Routing workloads and method thereof
Abstract
This record has no abstract on file.
Term
2.3 yearsleft in the term
Expires 28 January 2029.
- Priority
- Filed
- Granted
- Today
- Expires
8 claims: 2 independent, 6 dependent
- 1ワークロード・マネージャ(101)において 、ディスパッチャがアービトレータから受け取るワークロードのシェアを示す、 ディスパッチャのシェア(D)を計算する方法であって、 前記ワークロード・マネージャ(101)はアービトレータ(102)に結合され、前記アービトレータ(102)は複数のシステム(117~119)に結合され、各システムはディスパッチャ(103、104又は105)を備え、各ディスパッチャ(103、104又は105)は複数の実行ユニット(106~108、109~111又は112~114)に結合され、前記アービトレータ(102)はワークロード項目(115)のフローを受け取り且つこれを前記ディスパッチャ(103~105)に配分するのに適合しており、前記実行ユニット(106~114)は前記ワークロード項目(115)を実行するのに適合しており、前記ワークロード項目(115)は少なくとも第1~第3のワークロード・タイプを有しており、 前記方法は、 前記 複 数のシステム(117~119)にわたる各ワークロード・タイプ用のサービス単位の合計値(W)を 獲得する ステップを有し、前記サービス単位は時間間隔でCPU消費量を 示す 値であり、 前記 複 数のシステム(117~119)の各システム上の各ワークロード・タイプ用のキャパシティ値(cap)を 獲得する ステップを有し、前記キャパシティ値(cap)は前記時間間隔でシステムが実行することができる最大サービス単位を示し、 各システム上の各ワークロード・タイプの前記キャパシティ値(cap)を各ワークロード・タイプの前記サービス単位の合計値(W)で除算することにより、各システムのディスパッチャの相対シェア(R)を計算するとともに、各システムの当該ディスパッチャの相対シェア(R)のうち最小値を獲得するステップを有し、 各システム上の各ワークロード・タイプ用の前記ワークロード項目(115)のキュー長(q)を各システム上の各ワークロード・タイプの前記キャパシティ値(cap)で除算することにより、各システム用の相対キュー長(V)を計算するステップを有し、 前記最小値と(1+最大の前記相対キュー長(V))の逆数値との乗算 により、各システム用の前記ディスパッチャのシェア(D)を計算するステップを有する方法。
- 2前記相対キュー長(V)の前記第1の関数(427)は、前記相対キュー長(V)の最大値の逆数値を計算する、請求項1に記載の方法。
- 3前記算術演算は乗算である、請求項1又は請求項2に記載の方法。
- 4前記ディスパッチャのシェア(D)は、当該ディスパッチャのシェア(D)を全てのシステムの全てのディスパッチャの和で除算することにより正規化される、請求項1ないし請求項3の何れか1項に記載の方法。
- 5前記最小値は、当該最小値を全てのシステムの全ての最小値の和で除算することにより正規化される、請求項1ないし請求項4の何れか1項に記載の方法。
- 6前記アービトレータは、前記ワークロード項目(115)のフローを 前記ディスパッチャ(103、104又は105) に配分する、請求項1ないし請求項5の何れか1項に記載の方法。
- 7請求項1ないし請求項6の何れか1項に記載の方法の各ステップをコンピュータに実行させるためのコンピュータ・プログラム。
- 8ワークロード・マネージャ(101)において 、ディスパッチャがアービトレータから受け取るワークロードのシェアを示す、 ディスパッチャのシェア(D)を計算するためのデータ処理システムであって、 前記ワークロード・マネージャ(101)はアービトレータ(102)に結合され、前記アービトレータ(102)は複数のシステム(117~119)に結合され、各システムはディスパッチャ(103、104又は105)を備え、各ディスパッチャ(103、104又は105)は複数の実行ユニット(106~108、109~111又は112~114)に結合され、前記アービトレータ(102)はワークロード項目(115)のフローを受け取り且つこれを前記ディスパッチャ(103~105)に配分するのに適合しており、前記実行ユニット(106~114)は前記ワークロード項目(115)を実行するのに適合しており、前記ワークロード項目(115)は少なくとも第1~第3のワークロード・タイプを有しており、 前記データ処理システムは、 前記 複 数のシステム(117~119)にわたる各ワークロード・タイプ用のサービス単位の合計値(W)を 獲得する ためのコンポーネントを備え、前記サービス単位は時間間隔でCPU消費量を 示す 値であり、 前記 複 数のシステム(117~119)の各システム上の各ワークロード・タイプ用のキャパシティ値(cap)を 獲得する ためのコンポーネントを備え、前記キャパシティ値(cap)は前記時間間隔でシステムが実行することができる最大サービス単位を示し、 各システム上の各ワークロード・タイプの前記キャパシティ値(cap)を各ワークロード・タイプの前記サービス単位の合計値(W)で除算することにより、各システムのディスパッチャの相対シェア(R)を計算するとともに、各システムの当該ディスパッチャの相対シェア(R)のうち最小値を獲得するためのコンポーネントを備え、 各システム上の各ワークロード・タイプ用の前記ワークロード項目(115)のキュー長(q)を各システム上の各ワークロード・タイプの前記キャパシティ値(cap)で除算することにより、各システム用の相対キュー長(V)を計算するためのコンポーネントを備え、 前記最小値と(1+最大の前記相対キュー長(V))の逆数値との乗算 により、各システム用の前記ディスパッチャのシェア(D)を計算するためのコンポーネントを備える、データ処理システム。
Independent claims8
34 paragraphs, as filed
The present invention relates to a method and a computer program for calculating a routing workload in a workload manager.
Mainframes are computers used primarily by large organizations to run critical applications and process large amounts of data such as financial transaction processing. Mainframes have a high degree of redundancy to provide a reliable and secure system. Since mainframes can run or host multiple operating systems, replacing a large number of small servers in service with mainframes can reduce administrative burden and provide improved scalability. Recent mainframes include IBM's zSeries and system z9 servers.
A Parallel Sysplex is a cluster of multiple IBM mainframes that work together as a single system image using z / OS. Sysplex clusters up to 32 systems and spans these systems by allowing sharing of read / write data across multiple systems while using concurrency and maintaining complete data integrity. Share one workload. This workload can be dynamically distributed across multiple individual processors in a system, or distributed to any system in a cluster with available resources. Workload balancing also makes it possible to run a variety of applications across Parallel Sysplex clusters while maintaining critical response levels. If workload balancing or workload routing is not done correctly, overloading can occur in the system.
<p> Therefore, there is a need for a method for efficiently calculating the dispatcher's share in a workload manager, as well as a workload manager and computer program suitable for performing the method.</p>
<p> The present invention provides a method of calculating the dispatcher share (D) in a workload manager. The workload manager is combined into one arbitrator, the arbitrator is combined into multiple systems, each system has one dispatcher, each dispatcher is combined into multiple execution units, and the arbitrator is of a workload item. It is adapted to receive the flow and distribute it to the dispatcher, the execution unit is adapted to execute the workload item, and the workload item is at least the first to third workloads. -Has a type. The method comprises reading the total value (W) of service units for each workload type across the plurality of systems from memory, for the service units to measure CPU consumption during one time interval. The value of, which further has a step of reading the capacity (cap) value for each workload type on each system from said memory, said capacity value is the maximum service unit that one system can perform. Is shown.</p><p> Further, the method divides the capacity value of each workload type on each system by the total value of the service units of each workload type to obtain the relative share (R) of the dispatcher of each system. Along with the calculation, the step of obtaining the minimum value of the relative share (R) of the dispatcher of each system and the queue length (q) of the workload item for each workload type on each system are calculated on each system. To calculate the relative queue length (V) for each system by dividing by the capacity value for each workload type in, and using arithmetic operations to the minimum value and the relative queue length (V). ) To calculate the share (D) of the dispatcher for each system by combining the first function of).</p><p> The dispatcher share refers to the dispatcher share for all routes, that is, the workload share that all dispatchers receive from the arbitrator. The relative share of the dispatcher refers to the relative share of the dispatcher with the minimum relative execution capacity.</p><p> The first function of the relative queue length of the method comprises calculating the inverse of the maximum value of the relative queue length. The arithmetic operation is multiplication. The first to third workload types are general CPU (CP), z Application Assist Processor (zAAP) and z9 Integrated Information Processor (zIIP) workload types. The dispatcher's share is normalized by dividing the dispatcher's share by the sum of all dispatchers in all systems.</p><p> The minimum value of the method is normalized by dividing the minimum value by the sum of all the minimum values of all systems. The arbitrator distributes the flow of the workload item to other dispatchers in other systems.</p><p> In another aspect, the invention relates to a computer program for causing a computer to perform each step of the method.</p><p> In another aspect, the invention relates to a data processing system for calculating dispatcher share in a workload manager.</p>
<p> One of the advantages of the embodiment is insufficient capacity or outage of one execution unit as long as the total capacity of the sysplex still exceeds the total workload demand for all workload types. As a result, the queues in the dispatcher and execution unit do not grow indefinitely. In these cases, the dispatcher share of all routing to each system is relative to the queue length relative to its capacity on any processor and the sum of the service units of one workload type to that capacity. Depends on. When one execution unit on one system goes down, the execution unit is distributed by allocating the workload to the other system without resulting in overflow or congestion of resources in the other system. It is possible to compensate. Other advantages are that the dispatcher share calculation is independent of the initial workload allocation and that the algorithm concentrates on the optimal workload allocation.</p>
<figref num="1">It is a block diagram of Sysplex 100 according to the embodiment of this invention.</figref><figref num="2">FIG. 6 is a block diagram of a plurality of systems interconnected to form a sysplex 200 according to an embodiment of the present invention.</figref><figref num="3">It is a flowchart of the method of calculating the share of a dispatcher according to the embodiment of this invention.</figref><figref num="4">It is a block diagram of the calculation of the workload manager according to the embodiment of this invention.</figref>
The sysplex 100 of FIG. 1 comprises a first system 118, the first workload manager 101 of which is coupled to an arbitrator 102. The arbitrator 102 is coupled to the first to third dispatchers 103 to 105. The first workload manager 101 is combined with the other two workload managers 120 and 121. The first dispatcher 103 is provided in the second system 117 and is coupled to three execution units 106 to 108. The second dispatcher 104 is provided in the first system 118 and is coupled to the three execution units 109 to 111. The third dispatcher 105 is provided in the third system 119 and is coupled to the other three execution units 112 to 114. Memory 116 is coupled to workload manager 101.
The arbitrator 102 receives a plurality of incoming workload items 115 and sends them to dispatchers 103 to 105. Dispatchers 103-105 determine the workload types for the workload item 115 and send them to specialized execution units 106-114. The queue length of the plurality of workload items 115 can be measured for each dispatcher and each execution unit. For example, dispatcher 103 has a queue length of 5 workload items (q).<sub>D1</sub>), And the execution unit 106 has a queue length (q) of 3 workload items.<sub>E1,1</sub>). It is not necessary to measure the queue length for every dispatcher or every execution unit. The queue length must be measured at at least one stage of routing behind the arbitrator 102. This is the only requirement of the embodiments of the present invention.
Each workload item requires different CPU consumption of one execution unit during one time interval, which consumption is not known in advance to the arbitrator 102. Therefore, the allocation of the plurality of workload items 115 to the dispatchers 103 to 105 does not depend on the size of each workload item 115 or its CPU consumption. The workload type of each workload item 115 is identified by dispatchers 103-105, and the size of workload item 115 is identified by dispatchers 103-105 or execution units 106-114. As mentioned above, the arbitrator 102 does not know in advance the workload type of workload item 115 and the size of workload item 115, so the wrong algorithm can grow one queue indefinitely. The workload managers 101, 120 and 121 of all systems 117-119 communicate the capacity and workload values of the system by being coupled and interacting with each other on a regular basis. The memory 116 stores the total value of the service unit of all the workload types across the plurality of systems and the capacity value of all the workload types of each system.
The sysplex 200 illustrated in FIG. 2 is formed by interconnected first to fourth systems 201 to 204. Systems 201-204 include workload managers 205-208, all of which are coupled to each other. Workload managers 205 and 206 are coupled to arbitrators 209 and 210. Alternatively, it is sufficient to have one arbitrator within the sysplex 200. The arbitrators 209 and 210 are coupled to a plurality of dispatchers 213 to 216, while the dispatchers 213 to 216 are coupled to execution units 217 to 220, respectively. Not all systems 201-204 require one arbitrator. This is because the arbitrator can send the workload to any system in the sysplex 200.
Workload managers 205-208 handle routing algorithms, respectively. This routing algorithm calculates the dispatcher's share of the relative amount of workload items that one dispatcher should receive. Arbitrators 209 and 210 do not know the workload type of the workload item and the size of the workload item. Therefore, the arbitrators 209 and 210 receive routing recommendations from the workload managers 205-208 and allocate the workload items according to the result of the routing algorithm.
Dispatchers 213 to 216, also known as queue managers, receive workload items from arbitrators 209 and 210 and queue them until they are retrieved by execution units 217 to 220, also known as servers. Execution units 217-220 execute the workload item, read the workload type, and determine which processor can handle the workload item. The workload type can be three or more types, including the general CPU (CP), z Application Assist Processor (zAAP), and z9 Integrated Information Processor (zIIP) workload types. A different processor is used for each workload type. Workload managers 205-208 are coupled together to receive information related to dispatcher status across all systems 201-204.
FIG. 3 shows a flowchart 300 of how to calculate the dispatcher share (D) in the workload manager. The first step 301 obtains the total value (W) of service units for each workload type across multiple systems during the ongoing time interval. The service unit is the CPU consumption during one time interval.<u style="single">Show</u>The value. The total service unit value (W) is the first service unit value for the first workload type, the second service unit value for the second workload type, and the third workload type. Includes a third service unit value for the type. If more workload types are available in the system, the corresponding sum of service units (W) is obtained.
The second step 302 obtains one capacity value (cap) for each workload type, including first to third workload types on each system of the plurality of systems. The capacity value (cap) is the maximum amount of service units that one system can perform during one time interval.<u style="single">Show</u>.. Different capacity values (caps) are acquired for each workload type and for each system. Therefore, each system will contain one capacity value (cap) for each workload type.
The third step 303 calculates the relative share (R) of the dispatcher. This relative share (R) is obtained by dividing the capacity value (cap) for each workload type on each system by the total value (W) of service units for each workload type. .. The capacity value (cap) obtained for each workload type on each system is divided by the total value (W) of service units of the same workload type, and the minimum value is the dispatcher for each system. Corresponds to the relative share of. In the third step 303, a single value is obtained for each system.
The fourth step 304 calculates the relative queue length (V) from each system. This relative queue length (V) divides the queue length (q) of the workload item for each workload type on each system by the capacity value (cap) of each workload type on each system. Acquired by. Each system gets one value for each workload type.
Fifth step 305 calculates the dispatcher share (D) for each system by combining the first function of the minimum value and the relative queue length (V) using arithmetic operations.
Figure 4 illustrates the calculation of dispatcher share (D) to distribute the flow of workload items across multiple dispatchers. The first table 400 includes the first to third systems 401 to 403, the first to third workload types 404 to 406, the total value of service units (W) 407, and a plurality on each system. Capacity values (caps), i.e., multiple first capacity values 408 on the first system 401, multiple second capacity values 409 on the second system 402, and multiple second capacity values 409 on the third system 403. Includes multiple third capacity values of 410.
The second table 420 shows multiple queue lengths (q) 421-423 on each system and for each workload type and calculated relative queue lengths (V) 424 on each system and for each workload type. Includes ~ 426 and the relative queue length function 427 on each system and for each workload type.
One of the first steps in calculating the dispatcher share (D) in the workload manager is to get the total value (W) for each workload type service unit across all systems 401-403. Including doing. In this example, the total value of service units for the first workload type on the first system 401 is "300" and the total value of service units for the second workload type is "500". And the total value of service units for the third workload type is "10".
Other steps include obtaining capacity values (caps) for each system and for each workload type. For example, on the first system 401, the capacity value for the first workload type is "90", the capacity value for the second workload type is "100", and the third The capacity value for the workload type is "10". On the second system 402, the capacity value for the first workload type is "200", the capacity value for the second workload type is "400", and the third work. The capacity value for the load type is "100". After the capacity values have been obtained for all systems, the relative share (R) of the dispatcher for each workload type can be calculated on each system.
On the first system 401, the first workload that the system 401 can handle has a capacity value of "90" for the first workload type and a service unit of the first workload type. It can be obtained by dividing by the total value of "300" and deriving "0.3" as a result. Dividing the second capacity value "100" of the second workload type by the total value "500" of the service units of the second workload type gives system 401 a second that can be processed. The workload will be "0.2". This value is a percentage of the actual workload that System 401 can handle. Repeating the same process for the third workload type results in a "1". Also, repeating the same process across the other systems 402 and 403 yields three different values for each system.
Another step in the method of calculating the dispatcher share involves calculating the minimum value of the previously acquired relative share (R) of the dispatcher. For example, on the first system 401, there are three values of "0.3", "0.2", and "1" as the relative share (R) of the dispatcher. The minimum of these three values is "0.2". For the second system 402, the minimum relative share (R) of the dispatcher is "0.6667", and for the third system 403, the minimum relative share (R) of the dispatcher is "1".
The second table 420 illustrates the steps of the method of calculating the dispatcher share (D). That is, these steps calculate the relative queue lengths (V) 424-426 by dividing the queue lengths (q) 421-423 by the capacity values (cap) 408-410 for a particular workload type. Including doing. On the first system 401, the queue length 421 for the first workload type is "125", the queue length 421 for the second workload type is "89", and the third work. The queue length 421 for the load type is "67". On the first system 401, the first relative queue length 424 is obtained as "1.388" by dividing "125" by "90", and the second relative queue length 424 is "89" divided by "90". Obtained as "0.89" by dividing by "100". If the same process is repeated for all queue lengths of all workload types across all systems 401-403, the third relative queue length 424 of the first system 401 is acquired as "6.7".
In addition, the second table 420 contains the first function 427 with relative queue lengths (V) 424-426. This function is for getting the inverse of "1 + maximum relative queue length". On the first system 401, the maximum relative queue length is "6.7", so the inverse of that value + 1 is "0.1299". On the second system 402, the maximum relative queue length is "3.655", so the inverse of that value + 1 is "0.2148". On the third system 403, the maximum relative queue length is "4.3564", so the inverse of that value + 1 is "0.1867". Once all these values have been calculated, the minimum value of the dispatcher's relative share (R) for each system is multiplied by the first function value for each system in order to obtain the dispatcher's share (D). To. Therefore, for the first system 401, "0.02598" is obtained by multiplying "0.2" and "0.1299". For the second system 402, "0.1432" is obtained by multiplying "0.6667" and "0.2148".
Finally, for the third system 403, "0.1867" is obtained by multiplying "1" and "0.1867". These three values indicate the dispatcher's share for each system, allowing the arbitrator to send different amounts of workload units for each dispatcher. As a result, the best possible allocation of workload items can be obtained across all dispatchers and all systems. After a predetermined time interval, adjust the dispatcher share (D) to the new optimum by repeating the calculation, taking into account any changes in workload W, capacity and queue length. can do.
The present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment that includes both hardware and software elements. In a preferred embodiment, the invention is implemented in the form of software (including firmware, resident software, microcode, etc.).
Further, the present invention takes the form of a computer program accessible from a computer-enabled or computer-readable medium that provides program code for use by or in connection with a computer or any instruction execution system. Can be done. In this regard, computer-enabled or computer-readable media retains, stores, communicates, propagates or transports programs for use by or in connection with instruction execution systems, devices or devices. It can be any device that can be.
Such media can be electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems (devices or devices) or propagation media. Examples of computer-readable media include semiconductor or solid-state memory, magnetic tape, removable flexible disks, random access memory (RAM), read-only memory (ROM), fixed magnetic disks and optical disks. Current examples of optical discs include compact disc read-only memory (CD-ROM), compact disc read / write (CD-R / W) and DVD.
A suitable data processing system for storing and / or executing program code would include at least one processor directly or indirectly coupled to a memory element through the system bus. Such memory elements are used to reduce the number of times the local memory used during the actual execution of the program code and the mass storage, and the code must be retrieved from the mass storage during execution. It can include cache memory that provides temporary storage of code.
I / O devices (including keyboards, displays, pointing devices, etc.) can be coupled to the system either directly or via an I / O controller. Also, when the network adapter is coupled to a data processing system, the data processing system can be coupled to other data processing systems, remote printers, storage devices, etc. via a private network or a public network. Modems, cable modems and Ethernet® cards are currently available types of network adapters.
Although the specific embodiments of the present invention have been described above, it is clear that various changes and modifications can be made in these embodiments. The scope of the present invention is defined by the description of each claim.
101, 120, 121 ... Workload Manager 102 Arbitrator 103 ~ 105 Dispatcher 106 ~ 114 Execution unit 115 Workload items 200 ... Sysplex 201 ~ 204 System 205 ~ 208 Workload Manager 209 ~ 212 Arbitrator 213 ~ 216 Dispatcher
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| JP200648680A | Cites | Japan |
| JP200571031A | Cites | Japan |
| JP2000268012A | Cites | Japan |
| JP6243112A | Cites | Japan |
| JP2000259591A | Cites | Japan |
| JP200430663A | Cites | Japan |
| JP200847126A | Cites | Japan |
| JP4318655A | Cites | Japan |
14 members in 6 offices
Priority claims7
| Document | Office | Kind | Date |
|---|---|---|---|
| 08151917 | European Patent Office (EPO) | A | |
| 081519175 | European Patent Office (EPO) | – | |
| 2009050914 | European Patent Office (EPO) | W | |
| 200808151917 | – | – | – |
| 2009050914 | – | – | – |
| EP20080151917 | – | – | – |
| WO2009EP50914 | – | – | – |
Members14
| Document | Office | Kind | |
|---|---|---|---|
| US2009217288A1 | United States of America | A1 | |
| WO2009106398A1 | World Intellectual Property Organization (WIPO) | A1 | |
| EP2255286A1 | European Patent Office (EPO) | A1 | |
| KR20100138885A | Republic of Korea | A | |
| CN101960428A | China | A | |
| JP2011513807A | Japan | A | |
| JP4959845B2This record | Japan | B2 | |
| US8245238B2 | United States of America | B2 | |
| US2012291044A1 | United States of America | A1 | |
| CN101960428B | China | B | |
| US8875153B2 | United States of America | B2 | |
| US2015040138A1 | United States of America | A1 | |
| EP2255286B1 | European Patent Office (EPO) | B1 | |
| US9582338B2 | United States of America | B2 |
12 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on accelerated examinationJAPANESE INTERMEDIATE CODE: A971005A975 | A975 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 | |
| Explanation of circumstances concerning accelerated examinationJAPANESE INTERMEDIATE CODE: A871A871 | A871 |
Numbers
- Publication
- 4959845
- Publication, DOCDB
- 4959845
- Publication, EPODOC
- JP4959845B
- Application
- 2010547130
- Application, DOCDB
- 2010547130
- Application, EPODOC
- JP20100547130
Titles2
- Japanese
- ワークロード・マネージャにおいてディスパッチャのシェアを計算する方法、コンピュータ・プログラム及びデータ処理システム
- English
- How to calculate dispatcher share in workload managers, computer programs and data processing systems
Classification
- CPC, 6
- G06F9/5083
- G06F9/505
- H04L67/1008
- G06F9/4881
- H04L67/1001
- H04L67/1002
- IPC, 1
- G06F9 50