System and method for managing distributed system
Abstract
[Purpose] Even when the management device breaks down or the load is high and the functions cannot be fully exerted, the managed device performs management-related processing and enables the operation of the distributed system. [Constitution] Information about a computer capable of alternative processing is downloaded from the management device 100 to the managed device 200 in advance, and when the load cannot be properly distributed, such as when the management device 100 is out of order, the managed device 200 performs the alternative processing. Information such as load is collected from the managed device on a possible computer, the information is evaluated to determine an alternative computer, and the load is assigned to the computer.

Term
Term ended
Projected expiry passed 21 October 2014, 11.9 years ago.
- Priority and filed
- Published
- Projected expiry
- Today
8 claims: 8 independent, 0 dependent
- 1【特許請求の範囲】 【請求項1】 ネットワークに接続された各計算機のうち、少なくとも1つの計算機にシステム運用情報を保持する管理アプリケーションを有し、他の計算機に上記管理アプリケーションの指示に基づき計算機を管理する被管理アプリケーションを有した分散システム管理方式において、上記被管理アプリケーションは、代替動作が可能な他の計算機の情報を保持する代替システム管理手段と、この代替システム管理手段に保持された情報に基づき代替動作の候補となる各計算機の負荷情報を入手し評価して依頼先を決定する代替動作依頼先決定手段と、他の被管理アプリケーションからの依頼に基づき代替動作を行なう代替処理手段とを備えたことを特徴とする分散システム管理方式。
- 2【請求項2】 ネットワークに接続された各計算機のうち、少なくとも1つの計算機にシステム運用情報を保持する管理アプリケーションを有し、他の計算機に上記管理アプリケーションの指示に基づき計算機を管理する被管理アプリケーションを有した分散システム管理方式において、上記管理アプリケーションは、上記被管理アプリケーションから計算機の負荷情報を入手しタスクの配分先を決定し分配する負荷分配手段を備え、上記被管理アプリケーションは、計算機の稼動状況を調べて上記管理アプリケーションに報告すると共に上記管理アプリケーションからタスク配分決定の通知を受けるまで稼動状況をロックする管理対象モニタ制御手段を備えたことを特徴とする分散システム管理方式。
- 3【請求項3】 ネットワークに接続された各計算機のうち、少なくとも1つの計算機にシステム運用情報を保持する管理アプリケーションを有し、他の計算機に上記管理アプリケーションの指示に基づき計算機を管理する被管理アプリケーションを有した分散システム管理方式において、上記被管理アプリケーションは、他の上記被管理アプリケーションから計算機の負荷情報を入手しタスクの配分先を決定し分配する負荷分配手段と、計算機の稼動状況を調べて他の上記被管理アプリケーションに報告すると共に他の上記被管理アプリケーションからタスク配分決定の通知を受けるまで稼動状況をロックする管理対象モニタ制御手段とを備えたことを特徴とする分散システム管理方式。
- 4【請求項4】 ネットワークに接続された各計算機のうち、少なくとも1つの計算機にシステム運用情報を保持する管理アプリケーションを有し、他の計算機に上記管理アプリケーションの指示に基づき計算機を管理する被管理アプリケーションを有した分散システム管理方式において、上記被管理アプリケーションは、収集すべき管理情報の内容を定義した管理情報定義と、上記計算機の通信トラフィックを監視するトラフィック監視手段とを備え、上記通信トラフィックの量に応じて上記管理情報の内容を変更することを特徴とする分散システム管理方式。
- 5【請求項5】 ネットワークに接続された各計算機のうち、少なくとも1つの計算機にシステム運用情報を保持する管理アプリケーションを有し、他の計算機に上記管理アプリケーションの指示に基づき計算機を管理する被管理アプリケーションを有し、この被管理アプリケーションは以下のステップにより処理を行なうことを特徴とする分散システム管理方法。 (a)代替処理が必要となった場合、代替処理を行なうタスクと計算機を調べる。 (b)代替処理が可能な計算機上の他の被管理アプリケーションより代替処理をするための情報を入手する。 (c)上記代替動作をするための情報を評価して代替動作の依頼先を決定する。 (d)決定した依頼先の被管理アプリケーションにタスクの代替動作を依頼する。
- 6【請求項6】 ネットワークに接続された各計算機のうち、少なくとも1つの計算機にシステム運用情報を保持する管理アプリケーションを有し、他の計算機に上記管理アプリケーションの指示に基づき計算機を管理する被管理アプリケーションを有し、以下のステップにより処理を行なうことを特徴とする分散システム管理方法。 (a)上記管理アプリケーションは、負荷の再配分が必要になった場合、上記被管理アプリケーションに対し稼動状況を問い合わせる。 (b)上記被管理アプリケーションは、稼動状況を調べて上記管理アプリケーションに通知すると共に、上記管理アプリケーションより負荷の再配分の終了を通知されるまで稼動状況をロックする。 (c)上記管理アプリケーションは、上記各被管理アプリケーションからの稼動状況を評価し、負荷を割り当てる計算機を決定し、負荷割り当てが決定した計算機の上記被管理アプリケーションに通知する。 (d)負荷割り当てを受けた上記被管理アプリケーションは、上記管理アプリケーションに負荷割り当てを受けたことを通知する。 (e)管理アプリケーションは、稼動状況を問い合わせた上記被管理アプリケーションに負荷の割り当てが終了したことを通知する。 (f)上記被管理アプリケーションは、稼動状況のロックを解除し新たな負荷の受付を可能にする。
- 7【請求項7】 ネットワークに接続された各計算機のうち、少なくとも1つの計算機にシステム運用情報を保持する管理アプリケーションを有し、他の計算機に上記管理アプリケーションの指示に基づき計算機を管理する被管理アプリケーションを有し、この被管理アプリケーションは以下のステップにより処理を行なうことを特徴とする分散システム管理方法。 (a)代替処理が必要となった場合、依頼元の上記被管理アプリケーションは、代替処理を行なうタスクと計算機を調べ、代替処理が可能な計算機上の依頼先候補の被管理アプリケーションに対し稼動状況を問い合わせる。 (b)依頼先候補の被管理アプリケーションは、稼動状況を調べて上記依頼元の被管理アプリケーションに通知すると共に、上記依頼元の被管理アプリケーションから代替処理の終了を通知されるまで稼動状況をロックする。 (c)上記依頼元の被管理アプリケーションは、上記依頼先候補の各被管理アプリケーションからの稼動状況を評価し、負荷を割り当てる計算機を決定し、負荷割り当てが決定した計算機の依頼先の上記被管理アプリケーションに通知する。 (d)上記依頼先の被管理アプリケーションは、上記代替処理のタスクを起動し、依頼処理の終了を上記依頼元の被管理アプリケーションに通知する。 (e)上記依頼元の被管理アプリケーションは、稼動状況を問い合わせた上記依頼先候補の被管理アプリケーションに依頼処理の終了を通知する。 (f)上記依頼先候補の被管理アプリケーションは、稼動状況のロックを解除し新たな代替処理の受付を可能にする。
- 8【請求項8】 ネットワークに接続された各計算機のうち、少なくとも1つの計算機にシステム運用情報を保持する管理アプリケーションを有し、他の計算機に上記管理アプリケーションの指示に基づき計算機を管理する被管理アプリケーションを有し、この被管理アプリケーションは、収集すべき管理情報の内容を定義した管理情報定義と、この管理情報定義の内容を変更する管理情報量変更手順と、通信トラフィックの量に応じて実行される上記管理情報量変更手段との関係を記述した管理情報調整表とを備え、以下のステップにより処理を行なうことを特徴とする分散システム管理方法。 (a)通信トラフィックを監視しその変動を検出する。 (b)上記管理情報量調整表を検索し上記通信トラフィックの変動に対応した上記管理情報変更手順を決定する。 (c)決定した上記管理情報量変更手順により上記管理情報定義の内容を変更する。
Independent claims8
159 paragraphs, as filed
Description: TECHNICAL FIELD [Detailed description of the invention]
【0001】
[Industrial application field]
The present invention relates to a method and a method for managing a distributed processing system by a plurality of computers.
【0002】
[Conventional technology]
FIG. 16 shows the configuration of the load balancer and the transaction processing system disclosed in Japanese Patent Application Laid-Open No. 4-229356. In FIG. 16, the rear-end computers CP1 to CP3 (12 to 16) share the database 18 and execute the transaction business. The front-end computer 10 collects CPU usage information as feedback information from the rear-end computers CP1 to CP3 (12 to 16) at regular time intervals. The front-end computer 10 holds a destination specification table 11 internally, and can control the destination of transaction processing sent to the rear-end computers 12 to 16.
【0003】
From the above CPU utilization information, if the load on these rear-end computers becomes overloaded, the front-end computer modifies this destination table and redirects the transaction to another lightly-loaded rear-end computer. By allocating, the overload can be eliminated.
【0004】
[Problems to be Solved by the Invention]
Since the conventional distributed system management method is configured as described above, in order to distribute the load, it is necessary to centrally manage the CPU usage information, which is a prerequisite for the load, on the front-end computer 10. Therefore, when the front-end computer 10 is stopped due to a failure, there is a problem that important management work such as transaction allocation work is completely stopped.
【0005】
Also, if the information about the load status of the rear-end computers 12 to 16 and whether or not they are operating is delayed due to an increase in the network load, the load status and operating status of the rear-end computers 12 to 16 managed by the front-end computer 10 are delayed. Information about the actual load and operating status of the rear-end computers 12 to 16 may not be synchronized, and the load may be distributed to the wrong rear-end computer.
【0006】
Further, business-related communication between the rear-end computers 12 to 16 and the front-end computer 10 and business applications in the rear-end computers 12 to 16 are performed by the rear-end computers 12 to 16 and the front-end computer 10 used in distributed system management. Since it conflicts with communication between computers and monitoring processing in the rear end computers 12 to 16, conventional systems require distributed system management in response to an increase in communication traffic and an increase in the load on the rear end computers 12 to 16. It was not possible to change the load immediately, and it was difficult to obtain the optimum throughput.
【0007】
The present invention has been made to solve the above-mentioned problems, and even when the management device breaks down or the load is high and the functions cannot be sufficiently exhibited, the managed device side performs management-related processing. By doing so, the purpose is to enable operation without losing the management function of the distributed processing system.
【0008】
Another object of the present invention is to eliminate the time lag between the information received by the management device from the managed device and the information held by the managed device, and to provide the managed device with appropriate processing by the management device.
【0009】
Further, when the management is performed by the distributed management device, it is intended to enable accurate and efficient operation by having a mechanism for holding the management information synchronized between the distributed management devices.
【0010】
By having a mechanism to switch the level of detail of the information collected by the management device at an appropriate timing, the communication between the business application and the distributed system management and the processing of the management device are changed, and the load of the distributed system management is appropriate. The purpose is to reduce the amount to a large amount and improve the throughput of business applications.
【0011】
[Means for solving problems]
In the distributed system management method according to the present invention, the managed application can use an alternative system management means for holding information on other computers capable of alternative operation and a candidate for alternative operation based on the information held in the alternative system management means. It is provided with an alternative operation request destination determination means for obtaining and evaluating the load information of each computer to determine a request destination, and an alternative processing means for performing an alternative operation based on a request from another managed application.
【0012】
The distributed system management method according to the present invention includes a load distribution means in which a management application obtains computer load information from a managed application, determines a task allocation destination, and distributes the tasks, and the managed application checks the operating status of the computer. It is equipped with a managed monitor control means that locks the operating status until it reports to the management application and receives a notification of the task allocation decision from the management application.
【0013】
In the distributed system management method according to the present invention, the managed application obtains the load status of the computer from other managed applications, determines and distributes the task allocation destination, and examines the operating status of the computer. It is equipped with a managed monitor control means that locks the operating status until it reports to the managed application and receives a notification of the task allocation decision from another managed application.
【0014】
The distribution system management method according to the present invention includes a management information definition that defines the contents of management information to be collected by the managed application, and a traffic monitoring means for monitoring the communication traffic of the computer, according to the amount of communication traffic. The content of the management information is changed.
【0015】
In the distributed system management method according to the present invention, the managed application examines the task and the computer to perform the alternative processing, obtains information for performing the alternative processing from other managed applications of the computer capable of the alternative processing, and obtains the information. To determine the request destination of the alternative operation, and request the alternative operation of the task.
【0016】
In the distributed system management method according to the present invention, the management application inquires the managed application about the operating status, the managed application examines the operating status and notifies the management application, and the operating status is locked, and the managed application is sent from the managed application. It evaluates the operating status, determines the computer to allocate the load and notifies it, notifies the managed application that the notified application has received the load allocation to the management application, and informs the managed application that the management application inquired about the operating status. Notifies that the load allocation is complete, and the notified managed application unlocks the operating status.
【0017】
In the distributed system management method according to the present invention, the managed application of the requesting source inquires the managed application of the request destination candidate of the computer capable of alternative processing, and the managed application of the request destination candidate checks the operating status and requests. Notifies the original management application and locks the operation status, the requesting managed application evaluates the operation status from the request destination candidate managed application, determines the computer to allocate the load, notifies it, and receives the notification. The managed application of the request destination notifies the managed application of the request source that the load has been allocated, and the managed application of the request source inquires about the operating status of the managed application of the request destination. Notifies that this has been done, and the managed application that receives the notification unlocks the operating status.
【0018】
The distributed system management method according to the present invention includes a management information definition that defines the content of management information to be collected by the managed application, a procedure for changing the amount of management information that changes the content of this management information definition, and the amount of communication traffic. It is equipped with a management information amount adjustment table that describes the relationship with the management information amount change procedure executed in response, monitors communication traffic and detects the fluctuation, searches the management information amount adjustment table, and changes the communication traffic. Determine the corresponding management information amount change procedure, and change the contents of the management information definition according to the determined management information amount change procedure.
【0019】
[Action]
In the distributed system management method according to the present invention, the managed application holds information on other computers capable of alternative operation, and based on this information, obtains, evaluates, and requests the load information of each computer that is a candidate for alternative operation. Distribute the load on behalf of the management application by deciding ahead.
【0020】
In the distributed system management method according to the present invention, the management application obtains the load information of the computer from the managed application, determines the allocation destination of the task and distributes it, and the managed application checks the operating status of the computer and reports it to the management application. At the same time, by locking the operating status until the management application notifies the task allocation decision, the information of each computer is changed in synchronization.
【0021】
In the distributed system management method according to the present invention, the managed application obtains the load status of the computer from another managed application, determines the allocation destination of the task and distributes it, and the other managed application examines the operating status of the computer. By locking the operation status until the notification of the task allocation decision is received, the managed application distributes the load on behalf of the management application, and the information of each computer is changed in synchronization.
【0022】
In the distribution system management method according to the present invention, the managed application monitors the communication traffic of the computer and collects the management information according to the load by changing the content of the management information to be collected according to the amount of the communication traffic. To adjust.
【0023】
In the distributed system management method according to the present invention, the managed application examines the task and the computer to perform the alternative processing, obtains information for performing the alternative processing from other managed applications of the computer capable of the alternative processing, and obtains the information. Is evaluated to determine the request destination of the alternative operation, and the managed application distributes the load on behalf of the managed application by requesting the alternative operation of the task.
【0024】
In the distributed system management method according to the present invention, the management application inquires the managed application about the operating status, the managed application examines the operating status and notifies the management application, and the operating status is locked, and the managed application is sent from the managed application. After evaluating the operating status and deciding which computer to allocate the load to, the managed application unlocks the operating status, so that the information of each computer is changed in synchronization.
【0025】
In the distributed system management method according to the present invention, the managed application of the requesting source inquires the managed application of the request destination candidate of the computer capable of alternative processing, and the managed application of the request destination candidate checks the operating status and requests. After notifying the original management application and locking the operation status, the requesting managed application evaluates the operation status from the request destination candidate managed application and decides the computer to allocate the load, and then the managed application decides the computer to allocate the load. By unlocking the operating status, the managed application distributes the load on behalf of the managed application, and the information of each computer is changed in synchronization.
【0026】
In the distributed system management method according to the present invention, the managed application monitors the communication traffic, detects the fluctuation, searches the management information amount adjustment table, and determines the management information amount change procedure corresponding to the fluctuation of the communication traffic. By changing the content of the management information definition according to the determined management information amount change procedure, the collection of management information is adjusted according to the load.
【0027】
[Example]
Example 1. FIG. 1 shows an overall view of the distributed system management device according to the first embodiment. In the figure, 101 is a computer system connected by a network 400 to form a distributed system. The management device 100 provided in the computer system 101 is composed of a management application 110 and a management information communication means 300. The managed device 200 provided in the other computer system 101 is composed of the managed application 210 and the management information communication means 300. The management application 110 communicates with the managed application 210 by the management information communication means 300, collects the state of the computer system 101 in which the managed device 200 is operating, and uses the information to contact the managed device 200. And give instructions to change the state of the computer. The managed device 200 may communicate with another managed device 200 at the same time as communicating with the management device 100.
【0028】
FIG. 2 is a configuration diagram of the managed application 210 in the first embodiment. The managed application 210 normally refers to the management information definition 211, monitors the management target such as the computer system 101 and the network 400 by using the management target monitor control means 212, and the management data transmission means 213 is the management information communication means. The 300 is used to send the monitor results to another computer system 101. Here, the management information definition defines the type, frequency, detail, etc. of the information to be collected by the managed device. Further, the system operation information management means 218 downloads and holds the system operation information from the management application 110 of the management device 100, and the management target monitor control means 212 operates the task on the computer based on the download.
【0029】
Further, FIG. 2 has an alternative system management means 214 in which the managed application 210 itself holds information on another computer group capable of performing an alternative operation of the computer on which the managed application 210 is operating. In addition, there is a managed application-to-application communication means 215 for the managed application 210 to communicate with the managed application 210 operated by another computer group held by the alternative system management means 214. Further, there is an alternative operation request destination determining means 216 that obtains the state of another computer capable of performing the alternative operation by the managed application inter-communication means 215 and determines the managed application for which the alternative operation is requested based on the information. Further, there is an alternative processing means 217 that performs an alternative operation in response to a request from another managed application 210.
【0030】
Information about a group of computers capable of alternative operation is downloaded from the management device 100 to each managed device 200 by using the management information communication means 300 together with the system operation information. Figure 3 is an example of information about a computer that can perform alternative operations. In this embodiment, it is described that the application task "Business ABC" can be operated as an alternative with the computer names "Node3", "Node5", and "Node9". In addition, it is described that the addresses of the calculators are "131.141.51.10", "131.141.51.5", and "131.141.51.15". Figure 4 shows an example of system operation information. In this embodiment, it is described that the application task "Business ABC" operates with an operation plan such as a start time "8:00", an end time "18:00", and a date "Monday-Friday".
【0031】
FIG. 5 shows an example of an operation processing flow when the managed monitor control means 212 in the managed application 210 operates an application task on the computer system. First, based on the operation information downloaded to each managed device 200, the managed monitor control means 212 in the managed application 210 extracts the operation task and the operation time on the computer (procedure 501). Next, each information is set in the timer (step 502). Then, it enters the interrupt wait (step 503). When a certain interrupt occurs and exits the interrupt wait (procedure 503), the end interrupt judgment is entered (procedure 504). If it is an end interrupt, the process ends. If not, determine if it is a timer interrupt (step 505). If it is not a timer interrupt, other interrupt processing (procedure 506) is performed and the process returns to interrupt waiting (procedure 503). If it is a timer interrupt, start or stop the corresponding operation task (step 507). Then, the operation status of the task is monitored (procedure 508) to determine whether the operation plan can be maintained (procedure 509). If it can be maintained, return to waiting for an interrupt (step 503). If it cannot be maintained (for example, when the batch job is scheduled to be completed in a certain time at night but is likely to be exceeded), it is determined whether or not the management device 100 can be notified of the request for alternative processing (step 510). If notification is possible, the management device 100 is requested for alternative processing (step 511), and the process returns to interrupt waiting (step 503). If the management device 100 is out of order, communication with the management application 110 is impossible, or the management device 100 cannot perform the management function due to an overload, perform alternative processing of the task (step 512) and wait for an interrupt (step 512). Return to step 503).
【0032】
FIG. 6 is a flow of alternative processing in the managed application 210 in the first embodiment. First, the managed application 210 searches the alternative system management means 214 for information on the task to perform the alternative operation and the computer capable of the alternative operation (procedure 521). Next, the managed application 210 (alternative managed application) running on the computer capable of the alternative operation and the managed application communication means 215 communicate with each other, and the current status and the schedule change cost due to the alternative operation of the task are predicted. Receive parameters such as load increase (step 522). The information from the alternative managed application 210 is evaluated by the alternative operation request destination determination means 216 (step 523), and the request destination of the alternative operation is determined (procedure 524). Request the alternative managed application 210 of the determined request destination for the alternative operation of the task (step 525). Then, when the alternative processing means 217 in the alternative managed application 210 performs the alternative processing and the alternative task starts operating, the task of its own computer is stopped (procedure 526). Finally, if possible, the management device 100 is notified of the operation information regarding the task for which the alternative processing has been performed (procedure 527).
【0033】
According to the above embodiment, even if the management device 100 fails, the overload, the network failure, or the like makes it impossible to perform alternative processing of the task, the managed device 200 selects an appropriate alternative request destination from the alternative computer candidates. It is possible to select and perform alternative processing.
【0034】
In the above embodiment, the managed application 210 of the request source of the alternative process stops the task of its own computer when the alternative process of the task of the request destination starts, but the managed application 210 of the request destination It is also possible to stop the task of the own computer when the alternative processing is requested. Further, the managed application 210 of the requesting source notifies the management device 100 of the operation information of the task for which the alternative operation is performed, and this may be performed by the managed application 210 of the requesting destination.
【0035】
Example 2. The basic configuration of the distributed system management device of Example 2 is the same as that of FIG. 1 of Example 1. FIG. 7 is a configuration diagram showing the management application 110 in the second embodiment. The management application 110 usually includes a management information receiving means 116 for receiving management information from the managed application 210, and a management information storing / displaying means 117 for storing the management information and presenting it to the user. Further, in this example, a synchronous communication mechanism 111 that synchronizes and communicates with a plurality of managed applications, a synchronous information management unit 112 that performs synchronous communication and holds information synchronized with a plurality of managed applications, and synchronous information. It is composed of an information evaluation unit 113 that evaluates and determines the task distribution, a load distribution table 114 that manages the task distribution destination, and a load distribution unit 115 that distributes tasks based on the load distribution table 114.
【0036】
FIG. 8 is a configuration diagram showing the managed application 210 in the second embodiment. Within the managed application 210, there is a synchronous communication mechanism 219 for synchronizing and communicating with the management application 110. In addition, there is an operation status monitoring unit 221 that checks the operation status of the computer on which the managed application 210 is operating. These are controlled by the managed monitor control means 212, which receives inquiries about the operating status from the management application 110 via the synchronous communication mechanism 219, and examines and reports the operating status by the operating status monitoring unit 221.
【0037】
FIG. 9 shows a message flow between the management application 110 and the managed application 210. The processing flow of both applications will be described with reference to this figure. When the load (task) needs to be redistributed, the management application 110 inquires about the operation status of each managed application 210 (step 532). Upon receiving the inquiry, the managed application 210 uses its own operation status monitoring unit 221 to check the operation status (step 542) and notifies the management application 110 (step 543). Then, the managed application 210 locks the operating status without receiving a request for a new load until the management application 110 notifies that the load allocation work is completed (procedure 544). The management application 110 receives the operation status from the managed application 210 (procedure 533), determines the allocation for redistributing the load by the information evaluation unit 113 from the information (procedure 534), and loads the managed application 210. Submit the assignment (step 535). The assigned managed application 210 executes load distribution processing (task activation, etc.) (step 545), and sends a notification of load distribution completion to the management application when it is executed (step 546). The load distribution is completed when the load execution is correctly responded from the managed application 210 to the management application 110 to which the load is assigned (step 536). The management application 110 notifies all the inquired managed devices 210 that the load allocation work has been completed (step 537). When the management device 100 notifies that the load allocation work is completed, the managed application 210 unlocks the operating status (step 547) and makes it possible to accept a new load request (step). 548).
【0038】
When the management application 110 receives a response from the managed application 210 that tried to allocate the load that the load cannot be executed, the management application 110 determines another managed application 210 from the operating status of the remaining managed applications 210. And reallocate. Alternatively, it is conceivable to send the completion of the load allocation work to all the managed devices 200, end the work once, and restart from the operation status inquiry again.
【0039】
The management application 110 can exclude the managed application 210 that does not respond to the operation status inquiry even after a certain period of time from the load allocation target. In addition, if the managed application 210 that tried to allocate the load does not respond to whether the load can be executed within a certain period of time, it is possible to perform the same processing as receiving the response that the load cannot be executed. ..
【0040】
With such a management method, the operating status in response to the managed application 210 and the load distribution taken by the management application 110 can be surely synchronized. That is, when the operating status in which the managed application 210 responds is delayed due to a delay in the network 400 or an increase in the load on the management device 100, the operating status of the computer in the managed device 200 and the operating status of the management device 100 are conventionally used. The operating status to be grasped may be different, and the management device 100 may give an erroneous instruction, but this method will improve this problem.
【0041】
Example 3. The basic configuration of the distributed system management device of Example 3 is the same as that of FIG. 1 of Example 1. FIG. 10 is a configuration diagram showing the managed application 210 in the third embodiment. In this figure, there is an alternative system management means 214 in which the managed device 200 itself holds information on other computer groups that can perform alternative operations for the computer on which the managed application 210 is operating. Further, there is a synchronous managed application-to-application communication means 222 for synchronizing and communicating with the managed application 210 in which the managed application 210 operates in another computer group held by the alternative system management means 214. Further, there is an alternative operation request destination determining means 216 that obtains the state of the computer capable of performing the alternative operation by the synchronous inter-managed application communication means 222 and determines the managed application for which the alternative operation is requested based on the information.
【0042】
Information on the computer group capable of alternative operation is downloaded from the management application 110 to each managed application 210 together with the system operation information by using the management information communication means 300 as in the first embodiment. The operation processing flow when the managed application 210 operates the application task on the computer system is the same as that in the first embodiment.
【0043】
FIG. 11 is a flow of alternative processing in the third embodiment, and is composed of a flow of information exchange between a managed application of a request source and a request destination of alternative processing. First, the requested managed application 210 searches the alternative system management means 214 for information on the task to perform the alternative operation and the computer capable of the alternative operation (procedure 552). Next, it communicates with the managed application (alternative managed application) 210 running on the computer capable of the alternative operation by the synchronous inter-managed application communication means 222, and depends on the current status and the alternative operation of the task. Inquire about parameters such as schedule change cost and expected load increase. (Procedure 553).
【0044】
The managed application 210 receives an inquiry about the current operating status from the managed application 210 of a request source by the synchronous inter-managed application communication means 222, and acquires the operating status by the managed monitor control means 212 (procedure). 572) and respond to the managed application 210 that made the inquiry (step 573). Moreover, until the managed application 210 that made the inquiry notifies that the alternative processing has been completed, the lock is applied without accepting the alternative processing (response to the inquiry) from another managed application 210 (step 574). ).
【0045】
When the managed application 210 that issued the inquiry receives the operating status (step 554), the managed application 210 evaluates the information by the alternative operation request destination determination means 216 and determines the request destination of the alternative operation (procedure 555). Then, the determined request destination is requested to perform an alternative operation of the task (procedure 556). When the alternate managed application 210 receives an alternative request, it launches the alternate task (step 575), and when the alternate task launches correctly, it sends a completion notification to the requesting managed application 210 (step 576). When the requesting managed application 210 receives a notification from the request destination that the alternative request processing is completed, the alternative request processing is completed (step 557), and the alternative processing is performed for all the managed applications 210 that have made inquiries. Notify that it is complete (step 558). Then, the managed application notified of the completion of the alternative processing unlocks the operating status (procedure 577) and becomes in the alternative processing requestable state (procedure 578). In addition, the requested managed application 210 stops the alternative task (procedure 559), and notifies the management application 110 of the operation information regarding the task for which the alternative operation was finally performed (procedure 560).
【0046】
According to the above embodiment, communication between the managed applications 210 is delayed due to an increase in communication traffic, and information on the operating status between the managed application 210 of the request source of the alternative processing and the managed application 210 of the request destination. It is possible to eliminate the request for alternative processing to the wrong request destination due to the difference in the above.
【0047】
Example 4. The basic configuration of the distributed system management device of Example 4 is the same as that of FIG. 1 of Example 1. FIG. 12 is a configuration diagram showing the managed application 210 in the fourth embodiment. The managed application 210 normally refers to the management information definition 211, monitors a management target such as a computer system or a network by using the management target monitor control means 212, and the management data transmission means 213 uses the management information communication means 300. Use to send monitor results. Further, in this example, the managed application 210 includes a traffic monitoring unit 223 that monitors the communication traffic of the computer system 101 on which the managed application 210 operates, and a management information amount change procedure 224 that changes the contents of the management information definition 211. , Management information amount adjustment table 225 that describes the relationship with the management information amount change procedure 224 that is executed according to the amount of communication traffic, and management information amount adjustment table interpretation means that interprets and executes the management information amount adjustment table 225. It consists of 226.
【0048】
FIG. 13 shows the operation procedure of the management information amount adjustment table interpreting means 226 in the managed application 210. The management information amount adjustment table interpreting means 226 monitors the communication traffic by the traffic monitoring unit 223 (procedure 581). When the communication traffic fluctuates (step 582), the management information amount adjustment table interpreting means 226 searches the management information amount adjustment table 225 that associates the communication traffic with the management information amount change procedure 224 (step 583), and responds. Change the amount of management information to be performed Step 224 is executed to change the management information definition 211 (step 584).
【0049】
FIG. 14 shows an example of management information definition 211, and FIG. 15 shows an example of management information amount adjustment table 225. For example, when the traffic monitoring unit 223 notifies the management information amount adjustment table interpreting means 226 that the communication traffic has changed from 100 packets / sec to 500 packets / sec, the management information amount adjustment table interpreting means 226 manages the communication traffic. From the information amount adjustment table 225, start the procedure for changing the management information definition 211. In this case, the CPU utilization rate notification frequency is changed from once every 10 seconds to once every 30 seconds. Also, change the notification of process generation status from once every 30 seconds to once every 60 seconds. By changing the management information definition 211 in this way, the amount of management information sent from the managed application 210 to the management application 110 is reduced to 1/3 of the CPU utilization rate and the amount of management information related to the process generation status. Is reduced by half. Therefore, by reducing the communication of less urgent management information, it is expected to automatically avoid overwhelming the communication traffic of the more urgent business application and improve the throughput.
【0050】
Further, the above example avoids interference between the communication between the management application 110 and the managed application 210 and the communication of the business application, but the management application 110 monitors the state of the operating system and the management application 110. If you notify, there will be interference between the operating system monitor and the actual business application. Even in such a case, the monitor program that monitors the operating system status is set to automatically reduce the monitoring frequency and reduce the CPU usage rate by the monitor program when the CPU operating rate increases. Just do it.
【0051】
[Effect of the invention]
As described above, according to the present invention, the managed application holds information on other computers capable of alternative operation, and based on this information, obtains and evaluates load information of each computer that is a candidate for alternative operation. By deciding the request destination, the load can be distributed on behalf of the management application, and continuous operation becomes possible even when the computer having the management application is stopped due to a failure.
【0052】
Further, according to the present invention, the management application obtains the load information of the computer from the managed application, determines and distributes the task allocation destination, and the managed application examines the operating status of the computer and reports it to the management application. By locking the operating status until the management application notifies the task allocation decision, the information of each computer can be changed in synchronization, and the load can be distributed appropriately.
【0053】
Further, according to the present invention, the managed application obtains the load status of the computer from other managed applications, determines and distributes the task allocation destination, and the other managed application examines and reports the operating status of the computer. By locking the operation status until the notification of the task allocation decision is received, the managed application can distribute the load on behalf of the management application, and it continues even when the computer having the management application is stopped due to a failure. In addition to being able to operate, the information of each computer can be changed in synchronization, and the load can be distributed appropriately.
【0054】
According to the present invention, the managed application monitors the communication traffic of the computer and adjusts the collection of the management information according to the load by changing the content of the management information to be collected according to the amount of the communication traffic. It becomes possible to do. That is, while the load on the network or the computer is small, the items and frequency of management information are set to a large number by using the monitor result as a trigger, and detailed management information can be collected. In addition, when these loads increase, the items and frequency of management information are set to be small, which has the effect of suppressing the influence of the load of distributed system management on business applications and reducing the rate of reducing throughput.
【0055】
As described above, according to the present invention, the managed application investigates the task and the computer to perform the alternative processing, obtains the information for performing the alternative processing from the other managed application of the computer capable of the alternative processing, and obtains the information. By evaluating the information, deciding the request destination of the alternative operation, and requesting the alternative operation of the task, the managed application can distribute the load instead of the management application, and the computer having the management application fails. It will be possible to continue operation even when it is stopped.
【0056】
Further, according to the present invention, the management application inquires the managed application about the operating status, the managed application examines the operating status and notifies the managed application, and locks the operating status, so that the managed application operates from the managed application. After deciding which computer to allocate the load to, the managed application unlocks the operating status so that the information of each computer can be changed in synchronization, and the load can be distributed appropriately. It will be possible.
【0057】
Further, according to the present invention, the managed application of the requesting source inquires the managed application of the request destination candidate of the computer capable of alternative processing, and the managed application of the request destination candidate checks the operating status of the requesting source. After notifying the management application and locking the operation status, the managed application of the request source evaluates the operation status from the management application of the request destination candidate and decides the computer to allocate the load, and then the managed application operates the operation status. By unlocking, the managed application can distribute the load on behalf of the management application, and even if the computer with the management application is stopped due to a failure, it can be continuously operated and each computer can be operated continuously. The information can be changed in synchronization, and the load can be distributed appropriately.
【0058】
According to the present invention, the managed application monitors the communication traffic, detects the fluctuation thereof, searches the management information amount adjustment table, and determines and determines the management information amount change procedure corresponding to the fluctuation of the communication traffic. By changing the content of the management information definition according to the management information amount change procedure, it is possible to adjust the collection of management information according to the load. That is, while the load on the network or the computer is small, the items and frequency of management information are set to a large number by using the monitor result as a trigger, and detailed management information can be collected. In addition, when these loads increase, the items and frequency of management information are set to be small, which has the effect of suppressing the influence of the load of distributed system management on business applications and reducing the rate of reducing throughput.
[Simple explanation of drawings]
[Figure 1]
It is a block diagram which shows the whole system management system of Example 1 of this invention.
[Figure 2]
The internal configuration of the managed application of Example 1 of the present invention is shown.
[Fig. 3]
This is an example of information on an alternative calculator managed by the managed application of Example 1 of the present invention.
[Fig. 4]
This is an example of system operation information according to the first embodiment of the present invention.
[Fig. 5]
It is a flowchart of the operation process of the computer system performed by the managed application of Example 1 of this invention.
[Fig. 6]
This is a processing flow of alternative processing executed by the managed application according to the first embodiment of the present invention.
[Fig. 7]
This is an example of the internal configuration of the management application according to the second embodiment of the present invention.
[Fig. 8]
It is a configuration example of the managed application of Example 2 of this invention.
[Fig. 9]
The processing flow of the management application and the managed application of Example 2 of this invention is shown.
[Fig. 10]
The internal configuration of the managed application of Example 3 of the present invention is shown.
[Fig. 11]
The flow of the alternative processing of the managed application of Example 3 of this invention is shown.
[Fig. 12]
The internal configuration of the managed application of Example 4 of the present invention is shown.
[Fig. 13]
The processing flow of the managed application of Example 4 of this invention is shown.
[Fig. 14]
An example of the management information definition held in the managed application of Example 4 of the present invention is shown.
[Fig. 15]
An example of the management information amount adjustment table held by the managed application of Example 4 of the present invention is shown.
[Fig. 16]
It is a conventional system configuration diagram.
[Explanation of symbols]
100 Management device, 101 Computer system, 110 Management application, 111 Synchronous communication mechanism, 112 Synchronous information management unit, 113 Information evaluation unit, 114 Load distribution table, 115 Load distribution unit, 116 Management information receiving means, 117 Management information storage / display Means, 200 managed devices, 210 managed applications, 211 managed information definition, 212 managed monitor control means, 213 managed data transmission means, 214 alternative system management means, 215 managed application-to-application communication means, 216 alternative operation request destination determination Means, 217 Alternative processing means, 218 System operation information management means, 219 Synchronous communication mechanism, 221 Operation status monitoring unit, 222 Synchronous managed application-to-application communication means, 223 Traffic monitoring unit, 224 Management information amount change procedure, 225 Management information Amount adjustment table, 226 management information amount adjustment table interpretation means, 300 management information communication means, 400 networks.
17 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US9826031B2 | Cited by | United States of America | Applicant |
| US8819106B1 | Cited by | United States of America | Applicant |
| JP2007041763A | Cited by | Japan | Examiner |
| US11425194B1 | Cited by | United States of America | Applicant |
| US11263084B2 | Cited by | United States of America | Applicant |
| JP2012511784A | Cited by | Japan | Examiner |
| US10873623B2 | Cited by | United States of America | Applicant |
| US9207975B2 | Cited by | United States of America | Applicant |
| US8935404B2 | Cited by | United States of America | Applicant |
| JPH10312350A | Cited by | Japan | Search report |
| US10958716B2 | Cited by | United States of America | Applicant |
| JPH10161960A | Cited by | Japan | Search report |
| US9329909B1 | Cited by | United States of America | Applicant |
2 priority claims, no other members on record
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 25663194 | Japan | A | |
| JP19940256631 | – | – | – |
Numbers
- Publication
- 8-123768
- Publication, DOCDB
- H08123768
- Publication, EPODOC
- JPH08123768
- Application
- 6256631
- Application, DOCDB
- 25663194
- Application, EPODOC
- JP19940256631
Titles2
- Japanese
- 【発明の名称】分散システム管理方式及び分散システム管理方法
- English
- Description: Distributed system management method and distributed system management method
Classification
- IPC, 6
- G06F15 16
- G06F9 46
- G06F9 50
- G06F11 20
- G06F13 00
- G06F15 177