Cluster system and service control program
Abstract
Problem to be solved.To enable advanced and detailed service control to be realized in a cluster system.
Solution.Regarding the service provided by the cluster system, a service control part 122, arranged in a scenario management feature 12 inside the cluster system controls start or stop of the service, in order to match with an inter-service relationship indicated by the inter-service relationship information stored on a service related information DB 121a. In controlling the start of the service, the service control part 122 selects a computer which can provide service matching with the service relationship, indicated by the inter-service related information and instructs the start of the service to a service execution feature for its computer, and in controlling of the stoppage of the service, it instructs the stop of the service to the service execution feature for the computer providing the service.
Copyright (C)2005,JPO&NCIPI
Term
Term ended
Projected expiry passed 20 June 2023, 3.3 years ago.
- Priority and filed
- Published
- Projected expiry
- Today
15 claims: 3 independent, 12 dependent
- 1In a cluster system composed of a plurality of computers and capable of providing a plurality of types of services, the relationship between the services is defined in advance for each combination of two different types of services among the plurality of types of services. Regarding the inter-service relationship information storage means for storing information, the service execution means provided in each of the plurality of computers and controlling the start and stop of the service, and the service provided by the cluster system, the inter-service relationship information It is a service control means that controls the start or stop of the service so as to match the inter-service relationship shown, and in the start control of the service, the inter-service relationship that matches the inter-service relationship indicated by the inter-service relationship information of the service. A computer that can be provided is selected, and the service execution means of the computer is instructed to start the service. In the stop control of the service, the service execution means of the computer that has started the service is instructed. A cluster system including a service control means for instructing the stop of the service. 複数の計算機から構成され、複数種類のサービスが提供可能なクラスタシステムにおいて、前記複数種類のサービスのうちの相異なる2種類のサービスの組み合わせ毎に、そのサービス間の関係を予め定義したサービス間関係情報を記憶するサービス間関係情報記憶手段と、前記複数の計算機の各々に設けられ、サービスの開始及び停止を司るサービス実行手段と、前記クラスタシステムにより提供されるサービスに関し、前記サービス間関係情報の示すサービス間関係に合致するように、当該サービスの開始または停止を制御するサービス制御手段であって、当該サービスの開始制御では、前記サービス間関係情報の示すサービス間関係に合致する、当該サービスの提供が可能な計算機を選択して、その計算機の前記サービス実行手段に対して当該サービスの開始を指示し、当該サービスの停止制御では、当該サービスを開始している計算機の前記サービス実行手段に対して当該サービスの停止を指示するサービス制御手段とを具備することを特徴とするクラスタシステム。
- 10The number of computers constituting the cluster system is n + m (n is an integer of 2 or more, m is an integer of 1 or more), and n of the n + m computers are of a unique type. It is an operating system computer that provides services, and the remaining m units can take over the service started by the failure occurrence computer when a failure occurs in any of the n operating systems computers. This is a standby computer, and all the inter-service relationship information for each combination of two different types of services among the services provided by the n operating computers, which are stored in the inter-service relationship information storage means. The claim is characterized in that an exclusive relationship, a strong relationship, and a local relationship are set as the type of the relationship between the services, the attribute of the relationship between the services, and the scope of application of the relationship between the services, respectively. 8 Described cluster system. 前記クラスタシステムを構成する計算機の台数はn+m台(nは2以上の整数、mは1以上の整数)であり、当該n+m台の計算機のうちのn台はそれぞれ固有の種類のサービスを提供する稼働系計算機であり、残りのm台は前記n台の稼働系計算機のいずれかの計算機に障害が発生した場合に、当該障害発生計算機で開始されていたサービスを引き継ぐことが可能な待機系計算機であり、前記サービス間関係情報記憶手段に記憶される、前記n台の稼働系計算機がそれぞれ提供するサービスのうちの相異なる2種類のサービスの組み合わせ毎の全てのサービス間関係情報に、前記サービス間の関係の種類、前記サービス間の関係の属性及び前記サービス間の関係の適用範囲として、それぞれ排他関係、強い関係及び局所的関係が設定されていることを特徴とする請求項8記載のクラスタシステム。
- 13A service control program composed of a plurality of computers and executed to control the execution of services by the computers in a cluster system capable of providing a plurality of types of services, which can be provided to the computers by the cluster system. A step of detecting a service to be started among the services, a step of sequentially selecting a normal computer capable of providing the detected service to be started from the plurality of computers, and the computer being selected. Each time, by referring to the inter-service relationship information storage means in which the inter-service relationship information in which the relationship between the services is defined in advance is stored for each combination of two different types of services among the plurality of types of services. When it is determined that initiating the detected service on the selected computer matches the service-to-service relationship indicated by the corresponding service-to-service relationship information, and when it is determined to match the service-to-service relationship. A service control program for causing the selected computer to execute the step of starting the detected service. 複数の計算機から構成され、複数種類のサービスが提供可能なクラスタシステム内の前記計算機でサービスの実行を制御するために実行されるサービス制御プログラムであって、前記計算機に、前記クラスタシステムにより提供可能なサービスのうち開始すべきサービスを検出するステップと、前記複数の計算機の中から、前記検出された開始すべきサービスを提供可能な正常な計算機を順次選択するステップと、前記計算機が選択される都度、前記複数種類のサービスのうちの相異なる2種類のサービスの組み合わせ毎に、そのサービス間の関係を予め定義したサービス間関係情報が記憶されたサービス間関係情報記憶手段を参照することにより、前記選択された計算機で前記検出されたサービスを開始することが、対応する前記サービス間関係情報の示すサービス間関係に合致するかを判定するステップと、前記サービス間関係に合致すると判定された場合に、前記選択された計算機で前記検出されたサービスを開始させるステップとを実行させるためのサービス制御プログラム。
Independent claims3
278 paragraphs in 1 section, as filed
【0001】
[Technical field to which the invention belongs]
The present invention relates to a cluster system composed of a plurality of computers and capable of providing a plurality of types of services, and particularly relates to a cluster system and a service control program suitable for controlling the execution of services in the system.
【0002】
[Conventional technology]
It has been conventionally known that the purpose of a computer system is to execute an application program on the computer system and provide a service to a user. In recent years, as services have become more important, the computer systems on which services are executed are required to have high availability.
【0003】
Currently, a cluster system is attracting attention as a system with increased availability of a computer system (see, for example, Non-Patent Document 1). The cluster system is composed of a plurality of computers connected to the network, and is equipped with a cluster system management mechanism for managing those computers in a unified manner. The cluster system management mechanism controls the start processing and stop processing of the service and the status monitoring of the service according to the preset cluster system settings.
【0004】
For example, in a hot standby type cluster system, when the cluster system management mechanism detects a failure that occurs in the computer (operating computer) that is starting the service, the service is performed by the computer (standby computer) to be taken over. Is controlled to restart. The control to take over this service is called failover. This failover reduces service outages and increases the availability of cluster systems.
【0005】
In the conventional cluster system, when providing a plurality of services, a hot standby type cluster system composed of two computers is prepared as many as the number of services to be provided. However, in this form, if the number of services to be provided is n, 2 × n computers are required to build a cluster system, and computer resources are wasted. Therefore, in order to save computer resources, a cluster system in the form of n-to-1 backup with a common standby system is also known.
【0006】
[Non-Patent Document 1]
Tetsuo Kaneko, Yoshiya Mori, "Cluster Software", Toshiba Review, Vol.54 No.12 (1999), p.18-21 [0007]
[Problems to be Solved by the Invention]
As described above, in the conventional cluster system that provides a plurality of services, the number of services (n) provided by the hot standby type cluster system is prepared, or the form of n-to-1 backup is adopted. .. These forms are simply a combination of hot standby type cluster systems. Therefore, only services equivalent to those of a hot standby type cluster system can be controlled.
【0008】
The present invention has been made in consideration of the above circumstances, and an object of the present invention is to be optimal in terms of system availability, effective utilization of computer resources, and improvement of service processing capacity when providing a plurality of services. It is an object of the present invention to provide a cluster system and a service control program in which services can be arranged in a computer, thereby realizing more advanced service control than a hot standby type cluster system.
【0009】
[Means for solving problems]
According to one aspect of the present invention, a cluster system composed of a plurality of computers and capable of providing a plurality of types of services is provided. This cluster system includes a service-to-service relationship information storage means, a service execution means provided in each of the plurality of computers, and a service control means. The inter-service relationship information storage means stores inter-service relationship information in which the relationship between the services is defined in advance for each combination of two different types of services among the plurality of types of services. The service execution means starts and stops the service on the corresponding computer. The service control means controls the start or stop of the service provided by the cluster system so as to match the inter-service relationship indicated by the inter-service relationship information. In the service start control, this service control means selects a computer capable of providing the service that matches the service-to-service relationship indicated by the service-to-service relationship information, and provides the service to the service execution means of the computer. Instruct the start of. Further, in the service stop control, the service control means instructs the service execution means of the computer that has started the service to stop the service. Thus, in the present invention, the service control means is provided by the cluster system. Regarding services, by controlling the start or stop of the services so as to match the inter-service relationships indicated by the inter-service relationship information stored in the inter-service relationship information storage means, at an altitude that matches the inter-service relationships. Moreover, detailed service control can be realized.
【0010】
Here, a service permission status information storage means for storing service permission status information for managing the permission status of the service for each of the plurality of types of services is added, and the service permission status is added to the service control means. The service to be started and the computer to start the service to be started, or the service to be stopped and the service to be stopped so that the relationship between the services shown in the above-mentioned inter-service relationship information is matched according to the permission status indicated by the information. It is advisable to have an automatic control means for selecting the computer that has started the service to be serviced. In this way, it is possible to automatically determine which service is started or stopped by which computer so that the relationship between the services matches according to the permission status indicated by the service permission status information.
【0011】
Further, when the service is started by the service execution means, the service control means is provided with a batch job control means for disallowing the permission status indicated by the service permission status information corresponding to the service, and the service is permitted. When the permitted status indicated by the status information is disallowed, the automatic control means may be used to control the stop of the service corresponding to the service permitted status information. In this way, the batch job service can be easily realized.
【0012】
In addition, the service-to-service relationship information includes information indicating the type of the relationship between the services, the attribute of the relationship between the services, and the scope of application of the relationship between the services, and the service is defined as the type of the relationship between the services. It is possible to set a dependency that can start the service only when the service is started, an exclusive relationship that the predetermined service should not be started when the service is started, and an irrelevant service. As attributes of the relationship between services, when there are a plurality of computers that satisfy the strong relationship that the relationship between the services must be established and the strong relationship that is a candidate for selection by the automatic control means, the relationship between the services. It is possible to set a weak relationship in which the computer for which is established is preferentially selected, and as the scope of application of the relationship between the above services, the relationship between the services is applied only when the services are started on the same computer. It is preferable that the local relationship and the global relationship to which the relationship between the services is applied can be set regardless of which computer the service is started by.
【0013】
Here, when the above cluster system fails in any of the n operating computers that provide unique types of services and the n operating computers, the failure occurrence computer is started. It consists of m standby computers that can take over the services that have been provided, and is different from the services provided by the n operating computers that are stored in the inter-service relationship information storage means. For all service-to-service relationship information for each type of service combination, the types of relationships between services, the attributes of relationships between services, and the scope of application of relationships between services are exclusive relationships, strong relationships, and local relationships, respectively. If the configuration is such that the relationship is set, an n-to-m backup configuration can be easily realized.
【0014】
In addition, as the services provided by the cluster system, three types of services are defined: a first service corresponding to the data, a second service for backing up the data, and a third service for accessing the data. In the inter-service relationship information stored in the inter-service relationship information storage means, a strong dependency relationship between the second service and the third service is set as a local relationship with respect to the first service, and the above If a strong exclusive relationship between the second service and the third service is set as a global relationship, cold backup of data can be easily realized.
【0015】
Similarly, in the inter-service relationship information stored in the inter-service relationship information storage means, a strong dependency relationship between the second service and the third service is set as a local relationship with respect to the first service. If a strong dependency from the second service is set as a global relationship with respect to the third service and the configuration is configured, online backup of data can be easily realized.
【0016】
BEST MODE FOR CARRYING OUT THE INVENTION
Hereinafter, embodiments of the present invention will be described with reference to the drawings.
【0017】
[First Embodiment]
FIG. 1 is a block diagram showing a configuration of a cluster system according to the first embodiment of the present invention. In the figure, the cluster system 1 is a plurality of computers interconnected by a network (not shown), for example, N computers 10-1 (# 1), 10-2 (# 2), ... 10-N ( It is configured with #N). The cluster system 1 has a cluster system control mechanism 10. The cluster system control mechanism 10 exists over the computers 10-1 (# 1), 10-2 (# 2), ... 10-N (#N). That is, the cluster system 1 has a cluster system control mechanism 10 realized by using computers 10-1 to 10-N.
【0018】
The cluster system control mechanism 10 has a service execution mechanism 11-i realized on each computer 10-i (i = 1 to N). That is, the computer 10-i has a service execution mechanism 11-i. The service execution mechanism 11-i has a service starting means 111-i for starting the service and a service stopping means 112-i for stopping the service. The cluster system control mechanism 10 also has a scenario management mechanism 12. The scenario management mechanism 12 has communication paths 13-1 to 13-N between all service execution mechanisms 11-1 to 11-N. In the example of FIG. 1, the scenario management mechanism 12 exists across all the computers 11-1 to 11-N constituting the cluster system 1. The scenario management mechanism 12 is realized by the scenario management units 12-1 to 12-N that operate on each computer 11-1 to 11-N and communicate with each other. In the present embodiment, the service execution mechanism 11-i and the scenario management unit 12-i are realized by reading and executing the corresponding software programs by the computer 10-i (CPU not shown in the figure). To do.
【0019】
Here, before detailing each element in the system of FIG. 1, the "relationship between services" applied in the system will be described. The relationship between services defines what kind of relationship is between two different services. The elements that define the "relationship between services" applied in the present embodiment include: -type of relationship between services-attribute of relationship between services-applicable scope of relationship between services.
【0020】
"Types of relationships between services" include: -Dependencies-Exclusive relationships-Unrelated.
【0021】
A "dependency" is a relationship in which a service can be started only when a predetermined service is started. For example, if service # 1 has a dependency on service # 2, service # 1 can be started only when service # 2 has already started. In other words, if service # 1 and service # 2 are started and service # 2 is stopped, service # 1 is also stopped.
【0022】
The "exclusive relationship" is a relationship in which a predetermined service must not be started when the service is started. For example, if service # 1 has an exclusive relationship with service # 2, service # 1 can be started only when service # 2 is not started, and service # 2 can be started when service # 2 is started. # 1 is unstartable.
【0023】
"Unrelated" is a relationship in which there is no relationship between one service and another. "Dependency", "exclusive relationship" and "irrelevant" are independent of each other.
【0024】
Next, the "attributes of relationships between services" include: -strong relationships-weak relationships.
【0025】
"Strong relationship" means that the relationship between the target services must be established. Services cannot be started or stopped in a way that violates the strong relationship between services.
【0026】
"Weak relationship" is referred to by the scenario management mechanism 12 to automatically determine the computer to start the service when the user performs the service start operation without specifying the computer to start the service. It means that it is a relationship. Such service control is called automatic service control. In other words, for the "weak relationship", when there are multiple computers (options) that satisfy the "strong relationship" that is a candidate for selection in the automatic control of the service, the computer that holds the "weak relationship" is preferentially selected. It is a relationship. Here, if there are multiple computers that can start the service (computers that do not violate the relationship between services that have a strong relationship with starting the service on that computer), the service that has a weak relationship among them. The computer that does not violate the relationship is determined as the computer that starts the service.
【0027】
The "scope of application of relationships between services" includes: -local relationships-global relationships.
【0028】
"Local relationship" means that the relationship between the target services is a range that applies only when the services are started on the same computer. "Global relationship" means that the relationship between the services in question is within the scope of application regardless of which computer the service is started on.
【0029】
For example, assume that a strong exclusive relationship is set between service # 1 and service # 2, and service # 1 is started by computer # 1. In this case, if the above strong exclusive relationship is a global relationship, service # 2 cannot be started by any computer, and if it is a local relationship, service # 2 cannot be started by computer # 1.
【0030】
FIG. 2 is a block diagram showing the configuration of the scenario management mechanism 12 in FIG. The scenario management mechanism 12 is roughly divided into a database unit 121 and a service control unit 122. The database unit 121 is an information storage unit for storing a service-related information database (DB) 121a, a service status information database (DB) 121b, a computer status information database (DB) 121c, and a service permission status information database (DB) 121d. Is.
【0031】
Service-related information DB121a is a database for managing relationships between services. One data (record) of the service relationship information DB121a is the service name SN1, the service name SN2 of the service related to the service of the service name SN1, the type of the relationship between the services, the attribute of the relationship between the services, and the service between the services. It consists of 5 items of the scope of application of the relationship.
【0032】
The service status information DB121b is a database for managing the service status. One data (record) of service status information DB121b is composed of three items: service name, computer name, and status. The service name is the name of the service provided by the cluster system in Figure 1. The computer name is the name of the computer that last executed the service starting means 111-i. The state is information indicating whether or not the service has started normally, and has three states of "start", "stop", and "failure occurred".
【0033】
The computer status information DB121c is a database for managing the status of the computer. One data (record) of computer status information DB121c is composed of two items, computer name and status. The computer name is the name of the computer that constitutes the cluster system shown in Fig. 1. The state is information indicating whether or not the computer is operating normally, and has two states, "normal" and "abnormal".
【0034】
The service permission status information DB121d is a database for managing the service permission status. One data (record) of service permission status information DB121d consists of two items, service name and permission status. The service name is the name of the service provided by the cluster system in Figure 1. The permitted state indicates whether the service is started (that is, the service is in the started state) is permitted or not permitted, and has two states of "permitted" and "not permitted".
【0035】
The service control unit 122 has each processing unit of an operation processing unit 122a, a manual control unit 122b, an automatic control unit 122c, an inter-service relationship determination unit 122d, an execution control unit 122e, and a failure detection unit 122f. The operation processing unit 122a receives an operation content from a user given via a network from a client terminal (user terminal) (not shown), and controls a service corresponding to the operation content (a computer starts a service or starts a service). It is a user interface that requests the manual control unit 122b to stop).
【0036】
When the manual control unit 122b receives a service control request (starting or stopping a service on a computer) from the operation processing unit 122a, the control of the service is controlled by the service status, the computer status, and the service. Confirm that it is possible from the aspect of the relationship, and request the execution control unit 122e to control it.
【0037】
The automatic control unit 122c periodically refers to the service permission status information DB121d, and automatically executes service control according to a change in the information DB121d. That is, when the permission state of a certain service becomes "permission", the automatic control unit 122c starts the service (that is, transitions from the stop state to the start state), which is the relationship between the service state, the computer state, and the services. Determine the computer to start the service by confirming that it is possible in terms of. The automatic control unit 122c requests the execution control unit 122e to control the start of the service on the determined computer. In addition, the automatic control unit 122c identifies the computer that has started the service when the permitted status becomes "not permitted". The automatic control unit 122c requests the execution control unit 122e to control the service started by the specified computer.
【0038】
The inter-service relationship determination unit 122d determines whether or not executing the control of the service received from the manual control unit 122b or the automatic control unit 122c does not violate the inter-service relationship set in the database unit 121. The execution control unit 122e transmits the control of the service received from the manual control unit 122b or the automatic control unit 122c to the service execution mechanism 11-i of the corresponding computer 10-i. The fault detection unit 122f detects that a fault has occurred in the computer 10-i and updates the service status information DB121b and the computer status information DB121c.
【0039】
Next, the operation of the cluster system shown in FIG. 1 will be described by taking as an example the process executed by the scenario management mechanism 12 when the user starts the service. This process differs depending on whether or not the user specifies a computer to start the service.
【0040】
First, the processing mainly performed by the automatic control unit 122c of the scenario management mechanism 12 when the user does not specify the computer to start the service will be described with reference to the flowchart of FIG. When the operation processing unit 122a in the scenario management mechanism 12 receives the start operation content of a service (referred to as service #j) from the user from the client terminal, it accesses the service permission status information DB121d and the permission status of service #j. To "permit".
【0041】
The automatic control unit 122c periodically refers to the service permission status information DB121d (step S1). If the permission status of service #j becomes "permission" (step S2), the automatic control unit 122c refers to the computer status information DB121c and checks whether there is an undetermined computer (step S3). Here, in the first step S3, all the calculators are undecided. When there are no undetermined computers as a result of repeating step S3, the automatic control unit 122c ends the process associated with the start operation of service # j (step S4).
【0042】
On the other hand, when there is an undetermined computer, the automatic control unit 122c uses any one of the undetermined computers as a candidate for the computer to be started by the service #j (hereinafter referred to as a candidate computer). Select (step S5). Here, it is assumed that computer 10-i (computer #i) is selected as a candidate computer. Next, the automatic control unit 122c checks whether the status of the selected candidate calculator 10-i is "normal" based on the computer status information DB121c (step S6).
【0043】
If the state of the candidate calculator 10-i is not normal (step S6), the automatic control unit 122c considers the computer 10-i to be a determined computer and checks whether there is an undetermined computer described above (step). Return to S3). On the other hand, if the state of the candidate calculator 10-i is normal (step S6), the automatic control unit 122c inquires of the inter-service relationship determination unit 122d, and the start of the service #j in the computer 10-i is determined. Check for violations of relationships between services (step S7). The processing of the inter-service relationship determination unit 122d will be described later.
【0044】
If the start of service # j on the candidate computer 10-i violates the relationship between the services (step S7), the automatic control unit 122c considers the computer 10-i to be a determined computer and has not yet described it. Return to the process of checking if there is a judgment calculator (step S3). On the other hand, if the start of service # j on candidate computer 10-i does not violate the relationship between services (step S7), the automatic control unit 122c executes and controls the service control request for computer 10-i. Send to section 122e. In this case, the execution control unit 122e starts the service # j to the service execution mechanism 11-i of the computer 10-i (computer #i) specified in the service control request from the automatic control unit 122c. Send an instruction to execute the service start means 111-i to be performed (step S8).
【0045】
When the execution control unit 122e sends a command to the service execution mechanism 11-i of the computer 10-i to execute the service start means 111-i that starts the service # j, the execution control unit 122e sends an instruction to the automatic control unit 122c. Notify the completion of the service control request. When the automatic control unit 122c receives the completion notification of the service control request, it accesses the service status information DB121b and sets the status of the service #j to "start" (step S9). This completes the process associated with the service # j start operation when the user does not specify the computer to start the service (step S10).
【0046】
Next, the processing mainly performed by the manual control unit 122b of the scenario management mechanism 12 when the user specifies the computer to start the service will be described with reference to the flowchart of FIG. The operation processing unit 122a in the scenario management mechanism 12 receives a start operation of a service (service #j) from a computer (hereinafter referred to as a designated computer) 10-i (computer #i) from a user. A service control request based on this is sent to the manual control unit 122b. As a result, the manual control unit 122b starts the requested service control processing, that is, the processing associated with the service # j start operation (step S11).
【0047】
First, the manual control unit 122b refers to the computer status information DB121c and checks whether the status of the designated computer 10-i is normal (step S12). If the status of the designated calculator 10-i is not normal, the manual control unit 122b assumes that the service # j cannot be started by the designated calculator 10-i and ends the process associated with the start operation of the service # j (step S13). ). On the other hand, if the state of the designated calculator 10-i is normal, the manual control unit 122b inquires of the inter-service relationship determination unit 122d, and the start of the service #j in the computer 10-i is between services. Check for violations of the relationship (step S14).
【0048】
If the start of service # j on the designated calculator 10-i violates the relationship between services (step S14), the manual control unit 122b ends the process associated with the start operation of service # j (step S13). ). On the other hand, if the start of service # j on the designated computer 10-i does not violate the relationship between services (step S14), the manual control unit 122b executes and controls the service control request for the computer 10-i. Send to section 122e.
【0049】
The execution control unit 122e starts the service # j to the service execution mechanism 11-i of the computer 10-i specified in the service control request from the manual control unit 122b. Service start means 111-i Send an instruction to execute (step S15). Then, the execution control unit 122e notifies the manual control unit 122b of the completion of the service control request. Upon receiving the service control request completion notification, the manual control unit 122b accesses the service status information DB121b and sets the status of service #j to "started" (step S16). This completes the process associated with the service # j start operation when the user specifies the computer to start the service (step S17).
【0050】
Next, when a failure that occurs in a computer in cluster system 1 is detected, it is executed to take over (fail over) the service started in the failed computer to another computer in cluster system 1. The processing (failover processing) mainly performed by the automatic control unit 122c of the scenario management mechanism 12 will be described with reference to the flowchart of FIG.
【0051】
Now, it is assumed that the failure detection unit 122f in the scenario management mechanism 12 has detected a failure that has occurred in a computer in the cluster system 1, for example, calculator 10-k (computer #k). In this case, the fault detection unit 122f accesses the computer status information DB121c and sets the status of the fault occurrence computer 10-k (computer #k) to "abnormal". The failure detection unit 122f also accesses the service status information DB121b, and sets the status to "failure occurrence" for all services whose computer name is "computer #k" and whose current status is "started".
【0052】
The automatic control unit 122c in the scenario management mechanism 12 periodically refers to the service status information DB121b (step S21). If there is a service whose status is "failure occurred" (step S22), the automatic control unit 122c starts a process for failing over the service. Here, it is assumed that the status of service #j has been detected as "failure occurred".
【0053】
First, the automatic control unit 122c refers to the computer status information DB121c and checks whether there is an undetermined computer (step S23). Here, in the first step S23, all the calculators are undecided. When there are no undetermined computers as a result of repeating step S23, the automatic control unit 122c ends the process for failing over service # j (step S24).
【0054】
On the other hand, when there is an undetermined computer, the automatic control unit 122c uses any one of the undetermined computers as a candidate for the computer to be started (taken over) of the service #j (hereinafter referred to as the candidate computer). Select as (referred to) (step S25). Here, it is assumed that computer 10-i (computer #i) is selected as a candidate computer. Next, the automatic control unit 122c checks whether the status of the selected candidate calculator 10-i is "normal" based on the computer status information DB121c (step S26).
【0055】
If the state of the candidate calculator 10-i is not normal (step S26), the automatic control unit 122c considers the computer 10-i to be a determined computer and checks whether there is an undetermined computer described above (step). Return to S23). On the other hand, if the state of the candidate calculator 10-i is normal (step S26), the automatic control unit 122c inquires of the inter-service relationship determination unit 122d, and the start of the service #j in the computer 10-i is determined. Check for violations of relationships between services (step S27).
【0056】
If the start of service # j on the candidate computer 10-i violates the relationship between the services (step S27), the automatic control unit 122c considers the computer 10-i to be a determined computer and has not yet described it. The process returns to the process of checking whether there is a judgment calculator (step S23). On the other hand, if the start of service # j on candidate computer 10-i does not violate the relationship between services (step S27), the automatic control unit 122c executes and controls the service control request for computer 10-i. Send to department 122e. In this case, the execution control unit 122e starts the service # j to the service execution mechanism 11-i of the computer 10-i (computer #i) specified in the service control request from the automatic control unit 122c. Send an instruction to execute the service start means 111-i to be performed (step S28). Then, the execution control unit 122e notifies the automatic control unit 122c of the completion of the service control request.
【0057】
When the automatic control unit 122c receives the completion notification of the service control request, it accesses the service status information DB121b and sets the status of the service #j to "start" (step S29). With the above, the service #j started by the failure occurrence computer 10-k (failure occurrence computer #k) is failed over by another normal computer (here, computer 10-i (computer #i)) in the cluster system 1. The process for doing so ends (step S30). The automatic control unit 122c repeats the above processing as long as there is a service whose status is "failure occurred".
【0058】
Next, regarding the inter-service relationship determination process (that is, the process of checking whether the service of interest violates the inter-service relationship) executed by the inter-service relationship determination unit 122d in the scenario management mechanism 12, FIG. This will be described with reference to the flowchart.
【0059】
Now, the manual control unit 122b or the automatic control unit 122c has given an inquiry to the inter-service relationship determination unit 122d asking whether a service #j can be started on a computer 10-i (computer #i). To do. Here, the inquired service and the computer are referred to as a target service and a target computer, respectively. Here, it is assumed that the target service is service #j and the target computer is computer 10-i (computer #i).
【0060】
When the inter-service relationship determination unit 122d receives an inquiry from the manual control unit 122b or the automatic control unit 122c, the target service #j specified in the inquiry is set to the computer 10-i (computer #i) specified in the inquiry. The inter-service relationship determination process for checking whether the inter-service relationship is violated with respect to the start in step S41 is started (step S41). First, the inter-service relationship determination unit 122d refers to the service-related information DB121a, and acquires all the data (records) in which the service name SN1 represents the target service #j in the service-related information DB121a (step S42).
【0061】
Next, in the acquired data, the service-to-service relationship determination unit 122d has a type of relationship between services "exclusive relationship" and an attribute of the relationship between services "strong relationship", that is, "strong exclusion". Find out if there is a "relationship" and undecided (step S43). Here, in the first step S43, all "strong exclusive relationships" are undecided. When the undetermined "strong exclusive relationship" data disappears as a result of repeating step S43, the inter-service relationship determination unit 122d proceeds to the process of examining the dependency relationship (step S50), which will be described later.
【0062】
On the other hand, when there is undetermined "strong exclusive relationship" data, the inter-service relationship determination unit 122d selects any one of the undetermined data as candidate data for determining the inter-service relationship ( It is selected as candidate data (hereinafter referred to as candidate data) (step S44). Next, the inter-service relationship determination unit 122d refers to the service status information DB121b and acquires information about the service indicated by the service name SN2 in the selected candidate data in the service-related information DB121a (hereinafter referred to as exclusive target data). (Step S45).
【0063】
The inter-service relationship determination unit 122d checks whether the status of the acquired exclusive target data is started (step S46). If the status of the acquired exclusive target data is not started (step S46), the inter-service relationship determination unit 122d determines the data and returns to the process of checking whether there is any undetermined data (step S43). On the other hand, if the state of the acquired exclusive target data is the start, the inter-service relationship determination unit 122d examines whether the attribute of the relationship between the services of the data is a local relationship (step S47).
【0064】
If the attribute of the relationship between services of the acquired exclusive target data is not a local relationship (step S47), the service-to-service relationship determination unit 122d determines that the relationship between services is violated (step S49). On the other hand, if the attribute of the relationship between the services of the exclusive target data is a local relationship (step S47), the inter-service relationship determination unit 122d uses the computer name included in the exclusive target data as the target. Check if it represents a computer (step S48).
【0065】
If the computer name included in the exclusive target data represents the target computer (step S48), the inter-service relationship determination unit 122d can start the target service #j on the computers 10-i between services. It is determined that the relationship is violated (step S49). On the other hand, if the computer name included in the exclusive target data does not represent the target computer (step S48), the inter-service relationship determination unit 122d considers the data to have been determined and determines whether there is any undetermined data. The process returns to the checking process (step S43).
【0066】
When, as a result of repeating step S43, there is no undetermined "strong exclusive relationship" data, the inter-service relationship determination unit 122d includes the type of relationship between services in the data acquired in step S42. Is a "dependency" and the attribute of the relationship between services is a "strong relationship", that is, a "strong dependency", and it is examined whether there is an undetermined one (step S50). Here, in the first step S50, all "strong dependencies" are undecided. When there is no undetermined "strong dependency" data as a result of repeating step S50 (step S50), the inter-service relationship determination unit 122d may start the target service #j on the computer 10-i. Determine that the relationship between services is not violated (step S51).
【0067】
On the other hand, when there is undetermined "strong dependency" data, the inter-service relationship determination unit 122d determines any one of the undetermined "strong dependency" data as a target for determining the inter-service relationship. Select as candidate data for (step S52). Next, the inter-service relationship determination unit 122d refers to the service status information DB121b and acquires the information related to the service indicated by the service name in the selected candidate data in the service-related information DB121a as the dependency target data (step S53).
【0068】
The inter-service relationship determination unit 122d checks whether the status of the acquired dependent data is started (step S54). If the state of the acquired dependent target data is the start (step S54), the inter-service relationship determination unit 122d considers the data to have been determined and returns to the process of checking whether there is any undetermined data (step S50). .. On the other hand, if the state of the acquired dependent data is not the start, the inter-service relationship determination unit 122d examines whether the attribute of the relationship between the services of the data is a local relationship (step S55).
【0069】
If the attribute of the relationship between services of the acquired dependent data is not a local relationship (step S55), the inter-service relationship determination unit 122d considers the data to have been determined and checks whether there is any undetermined one. Return to (step S50). On the other hand, if the attribute of the relationship between the services of the dependent target data is a local relationship (step S55), the inter-service relationship determination unit 122d uses the computer name included in the exclusive target data as the target. Check if it represents a computer (step S56).
【0070】
If the computer name included in the dependent target data does not represent the target computer (step S56), the inter-service relationship determination unit 122d may start the target service #j on the computers 10-i between services. It is determined that the relationship is violated (step S57). On the other hand, if the computer name included in the dependent target data represents the target computer (step S56), the inter-service relationship determination unit 122d considers the data to have been determined and determines whether there is any undetermined data. Return to the checking process (step S50).
【0071】
As described above, in the first embodiment of the present invention, advanced and detailed service control can be realized by setting the relationship between services by the service-related information DB121a. Specifically, by setting services that must be started in order and services that must not be started at the same time, advanced and detailed service control can be controlled from the scenario management mechanism 12 in the cluster system control mechanism 10. It becomes. Further, in the first embodiment of the present invention, by mainly using the automatic control of the service realized by the automatic control unit 122c, the judgment required of the user is reduced and the control of the cluster system is facilitated. Can be done. Further, in the first embodiment of the present invention, a mechanism for failing over a service can be easily realized.
【0072】
(First Modified Example of First Embodiment) FIG. 7 is a block diagram showing a configuration of a cluster system according to a first modified example of the first embodiment of the present invention. In the configuration of FIG. 7, the same reference numerals are given to the parts equivalent to those of FIG.
【0073】
The feature of the cluster system 1'shown in FIG. 7 is that the scenario management mechanism 12'corresponding to the scenario management mechanism 12 in FIG. 1 is any of computers 10-1 (# 1) to 10-N (#N). It is in the point that it exists on one computer, for example, computer 10-1 (# 1).
【0074】
As is clear from the first embodiment of the present invention and the first modification thereof, the scenario management mechanism in the cluster system exists only on any one of the plurality of computers constituting the cluster system. It may be present across all computers, or it may be present on some of two or more computers.
【0075】
(Second Modification Example of the First Embodiment) Next, a second modification example of the first embodiment will be described. The types of relationships between services (dependencies, exclusive relationships, irrelevant) and the attributes of relationships between services (strong relationships, weak relationships) can be combined. The number of combinations is as follows: Strong dependency Weak dependency Strong exclusion relationship Weak exclusion relationship Irrelevant.
【0076】
It is also possible to combine the above five combinations with the scope of application of relationships between services (local relationships, global relationships). The number of combinations is 10 (5 × 2) shown below, that is, (1) local strong dependency (2) local weak dependency (3) local strong exclusion (4) local Weak exclusive relationship (5) Local irrelevance (6) Global strong dependency (7) Global weak dependency (8) Global strong exclusive relationship (9) Global weak exclusive relationship (10) Global weak exclusive relationship Is irrelevant.
【0077】
Therefore, when setting the relationship between services, it is possible to select some of the above 10 combinations and set them at the same time for the same service. However, some of the selected combinations (selection patterns) that are set at the same time are logically inconsistent or meaningless. For example, if the combination of (6) and (8) are set at the same time, a logical contradiction occurs. Also, considering the exclusive relationship, it is meaningless to set the combination of (3) when the combination of (8) is set. Also, considering the dependency, it is meaningless to set the combination of (6) when the combination of (1) is set.
【0078】
When such selection patterns that cause logical inconsistency and meaningless selection patterns are excluded from the selection patterns that can be set at the same time among the above 10 combinations, the selection patterns that can be set are shown below. There are 15 ways: (1) local strong dependency and global strong dependency (2) local weak dependency and global strong dependency (3) local weak dependency and global weak dependency. Dependencies (4) Strong global dependencies only (5) Global weak dependencies only (6) Global strong dependencies and local strong exclusive relationships (7) Global weak dependencies and local Strong global exclusive relationship (8) Global strong dependency and local weak exclusive relationship (9) Global weak dependency and local weak exclusive relationship (10) Local weak exclusive relationship only (11) Local Strong local exclusive relationship only (12) Local weak exclusive relationship and global weak exclusive relationship (13) Local strong exclusive relationship and global weak exclusive relationship (14) Local strong exclusive relationship and global weak exclusive relationship Strong exclusive relationship (15) Only irrelevant.
【0079】
As described above, in the second modification of the first embodiment of the present invention, some of the combinations of the types of relationships between services, their attributes, and the scope of application thereof are selected to provide service-related information. When setting to DB121a, it is possible to efficiently set the relationship between logically correct services by excluding in advance the selection pattern that causes a logical contradiction and the selection pattern that is logically meaningless.
【0080】
[Second Embodiment]
Next, a second embodiment of the present invention will be described. FIG. 8 is a block diagram showing the configuration of the scenario management mechanism 120 used in place of the scenario management mechanism 12 of the configuration of FIG. 2 in the cluster system of FIG. In the configuration of FIG. 8, the same reference numerals are given to the parts equivalent to those of FIG.
【0081】
The feature of the scenario management mechanism 120 having the configuration shown in FIG. 8 is that it has a configuration in which batch jobs can be realized. A batch job is a job of a method in which a predetermined series of processes are executed and the job ends when the processes are completed. In a conventional cluster system, it is difficult to realize an automatically terminated process such as a batch job as a service because an operation by a user is always required to stop the service.
【0082】
The scenario management mechanism 120 of FIG. 8 has a batch job control unit 122g. The configuration of this scenario management mechanism 120 is equivalent to the configuration in which the batch job control unit 122g is added to the scenario management mechanism 12 of FIG. The batch job control unit 122g has the batch job type service information DB121e. This batch job type service information DB121e is a database of services (batch job type services) handled as batch jobs. One data (record) of batch job type service information DB121e is composed of two items, service name and status. The service name is the name of the batch job type service provided by the cluster system. The state is information indicating whether or not the batch job type service is running, and has two states, "running" and "stopped".
【0083】
Next, regarding the processing mainly performed by the batch job control unit 122g of the scenario management mechanism 120 having the configuration shown in FIG. 8, refer to the flowchart of FIG. 9 by taking as an example the processing of handling a certain batch job type service (referred to as service #j). I will explain. First, it is assumed that the batch job type service information DB121e is set with data (record) whose service name is service # j and whose status is "stopped". In this state, it is assumed that the operation processing unit 122a receives the service # j start operation from the user. Then, the operation processing unit 122a accesses the service permission status information DB121d and sets the permission status of the service #j to "permission".
【0084】
The batch job control unit 122g periodically refers to the service permission status information DB121d and acquires the data related to the service #j each time (step S61). If the permission status of the service #j acquired from the service permission status information DB121d becomes "permission" (step S62), the batch job control unit 122g accesses the batch job type service information DB121e and of the service #j. Set the state to "Running" (step S63).
【0085】
Next, the batch job control unit 122g periodically refers to the service status information DB121b to acquire data (records) related to the service #j. On the other hand, the automatic control unit 122c executes the start processing of the service #j when the permission status of the service #j becomes "permission" as in this example, and accesses the service status information DB121b at the end. And set the status of service #j to "started".
【0086】
When the status of service #j in the service status information DB121b becomes "started" (step S64), the batch job control unit 122g accesses the service permission status information DB121d and "disallows" the permission status of the service #j. (Step S65). When the permitted status of the service #j becomes "not permitted", the automatic control unit 122c executes a process of stopping the service #j, and finally accesses the service status information DB121b to access the status of the service #j. To "stop".
【0087】
When the status of service #j in the service status information DB121b becomes "stopped" (step S66), the batch job control unit 122g accesses the batch job type service information DB121e and stops the status of the service #j. (Step S67). This completes the batch job processing for service #j (step S67).
【0088】
As described above, in the second embodiment of the present invention, the service can be treated as a batch job by using the scenario management mechanism 120 provided with the batch job control unit 122g. As a result, the scenario management mechanism 120 can control a process suitable for handling as a batch job such as data backup.
【0089】
[Third Embodiment]
Next, a third embodiment of the present invention will be described. FIG. 10 is a block diagram showing a configuration of a cluster system (n vs. m backup configuration cluster system) according to the third embodiment of the present invention. The feature of the n-to-m backup configuration cluster system of FIG. 10 is that the conventionally known n-to-m (n is an integer of 2 or more and m is an integer of 1 or more) backup configuration is applied in the first embodiment. The point is that it is realized by applying the applied cluster system.
【0090】
The n-to-m backup configuration cluster system in Fig. 10 consists of n 1-to-m backup configuration cluster systems 1-1 (# 1), 1-2 (# 2), ... 1-n (#n). Has been done. Cluster systems 1-1 (# 1), 1-2 (# 2), ... 1-n (#n) are computers 100-1 (# 1) and 100-2 used as operating computers, respectively. It has (# 2), ... 100-n (#n). Computers 100-1 (# 1), 100-2 (# 2), ... 100-n (#n) provide services # 1, # 2, ... # n, respectively. The cluster system 1-1 (# 1), 1-2 (# 2), ... 1-n (#n) is the cluster system 1-1 (# 1), 1-2 (# 2) ,. ..1-N (#n) is a working computer in 100-1 (# 1), 100-2 (# 2), ... 100-n (#n), which is commonly provided and stands by. M computers used as system computers 100- (n + 1) (# n + 1), 100- (n + 2) (# n + 2), ... 100- (n + m) (# n It has + m). Operating computers 100-1 (# 1) to 100-n (# n) in cluster system 1-1 (# 1) to 1-n (# n), and cluster system 1-1 (# 1) to 1 The standby computers 100- (n + 1) (# n + 1) ~ 100-(n + m) (# n + m) common to -n (#n) are interconnected by a network (not shown). There is.
【0091】
In this way, the n-to-m backup configuration cluster system shown in Fig. 10 has m computers 100- (n + 1) (# n + 1) to 100- (n + m) in order to save computer resources. It consists of n 1-to-m backup configuration cluster systems 1-1 (# 1) to 1-n (#n) that use (# n + m) in common as a standby computer. In this n-to-m backup configuration cluster system, if any of the n active computers 100-1 to 100-n that provide services # 1 to # n fails, that computer provides them. It is configured so that the existing service can be taken over by any of m computers 100- (n + 1) to 100- (n + m). The configuration of the n-to-m backup configuration cluster system in FIG. 10 is the above-mentioned first, except that it includes n 1-to-m backup configuration cluster systems 1-1 (# 1) to 1-n (#n). It is the same as the cluster system in the first embodiment. That is, the n-to-m backup configuration cluster system shown in Fig. 10 has a service execution mechanism provided for each computer 100-1 (# 1) to 100- (n + m) (# n + m) and scenario management of the configuration shown in Fig. 2. It has a cluster system control mechanism (neither is shown) including a mechanism. Therefore, in the third embodiment, FIG. 2 is incorporated as necessary.
【0092】
In the conventional technology, in order to realize a computer system with an n-to-m backup configuration, it is the same which computer starts with which computer for all n services # 1 to # n and which computer takes over in the event of a failure. It is necessary to set it carefully so that services are not concentrated on the computer. On the other hand, in the third embodiment of the present invention, by using the cluster system applied in the first embodiment, this kind of setting can be easily performed as described below.
【0093】
FIG. 11 shows an example of the service-related information DB121a possessed by the scenario management mechanism in the n-to-m backup configuration cluster system of FIG. In the service-related information DB121a shown in FIG. 11, service # i (i = 1 to n-1) and other services # j among the n services provided by the n-to-m backup configuration cluster system shown in FIG. 10 are provided. For all combinations with (j = i + 1 ~ n), the service name (SN1, SN2) representing the service of that combination, the type of relationship between the services (#i, #j), and the service (#i, #) Information on the attributes of the relationship between j) and the scope of application of the relationship between services (#i, #j) is registered. Here, the types of relationships between services, the attributes of relationships between services, and the scope of relationships between services are irrespective of the combination of services: -Types of relationships between services ... Exclusive relationships-Exclusive relationships between services Relationship attributes ... Strong relationships-The scope of relationships between services ... Set as local relationships.
【0094】
Next, the operation of the n vs. m backup configuration cluster system of FIG. 10 to which the service-related information DB121a shown in FIG. 11 is applied will be described with reference to the configuration of FIG. 2 and the flowchart of FIG. When the operation processing unit 122a in the scenario management mechanism 12 shown in FIG. 2 receives the start operation of n services # 1 to # n by the user, it accesses the service permission status information DB121d and n services # 1 ~. Set the permission status of #n to "permission". When the permission status of services # 1 to #n in the service permission status information DB121d becomes "permitted", the automatic control unit 122c sets the services # 1 to #n between the services indicated by the service-related information DB121a. According to the relationship, n vs. m backup configuration Of the n + m computers that make up the cluster system, n active computers 100-1 (# 1) to 100-n (# n) start one service at a time. Here, it is assumed that the service #i is started on the computer 100-i (#i).
【0095】
In this state, for example, due to a failure in computer 100-1 (# 1), service # 1 started on computer 100-1 shall be failed over. This failover process is executed by the automatic control unit 122c according to the flowchart of FIG. 5 as in the first embodiment. That is, when the automatic control unit 122c detects that a failure has occurred in the computer 100-1 that started the service # 1 by referring to the service status information DB121b (Fig. 5, steps S21 and S22), the service # The computer that inherits 1 is determined by a series of processes (steps S23, S25 to S27) starting from step S23 in FIG. Here, based on the relationship between the computer status indicated by the computer status information DB121c and the services indicated by the service-related information DB121a, the service is selected from the n + m computers in the n-to-m backup configuration cluster system shown in FIG. One normal computer is selected whose start of # 1 does not violate the relationship between services. In the example of the service-related information DB121a shown in FIG. 11, the computers whose start of service # 1 does not violate the relationship between services are the standby computers 100- (n + 1) to 100- (n) which have not started the service. + M), and it is assumed that the computer 100- (n + 1) is selected. In this case, service # 1 is started (taken over) by computer 100- (n + 1).
【0096】
Further, it is assumed that a failure occurs in the computer 100-2 in this state, and the service # 2 started in the computer 100-2 is failed over. In this case as well, the automatic control unit 122c determines the calculator to be taken over in the same manner as described above. Here, m-1 computers 100- (n + 2) to 100- (n) excluding the operating computers 100-1 to 100-n and the computers 100- (n + 1) that inherited service # 1. One of + m), for example, computer 100- (n + 2) is selected as the takeover computer, and service # 2 is started (taken over) on the computer 100- (n + 2). To do.
【0097】
Hereinafter, by the same control, when n m, even if a failure occurs in m of computers 100-1 to 100-n, the service started by the m computers will be provided. , M standby computers 100- (n + 1) ~ 100- (n + m) can be automatically failover. In other words, in the n-to-m backup configuration cluster system shown in Fig. 10, when n m, up to m services can be provided to m standby computers 100- (n + 1) to 100- (n + m). ) Can be failover. If n <m, up to n services can be failover to n out of m standby computers 100- (n + 1) to 100- (n + m).
【0098】
As described above, in the third embodiment of the present invention, the service-related information DB121a is used, and the relationships between all the services (combinations of services) are described in the service-related information DB121a, both of which are between services. An n-to-m backup configuration can be realized by defining the type of relationship = exclusive relationship, the attribute between services = strong relationship, and the scope of application between services = local relationship. This third embodiment combines the relationship between services applied in the first embodiment and the automatic control of services by the automatic control unit 122c, and is a conventional cluster system with an n-to-m backup configuration. It can be said that the setting is simple and unified.
【0099】
[Fourth Embodiment]
Next, a fourth embodiment of the present invention will be described. The feature of this fourth embodiment is that the cluster system applied in the second embodiment realizes cold backup of data by utilizing the relationship between services. Therefore, Fig. 1 and Fig. 8 are used for the configuration of the cluster system.
【0100】
Cold backup of data refers to the operation of backing up data by stopping the service that accesses the data when there is data and the service that accesses the data. That is, a backup form in which access to data does not occur during the backup process is called a cold backup of data.
【0101】
In order to realize cold backup in a conventional cluster system, the user first stops the service that accesses the data, then executes the process of backing up the data, and then starts the service that accesses the data again. There is a need to do. Also, in a conventional cluster system, the service to access the data and the process to back up the data are executed on the computer where the data exists, or the service to access the data is started after the data becomes accessible. It is necessary to control such services and processes by user operations. In the fourth embodiment of the present invention, by using the cluster system applied in the second embodiment, control of this kind of service and processing can be automatically performed as described below.
【0102】
FIG. 12 shows service-related information included in the scenario management mechanism 120 having the configuration of FIG. 8 (used in place of the scenario management mechanism 12 in the cluster system 1 of FIG. 1), which is applied in the fourth embodiment of the present invention. An example of DB121a is shown. In this fourth embodiment, as the service registered in the service-related information DB121a of FIG. 12, the service data corresponding to the following three service data (data itself) and the service data for accessing the data are backed up. Service is prepared. The "data backup service" is a batch job type service applied in the second embodiment.
【0103】
In the service relationship information DB121a shown in FIG. 12, three types of relationships # 1 to # 3 are defined as relationships between services. The contents of the data (record) of the service relationship information DB121a for each of the relationships # 1 to # 3, that is, the service names SN1 and SN2, the types of relationships between services, the attributes of the relationships between services, and the services. Information on the scope of application of the relationship between them is as follows.
【0104】
(Relationship between services # 1) First, the information about the relationship # 1 between services is as follows: Service name SN1 ... Service that accesses data Service service name SN2 ... Relationship between services that correspond to data Types of ... Dependencies / Attributes of relationships between services ... Strong relationships / Scope of relationships between services ... Local relationships.
【0105】
(Relationship between services # 2) Next, the information about the relationship # 2 between services is as follows: Service name SN1 ... Service for backing up data Service name SN2 ... Between services and services corresponding to data Relationship types ... Dependencies / Attribute of relationships between services ... Strong relationships / Scope of relationships between services ... Local relationships (relationships between services # 3) Finally, relationships between services # Information about 3 is: -Service name SN1 ... Service that accesses data Service-Service name SN2 ... Services that back up data Types of relationships between services ... Exclusive relationships-Attributes of relationships between services. .. Scope of strong relationships / relationships between services ... Global relationships.
【0106】
Next, the cold backup of the data according to the fourth embodiment of the present invention will be described with reference to the configurations of FIGS. 1 and 8 and the flowcharts of FIGS. 3 and 9. However, it is assumed that the scenario management mechanism 120 having the configuration shown in FIG. 8 is used instead of the scenario management mechanism 12 in FIG.
【0107】
It is assumed that the "service corresponding to the data" and the "service for accessing the data" have already been started. Here, due to the relationship # 1 between services, the "service corresponding to data" is started first. In addition, both the "data-equivalent service" and the "data-accessing service" are started on the same computer due to the relationship # 1 between the services. Here, it is assumed that both of the above services are started by the computer 10-1 in the cluster system shown in FIG.
【0108】
In such a state, it is assumed that the user has performed an operation to start a "data backup service" on the client terminal for data backup. Then, the operation processing unit 122a in the scenario management mechanism 120 shown in FIG. 8 accesses the service permission status information DB121d and sets the permission status of the service for backing up data to permission.
【0109】
The automatic control unit 122c periodically refers to the service permission status information DB121d, and when the permission status of the "service for backing up data" is set to "permission", the "permission" is performed according to the relationship # 2 between services. The "data backup service" that has become "data backup service" is started on the computer 10-1 in which the "data equivalent service" is started (Fig. 3, steps S2 to S10). In this state, the automatic control unit 122c stops the "service for accessing data" started by the computer 10-1 according to the relationship # 3 between services.
【0110】
The "data backup service" is a batch job type service. Therefore, as in the above example, when the "data backup service" is started on the calculator 10-1, the permission status of the "data backup service" is immediately "disabled" by the batch job control unit 122g. It is "permitted" (Fig. 9 steps S62 to S65). Then, the automatic control unit 122c executes a process of stopping the "data backup service" on the computer 10-1. The process of stopping the "data backup service" is performed after the process associated with the start of the service, that is, the process for backing up the data is completed.
【0111】
By the way, when the "data backup service" on the calculator 10-1 is stopped, the fact that the "data access service" is started does not violate the relationship # 3 between services. Therefore, the automatic control unit 122c restarts the "service for accessing data" (steps S7 to S9 in FIG. 3).
【0112】
As described above, in the fourth embodiment of the present invention, the service-related information DB121a is used to define "a service corresponding to data", "a service for backing up data", and "a service for accessing data". For the "data equivalent service", a strong dependency is set as a local relationship from the "data backup service" and the "data access service", respectively, and the "data backup service" and "data backup service" and "data are backed up". Cold backup of data can be realized by setting a strong exclusive relationship with the "service that accesses data" as a global relationship. This fourth embodiment is applied by combining the relationship between services applied in the second embodiment and the automatic control of services by the automatic control unit 122c, and data is used by using a conventional cluster system. Compared to implementing a cold backup of, the user does not have to start or stop the service. Further, in the fourth embodiment, when the "service for backing up data" is started, it is explicitly specified on which computer the "service corresponding to the data" is started, and the order in which the services are started is specified. There is no need to consider it. Therefore, it can be said that the cluster system according to the fourth embodiment is easier and more unified than the conventional cluster system.
【0113】
[Fifth Embodiment]
Next, a fifth embodiment of the present invention will be described. The feature of the fifth embodiment is that the cluster system applied in the second embodiment realizes online backup of data by utilizing the relationship between services. Therefore, Fig. 1 and Fig. 8 are used for the configuration of the cluster system.
【0114】
Online data backup refers to the operation of backing up data when there is data and a service that accesses the data, without stopping the service that accesses the data.
【0115】
In order to realize online backup in a conventional cluster system, the service to access the data and the process of backing up the data should be executed on the computer where the data exists, or the data should be accessed after the data becomes accessible. It is necessary to control the service and the process such as starting the service to be performed by the user's operation. In a fifth embodiment of the present invention, by using the cluster system applied in the second embodiment, control of this kind of service and processing can be automatically performed as described below.
【0116】
FIG. 13 shows service-related information included in the scenario management mechanism 120 having the configuration of FIG. 8 (used in place of the scenario management mechanism 12 in the cluster system 1 of FIG. 1), which is applied in the fifth embodiment of the present invention. An example of DB121a is shown. In the fifth embodiment, as the service registered in the service-related information DB121a of FIG. 13, a service for accessing service data corresponding to the following three service data, as in the fourth embodiment, is performed. -A service to back up data will be prepared. The "data backup service" is a batch job type service applied in the second embodiment.
【0117】
In the service relationship information DB121a shown in FIG. 13, three types of relationships # 1 to # 3 are defined as relationships between services. Of the relationships # 1 to # 3 between services, the relationships # 1 and # 2 between services are the same as the relationships # 1 and # 2 between services in the fourth embodiment, and thus description thereof will be omitted. On the other hand, the relationship # 3 between services is different from the relationship # 3 between services in the fourth embodiment in the type of relationship between services.
【0118】
That is, the relationship # 3 between services applied in the fifth embodiment is as follows: -Service name SN1 ... Service for backing up data Service name SN2 ... Type of relationship between services for accessing data ... Attributes of dependencies / relationships between services ... Strong relationships / scope of relationships between services ... Global relationships. In this way, in the relationship # 3 between services applied in the fifth embodiment, the service names SN1 and SN2 are opposite to those in the fourth embodiment, and the "type of relationship between services" becomes "dependency". , The feature is that "the scope of application of the relationship between services" is a global relationship.
【0119】
Next, the cold backup of the data according to the fifth embodiment of the present invention will be described with reference to the configurations of FIGS. 1 and 8 and the flowcharts of FIGS. 3 and 9. However, it is assumed that the scenario management mechanism 120 having the configuration shown in FIG. 8 is used instead of the scenario management mechanism 12 in FIG.
【0120】
It is assumed that the "service corresponding to the data" and the "service for accessing the data" have already been started. Here, due to the relationship # 1 between services, the "service corresponding to data" is started first. In addition, both the "data-equivalent service" and the "data-accessing service" are started on the same computer due to the relationship # 1 between the services. Here, both of the above services are assumed to be started by the computer 10-1 in the cluster system of FIG. 1 as in the fourth embodiment.
【0121】
In such a state, it is assumed that the user has performed an operation to start a "data backup service" on the client terminal for data backup. Then, the operation processing unit 122a in the scenario management mechanism 120 shown in FIG. 8 accesses the service permission status information DB121d and sets the permission status of the service for backing up data to permission.
【0122】
The automatic control unit 122c periodically refers to the service permission status information DB121d, and when the permission status of the "service for backing up data" is set to "permission", the relationship between services # 2 and # 3 is followed. The "data backup service" is started on the computer 10-1 in which the "data equivalent service" and the "data access service" are started (Fig. 3, steps S2 to S10).
【0123】
The "data backup service" is a batch job type service. Therefore, when the "data backup service" is started on the computer 10-1, the permission status of the "data backup service" is immediately set to "disapproval" by the batch job control unit 122g (Fig.). 9 steps S62 ~ S65). Then, the automatic control unit 122c executes a process of stopping the "data backup service" on the computer 10-1. The process of stopping the "data backup service" is performed after the process associated with the start of the service, that is, the process for backing up the data is completed.
【0124】
As described above, in the fifth embodiment of the present invention, the service-related information DB121a is used to define "a service corresponding to data", "a service for backing up data", and "a service for accessing data". For the "data equivalent service", a strong dependency is set as a local relationship from the "data backup service" and the "data access service", respectively, and for the "data access service". By setting a strong dependency as a global relationship from the "data backup service", online data backup can be realized. This fifth embodiment is applied by combining the relationship between services applied in the second embodiment and the automatic control of services by the automatic control unit 122c, and is the same as the fourth embodiment. In addition, it can be said that the setting is simpler and more unified than the case of realizing online backup of data using a conventional cluster system.
【0125】
The present invention is not limited to the above embodiment as it is, and at the implementation stage, the components can be modified and embodied within a range that does not deviate from the gist thereof. In addition, various inventions can be formed by appropriately combining a plurality of components disclosed in the above-described embodiment. For example, some components may be removed from all the components shown in the embodiments. Further, the components of different embodiments may be combined as appropriate.
【0126】
[Effect of the invention]
As described in detail above, according to the present invention, with respect to the service provided by the cluster system, the service is provided so as to match the inter-service relationship indicated by the inter-service relationship information stored in the inter-service relationship information storage means. By controlling the start or stop, it is possible to realize advanced and detailed control of services that matches the relationship between services. Therefore, when multiple services are provided by a cluster system, the services can be arranged in the computer so as to be optimal in terms of system availability, effective use of computer resources, and improvement of service processing capacity, and a hot standby type cluster. Higher level of service control than the system can be realized.
[Simple explanation of drawings]
FIG. 1 is a block diagram showing a configuration of a cluster system according to a first embodiment of the present invention.
FIG. 2 is a block diagram showing a configuration of a scenario management mechanism 12 in FIG.
FIG. 3 is a flowchart for explaining a process mainly performed by the automatic control unit 122c of the scenario management mechanism 12 when the user does not specify a computer to start the service in the first embodiment.
FIG. 4 is a flowchart for explaining a process mainly performed by the manual control unit 122b of the scenario management mechanism 12 when a user specifies a computer to start a service in the first embodiment.
FIG. 5 is a flowchart for explaining failover processing mainly by the automatic control unit 122c of the scenario management mechanism 12 in the first embodiment.
FIG. 6 is a flowchart for explaining an inter-service relationship determination process according to the first embodiment.
FIG. 7 is a block diagram showing a configuration of a cluster system according to a first modification of the first embodiment.
FIG. 8 is a block diagram showing a configuration of a scenario management mechanism applied in a second modification of the first embodiment.
FIG. 9 is a flowchart for explaining processing mainly performed by the batch job control unit 122g of the scenario management mechanism 120 in the second modification.
FIG. 10 is a block configuration diagram of an n-to-m backup configuration cluster system according to a third embodiment of the present invention.
FIG. 11 is a diagram showing an example of service-related information DB121a possessed by the scenario management mechanism in the n-to-m backup configuration cluster system of FIG.
FIG. 12 is a diagram showing an example of service-related information DB121a for realizing cold backup of data, which is applied in the fourth embodiment of the present invention.
FIG. 13 is a diagram showing an example of service-related information DB121a for realizing online backup of data, which is applied in the fifth embodiment of the present invention.
[Explanation of symbols]
1 ... Cluster system, 10 ... Cluster system control mechanism, 10-1 ~ 10-N ... Computer, 11-1 ~ 11-N ... Service execution mechanism, 12, 12'... Scenario Management mechanism, 100-1 ~ 100-n ... Working system computer, 100- (n + 1) ~ 100- (n + m) ... Standby system computer, 121 ... Database part, 121a ... Service-related information DB, 121b ... Service status information DB, 121c ... Computer status information DB, 121d ... Service permission status information DB, 121e ... Batch job type service information DB, 122 ... Service control Department, 122a ... Operation processing unit, 122b ... Manual control unit, 122c ... Automatic control unit, 122d ... Inter-service relationship judgment unit, 122e ... Execution control unit, 122g ... Batch job Control unit.
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| JP2007172334A | Cited by | Japan | Examiner |
| JP2010198060A | Cited by | Japan | Examiner |
| US8713352B2 | Cited by | United States of America | Applicant |
| JP2011138454A | Cited by | Japan | Search report |
| US9703653B2 | Cited by | United States of America | Applicant |
2 priority claims, no other members on record
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 2003176959 | Japan | A | |
| JP20030176959 | – | – | – |
16 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of completion of termEXPY | EXPY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Written notification of patent or utility model registrationJAPANESE INTERMEDIATE CODE: R151R151 | R151 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Re-examination (zenchi) completed and case transferred to appeal boardAppealJAPANESE INTERMEDIATE CODE: A912A912 | A912 | |
| Transfer to examiner for re-examination before appeal (zenchi)AppealJAPANESE INTERMEDIATE CODE: A911A911 | A911 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Decision of refusalJAPANESE INTERMEDIATE CODE: A02A02 | A02 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 |
Numbers
- Publication
- 2005011237
- Publication, DOCDB
- 2005011237
- Publication, EPODOC
- JP2005011237
- Application
- 176959
- Application, DOCDB
- 2003176959
- Application, EPODOC
- JP20030176959
Titles2
- English
- CLUSTER SYSTEM AND SERVICE CONTROL PROGRAM
- Japanese
- クラスタシステム及びサービス制御プログラム
Classification
- IPC, 2
- G06F11 20
- G06F15 177