Method for isolating a defective computer in a fault-tolerant multiprocessor system
Abstract
The multi processor system executes a program to identify a faulty computer (101) and a fail safe process is initiated . This consists of generating a command to control the generation of outputs. In the event that a fault does occur the processing function is switched to a non defective computer.

Term
Term ended
Projected expiry passed 28 August 2018, 8.1 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
5 claims: 5 independent, 0 dependent
- 1Method for isolating an identified defective computer by not defective computers in a computer system having at least three computers (R1, R2, R3 in Fig. 4), wherein the at least three computers via system internal communication network (KV12, KV13, KV23) together exchange data and from them calculated results to a common, independent of the system-internal communication network Input / output channel (EAK) spend characterized by the following steps:a) the defective computer is sent a shutdown command (102 in Fig. 1),b) it is checked whether the defective computer results to the common I / O channel outputs (103)c) if the defective computer after a predetermined period of time after the shutdown command still results to the outputs common I / O channel, received a non-defective Calculator also shutdown commands (104). Verfahren zur Isolation eines als defekt identifizierten Rechners durch nicht defekte Rechner in einem Rechnersystem mit mindestens drei Rechnern (R1, R2, R3 in Fig. 4), bei dem die mindestens drei Rechner über ein system internes Kommunikationsnetz (KV12, KV13, KV23) miteinander Daten austauschen und von ihnen berechnete Ergebnisse an einen gemeinsamen, vom system internen Kommunikationsnetz unabhängigen Ein-/Ausgabekanal (EAK) ausgeben, gekennzeichnet durch folgende Schritte:a) dem defekten Rechner wird ein Herunterfahrkommando übermittelt (102 in Fig. 1),b) es wird überprüft, ob der defekte Rechner Ergebnisse an den gemeinsamen Ein-/Ausgabekanal ausgibt (103),c) falls der defekte Rechner nach einer vorbestimmten Zeitspanne nach dem Herunterfahrkommando noch immer Ergebnisse an den gemeinsamen Ein-/Ausgabekanal ausgibt, erhalten die nicht defekten Rechner ebenfalls Herunterfahrkommandos (104).
- 2The method of claim 1, wherein at least one of the non-defective the defective computer transmits the computer shutdown command and wherein at least one of the non-defective computer checks whether the defective computer results to the common I / O channel outputs, and wherein at least one of the non-defective computers to non-defective computers gives the shut-down command, if the defective Computer after a predetermined time interval after the Shutdown command still results in the joint I / O channel outputs. Verfahren nach Anspruch 1, bei dem wenigstens einer der nicht defekten Rechner dem defekten Rechner das Herunterfahrkommando übermittelt und bei dem wenigstens einer der nicht defekten Rechner überprüft, ob der defekte Rechner Ergebnisse an den gemeinsamen Ein-/Ausgabekanal ausgibt, und bei dem wenigstens einer der nicht defekten Rechner den nicht defekten Rechnern das Herunterfahrkommando gibt, falls der defekte Rechner nach einer vorbestimmten Zeitspanne nach dem Herunterfahrkommando noch immer Ergebnisse an den gemeinsamen Ein-/Ausgabekanal ausgibt.
- 3The method of claim 1, wherein the non-defective computer after the defective computer has received the shutdown command, no More data with this exchange (204). Verfahren nach Anspruch 1, bei dem die nicht defekten Rechner, nachdem der defekte Rechner das Herunterfahrkommando erhalten hat, keine Daten mehr mit diesem austauschen (204).
- 4Mehrrechnersytem comprisinga) at least three computers (R1, R2, R3)b) a system-internal communication network (V1, V2, V3 in Fig. 3;KV12, KV13, KV23 in Fig. 4) for exchanging data between the computers,c) computer interfaces (EA1, EA2, EA3), on the computer of them calculated results to a common I / O channel (EWC) spendd) and identification means (IM;M1, IM2, IM3) for identifying a defective computer, characterized,e) that command means (KM, KM1, KM2, KM3) are present, the on the system internal communication network to a defective identified computers transmit a shutdown command,f) that testing means (PM;PM1, PM2, PM3) are present, the check, whether identif ed by the identification means as broken computer Results to the common I / O channel outputs, and the cause that the non-defective computers a is received shutdown command, if the defective computer after a predetermined period of time of the transmission of Shutdown commands to the defective computer starts, yet always results to the common I / O channel outputs. Mehrrechnersytem, umfassend a) mindestens drei Rechner (R1, R2, R3),b) ein systeminternes Kommunikationsnetz (V1, V2, V3 in Fig. 3;KV12, KV13, KV23 in Fig. 4) zum Austausch von Daten zwischen den Rechnern,c) Rechnerschnittstellen (EA1, EA2, EA3), über die die Rechner von ihnen berechnete Ergebnisse an einen gemeinsamen Ein-/Ausgabekanal (EAK) ausgeben,d) und Identifikationsmittel (IM;M1, IM2, IM3) zum Identifizieren eines defekten Rechners, dadurch gekennzeichnet,e) daß Kommandomittel (KM;KM1, KM2, KM3) vorhanden sind, die über das systeminterne Kommunikationsnetz an einen als defekt identifizierten Rechner ein Herunterfahrkommando übermitteln,f) daß Prüfmittel (PM;PM1, PM2, PM3) vorhanden sind, die überprüfen, ob der von den Identifikationsmitteln als defekt identifzierte Rechner Ergebnisse an den gemeinsamen Ein-/Ausgabekanal ausgibt, und die veranlassen, daß den nicht defekten Rechnern ein Herunterfahrkommando übermittelt wird, falls der defekte Rechner nach einer vorbestimmten Zeitspanne, die mit der Übermittlung des Herunterfahrkommandos an den defekten Rechner beginnt, noch immer Ergebnisse an den gemeinsamen Ein-/Ausgabekanal ausgibt.
- 5A multicomputer system as claimed in claim 4, wherein the intrinsic Communications network is designed such that the computers with each other completely apart physically independent lines are intermeshed. Mehrrechnersystem nach Anspruch 4, bei dem das systeminterne Kommunikationsnetz so ausgelegt ist, daß die Rechner untereinander vollständig durch voneinander physikalisch unabhängige Leitungen vermascht sind.
Independent claims5
19 paragraphs, as filed
The invention relates to a method for isolating an identified as defective Computer in a fault tolerant multiprocessor system according to the preamble of claim 1. The invention further relates to a fault-tolerant Multicomputer system according to the preamble of claim 3rd
Fault-tolerant multiprocessor systems - mostly 2-of-3 computer systems- especially such processes are used for control, in which particularly high demands in terms of security and availability the controller are provided. Examples of such processes include the Assurance of railway infrastructure for railway or monitoring Nuclear power plants. A computer system is called fault tolerant if even in the presence of a limited number of hardware and / or Software errors, the system still works. From such Data processing equipment can also request that they rest with the required for process control devices according to the "fail-safe" principle cooperate. This means that even with a total loss the data processing system of the process under any circumstances in a dangerous condition may pass.
A method according to the preamble of claim 1 is known from EP-B1-0 246 218 known. There, however, is not isolation but rather the necessary before an isolation identification of a faulty computer in a redundant data processing system described. The known data processing system consists of several Computer nodes, in turn, as a multi-computer systems - preferably as a 2-of-3 computer systems - Are executed. The nodes are over serial point-to-point connections interconnected. The computers within a each node are also fully meshed with each other. The a Majority decision required reconciliation of each of the computers within a node supplied results ( "Voting") is not central performed, but is distributed over the individual computers.
Furthermore, from DE-A1-41 35 640 a 2-from-3-computer system known in the computer also has its own control bus to each other can exchange data. The computers communicate with the process on one of them independent I / O data bus. If a computer is defective is identified, cause there the two non-defective computers that the Data port is blocked to the I / O bus. Depending on the nature of the defect, however, it may happen that the non-defective computers such an intervention in the Data port of the defective computer fails. It is then possible that the defective computer continues to output erroneous data to the I / O bus. If another computer is broken and also errored data output to the I / O bus, so these data could match by chance. A receiver could these two results then because of their maintain compliance for properly, so they accept and pursue tax acts that result in a non-safe state could.
Finally, from DE-C2-32 08 573 a selection device for a Three computer system is known in which an identified as defective computer via a relay switch from I / O channel is irreversibly separated. Of the defective computer can not there again after a restart without any intervention an operator will be integrated into the system. This solution is comparatively expensive since the relay switch dedicated to each Computer system needs to be developed. Further, control lines for provide control of Reloisschalters.
It is therefore an object of the invention to provide a method by means of which a identif ed as broken computers effectively in a multi-computer system not defective considering the ,, failsafe "principle of the computers can be isolated. The process should not under any circumstances unsafe process conditions lead. Moreover, the method is not intended to the use of additional hardware such as relay switches require.
The invention solves this problem by the teaching specified in claim. 1 Since the invention of the defective computer is prompted shut down, a subsequent re-integration of the computer - unless a reboot is successful - even without the intervention of an operator possible. The security of the system is ensured by the fact that the non- Shutdown defective computer even if the defective computer the Shutdown command has not complied with successfully. namely, if the defective computer to spite shutdown commands further results outputs the input / output channel, then only by a shutdown of the non-defective computers are reliably prevented it up to the aforementioned random match results from two defective computers and thereby comes to an unsafe condition.
In an advantageous embodiment of claim 2 is the Task, the defective computer to send a shutdown command, taken from the non-defective computers. The non-defective computers Check also whether the defective computer results common to the I / O channel outputs. Further, the non-defective computer boots up itself down if the defective computer after a predetermined period after the shutdown command still results to the common I / O channel outputs. Thus, the computer system is in a safe state, because it is - as in secure computer systems usual - provided that a safe output always two computers must participate. A single computer alone can not under any circumstances effectively engage the consumers to due process. Additional hardware devices are not necessary in this embodiment.
filters In a further advantageous embodiment according to claim 3 the non-defective computers to exchange data with the defective Computer on after it has received the shutdown command. This is the defective computer signaled again that he all output Set to the common I / O channel and Shutdown should.
The invention is described with reference to embodiments and the Drawings in detail. Show it:<sl><li>Fig. 1: A representation of the method according to claim 1 in the form of a flow chart,</li><li>Fig. 2: A representation of the method according to claim 3 in the form of a flow chart,</li><li>Fig. 3: A schematic representation of one embodiment of a Multicomputer system according to the invention of claim 4,</li><li>Fig. 4: A schematic representation of another embodiment a multi-computer system according to the invention of claim 4.</li></sl>
Fig. Figure 1 illustrates the inventive method 100 according to claim 1 in the form of of a flow chart. A multi-computer system in which the method 100 may be advantageously applied, Fig. 4. This multi-computer system consists in this example of three computers R1, R2, R3, via a system internal communication network are intermeshed, ie each Computer can have its own communication link with each Replacing other computer data. Thus, between the computers R1 and R2 communication connection KV12, between the computers R2 and R3 communication connection KV23 and between the computers R1 and R3 communication connection KV13. The exchange of data between the multicomputer system, and the process to be controlled via a common, independent of the system-internal communication network I / O channel EAK.
In a first step 101, by the inventive method 100 a defective computer identified. In general, this identification is performed characterized in that the three computers R1, R2, R3 via the intrinsic Communication network, the calculated results of them with each other change. If all the results agree with each other, so it goes System assumes that no host is defective. however soft the result a computer of the results of the other two computers from, so is this result as flawed and therefore the corresponding computer as defective viewed. Because of the complete meshing of the computer each other is always a clear allocation of results ensured so that the defective computer can be clearly identified. A particularly advantageous method for error identification is already in the cited patent EP-B1-0 246 218 describes, on to this Reference is made.
In a second step 102, in accordance with the invention to be defective identified computer a shutdown command received. This Command can, for example, by external command means are, as in the description of the invention Multicomputer system is described in detail below. Especially then However, if the error ID decentralized, ie distributed to all computers, is performed, it is advisable, broken this command not to Computers or start from one of the non-defective computers allow. Under shutdown is understood here that the computer be concerned Application program, as well as its operating system with all Interface drivers properly terminated. Such proper Termination includes, for example conducting diagnostics and Storing data on a non-volatile carriers. After this Shutdown, the computer will not back more, either through the I / O channel still on the system internal communication network. If no errors were detected and it allows the security concept, will the computer then rebooted. This typically includes extensive hardware testing a with.
In a next step 103 it is checked whether the detected as defective computer continue outputting results to the common I / O channel. This amounts to a review of whether the defective computer the Shutdown command is actually complied with. As already mentioned, a computer can after shutdown no issues to the common I / O channel make more. yet Can such demonstrate expenditures, then it can be assumed that the defect such is severe, that a shutdown was no longer possible. If he makes computer output to the input / output channel, can the I / O channel be even detected. Again, it is again possible, this task by external test equipment or defective of not the carry computers even allow.
If it is found during this check that the broken computer to a predetermined period still issues to the common I / O channel makes, so go the non-defective computers in one step 104 down to. This shutdown will of external test equipment are caused or of the non-defective computers themselves. By Shutdown is, as explained above, the system to a safe state over because the defective computer can alone make any expenditure on the side of expenditure receiving recipient an effect could unfold. That task must always - in the case of a 2-of-3 computer system - Correspond at least two results.
If on the other hand found in this test 103 that the defective computer after a predetermined period of time no issues to the common makes I / O channel more, so go the non-defective computers in the of conventional 2-of-3 computer systems known manner with their Computing activity continued. The predetermined time period should be at least as long as be like a broken computer needs to shut down. Otherwise could shut down the non-defective computers, although the defective Computer is shutting down successfully and thus a dangerous condition not may occur.
Another embodiment of the method according to Claim 3 is shown in Fig. 2. Steps 201 and 202 correspond to the Steps 101 to 102. In addition to the steps shown in Fig. 1 is here However, provided that, in a step 204, the non-defective computers of from the data exchange with the defective computer via the Setting the internal system communications network. This is the defective Computer again signaled that it is no longer as an equal calculator is recognized in the system and should therefore shut down. The non defective computers can exchange data in each case after the Setting the transmission of the shutdown commands 202 or, as in shown Fig. 2, in step 203, do this depends on whether the defective Calculator adjusts on its own data exchange. The additional Step 204 reduces the likelihood that the defective computer not shutting down and therefore to a shutdown and the not defective computer in step 207 comes.
An embodiment of an inventive multi-computer system is in shown FIG. 3. The multiprocessor system includes three computer R1, R2, R3, the via input / output ports EA1, EA2, EA3 with a common I / O channel EAK are connected. According to the invention are further provided Identification means IM for identifying defective computers and Command means KM to issue shutdown commands. This means are connected to the computer via a system internal communication network, comprising the compounds V1, V2 and V3. In this way, the Identification means IM and the command means KM with the computers R1, R2, R3 communicate. In addition, test equipment PM are present, both with the I / O channel EWC as well as with the three computers R1, R2, R3 are connected. These agents have the task of reviewing whether the of the identification means as defective recognized calculator results to the common I / O channel outputs. If the test equipment to determine that the defective computer after a predetermined period still cause outputting results to the common input / output channel, the test means that the non-defective computers shut down and thus a safe state is achieved.
In the embodiment shown in Fig. 4, the Identification means IM, the command means KM and test equipment not PM externally arranged, but divided among the three computers. Preferably is it in this means Im1 ... IM3, KM1 KM3 ... and PM1 PM3 to ... Software modules, which take over the corresponding functions.
4 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| WO0182010A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| EP1205834A2 | Cited by | European Patent Office (EPO) | Search report |
| US6574744B1 | Cited by | United States of America | Applicant |
| EP2352109A1 | Cited by | European Patent Office (EPO) | Search report |
| EP1148396A1 | Cited by | European Patent Office (EPO) | Search report |
| EP1205834A3 | Cited by | European Patent Office (EPO) | Search report |
| US8745735B2 | Cited by | United States of America | Applicant |
| EP2352109A4 | Cited by | European Patent Office (EPO) | Search report |
| DE4135640A1 | Cites | Germany | Search report |
| WO9203787A1 | Cites | World Intellectual Property Organization (WIPO) | Search report |
8 members in 4 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 19740136 | Germany | A | |
| 19740136 | Germany | A | |
| 19740136 | Germany | – | |
| 19740136 | – | – | – |
| DE1997140136 | – | – | – |
Members8
| Document | Office | Kind | |
|---|---|---|---|
| EP0902369A2This record | European Patent Office (EPO) | A2 | |
| DE19740136A1 | Germany | A1 | |
| EP0902369A3 | European Patent Office (EPO) | A3 | |
| EP0902369B1 | European Patent Office (EPO) | B1 | |
| AT230132T | Austria | T | |
| ATE230132T1 | Austria | T1 | |
| DE59806695D1 | Germany | D1 | |
| ES2185131T3 | Spain | T3 |
50 legal events, as 7 offices reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | Office | |
|---|---|---|---|
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Announcement of lapse in spainLapsedFD2A | FD2A | ES | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| ExpiryMK07 | MK07 | AT | |
| Patent expired after termination of 20 yearsExpiredPE20 | PE20 | GB | |
| Expiry of rightR071 | R071 | DE | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Fee paymentPLFP | PLFP | FR | |
| Fee paymentPLFP | PLFP | FR | |
| Patent ceasedCeasedPL | PL | CH | |
| Be: lapsedLapsedBERE | BERE | EP | |
| No opposition filedOpposition26N | 26N | EP | |
| No opposition filed within time limitOppositionORIGINAL CODE: 0009261PLBE | PLBE | EP | |
| Information on the status of an ep patent application or granted ep patentGrantedSTATUS: NO OPPOSITION FILED WITHIN TIME LIMITSTAA | STAA | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Fr: translation filedET | ET | EP | |
| Definitive protectionFG2A | FG2A | ES | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Corresponds to:REF | REF | EP | |
| Corresponds to:REF | REF | EP | |
| Gb: translation of ep patent filed (gb section 77(6)(a)/1977)GBT | GBT | EP | |
| European patent takes effect as a national patent in ch/liEP | EP | CH | |
| Designated contracting statesAK | AK | EP | |
| European patent grantedGrantedNOT ENGLISHFG4D | FG4D | GB | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Corresponds to:REF | REF | EP | |
| (expected) grantORIGINAL CODE: 0009210GRAA | GRAA | EP | |
| Despatch of communication of intention to grant a patentORIGINAL CODE: EPIDOS IGRAGRAH | GRAH | EP | |
| Despatch of communication of intention to grantORIGINAL CODE: EPIDOS AGRAGRAG | GRAG | EP | |
| Despatch of communication of intention to grant a patentORIGINAL CODE: EPIDOS IGRAGRAH | GRAH | EP | |
| Despatch of communication of intention to grantORIGINAL CODE: EPIDOS AGRAGRAG | GRAG | EP | |
| First examination report despatched17Q | 17Q | EP | |
| Designation fees paidAT BE CH DE DK ES FI FR GB IT LI NL PT SEAKX | AKX | EP | |
| Request for examination filed17P | 17P | EP | |
| Designated contracting statesAK | AK | EP | |
| Request for extension of the european patentAL;LT;LV;MK;RO;SIAX | AX | EP | |
| Search report despatchedORIGINAL CODE: 0009013PUAL | PUAL | EP | |
| Designated contracting statesAK | AK | EP | |
| Request for extension of the european patentAL;LT;LV;MK;RO;SIAX | AX | EP | |
| Public reference made under article 153(3) epc to a published international application that has entered the european phaseORIGINAL CODE: 0009012PUAI | PUAI | EP |
Numbers
- Publication
- 0902369
- Publication, DOCDB
- 0902369
- Publication, EPODOC
- EP0902369
- Application
- 98440187
- Application, DOCDB
- 98440187
- Application, EPODOC
- EP19980440187
Titles3
- German
- Verfahren zur Isolation eines defekten Rechners in einem fehlertoleranten Mehrrechnersystem
- English
- Method for isolating a defective computer in a fault-tolerant multiprocessor system
- French
- Méthode pour l'isolation d'un ordinateur défectueux dans un système à multiprocesseur à tolérance de fautes
Classification
- CPC, 3
- G06F11/181
- G06F11/0796
- G06F11/182
- IPC, 4
- G06F11 00
- G06F11 16
- G06F11 18
- G06F15 16
Designated states25
- Contracting states, 19
- Austria
- Belgium
- Switzerland
- Cyprus
- Germany
- Denmark
- Spain
- Finland
- France
- United Kingdom
- Greece
- Ireland
- Italy
- Liechtenstein
- Luxembourg
- Monaco
- Netherlands (Kingdom of the)
- Portugal
- Sweden
- Extension states, 6
- Albania
- Lithuania
- Latvia
- North Macedonia
- Romania
- Slovenia