Computer system.
Abstract
Computer system having at least two computer units, the computer system being equipped with at least one means for detecting faulty conditions of the computer units and, upon detecting faulty conditions in an active computing system, the activity is withdrawn from this and allocated to another intact standby computer unit, characterised in that an independent (autarchical) diagnostic computer unit is provided as means for monitoring and detecting faulty conditions in the individual computer units, in that, upon detecting a faulty condition in an active computer unit, this diagnostic computer unit decides to which other computer unit the activity is switched, and in that, by means of a highly reliable hardware component, exactly one computer unit is always selected as active computer unit. Computer system with a very high availability at minimum hardware cost. <IMAGE>

Term
Term ended
Projected expiry passed 27 November 2013, 12.8 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
9 claims: 6 independent, 3 dependent
- 1Rechnersystem mit wenigstens zwei Rechnereinheiten,zu wobei das Rechnersystem mit wenigstens einem Mittel zur Erkennung von fehlerhaften Zuständen der Rechnereinheiten ausgestattet ist und wobei die Erkennung von fehlerhaften Zuständen bei einem aktiven Rechensystem diesem die Aktivität entzogen und einer anderen intakten stand-by Rechnereinheit zugeteilt wird, dadurch gekennzeichnet, daß als Mittel zur Überwachung und Erkennung von fehlerhaften Zuständen bei den einzelnen Rechnereinheiten eine autarke Diagnose-Rechnereinheit vorgesehen ist, daß bei Erkennung eines fehlerhaften Zustandes bei einer aktiven Rechnereinheit diese Diagnoserechnereinheit entscheidet, auf welche andere Rechnereinheit die Aktivität umgeschaltet wird und daß mittels einer hoch zuverlässigen Logikschaltung stets genau eine Rechnereinheit als aktive Rechnereinheit selektiert wird.
- 2Rechnersystem nach Anspruch 1, dadurch gekennzeichnet, daß die Erkennung von fehlerhaften Zuständen bei einer Rechnereinheit von der Diagnoserechnereinheit als Alarm an eine übergeordnete Stelle gemeldet wird.
- 3Rechnersystem nach einem der vorhergehenden Ansprüche, dadurch gekennzeichnet, daß für Prüfzwecke und Wartungsmaßnahmen die Diagnoserechnereinheit eine manuelle Umschaltung der Aktivität von einer Rechnereinheit auf eine andere ermöglicht.
- 4Rechnersystem nach einem der vorhergehenden Ansprüche, dadurch gekennzeichnet, daß die zur Erkennung von fehlerhaften Zuständen erforderlichen Informationen (Report) von den einzelnen Rechnereinheiten jeweils über eine Leitung an die Diagnoserechnereinheit übertragen wird.
- 5Rechnersystem nach Anspruch 4, dadurch gekennzeichnet, daß die Diagnoserechnereinheit mit den einzelnen Rechnereinheiten jeweils über eine Leitung verbunden ist, über die die empfangenen Informationen (Report R) von der Diagnoserechnereinheit quittiert (Quittung Q) wird.
- 6Rechnersystem nach einem der vorhergehenden Ansprüche, dadurch gekennzeichnet, daß die Diagnoserechnereinheit mit den einzelnen Rechnereinheiten jeweils über eine Leitung verbunden ist, über die eine Rechnereinheit aktivierbar ist.
- 7Rechnersystem nach Anspruch 6, mit zwei Rechnereinheiten, dadurch gekennzeichnet, daß zur Selektion des aktiven Rechners als Logikschaltung ein Inverter eingesetzt wird, der in die eine Aktivierungsleitung eingefügt ist.
- 8Rechnersystem nach einem der vorhergehenden Ansprüche, dadurch gekennzeichnet, daß die Rechnereinheiten über wenigstens ein Leitungssystem zum Austausch von Daten miteinander verbunden sind.
- 9Rechnersystem nach einem der vorherghenden Ansprüche, dadurch gekennzeichnet, daß die Diagnoserechnereinheit ein Single-Chip-Mikroprozessor ist.
Independent claims9
10 paragraphs, as filed
0001The invention relates to a computer system according to the preamble of claim 1. Such computer systems are known, for example from German published patent application DE 40 04 709 A1. In this system, for the mutual monitoring of the two processors, each processor is equipped with a so-called watchdog, which sends control signals to the other computer, by means of which the other computer can check the functionality of the computer emitting the control signal. The disadvantage of such a system is the high hardware expenditure, but also the high software expenditure and relatively low availability. A two-of-three computer system provides higher availability, which, however, has to be paid for with a very high expenditure of hardware and software.
0002The present invention was based on the object of specifying a computer system of the type mentioned at the outset which achieves a very high level of system availability while at the same time requiring relatively little hardware and software.
0003This object is achieved by the means of claim 1. Advantageous refinements result from the subclaims.
0004The 1-out-of-n-computer system according to the invention has a very high availability with low hardware costs at the same time. This is achieved by minimizing the linear component of the system unavailability that is always present (it is always a single instance that makes the selection). In particular, the exemplary embodiment 1-of-2 computer system is very economical, since the more complex computer units are only present twice and the diagnostic computer as a single-chip processor is very inexpensive.
0005There now follows the description of the invention with reference to the figures. FIG. 1 shows a 1 of 2 computer system as a block diagram. FIG. 2 shows an equivalent circuit diagram for calculating the availability. FIG. 1 shows at the top the diagnostic computer DR, which is connected to the two computer units R1 and R2 each via a line report R and a return line acknowledgment Q. The main task of the diagnostic computer is the selection of the active computer unit. If the arbitration were left to the two computer units themselves, there would be a risk that a defective computer unit would not relinquish control or falsely force control. Furthermore, an activation line is drawn, which leads from the diagnostic computer DR to the left computer R1. An inverter 1 is connected to this activation line, the output of which leads to the activation / activation input of the right-hand computer R2. This inverter ensures that exactly one computer unit is always active. In principle, there would be the possibility to control the activation using (different) ports present in the diagnostic computer. A defective diagnostic computer could then select both computer units or no computer unit. As will be demonstrated arithmetically below, the isolated inverter circuit in particular causes the high unavailability of the system described.
0006In the case of more than two mutually redundant computer units, instead of the inverter, a simple logic circuit ensures that exactly one computer unit is always active. Both computer units R1, R2 continuously send the results of their own self-tests to the diagnostic computer, which acknowledges them (R / Q interface). The results are evaluated by the diagnostic computer, which switches over to the intact (better) computer via the active line already mentioned. The switchover occurs either by isolating or resetting the active line exactly the defective computer from its outside world by switching the driver modules to all external interfaces with high resistance, or by stopping the defective computer itself. According to FIG. 2, the overall unreliability or overall unavailability results<maths id="math0001" num=""><math display="inline"><mrow><mtext>U total = Ui Inverter + Udr * Ur1 + Udr * Ur2 + Ur1 * Ur2 -Ur1 * Ur2 * Udr - Ui * Ur1 * Ur2 - Ui * Ur1 * Udr -Ui * Ur2 * Udr - Ui * Ur1 * Ur2 * Udr</mtext></mrow></math><img file="EP0601424A2_D0001.tif" /></maths><dl id="dl0001"><dt>Ui:</dt><dd>Component unavailability of the inverter</dd><dt>Ur1:</dt><dd>Component unavailability of the first computer unit</dd><dt>Ur2:</dt><dd>Component unavailability of the second computer unit</dd><dt>Udr:</dt><dd>Component availability of the diagnostic computer</dd></dl> Under the conditions<maths id="math0002" num=""><math display="inline"><mrow><mtext>Ur1 = Ur2 = Ur and Udr << Ur and 0 <Ur << 1</mtext></mrow></math><img file="EP0601424A2_D0002.tif" /></maths> From a practical point of view, the total unavailability results:<maths id="math0003" num=""><math display="inline"><mrow><mtext>Utotal = Ui + Ur²</mtext></mrow></math><img file="EP0601424A2_D0003.tif" /></maths> The linear component of the system unavailability and thus the determining factor essentially consists, as the equation above shows, of the inverter (a gate), which inverts the active line. When considering a single failure, even a failure of the diagnostic computer does not lead to a system failure, since an intact computer then necessarily remains or is activated. That is the central idea of the invention.
0007The redundancy manager RM is the part of the software in the computer units which has the task of supplying the diagnostic computer with the information which enables the two computer units to be monitored by the diagnostic computer. In addition, the redundancy manager monitors the messages which are sent from the diagnostic computer to the computer unit, so that the diagnostic computer and the computer unit check each other. Another task of the redundancy manager is to create the conditions to enable an orderly switch from the active computer unit to the stand-by computer unit. To do this, the application software on both computer units must be prepared for the switchover. The switchover preparation, the actual switchover by the diagnostic computer and the subsequent re-installation of the two computer units are carried out automatically without external control. Communication between the diagnostic computer and the computer unit takes place on the application level between the diagnostic computer software and the redundancy manager. Message frames are constantly being exchanged, namely message frames that are sent from a computer unit to the diagnostic computer, reports on line R and message frames from the diagnostic computer to a computer unit, called a receipt, on line Q. Both the active and the stand-by redundancy manager send reports cyclically to the diagnostic computer, which contain information about self-test results, commands to the diagnostic computer and states of the computer unit. This information is evaluated by the diagnostic computer software and answered with a receipt to the respective redundancy manager. Analogous to the content of an RM report, this receipt contains self-test results from the diagnostic computer and commands to the redundancy manager. In systems in which transmission errors between diagnostic computers and computer units cannot be ruled out, communication between these communication partners must be secured using a Layer 2 protocol. In addition, both the reports in the diagnostic computer and the receipts in the computer unit are checked for their plausibility. If the content of a report is incorrect, this is interpreted as an error in the computing unit. Appropriate error handling measures may be initiated by the diagnostic computer. Incorrect acknowledgment contents that occur are classified by the redundancy manager as errors of the diagnostic computer. The diagnostic computer is also monitored by transferring the self-test results to the redundancy manager when the diagnostic computer is acknowledged. The task of the redundancy manager is to evaluate these results and, if necessary, to issue alarm messages. This is also done if implausible acknowledgments are received. This enables maintenance measures on the diagnostic computer.
0008In the first instance, a computer unit RE is switched over due to an external command on the part of the maintenance personnel or due to an error in the active computer. Only the diagnostic computer can set and change the states of the activation line to the two computer units. In the case of an orderly switchover, there is a controlled transition of the activities from the active computer unit to the stand-by computer unit, which in turn becomes active. The rough timing of a switchover from the perspective of the redundancy manager can be outlined as follows:<ul id="ul0001" list-style="none"><li>1. Arrival of a toggle command</li><li>2nd Switchover preparation phase of the passive computing unit</li><li>3rd Switchover preparation phase for the active computer</li><li>4th Switching the state of the computer unit - active line on both computer units through the diagnostic computer</li><li>5. Changeover postprocessing phase on both computer units RE A changeover (panic switchover) forced by the diagnostic computer DR is not realized by an abrupt system transition, but takes place in an orderly manner if possible. The forced switchover to the redundant computer unit is characterized in that it is only initiated by the diagnostic computer when the diagnostic computer has identified an error in the active computer due to a report. The prerequisite for the diagnostic computer is that the passive computer sends correct reports and that the switchover has been prepared by its redundancy manager. Then, analogously to the orderly switchover, the diagnostic computer tries to move the defective computer unit to orderly deliver its activities. If the activities are still submitted to the defective computer unit within a time-out period set by the diagnostic computer, then an orderly switchover is carried out. In general, however, it cannot be assumed that a defective computer unit will send a confirmation of correct preparation to the diagnostic computer within the specified time period. In this case it switches automatically after 'time-out'. With the report of a computing unit, an error of the computing unit recognized in the self-tests is displayed to the diagnostic computer. To do this, the results of the self-tests must be included in the report. These self-tests are carried out cyclically by a diagnostic software (diagnostic manager) implemented on the computing unit and their results are regularly transmitted to the redundancy manager in a defined period.</li></ul>
0009In order to check the functionality of the diagnosis manager and the communication between the diagnosis manager and the redundancy manager, the arrival of the messages of the diagnosis manager is monitored for regularity.
0010It should also be noted that the diagnostic computer must ensure that the joint switching of the two computer unit active lines takes place atomically, that is to say that the two switching times between the active and the passive computer cannot drift apart in time. Regarding the functionality between the redundancy and diagnostics manager, the following should also be noted: The functional properties between the redundancy manager and diagnostics manager are identical for both the active and the passive computer. The redundancy manager does not evaluate the diagnostic results. This only happens in the diagnostic computer. After system start, the diagnostic manager sends a complete set of self-tests or Diagnostic results (power-up diagnostics) to the redundancy manager. The diagnosis manager cyclically sends the diagnosis results of the respective computer unit to the corresponding redundancy manager. These results are sent in closed form and then integrated into the reports by the redundancy manager and passed on to the diagnostic computer without being seen. The communication between diagnosis and redundancy manager is monitored over time, ie If no report is received within a specified period of time, an error by the diagnostic manager is accepted and the diagnostic computer is informed accordingly.
5 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| CN106170416A | Cited by | China | Search report |
| WO2015128639A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| US10447196B2 | Cited by | United States of America | Applicant |
| GB2084770A | Cites | United Kingdom | Search report |
| GB2085205A | Cites | United Kingdom | Search report |
| DE2108836A1 | Cites | Germany | Search report |
| FR2448192A1 | Cites | France | Search report |
| FR2490366A1 | Cites | France | Search report |
| FR2492132A1 | Cites | France | Search report |
| DE3700986A1 | Cites | Germany | Search report |
| US3959638A | Cites | United States of America | Search report |
| DE4134207A | Cites | Germany | Search report |
4 priority claims, no other members on record
Priority claims4
| Document | Office | Kind | Date |
|---|---|---|---|
| 4241319 | Germany | A | |
| 4241319 | Germany | – | |
| DE19924241319 | – | – | – |
| 4241319 | – | – | – |
9 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Application deemed to be withdrawnWithdrawn18D | 18D | |
| Information on the status of an ep patent application or granted ep patentGrantedSTATUS: THE APPLICATION IS DEEMED TO BE WITHDRAWNSTAA | STAA | |
| First examination report despatched17Q | 17Q | |
| Party data changed (applicant data changed or rights of an application transferred)RAP1 | RAP1 | |
| Request for examination filed17P | 17P | |
| Designated contracting statesAK | AK | |
| Search report despatchedORIGINAL CODE: 0009013PUAL | PUAL | |
| Designated contracting statesAK | AK | |
| Public reference made under article 153(3) epc to a published international application that has entered the european phaseORIGINAL CODE: 0009012PUAI | PUAI |
Numbers
- Publication
- 0601424
- Publication, DOCDB
- 0601424
- Publication, EPODOC
- EP0601424
- Application
- 93119147
- Application, DOCDB
- 93119147
- Application, EPODOC
- EP19930119147
Titles6
- German
- Rechnersystem.
- English
- Computer system.
- French
- Système d'ordinateur.
- German
- Rechnersystem
- English
- Computer system
- French
- Système d'ordinateur
Classification
- CPC, 2
- G06F11/20
- G06F11/22
- IPC, 2
- G06F11 20
- G06F11 22
Designated states6
- Contracting states, 6
- Austria
- Switzerland
- Spain
- United Kingdom
- Italy
- Liechtenstein