Handling errors in an error-tolerant distributed computer system
Abstract
Method of dealing with faults in a fault-tolerant distributed computer system, as well as such a system, with a plurality of node computers (K 1 . . . K 4 ) which are connected by means of communication channels (c 11 . . . c 42 ) and access to the channels takes place according to a cyclic time slicing method. Messages leaving node computers (K 1 . . . K 4 ) are checked by independently formed guardians (GUA) which convert a message burdened with an SOS ("slightly off specification") fault either into a correct message or into a message which can be recognised by all node computers as clearly incorrect.

Term
Term ended
Expired 10 October 2020, 6 years ago.
- Priority and filed
- Granted
- Expired
- Today
6 claims: 6 independent, 0 dependent
- 1Method for tolerating slightly-off-specification (SOS) errors in a fault-tolerant distributed computer system in which a large number of node computers 111, 112, 113 and 114 are connected via one or more communication channels 121 and where each node computer has an autonomous communication control unit with the corresponding connections to the communication channels and where the communication channels are accessed according to a cyclical time slice method and 1. Verfahren zur Tolerierung von slightly-off-specification (SOS) Fehlern eines fehlertoleranten verteilten Computersystems, in dem eine Vielzahl von Knotenrechnern 111, 112, 113 und 114 über einen oder mehrere Kommunikationskanäle 121 verbunden sind und wo jeder Knotenrechner über eine autonome Kommunikationskontrolleinheit mit den entsprechenden Anschlüssen an die Kommunikationskanäle verfügt und wo der Zugriff auf die Kommunikationskanäle entsprechend einem zyklischen Zeitscheibenverfahren erfolgt und A A. AT 41 0 490 B where the correctness of each outgoing message from a node computer is checked by an independent BusGuardian 122, characterized in that a correct, independent BusGuardian converts a wrong SOS message either into a correct message or into a message sent by all recipients can be identified as clearly incorrect. AT 41 0 490 B wo die Korrektheit jeder von einem Knotenrechner ausgehenden Nachricht durch einen unabhängigen BusGuardian 122 überprüft wird, dadurch gekennzeichnet, daß ein korrekter unabhängiger BusGuardian eine SOS-falsche Nachricht entweder in eine korrekte Nachricht oder in eine Nachricht umformt, die von allen Empfängern als eindeutig inkorrekt erkannt werden kann.
- 2Method according to Claim 1, characterized in that the independent BusGuardian checks with reference to its independent time base whether the start of each message sent by the communication controller falls within the start time window of the message known a priori to the BusGuardian and, if this is not the case, the Communication channel shoots immediately, so that an incomplete message is created that is recognized as incorrect by all recipients. 2. Verfahren nach Anspruch 1, dadurch gekennzeichnet, daß der unabhängige BusGuardian unter Bezugnahme auf seine unabhängige Zeitbasis überprüft, ob der Beginn jeder vom Kommunikationskontroller gesendeten Nachricht innerhalb des dem BusGuardian a priori bekannten Beginnzeitfensters der Nachricht fällt und der, falls dies nicht der Fall ist, den Kommunikationskanal sofort schießt, damit eine unvollständige Nachricht entsteht, die von allen Empfängern als fehlerhaft erkannt wird.
- 3Method according to Claim 1, characterized in that the BusGuardian 122 regenerates the incoming analog signal of each message in the time and value domain taking into account the coding rules known to the BusGuardian and with reference to its local time base and its local power supply. 3. Verfahren nach Anspruch 1, dadurch gekennzeichnet, daß der BusGuardian 122 das eingehende Analogsignal jeder Nachricht im Zeit- und Wertebereich unter Berücksichtigung der dem BusGuardian bekannten Codierungsvorschriften und unter Bezugnahme auf seine lokale Zeitbasis und seine lokale Stromversorgung regeneriert.
- 4Verfahren nach einem oder mehreren der Ansprüche 1 bis 3, dadurch gekennzeichnet, daß aufgrund der Konstruktion des BusGuardians ausgeschlossen werden kann, daß ein fehlerhafter BusGuardian Nachrichten mit einer korrekten CRC und einer korrekten Länge aus sich heraus generieren kann. 4th Method according to one or more of Claims 1 to 3, characterized in that, due to the construction of the BusGuardian, it can be ruled out that a faulty BusGuardian can generate messages with a correct CRC and a correct length on its own.
- 6Computerarchitektur, dadurch gekennzeichnet, daß die BusGuardians eines jeden Broadcastkanals, die das Verfahren nach einem der Ansprüche 1 bis 4 durchführen, in eine pro Broadcastkanal zentralen Verteilereinheit, die über eine unabhängige Stromversorgung und über eine unabhängige fehlertolerante verteilte Uhrensynchronisation verfügt, integriert werden. 6th Computer architecture, characterized in that the BusGuardians of each broadcast channel, which carry out the method according to one of Claims 1 to 4, are integrated into a central distribution unit for each broadcast channel, which has an independent power supply and an independent, fault-tolerant distributed clock synchronization.
Independent claims6
28 paragraphs in 4 sections, as filed
TECHNICAL ENVIRONMENT
This invention relates to a distributed, time-controlled computer architecture for reliable real-time applications, such as is used in the field of control of embedded systems.
BACKGROUND TO THIS INVENTION
Safety-critical technical applications, ie applications where an error can lead to a catastrophe, are increasingly being managed by distributed, fault-tolerant real-time computer systems.
In a distributed, fault-tolerant real-time computer system, consisting of a number of node computers and a real-time communication system, each individual failure of a node computer should be tolerated. At the core of such a computer architecture is a fault-tolerant real-time communication system for the predictably fast and secure exchange of messages.
A communication protocol that meets these requirements is described in the cited European Patent 0 658 257 v. December 18, 1996 as well as in the international patent application PCT / AT 93/00138. The protocol has become known under the name Time-Triggered Protocol / C (TTP / C). It is based on the well-known cyclical time slice method (TDMA time-division multiple access) with time slices defined a priori. TTP / C uses a fault tolerant clock synchronization technique disclosed in US Patent 4,866,606.
TTP / C assumes that the communication system supports a logical broadcast topology and that the node computers show fail-silence behavior, ie either the node computers function correctly in the value domain and in the time domain, or they are quiet. The prevention of errors in the time domain, ds, so-called Babbling Idiot errors, is achieved in TTP / C by an independent error detection unit, the BusGuardian, which has an independent time base and continuously checks the time behavior of the node computer. In order to implement the fault tolerance, several fail-silent node computers are combined to form a fault-tolerant unit (FTU) and the communication system is replicated. As long as a node computer of an FTU and a replica of the communication system are functioning, the services of the FTU in the time and value domain are provided in good time.
A logical broadcast topology of communication can be physically established either by a distributed bus system, a distributed ring system or by a central distribution unit (eg a star coupler) with point-to-point connections to the node computers. If a distributed bus system or a distributed ring system is set up, each node computer must have its own BusGuardian. If, on the other hand, a central distribution unit is used, all BusGuardians can be integrated into this distribution unit, which can effectively enforce regular transmission behavior in the time domain based on the global observation of the behavior of all nodes (see PCT / AT 2000/174).
In a distributed computer system, errors that can lead to an inconsistent system state are particularly critical. An example is a “brake-by-wire” application in a car, in which a central brake computer sends braking messages to the four bike computers on the wheels. If a braking message is correctly received by two bike computers and the other two bike computers do not receive the message, this results in an inconsistent state. If two wheels that are on the same side of the vehicle are braked, the vehicle can get out of control. The type of error described here is also referred to in the literature as a Byzantine error (Kopetz, p.133). The quick detection and correct handling of Byzantine errors is one of the difficult problems of computer science.
A subclass of the Byzantine errors is formed by the Slightly-Off-Specification (SOS) errors. An SOS error can occur at the interface between analog technology and digital technology. In realizing a data transmission, each logical bit on the line can be represented by a signal value (eg voltage from a specified voltage tolerance interval) during a specified time interval. A correct sender must have its o
AT 41 0 490 B
Generate analog signals within the specified tolerance intervals to ensure that all correct receivers interpret these signals correctly. If a sender of a broadcast message generates a signal just outside the specified interval (slightly off specification) (in the range of values, in the time range, or in both), it can happen that some receivers interpret this signal correctly, while other receivers cannot interpret the signal correctly. We falsely label such a broadcast message as SOS. As a result, a Byzantine fault, as described above in the braking system, can occur. Such an error can be caused by a faulty voltage supply, a faulty clock generator or a component that has deteriorated due to aging. SOS errors cannot prevent the transmission of a message on two communication channels if the cause of the error, for example a faulty clock in the computer node that generates the bit sequence, affects both channels.
It is a principle of safety technology to recognize occurring errors at the earliest possible point in time, in order to be able to take countermeasures before subsequent errors cause further damage. This principle is complied with in the cited TTP / C protocol (European Patent 0 658 257) in that SOS errors are consistently recognized via the membership algorithm of the TTP / C protocol within a maximum of two TDMA rounds. Since SOS errors are typically very rarely occurring transient errors, in the existing prototype implementation of TTP / C SOS errors are assigned to the class of near-coincident multiple errors, which also occurs very rarely, and how these are treated.
The present invention relates to an innovative method of how the class of SOS errors can be tolerated in a time-controlled architecture by means of suitable architectural measures and a construction according to the invention of the BusGuardian.
DESCRIPTION OF A REALIZATION
The objective described above and other new properties of the present invention are explained with reference to the following figures using the example of the fault-tolerant protocol TTP / C.
1 shows a system of four node computers 111, 112, 113 and 114. A node computer forms an exchangeable unit. Each node computer is connected by a point-to-point connection 121 to one of the replicated central distribution units 101 or 102. Between each output of a node computer and each input of the distribution unit there is a BusGuardian 122, which is either implemented independently or can be integrated into the distribution unit. The two connections 131 and 132 between the distribution units 101 and 102 are used for mutual monitoring and the exchange of information between the central distribution units.
2 shows a node computer 111, the two connections 121 and the two BusGuardians 122 and 132, as well as the connections 123 and 133 to the other node computers of the distributed computer system. As shown in Fig. 1, the BusGuardians can be integrated into two independent central distribution nodes. From a logical point of view, the three subsystems, the node computer 111 and the two BusGuardians 122 and 132 form the Fault Containment Unit (FCU) 140, even if the BusGuardians are physically integrated in the central distribution units 101 and 102.
We designate any fault of an active subsystem, for example the node computer 111, as being active as required (unconstrained active). We denote a fault in a passive subsystem, e.g. of a BusGuardian or a connection line 121 or 122 as arbitrarily passive (unconstrained passive), if the construction of the passive subsystem ensures that this subsystem cannot generate a bit sequence by itself, that is, without an input from an active subsystem Recipient can be interpreted as a syntactically correct message. A message is syntactically correct if the CRC check does not indicate an error, it has the expected correct length, complies with the coding rules and arrives within the expected time interval.
If a passive subsystem does not have the knowledge of how to generate a correct CRC (has no access to the CRC generation algorithm) and how long a correct message must be, the probability is that this is due to random statistical processes
AT 41 0 490 B (malfunctions) a syntactically correct message arises, negligibly small.
The FCU 140 can convert any active error of the node computer 111 or any passive error of one of the two BusGuardians 122 into an error that is not a Byzantine error, if the following assumptions are met:
(i) a correct node computer 111 sends the same syntactically correct message on both channels 122 and 132 and (ii) a correct BusGuardian 122 or 123 converts an SOS incorrect message from node computer 111 either into a syntactically correct message or into a message, which can be recognized as clearly incorrect by all recipients (not SOS message) and (iii) while a message is being sent, a maximum of one of the listed subsystems is defective.
Due to the error assumption (iii), only one of the three listed subsystems (111, 122, 123) can be defective. If the node computer 111 is faulty in any way, then the two BusGuardians 122 and 123 are not faulty and, according to assumption (ii), generate non-SOS messages. If one of the two BusGuardians is passively defective in any way, the node computer 111 generates a syntactically correct message and transmits this syntactically correct message to both BusGuardians 122 and 123 (assumption i). The correct BusGuardian now correctly transmits the message to all recipients. Due to the reception logic and the self-confidence principle of TTP / C, in this case all correct recipients will select the correct message and classify the sending node computer as correct. In order to tolerate SOS errors, no change in the TTP / C protocol is necessary.
Any message can be SOS incorrect for three reasons:
(i) the message has an SOS error in the range of values and / or (ii) the message has an internal SOS error in the time range (e.g. timing error within the code) and / or (iii) the transmission of the message is just outside of the specified Transmission interval started.
A correct BusGuardian converts these causes of errors into non-SOS errors as follows:
(i) The output values of the message are regenerated by the BusGuardian with the independent power supply of the BusGuardian.
(ii) The coding of the message is regenerated by the BusGuardian with the independent timer of the BusGuardian.
(iii) The BusGuardian blocks the transmission channel as soon as it detects that the transmission has started outside the specified time interval. This means that all recipients receive severely garbled messages that are recognized as faulty.
Blocking the transmission channel immediately after the specified end of the transmission time of a message is generally not sufficient to prevent SOS errors, since it cannot be ruled out that a message that has been weakly garbled by the blocking can be the reason for an SOS error of an inherently error-free BusGuardian . If both BusGuardians mutilate the message weakly in the same way, an SOS error can occur at system level.
Finally, it should be noted that this invention is not limited to the implementation described with four node computers, but can be expanded as required. It can be used not only with the TTP / C protocol, but also with other time-controlled protocols.
Contents4
1 sheet
Sheet 1
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| WO0113230A1 | Cites | World Intellectual Property Organization (WIPO) | Search report |
| EP0658257A1 | Cites | European Patent Office (EPO) | Search report |
| US4866606A | Cites | United States of America | Search report |
| US5694542A | Cites | United States of America | Search report |
| WO9406080A1 | Cites | World Intellectual Property Organization (WIPO) | Search report |
16 members in 8 offices
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 17232000 | Austria | A | |
| AT20000001723 | – | – | – |
Members16
| Document | Office | Kind | |
|---|---|---|---|
| WO0231656A2 | World Intellectual Property Organization (WIPO) | A2 | |
| AU9146701A | Australia | A | |
| WO0231656A3 | World Intellectual Property Organization (WIPO) | A3 | |
| ATA17232000A | Austria | A | |
| AT410490BThis record | Austria | B | |
| KR20030048430A | Republic of Korea | A | |
| EP1325414A2 | European Patent Office (EPO) | A2 | |
| US2004030949A1 | United States of America | A1 | |
| JP2004511056A | Japan | A | |
| EP1325414B1 | European Patent Office (EPO) | B1 | |
| AT265063T | Austria | T | |
| ATE265063T1 | Austria | T1 | |
| DE50102075D1 | Germany | D1 | |
| US7124316B2 | United States of America | B2 | |
| JP3953952B2 | Japan | B2 | |
| KR100848853B1 | Republic of Korea | B1 |
2 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| ExpiryMK07 | MK07 | |
| ExpiryMK07 | MK07 |
Numbers
- Publication, DOCDB
- 410490
- Publication, EPODOC
- AT410490B
- Application
- 172300
- Application, DOCDB
- 17232000
- Application, EPODOC
- AT20000001723
Titles2
- German
- VERFAHREN ZUR TOLERIERUNG VON ''SLIGHTLY-OFF- SPECIFICATION'' FEHLERN IN EINEM VERTEILTEN FEHLERTOLERANTEN ECHTZEITCOMPUTERSYSTEM
- English
- METHOD FOR tolerating '' SLIGHTLY OFF-SPECIFICATION 'ERRORS IN A DISTRIBUTED FAULT TOLERANT REAL-TIME COMPUTER SYSTEM
Classification
- CPC, 4
- G06F11/16
- G06F11/00
- G06F11/0757
- G06F2201/805
- IPC, 2
- G06F13 00
- G06F11 00