Fault-tolerant distributed computer system
Abstract
Method for ensuring the fail-silent property in the time domain of remote communication computers (111, . . . 114) of a fault-tolerant distributed computer system, in which a plurality of remote computers are connected via a distributor unit (101, 102), each remote computer has an independent communications controller unit with the corresponding connections to the communication channels (121), and the access to the communication channels occurs by a cyclical time-division multiple access method. The at least one distributor unit makes sure, by virtue of the correct sending behavior of the remote computer that is known a priori by it, that a remote computer can only send to the other remote computers within its statically assigned time slice.

Term
Term ended
Expired 13 August 2019, 7.1 years ago.
- Priority and filed
- Granted
- Expired
- Today
7 claims: 7 independent, 0 dependent
- 1Method for enforcing the fail-silent property in the time domain of node computers of a fault-tolerant distributed computer system in which a large number of node computers are connected via one or more distribution units and where each node computer has an autonomous communication control unit with the appropriate connections to the communication channels and where access is available takes place on the communication channels according to a cyclical time slice procedure, characterized in that a central distribution unit, on the basis of the regular transmission behavior of the node computers known to it a priori, forces a node computer to be able to send to the other node computers only within its statically assigned time slice. 1. Methode zur Erzwingung der Fail-silent Eigenschaft im Zeitbereich von Knotenrechnern eines fehlertoleranten verteilten Computersystems, in dem eine Vielzahl von Knotenrechnern über einen oder mehrere Verteilereinheiten verbunden sind und wo jeder Knotenrechner über eine autonome Kommunikationskontrolleinheit mit den entsprechenden Anschlüssen an die Kommunikationskanäle verfügt und wo der Zugriff auf die Kommunikationskanäle entsprechend einem zyklischen Zeitscheibenverfahren erfolgt, dadurch gekennzeichnet, daß eine zentrale Verteilereinheit aufgrund des ihr a priori bekannt regulären Sendeverhaltens der Knotenrechner erzwingt, daß ein Knotenrechner nur innerhalb seiner statisch zugewiesenen Zeitscheibe an die anderen Knotenrechner zu senden vermag.
- 2Method according to Claim 1, characterized in that the central distribution unit changes from the unsynchronized state, in which all input ports can be received, to the synchronized state after receiving a correct initialization message, in which received via an input port only during the time slice statically assigned to this input port can be. 2. Methode nach Anspruch 1 dadurch gekennzeichnet, daß die zentrale Verteilereinheit vom Zustand unsynchronisiert, in dem über alle Eingangsports empfangen werden kann, nach dem Empfang einer korrekten Initialierungsnachricht in den Zustand synchronisiert wechselt, in dem über einen Eingangsport nur während der diesem Eingangsport statisch zugewiesenen Zeitscheibe empfangen werden kann.
- 3Methode nach den Ansprüchen 1 oder 2 dadurch gekennzeichnet, daß eine zentrale Verteilereinheit vom Zustand synchronisiert in den Zustand unsynchronisiert wechselt, wenn an keinem ihrer Eingangports innerhalb eines a priori vorgegebenen Zeitintervalls dfaUit-i eine korrekte Initialisierungsnachricht empfangen wird. 3rd Method according to Claims 1 or 2, characterized in that a central distribution unit changes from the synchronized state to the unsynchronized state if at none of its input ports within an a priori predetermined time interval dfaUit-i a correct initialization message is received.
- 4Methode nach einem oder mehreren der Ansprüche 1 bis 3 dadurch gekennzeichnet, daß die zentrale Verteilereinheit vom Zustand synchronisiert in den Zustand unsynchronisiert wechselt, wenn an keinem ihrer Eingangports innerhalb eines a priori vorgegebenen Zeitintervalls dfaUit-2 eine Nachricht, die den Codierungsvorschiften des gewählten Codierungssystems entspricht, empfangen wird. 4th Method according to one or more of Claims 1 to 3, characterized in that the central distribution unit changes from the synchronized state to the unsynchronized state if d at none of its input ports within an a priori predetermined time intervalfaUit-2 a message is received that complies with the coding rules of the selected coding system.
- 5Method according to one or more of Claims 1 to 4, characterized in that a powerful central distribution unit stores the content of initialization messages 5. Methode nach einem oder mehreren der Ansprüche 1 bis 4 dadurch gekennzeichnet, daß eine leistungsfähige zentrale Verteilereinheit den Inhalt von Initialisierungsnachrichten AT 407 582 B auswertet, um eine zusätzliche Fehlererkennung durchzuführen. AT 407 582 B evaluates in order to carry out an additional error detection.
- 6Methode nach einem oder mehreren der Ansprüche 1 bis 5 dadurch gekennzeichnet, daß die zentrale Verteilereinheit nach Power-up den Zustand ''unsynchronisieif' einnimmt. 6th Method according to one or more of Claims 1 to 5, characterized in that the central distribution unit assumes the "unsynchronized" state after power-up.
- 7Verteilereinheit mit integriertem Guardian zur Erzwingung der Fail-silent Eigenschaft im Zeitbereich von Knoten rechnern eines fehlertoleranten verteilten Computersystems, in dem eine Vielzahl von Knotenrechnern über einen oder mehrere Verteilereinheiten verbunden sind und wo jeder Knotenrechner über eine autonome Kommunikationskontrolleinheit mit den entsprechenden Anschlüssen an die Kommunikationskanäle verfügt und wo der Zugriff auf die Kommunikationskanäle entsprechend einem zyklischen Zeitscheibenverfahren erfolgt, dadurch gekennzeichnet, daß die zentrale Verteilereinheit aufgrund des ihr a priori bekannten reguläre Sendeverhaltens der Knotenrechner erzwingt, daß ein Knotenrechner nur innerhalb seiner statisch zugewiesenen Zeitscheibe an die anderen Knotenrechner zu senden vermag und wo die Verteilereinheit eine oder mehrere der in Ansprüchen 2 bis 6 beschriebenen Methoden realisiert. 7th Distribution unit with integrated guardian for enforcing the fail-silent property in the time domain of node computers of a fault-tolerant distributed computer system, in which a large number of node computers are connected via one or more distribution units and where each node computer has an autonomous communication control unit with the corresponding connections to the communication channels and where the communication channels are accessed according to a cyclical time-slice method, characterized in that the central distribution unit is based on enforces the regular transmission behavior of the node computers known a priori, that a node computer is only able to send to the other node computers within its statically assigned time slice and where the distribution unit implements one or more of the methods described in claims 2 to 6.
Independent claims7
37 paragraphs in 8 sections, as filed
MESSAGE DISTRIBUTION UNIT WITH INTEGRATED GUARDIAN TO PREVENT BABBLING IDIOT ERRORS
The main aim of the present invention is to increase the fault tolerance of a distributed, time-controlled computer system and to reduce costs by integrating the guardian in a central distribution unit.
This goal is achieved in that an intelligent distribution unit with an integrated guardian is used in a time-controlled computer system to prevent babbling idiot errors in the node computers. The function of this distribution unit is based on the evaluation of a combination of static a priori information about the temporal transmission authorization of the individual node computers with a dynamic synchronization of the Guardian through the initialization messages from TTP / C.
<img file="AT407582B_D0001.tif" />
FIG. 3
DVR 0078018
AT 407 582 B
TECHNICAL ENVIRONMENT
This invention relates to a central distribution unit with an integrated error detection unit for detecting and preventing errors in the time domain for a time-controlled communication method for the transmission of messages in a distributed, fault-tolerant real-time computer system.
BACKGROUND TO THIS INVENTION
Safety-critical technical applications, ie applications where an error can lead to a catastrophe, are increasingly being managed by distributed, fault-tolerant real-time computer systems.
In a distributed fault-tolerant real-time computer system consisting of a number of node computers and a real-time communication system, the failure of one node computer must be tolerated. At the core of such a computer architecture is a fault-tolerant real-time communication system for the predictably fast and secure exchange of messages.
A communication protocol that meets these requirements is cited in European Patent 0 658 257 v. December 18, 1996 as well as in the international patent application PCT / AT 93/00138. The protocol has become known under the name Time-Triggered Protocol / C (TTP / C). It is based on the well-known cyclical time slice method (TDMA time-division multiple access) with time slices defined a priori. TTP / C uses a fault tolerant clock synchronization method disclosed in US Patent: 4,866,606.
TTP / C assumes that the communication system supports a logical broadcast topology and that the node computers show fail-silence behavior, ie either the node computers function correctly in the value domain and in the time domain or they are quiet. The prevention of errors in the time domain, ds, so-called babbling idiot errors, is achieved in TTP / C by an independent error detection unit, the Guardian, which has an independent time base and continuously checks the time behavior of the node computer. In order to implement the fault tolerance, several fail-silent node computers are combined to form a fault-tolerant unit (FTU) and the communication system is replicated. As long as a node computer of an FTU and a replica of the communication system are functioning, the services of the FTU in the time and value domain are provided in good time.
The problem of the structure of node computers that exhibit fail-silence failure behavior has been dealt with several times in the specialist literature. It is suggested to duplicate the node computers and to compare the results (Rostamzadeh, B., Lonn, H., Snedsbol, R., Torin, J., DACAPO, A distributed Computer architecture for safety-critical control applications, Proc. Of the Intelligent Vehicies 95, IEEE Press, pp. 376-385, New York 1995; Levin, LS, MacLeod, IM, Implementation strategy for a fail-silent Processing node for distributed Computer control, Transactions of the South African Institute of Electrical Engineers, VI. 85, no. 3, p. 80-87, September 1994) or to provide a largely independent bus guardian in each node computer (Kopetz, H., Real-Time Systems, Design Principles for Distributed Embedded Applications; ISBN: 0-7923-9894-7. Boston. Kluwer Academic Publishers 1997). However, neither in these publications nor anywhere else is a method known that enforces fail-silence in the time domain by integrating a central distribution unit with a central bus guardian.
A logical broadcast topology of communication can be physically established either by a distributed bus system, a distributed ring system, or by a central distribution unit (eg a star coupler) with point-to-point connections to the node computers. If a distributed bus system or a distributed ring system is set up, each node computer must have its own guardian. If, on the other hand, a central distribution unit is used, then all guardians can be integrated into this distribution unit, which can effectively enforce regular transmission behavior in the time domain based on the global observation of the behavior of all nodes. The present invention relates to a method and an apparatus for integrating the guardians in such a central distribution unit.
Such central distribution units with integrated Guardian have the following advantages:
AT 407 582 B (i) The fault containment region for globally critical errors is reduced by the point-to-point connections between the node computers and the central distribution unit, ie errors caused by EMI (electromagnetic immission) in this point-to-point. Point connections are interspersed, can be clearly assigned to a node computer and have no global effect.
(ii) The replicated, globally critical distribution units can be installed spatially separated in protected areas and designed to be physically compact. This significantly reduces the probability that one cause of failure will destroy all globally critical distribution units.
(iii) The central guardian of the distribution unit replaces the decentralized guardians in the node computer. This saves hardware (e.g. the Guardian oscillators) on the node computers.
(iv) Physical point-to-point connections are well suited for the introduction of fiber optics and also bring advantages in impedance matching with twisted cables.
SUMMARY
The main aim of the present invention is to increase the fault tolerance of a distributed, time-controlled computer system and to reduce costs by integrating the guardian in a central distribution unit.
This goal is achieved in that an intelligent distribution unit with an integrated guardian is used in a time-controlled computer system to prevent babbling idiot errors in the node computers. The function of this distribution unit is based on the evaluation of a combination of static a priori information about the temporal transmission authorization of the individual node computers with a dynamic synchronization of the Guardian through the initialization messages from TTP / C.
BRIEF DESCRIPTION OF THE FIGURES
The above-described object and other novel features of the present invention are illustrated in the accompanying figures.
1 shows the structure of a distributed computer system with four node computers which are connected via two replicated central distribution units.
2 shows the structure of a node computer, consisting of a communication control unit and a host computer, which communicate via the Communication Network Interface (CNI).
3 shows the structure of a central distribution unit with an integrated guardian.
4 shows the data structure which contains the a priori information for the central guardian.
Fig. 5 shows the structure of an initialization message
Fig. 6 shows the internal states of the central distribution unit.
DESCRIPTION OF A REALIZATION
In the following section, an implementation of the new method is shown using an example with four node computers that are connected via two replicated distribution units. The objects in the images are numbered in such a way that the first of the three-digit object numbers always indicates the image number.
1 shows a system of four node computers 111, 112, 113 and 114. A node computer forms an exchangeable unit. Each node computer is connected to the replicated central distribution units 101 and 102 by a point-to-point connection 121.
Fig. 2 shows the internal structure of a node calculator. It consists of two subsystems, the communication controller 210, which is connected to the replicated communication channels 201 and 202, and the host computer 220 on which the application programs of the node computer are executed. These two subsystems are connected via the Communication Network Interface (CNI) 241 and a signal line 242. The CNI consists of a memory
AT 407 582 B (Dual Ported RAM, DPRAM) 241 which both subsystems can access. The two subsystems exchange the communication data via this shared memory 241. The signal line 242 is used to transmit the synchronized time signals. This signal line is described in detail in the referenced U.S. Patent 4,866,606. The communication controller 210, which operates autonomously, has a communication control unit 211 and a data structure 212 which specifies the times at which messages must be sent and received. The data structure 212 is referred to as a Message Descriptor List (MEDL).
3 shows the structure of a central distribution unit with an integrated guardian. Such a central distribution unit consists of the input ports 311, the output ports 312, a data distributor 330 and a control computer 340. The data connections from the distribution unit to the node computer 301 are routed to an input port 311 and an output port 312 of the distribution unit. The same applies to the node computers 302, 303 and 304. In the case of a unidirectional communication line, these two ports 311 and 312 can also be connected separately to the node computer 301. In each input port 311 there is - in addition to the usual filters and, if necessary, electrical isolation - a switch 313 which can be controlled by the control computer 340 of the distribution unit via the signal line 314 and which informs the control computer 340 when this port is receiving. The data arriving at the input port 311 are forwarded via the data distributor 330 to the output lines 312 and to the control computer 340 (via the data line 331). The control computer 340 also has a serial I / O channel 341 via which the static data structure according to FIG. 4 can be loaded and which periodically sends a diagnostic report on the status of the control computer 340 to a maintenance unit. If necessary, the data can be amplified in front of the output. These state-of-the-art amplifiers are not shown in Figure 3.
4 shows the data structure which is made available to the central control computer 340 a priori, ie before the runtime. This data structure has its own data record for each port of the distribution unit (411 for node 111, 412 for node 112, 413 for node 113 and 414 for node 114). The first field of this data record 401 contains the port number to which this data record relates. The second field 402 contains the transmission duration of the node connected to the port according to the entry in the MEDL 212. The third field 403 contains the duration of the time interval between the end of the current transmission and the start of the next transmission by the node connected to the port. The fourth field 404 contains the number of the next port in time. The fifth field 405 contains the duration of the time interval between the end of the current transmission and the beginning of the transmission of the node at the port that is next in time. Field 406 contains the length of an initialization message that can be received on the current port. The content of the data structure of FIG. 4 is created by a development tool in coordination with the MEDLs 212 and is loaded into the control computer 340 before the runtime.
Fig. 5 shows the structure of an initialization message. The initialization message must contain a distinctive bit 510 in the header 501, which identifies the message as an initialization message. The data field 502 of the initialization message contains further information that is irrelevant for the function of a simple distribution unit. The CRC field 503 is located at the end of the initialization message. More powerful distribution units can evaluate the information in data field 502 of an initialization message in addition to increasing the error detection probability. For example, such more powerful distribution units can evaluate the time field of a TTP / C initialization message in order to be able to compare the clock status of the transmitter with its own clock.
6 shows the two most important internal states of the control computer 340, unsynchronized 601 and synchronized 602. After power-up 610, the control computer 340 goes into the unsynchronized state. In this state, all input ports 311 are connected to the data distributor 330. As soon as an initialization message with the correct CRC is received on an input port via the data line 331 from the control computer 340, the control computer 340 determines via the signal line 314 via which port was received, saves the time of receipt, checks the length of the message by comparing it with the stored length 406 and, in the event of a positive outcome of the test, goes to the synchronized state 602, wherein the stored time of receipt of the initialization message represents the synchronization event. In condition
AT 407 582 B synchronized 602, the control computer 340 establishes a connection at the corresponding input port only during the time period 403. If any message arrives at approximately the right time that complies with the coding rules of the selected coding system, then the control computer uses the measured time difference between the observed and expected arrival time of the message to check its clock using a known fault-tolerant algorithm (e.g. Kopetz 1997) to synchronize. If during an a priori fixed time interval d<sub>stupid</sub>it-i no further correct initialization message arrives at any input port of the distribution unit, then the control computer 340 goes to the unsynchronized state 601. In the synchronized state 602, an initialization message is correct if it arrives at the input port approximately at the expected time, has a correct CRC field 503 and is the correct length according to 406. The control computer 340 communicates its internal status via the serial I / O line 341.
In order to accelerate the restart, the control computer 340 can change from the synchronized state 602 to the unsynchronized state 601 if d<sub>fa</sub>uit-2 no message, which corresponds to the coding regulations of the selected coding system, arrives at any input port at approximately the right time.
It is an important property of this invention that the control computer 340 can only cause the switches 313 to open and close, but cannot change the messages sent or insert new messages. The only type of failure of the distribution unit is therefore a faiisilent failure of a communication channel. In a fault-tolerant configuration, however, there is a second independent communication channel.
Finally, it should be noted that this invention is not limited to the implementation described with four node computers and two star couplers, but can be expanded as required. It can be used not only with the TTP / C protocol, but also with other time-controlled protocols.
Contents8
4 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4
Every citation, both waysCites: the store holds 0 of 1
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US7502334B2 | Cited by | United States of America | Applicant |
| US7606179B2 | Cited by | United States of America | Applicant |
| US7912094B2 | Cited by | United States of America | Applicant |
| US7656881B2 | Cited by | United States of America | Applicant |
| US7698395B2 | Cited by | United States of America | Applicant |
| AT411948B | Cited by | Austria | Search report |
| US8301885B2 | Cited by | United States of America | Applicant |
| US7839868B2 | Cited by | United States of America | Applicant |
| US7649835B2 | Cited by | United States of America | Applicant |
| US7778159B2 | Cited by | United States of America | Applicant |
| US8315274B2 | Cited by | United States of America | Applicant |
| US7729297B2 | Cited by | United States of America | Applicant |
| US8817597B2 | Cited by | United States of America | Applicant |
| US7505470B2 | Cited by | United States of America | Applicant |
| US7372859B2 | Cited by | United States of America | Applicant |
| US7668084B2 | Cited by | United States of America | Applicant |
| US7889683B2 | Cited by | United States of America | Applicant |
14 members in 8 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 139599 | Austria | A | |
| 7199102 | United States of America | A | |
| 7199102 | United States of America | A | |
| AT19990001395 | – | – | – |
| US20020071991 | – | – | – |
Members14
| Document | Office | Kind | |
|---|---|---|---|
| ATA139599A | Austria | A | |
| WO0113230A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU5952400A | Australia | A | |
| AT407582BThis record | Austria | B | |
| KR20020029914A | Republic of Korea | A | |
| EP1222542A1 | European Patent Office (EPO) | A1 | |
| JP2003507790A | Japan | A | |
| EP1222542B1 | European Patent Office (EPO) | B1 | |
| AT237841T | Austria | T | |
| ATE237841T1 | Austria | T1 | |
| DE50001819D1 | Germany | D1 | |
| US2003154427A1 | United States of America | A1 | |
| KR100433649B1 | Republic of Korea | B1 | |
| JP4099332B2 | Japan | B2 |
2 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| ExpiryMK07 | MK07 | |
| ExpiryMK07 | MK07 |
Numbers
- Publication, DOCDB
- 407582
- Publication, EPODOC
- AT407582B
- Application
- 139599
- Application, DOCDB
- 139599
- Application, EPODOC
- AT19990001395
Titles2
- German
- NACHRICHTENVERTEILEREINHEIT MIT INTEGRIERTEM GUARDIAN ZUR VERHINDERUNG VON ''BABBLING IDIOT'' FEHLERN
- English
- NEWS DISTRIBUTION UNIT WITH INTEGRATED GUARDIAN TO PREVENT '' babbling idiot '' ERROR
Classification
- CPC, 3
- G06F11/2005
- H04L12/40026
- H04L12/66
- IPC, 5
- G06F13 00
- G06F11 00
- G06F11 07
- G06F11 20
- G06F11 30