Method for exchanging data packets in a safe multicomputer system and multicomputer system for carrying out the same
10 claims: 10 independent, 0 dependent
- 1Method for exchanging data packets (A, B, C in Fig. 2b) between the computers of a multicomputer system (MRS), the computers processing the same tasks in parallel and independently of one another, characterized in that the exchange of data packets occurs in the form of successive transmission rounds, the data packets being brought into a uniform order in all computers, and the transmission rounds comprising the following steps:a) any computer (R1) sends a first data packet (A) on its own initiative to all other computers (R2, R3),b) each of the other computers (R2, R3) reacts to the receipt of the first data packet (A) within a predefinable time by sending in each case a data packet of their own (B, C) to all other computers respectively in the multicomputer system. Procédé pour l'échange de paquets de données (A, B, C dans la figure 2b) entre les ordinateurs d'un système multi-ordinateurs (MRS), les ordinateurs accomplissant les mêmes tâches en parallèle et indépendamment les unes des autres, caractérisé en ce que l'échange des paquets de données s'effectue sous la forme de tournées de transmission successives, les paquets de données étant amenés dans tous les ordinateurs selon une séquence homogène, lequel comprend les étapes suivantes : a) un ordinateur quelconque (R1) envoie sur sa propre initiative un premier paquet de données (A) à tous les autres ordinateurs (R2, R3),b) chacun des autres ordinateurs (R2, R3) réagit à la réception du premier paquet de données (A) dans un intervalle de temps pouvant être prédéfini avec l'envoi à chaque fois d'un paquet de données (B, C) propre respectivement à tous les autres ordinateurs dans le système multi-ordinateurs. Verfahren zum Austausch von Datenpaketen (A, B, C in Fig. 2b) zwischen den Rechnern eines Mehrrechnersystems (MRS) wobei die Rechner die gleichen Aufgaben parallel und unabhängig voneinander bearbeiten, dadurch gekennzeichnet,daß der Austausch der Datenpakete in Form von aufeinanderfolgenden Übertragungsrunden erfolgt wobei die Datenpackete in allen Rechnern in eine einheitliche Reihenfolge gebracht werden, welche folgende Schritte umfassen: a) ein beliebiger Rechner (R1) sendet auf eigene Initiative ein erstes Datenpaket (A) an alle anderen Rechner (R2, R3),b) jeder der anderen Rechner (R2, R3) reagiert auf den Empfang des ersten Datenpakets (A) innerhalb einer vorgebbaren Zeitspanne mit dem Aussenden jeweils eines eigenen Datenpakets (B, C) an jeweils alle anderen Rechner im Mehrrechnersystem.
- 2Method according to Claim 1, in which one of the other computers (R2, R3) in a transmission round does not react to the receipt of the first data packet (A), if it has already in this transmission round sent a data packet itself on its own initiative to all other computers. Procédé selon la revendication 1, avec lequel l'un des autres ordinateurs (R2, R3) ne réagit pas à la réception du premier paquet de données (A) dans une tournée de transmission s'il a lui-même déjà envoyé un paquet de données de sa propre initiative à tous les autres ordinateurs dans cette tournée de transmission. Verfahren nach Anspruch 1, bei dem einer der anderen Rechner (R2, R3) in einer Übertragungsrunde nicht auf den Empfang des ersten Datenpakets (A) reagiert, wenn er in dieser Übertragungsrunde bereits selbst auf eigene Initiative ein Datenpaket an alle anderen Rechner gesendet hat.
- 3Method according to Claims 1 or 2, in which the data packets (DP in Fig. 9) contain at least one of the following items of information:a) user data (ND),b) Data (W) to identify which path the data packet has taken through the multicomputer system,c) Data (LS), which identifies the sending computer's local view of the system status in the preceding transmission round,d) Data (LZ), from which the local time of the computer sending the data packet can be determined,e) Data (RN) to identify the current transmission round,f) Control data (K) for checking whether all or part of the data mentioned above has been correctly received. Procédé selon la revendication 1 ou 2, avec lequel les paquets de données (DP dans la figure 9) contiennent au moins l'une des informations suivantes : a) Données utiles (ND),b) Données (W) destinées à identifier le trajet parcouru par le paquet de données à travers le système multi-ordinateurs,c) Données (LS) qui identifient la vue locale de l'ordinateur émetteur de l'état du système dans la tournée de transmission précédente,d) Données (LZ) à partir desquelles il est possible de déterminer l'heure locale de l'ordinateur ayant envoyé le paquet de données,e) Données (RN) pour identifier la tournée de transmission actuelle,f) Données de contrôle (K) destinées à vérifier si la totalité ou une partie des données mentionnées précédemment a été reçue correctement. Verfahren nach Anspruch 1 oder 2, bei dem die Datenpakete (DP in Fig. 9) wenigstens eine der folgenden Informationen enthalten: a) Nutzdaten (ND),b) Daten (W) zur Kennzeichnung, welchen Weg das Datenpaket durch das Mehrrechnersystem genommen hat,c) Daten (LS), die die lokale Sicht des sendenden Rechners vom Systemstatus in der vorhergehenden Übertragungsrunde kennzeichnen,d) Daten (LZ), aus denen die lokale Zeit des das Datenpaket sendenden Rechners ermittelbar ist,e) Daten (RN) zur Kennzeichnung der aktuellen Übertragungsrunde,f) Kontrolldaten (K) zum Überprüfen, ob alle oder ein Teil der vorstehend genannten Daten korrekt empfangen worden sind.
- 4Method according to Claim 3, in which each computer determines, from the individual computers' local views (LSk) of the system status, a global view (GS) of the system status identical for all computers. Procédé selon la revendication 3, avec lequel chaque ordinateur détermine à partir des vues locales (LSk) individuelles des ordinateurs de l'état du système une vue globale (GS) de l'état du système identique pour tous les ordinateurs. Verfahren nach Anspruch 3, bei dem jeder Rechner aus den einzelnen lokalen Sichten (LSk) der Rechner vom Systemstatus eine für alle Rechner gleiche globale Sicht (GS) vom Systemstatus ermittelt.
- 5Method according to Claim 3, in which with the help of the data (LZ), from which the local time of the computer sending the data packet can be determined, a synchronized global time is determined. Procédé selon la revendication 3, avec lequel une heure globale synchronisée est déterminée à l'aide des données (LZ) à partir desquelles peut être déterminée l'heure locale de l'ordinateur ayant envoyé le paquet de données. Verfahren nach Anspruch 3, bei dem mit Hilfe der Daten (LZ), aus denen die lokale Zeit des das Datenpaket sendenden Rechners ermittelbar ist, eine synchronisierte globale Zeit ermittelt wird.
- 6Method according to one of the preceding claims, in which a data packet (A) sent from any one computer (R1 in Fig. 5) to the other computers (R2, R3) is forwarded by these other computers in such a way that each of the other computers receives this data packet at least twice in the absence of any faults. Procédé selon l'une des revendications précédentes, avec lequel un paquet de données (A) envoyé par un ordinateur quelconque (R1 dans la figure 5) aux autres ordinateurs (R2, R3) est retransmis par ces autres ordinateurs de telle sorte que chacun des autres ordinateurs, en l'absence de défaut, reçoive ce paquet de données au moins deux fois. Verfahren nach einem der vorhergehenden Ansprüche, bei dem ein von einem beliebigen Rechner (R1 in Fig. 5) an die anderen Rechner (R2, R3) gesendetes Datenpaket (A) so von diesen anderen Rechnern weitergeleitet wird, daß jeder der anderen Rechner dieses Datenpaket im fehlerfreien Fall wenigstens zweimal empfängt.
- 7Method according to one of the preceding claims, in which a computer on its own initiative sends a data packet to all other computers of the multicomputer system, if one or more of the following conditions are satisfied:a) an application program executed by the computer gives a send command;b) a predefinable trigger interval has elapsed since the last sending of a data packet;c) the quantity of user data waiting in the computer for transmission exceeds a predefinable size;d) the number of messages waiting in the respective computer for transmission exceeds a predefinable figure;e) a message whose priority value is above a predefinable threshold is waiting in the computer for transmission. Procédé selon l'une des revendications précédentes, avec lequel un ordinateur envoie de sa propre initiative un paquet de données à tous les autres ordinateurs du système multi-ordinateurs si une ou plusieurs des conditions suivantes sont remplies : a) un programme d'application exécuté par l'ordinateur délivre une instruction d'émission ;b) un intervalle de temps de déclenchement pouvant être prédéfini s'est écoulé depuis la dernière émission d'un paquet de données ;c) le volume d'informations utiles prêtes à être transmises dans l'ordinateur dépasse une valeur pouvant être prédéfinie ;d) le nombre d'informations à transmettre présentes dans l'ordinateur correspondant dépasse un nombre pouvant être prédéfini ;e) il existe dans l'ordinateur une information à transmettre dont la valeur de priorité est supérieure à un seuil pouvant être prédéfini. Verfahren nach einem der vorhergehenden Ansprüche, bei dem ein Rechner auf eigene Initiative allen anderen Rechnern des Mehrrechnersystems ein Datenpaket sendet, wenn eine oder mehrere der folgenden Bedingungen erfüllt sind: a) ein von dem Rechner ausgeführtes Applikationsprogramm gibt einen Sendebefehl;b) seit dem letzten Senden eines Datenpakets ist eine vorgebbare Auslösezeitspanne verstrichen;c) die im Rechner zur Übermittlung anstehende Nutzdatenmenge überschreitet ein vorgebbares Maß;d) die Anzahl von im jeweiligen Rechner zur Übermittlung anstehenden Nachrichten überschreitet eine vorgebbare Zahl;e) im Rechner steht eine Nachricht zur Übermittlung an, deren Prioritätswert über einer vorgebbaren Schwelle liegt.
- 8Mehrrechnersystem mit wenigstens zwei Rechnern, die Mittel zum Austausch von Datenpaketen umfassen, wobei die Rechner die gleichen Aufgaben parallel und unabhängig voneinander bearbeiten, dadurch gekennzeichnet,daß der Austausch der Datenpakete in Form von aufeinanderfolgenden Übertrogungsrunden erfolgt wobei die Datenpackete in allen Rechnern in eine einheitliche Reihenfolge gebracht werden, und wobei a) ein beliebiger Rechner (R1) derart ausgestaltet ist, auf eigene Initiative ein erstes Datenpaket (A) an alle anderen Rechner (R2, R3) zu senden, undb) jeder der anderen Rechner (R2, R3) Auslösemittel (ATVM, AM in Fig. 8) umfasst, die derart ausgestaltet sind, bei Empfang des ersten Datenpakets (A) innerhalb einer vorgebbaren Zeitspanne das Aussenden jeweils eines eigenen Datenpakets (B, C) an jeweils alle anderen Rechner im Mehrrechnersystem auszulösen. Multicomputer system with at least two computers, which include means of exchanging data packets, the computers processing the same tasks in parallel and independently of one another, characterized in that the exchange of data packets occurs in the form of successive transmission rounds, the data packets being brought into a uniform order in all computers, and where a) any computer (R1) is developed in such a way that it sends a first data packet (A) on its own initiative to all other computers (R2, R3), andb) each of the other computers (R2, R3) includes triggering means (ATVM, AM in Fig. 8), which are developed such that upon receipt of the first data packet (A) they trigger the sending in each case of a data packet of their own (B, C) within a predefinable time to all other computers respectively in the multicomputer system. Système multi-ordinateurs comprenant au moins deux ordinateurs, lesquels comprennent des moyens pour échanger des paquets de données, les ordinateurs accomplissant les mêmes tâches en parallèle et indépendamment les unes des autres, caractérisé en ce que l'échange des paquets de données s'effectue sous la forme de tournées de transmission successives, les paquets de données étant amenés dans tous les ordinateurs selon une séquence homogène, a) un ordinateur quelconque (R1) étant configuré de telle sorte à envoyer sur sa propre initiative un premier paquet de données (A) à tous les autres ordinateurs (R2, R3), etb) chacun des autres ordinateurs (R2, R3) comprenant des moyens de déclenchement (ATVM, AM dans la figure 8) qui sont configurés de manière à déclencher l'envoi à chaque fois d'un paquet de données (B, C) propre respectivement à tous les autres ordinateurs dans le système multi-ordinateurs lors de la réception du premier paquet de données (A) dans un intervalle de temps pouvant être prédéfini.
- 9Mehrrechnersystem nach Anspruch 8, bei denen jeder Rechner Mittel zum Umsetzen von Nachrichten in Datenpakete umfaßt, welche wenigstens eine der folgenden Informationen enthalten:a) Nutzdaten (ND in Fig. 9),b) Daten (W) zur Kennzeichnung, welchen Weg das Datenpaket durch das Mehrrechnersystem genommen hat,c) Daten (LS), die die lokale Sicht des sendenden Rechners vom Systemstatus in der vorhergehenden Übertragungsrunde kennzeichnen,d) Daten (LZ), aus denen die lokale Zeit des das Datenpaket sendenden Rechners ermittelbar ist,e) Daten (RN) zur Kennzeichnung der aktuellen Übertragungsrunde,f) Kontrolldaten (K) zum Überprüfen, ob alle oder ein Teil der vorstehend genannten Daten korrekt empfangen worden sind. Multicomputer system according to Claim 8, in which each computer includes means of converting messages into data packets, which contain at least one of the following items of information: a) user data (ND in fig. 9),b) Data (W) to identify which path the data packet has taken through the multicomputer system,c) Data (LS), which identifies the sending computer's local view of the system status in the preceding transmission round,d) Data (LZ), from which the local time of the computer sending the data packet can be determined,e) Data (RN) to identify the current transmission round,f) Control data (K) for checking whether all or part of the data mentioned above has been correctly received. Système multi-ordinateurs selon la revendication 8, avec lequel chaque ordinateur comprend des moyens pour convertir les informations en paquets de données qui contiennent au moins l'une des informations suivantes : a) Données utiles (ND dans la figure 9),b) Données (W) destinées à identifier le trajet parcouru par le paquet de données à travers le système multi-ordinateurs,c) Données (LS) qui identifient la vue locale de l'ordinateur émetteur de l'état du système dans la tournée de transmission précédente,d) Données (LZ) à partir desquelles il est possible de déterminer l'heure locale de l'ordinateur ayant envoyé le paquet de données,e) Données (RN) pour identifier la tournée de transmission actuelle,f) Données de contrôle (K) destinées à vérifier si la totalité ou une partie des données mentionnées précédemment a été reçue correctement.
- 10Ein oder mehrere Datenträger mit einem darauf gespeicherten Datenverarbeitungsprogramm, welches bei Einlesen in ein Mehrrechnersystem mit wenigstens zwei Rechnern das Verfahren nach einem der Ansprüche 1 bis 7 steuert. One or more data carriers with, stored on them, a data processing program which, for reading into a multicomputer system with at least two computers, controls the method according to one of the claims 1 to 7. Un ou plusieurs supports de données sur lequel est enregistré un programme de traitement de données qui, lorsqu'il est chargé dans un système multi-ordinateurs comprenant au moins deux ordinateurs, commande le procédé selon l'une des revendications 1 à 7.
Independent claims10
48 paragraphs, as filed
The invention relates to a method for exchanging data packets within a secure multi-computer system. The invention further relates to a secure multi-computer system for performing the method.
Multi-computer systems are generally used to control and monitor safety-critical systems, for example in the field of railway signaling or aerospace. Such multicomputer systems consist of at least two computers, which process pending tasks largely in parallel and independently of one another. The computers of most multicomputer systems communicate their results to an internal or external comparator, which leads to a majority decision. The comparator ensures that only those results are accepted that have been determined by the majority of the participating computers in agreement. A result determined by only one computer can therefore never affect the process to be controlled, since a computer in a multicomputer system can under no circumstances have the majority. Such multi-computer systems are often designed as 2-out-of-3 computer systems, in which the agreement of the results is required by at least two computers. A failure of one of the three computers is tolerated by the system, since two computers can still determine matching results.
Since, as already mentioned at the beginning, the computers in such multi-computer systems largely process pending tasks in parallel, it must be ensured that the same data stream is also supplied to all computers. If, for example, one of the computers receives data packets in the order a → b → c and another computer in the order a → c → b, the computers will usually come to different results in their calculations despite the same programming, even though they do work flawlessly. A 2-out-of-2 computer system becomes unable to output a result due to such an error. The occurrence of a further error can already lead to two incorrect results matching and thus being recognized by the comparator as "correct". Catastrophic consequences such as train crashes and plane crashes cannot be ruled out.
Multi-computer systems are known which operate according to a fixed, system-uniform synchronization cycle. The data are exchanged in the form of data packets between the computers of the multi-computer system. Each data packet receives a cycle number from the sending computer, with the aid of which a receiving computer can put incoming data packets in the correct order. This ensures a data stream that is uniform for the entire multicomputer system. A dedicated synchronization network is used to maintain the system-wide synchronization clock. A selected computer (so-called "master computer") sends synchronization signals to the other computers at short intervals via this synchronization network. However, such solutions require special hardware, for example for the synchronization network, and are therefore expensive.
US-A-5 506 962 discloses a distributed processing system having a plurality of processors for performing a sequence of processing steps. Information to be processed is transmitted together with a content code. An event number is used to compare a received message with a previously received message in order to detect repeatedly sent messages.
It is therefore an object of the invention to provide an inexpensive method for exchanging data packets within a secure multi-computer system, which is intended to ensure that all computers process the same data stream. The method should not require the provision of a separate synchronization network.
The invention solves this problem with the aid of the features specified in independent claims 1, 8 and 10. Further advantageous embodiments of the invention can be found in the subclaims.
The invention is explained in detail below using the exemplary embodiments and the drawings. Show it:<dl id="dl0001"><dt>Fig. 1:</dt><dd>A 2-out-of-3 computer system MRS as an example of a multi-computer system in a schematic representation to explain the method according to the invention;</dd><dt>Fig. 2a:</dt><dd>The 2-out-of-3 computer system MRS from FIG. 1, in which the computer R1 sends data packets A to the other two computers R2 and R3;</dd><dt>Fig. 2b:</dt><dd>The 2-out-of-3 computer system MRS from FIG. 1, in which the two computers R2 and R3 in turn send data packets B and C to the computer R1;</dd><dt>Fig. 3a:</dt><dd>The 2-out-of-3 computer system MRS from FIG. 1, in which the computers R1 ... R3 are provided with status characters S;</dd><dt>3b:</dt><dd>The 2-out-of-3 computer system MRS from FIG. 1, in which the local views LS for one exemplary embodiment of the invention<sub>k</sub> and the global view GS of the computers R1 ... R3 is specified;</dd><dt>Fig. 4:</dt><dd>Schematic representation to explain how a global view can be determined from local views;</dd><dt>Fig. 5:</dt><dd>The 2-out-of-3 computer system MRS from FIG. 1 to explain an exemplary embodiment according to claim 5;</dd><dt>Fig. 6:</dt><dd>The 2-out-of-3 computer system MRS from FIG. 1, in which the communication link KV12 is disrupted;</dd><dt>Fig. 7:</dt><dd>The 2-out-of-3 computer system MRS from FIG. 1, in which the computer R1 is faulty;</dd><dt>Fig. 8:</dt><dd>Highly schematic representation of a computer in a multi-computer system according to the invention to explain the data exchange within the computer and between the computers;</dd><dt>Fig. 9:</dt><dd>Structure of a data packet;</dd><dt>Fig. 10:</dt><dd>3-of-5 computer system with two transmission buses.</dd></dl>
Example 1:
As an example of a multi-computer system, FIG. 1 shows a 2-out-of-3 computer system MRS, which comprises the computers R1, R2 and R3. The three computers are completely meshed with each other via communication links KV12, KV13 and KV32. The communication connections KV12, KV13 and KV32 enable bidirectional communication between the computers. These communication connections can be implemented, for example, as physical point-to-point connections or via a local area network (LAN), in which the free addressability of the computers is guaranteed. A message a is fed to the computer R1 via communication connections (not shown). In the same way, computer R2 receives a message b and computer R3 receives a message c. The message sources that deliver the messages to the computers can be, for example, other multicomputer systems or a single computer that does not belong to the MRS multicomputer system. It is also possible that the messages are not supplied by external message sources, but are generated in the computers themselves by application programs running there.
After receiving the messages, the computers convert messages a, b, c into corresponding data packets A, B and C according to an agreed protocol. If there are no messages on one or more computers, data packets are created whose user data area is empty. It is also possible to arrange several pending messages together in one data packet. In Fig. 2a shows how one of the computers, here the computer R1, begins with the transmission of the data packet A created by it to the two neighboring computers R2 and R3. Which computer takes the initiative at this moment will be explained in more detail later. After receiving the data packet A, the two other computers R2 and R3 in turn send out the data packets B and C they have created. This is shown in Fig. 2b with the solid arrows. The computers R1 and R3 adjacent to the computer R2 each receive a data packet B in this way, the computers R1 and R2 adjacent to the computer R3 each receive a data packet C. After the transmission, each of the computers R1, R2 and R3 now has from every other computer in the Multi-computer system receive exactly one data packet. In Fig. 2b this can be seen from the fact that exactly two arrows (solid or interrupted) are directed at each of the computers. The messages a, b and c present in the individual computers have thus been distributed to the computers after carrying out the method according to the invention in such a way that all computers can have all the messages supplied to the overall system. In the following, the process just described is referred to as a "broadcast round". The decisive factor for the establishment of such a transmission round is the fact that the two computers R2 and R3, after receiving the data packet A sent by the computer R1, in turn transmit their own data packets to all other computers within a predeterminable time period.
In order to ensure a uniform data flow in all computers, the data packets must now be brought into a uniform order. Here, for. B. the fact can be exploited that each computer can determine the transmitter when receiving a data packet. Determining the transmitter is possible both with physical point-to-point connections and with networks and does not require any special measures. Based on the sender determination, the received data packets can be clearly assigned to the sending computer. An agreement can then stipulate that the data packets are to be sorted in such a way that first the data packet A originating from the computer R1, then the data packet B originating from the computer R2 and finally the data packet C originating from the computer R3 are fed to the application processes running in the computers becomes. In the example shown above, this would result in a sequence A → B → C. In this way, however, the messages fed to the computer R1 are processed preferentially. Therefore, it can make sense to set the order in a different way. So z. B. conceivable to permute the sequence from transmission round to transmission round cyclically. In round 1, the sequence A → B → C is then chosen, in the following round 2 the sequence B '→ C' → A 'and then in round 3 the sequence C "→ A" → B "etc. With this one The procedure is the same for all computers when determining the sequence.
With the implementation of the method according to the invention, a uniform data stream is thus made available to all computers in the multi-computer system. In contrast to known solutions, a separate synchronization network is not required for this. Another advantage of the method according to the invention is that, due to the chronological sequence of the transmission and reception processes, the computers are able to determine and also localize certain errors in the transmission. For example, in the example shown above with reference to FIG. 2b, if the computer R1 does not receive the data packet B from the computer R2, the computer R1 can assume that one of the following faults is present:<ul id="ul0001" list-style="bullet"><li>the computer R2 has not received the data packet A;</li><li>the computer R2 has received the data packet A, but has not sent a data packet B due to an internal error in the computer;</li><li>The computer R2 has sent out a data packet B, but this could not be delivered due to a defective communication connection.</li></ul>
The computer R1 can z. B. set the communication with the computer R2 and still only communicate with the computer R3.
Embodiment 2
:
It is u. Depending on the configuration of the method according to the invention, it cannot be ruled out (see the explanations further below) that, in addition to the computer R1, another computer initiates a transmission round almost simultaneously. For example, computer R2 could initiate a further transmission round by sending its data packet B after a transmission round has been initiated by computer R1, but before the data packet A has been received. In this case, the computer R2 would send its data packet B again after receiving the data packet A from the computer R1, although it had initiated a round of transmission with this data packet B shortly before. The computers R1 and R3 would accordingly receive the data packet B twice. There is also the danger that there will be a kind of "ping-pong effect" in which the computers transmit identical data packets to one another in an endless sequence.
To rule this out, it is provided in one embodiment of the invention according to claim 2 that a computer in a transmission round does not respond to the receipt of a data packet by sending out its own data packet, if in this transmission round it itself already sends a data packet to everyone else on its own initiative Computer sent. This presupposes that the data packets comprise data RN for identifying the transmission round, so that each computer can assign received data packets to a specific transmission round at any time. If, for example, a transmission round i is completed and the two computers R1 and R2 independently of one another a transmission round 1 + 1 by sending out data packets A or Trigger B, the computer R2 does not send its data packet B again in this transmission round, since the other computers have already received this once or will soon receive it. It also does not respond by sending another self-created data packet B ', so that the "ping-pong effect" mentioned above is avoided. The same applies accordingly to the computer R1.
Embodiment 3
:
In the method according to the invention, the communication between the computers takes place in the form of rounds of transmission. At the end of such a round of transmission, each computer received at least one data packet from all other computers. This form of data packet exchange is particularly advantageous if the data packets contain status information on the individual computers. In this way, each individual computer can determine a global view of the system status for the past transmission round. What this means in detail is described in more detail below.
3a shows the multicomputer system from FIG. 1 and additionally the status S for each computer<sub>k</sub>. In the example shown, the two computers R1 and R2 work error-free and are therefore fully-fledged "members" in the multi-computer system. They are therefore marked with the status symbol "m". It is agreed that only those data are processed or evaluated (for example in a comparator) that originate from a computer in the "member" status. The computer R3 is defective in this example and is therefore referred to as a "non-member". For this reason, the status symbol "n" is assigned to it. Data originating from this computer R3 are therefore discarded.
After carrying out a round of transmission, each computer determines its individual local view LS<sub>k</sub> system status by evaluating the received data packets. For example, if the computer R1 receives no or only an incorrect data packet from the computer R2, it looks at this computer in its local view LS<sub>1</sub> as a "non-member". Depending on the type of error, it may well happen that computer R2, which is defective from the perspective of computer R1, regards itself as error-free and therefore as a "member". In the local view LS<sub>2</sub> of the computer R2, all computers may therefore be "members". The local views LS<sub>k</sub> can therefore differ from the system status.
For certain purposes, however, it is advantageous to also determine a global view GS of the system status that is uniform for all computers. The global view of the system status is especially needed<ul id="ul0002" list-style="bullet"><li>so that there is a uniform basis for decision-making at system level in order to be able to carry out the same processing steps in the same order for the fault-free computers ("members"),</li><li>so that a central error handling authority can identify errors and, if necessary, initiate remedial measures,</li><li>to integrate defective computers back into the multi-computer system</li></ul>
Determining the global view is the subject of the following sections.
After completion of a transmission round i, each computer determines its local view LS<sub>k</sub> from system status. In the following transmission round 1 + 1, each computer transmits its local view LS<sub>k</sub> from the system status in transmission round i to the other computers. After completion of the transmission round i + 1, each computer thus has the local views of all the computers in the multicomputer system from the previous transmission round i. From these local views, each computer can now determine a global view of the system status in transmission round i according to an agreed rule. Such a rule can e.g. B. include the following steps:<ul id="ul0003" list-style="bullet"><li>Use the local views LS<sub>k</sub> from the system status in transmission round i of those computers which had the status "member" in the global view of the system status in this transmission round i;</li><li>Guide LS for these local views<sub>k</sub> a majority vote;</li><li>Only accept a computer designated as a "member" by majority vote if it has registered itself as a "member" in the current transmission round.</li></ul>
The global view GS for a broadcast round is identified by the symbols “MMN” in FIG. 3b. This means that computers R1 and R2 are globally recognized as members, while computer R3 is considered a non-member. The local views LS are also shown<sub>k</sub> the computer R1 ... R3 from the system status of the previous transmission round. In this simple example, all computers R1 ... R3 have the same local view of the system status of the previous transmission round, namely "mmn". This means that all computers have computers R1 and R2 as members and computer R3 as "non-members".
The determination of a global view is explained again in FIG. 4 using a simple example. The global view GS in the transmission round i is "MMN", ie the two computers R1 and R2 are recognized as members, while the computer R3 is regarded as defective. In accordance with the rule mentioned above, only the local views of the computers R1 and R2 are taken into account in the majority decision ME in the following transmission round i + 1. In the example shown, these local views match, so that the majority decision ME delivers the result "mmn". The local view of the defective computer R3 is u. U. nonsensical, incorrect or incomplete, which is shown in FIG. 4 by the symbols "<img file="EP0905623B1_D0001.tif" />"Since the two computers R1 and R2 specified as members see themselves in their local view LS as members, the global view of the system status is equal to the result of the majority decision, ie" MMN ".
The concept of adding status information to the data packets can be expanded depending on the requirements of the overall system. In addition to the listed status options "member" and "non-member", intermediate states such as "provisional member" or "candidate for member" can also be defined. This is useful, for example, if defective computers are to be integrated back into the multi-computer system without the intervention of an operator. In the case of such an integration, it is expedient for security reasons to have the function of the computer or computers to be integrated first observed by the non-defective computers. Global recognition as a member only occurs when the computer to be integrated has successfully passed a type of "trial period".
In a further advantageous exemplary embodiment, the data packets contain data from which a local time of the computer sending the data packet can be determined. Local time is understood to mean the time that the sending computer taps from an internal or external clock at the time the data packet is created. After the completion of a transmission round, each computer in the multi-computer system is able to determine a global time from this data. This can be done, for example, by simply forming the median value (= central value) from the individual local times. The application programs executed by the multicomputer system or also certain external modules such as the comparator mentioned above, which subjects the results of the individual computers to a majority decision, often need a global time.
Furthermore, it can be provided to add control data to the data packets, which are used to check whether all or part of the data contained in the data packet has been received correctly. Such measures, which are known per se, can ensure that only those data packets are evaluated - for example for determining a global view of the system status - which have actually been transmitted without errors.
Example 4:
In a particularly advantageous exemplary embodiment according to claim 6, the computers R1 ... R3 forward received data packets to the neighboring computers in such a way that each computer receives at least twice each circulating data packet in the error-free case. This is explained below with reference to FIG. 5. In this example, the computer R1 sends out a data packet A. The two computers R2 and R3 receive this data packet A and pass it on to the computers R3 and R2. In this way, the computer R2 receives the data packet A once directly from the computer R1 and once indirectly via the computer R3. The same applies accordingly to the computer R3. Data packets B and C sent by computers R2 and R3 are also forwarded in the manner described. Thanks to the double reception, the tolerance of the multicomputer system against errors and failures is significantly improved. Furthermore, it is possible to precisely localize a larger class of errors, especially if the data packets contain control data as described above. Two error scenarios which are possible in a multicomputer system are explained below in order to demonstrate the advantages of this embodiment of the invention.
FIG. 6 shows the multi-computer system from FIG. 1, in which the communication connection KV12 between the computers R1 and R2 is interrupted in both transmission directions. The computer R2 can therefore not receive data packets directly from the computer R1, but only indirectly via the computer R3. The same applies to data packets that computer R2 sends to computer R1. Due to the forwarding of the data packets according to the invention, the interruption of the communication link KV12 is tolerated. In an advantageous variant of the invention, the data packets contain data which identify the route which the data packets have taken through the multicomputer system. This makes it possible for every computer to precisely localize the error, because computer R1 then knows, for example, that it has only received one data packet from computer R2, which has taken the route via computer R3. The cause of this can only be an interruption in the communication link KV12. In an analogous manner, the other computers R2 and R3 can also determine from the transmission path of the received data packets at which point in the multi-computer system a communication link is interrupted. This status information, which is determined locally for each individual computer, is preferably exchanged between the computers by transmission rounds and, as described above, is used to determine a global view of the system status. A central error handling entity then takes suitable measures. In the example shown in FIG. 6, a possible measure could be e.g. B. consist in permanently sending or receiving over this communication link KV12 permanently or for a predetermined time.
In the multi-computer system MRS shown in FIG. 7, the computer R1 does not send identical data packets A to both computers R2 and R3, as actually provided, but rather different data packets A and A '. The cause of such an error scenario, referred to as "Byzantine", can be, for example, that the transmitter unit is disturbed for a communication connection. Without forwarding the data packets according to claim 6, the computers R2 and R3 would not be able to determine that they have received different data packets. In their view, an error has not occurred. You would therefore process the data packets A or A 'further, however, since the data packets differ from one another, the results may be different. Since at least one of the results must be incorrect, the case could arise that this incorrect result coincides with the result determined by the defective computer R1. A simple error, namely the disturbance of the computer R1, could therefore already lead to a correspondence between two incorrect results and thus to a dangerous state.
If, on the other hand, the data packets are forwarded, the computers R2 and R3 in this case each receive two data packets from the computer R1 (one directly and one indirectly), which, however, differ from one another. If the data packets contain control data, the computers R2 and R3 can determine whether the data packets have been tampered with on the transmission path. If corruption has not occurred, computers R2 and R3 can reliably determine that computer R1 has sent different data packets A and A '. However, since computers R2 and R3 cannot see which of the two data packets is the right one, both computers discard the data packets received directly or indirectly from computer R1. In the local view of the two computers R2 and R3 and also in the global view, the computer R1 is therefore regarded as a "non-member".
The following explanations deal with the question already mentioned above, from which computers in the multi-computer system the transmission rounds are initiated. According to the invention, various criteria are provided here that can be used individually or in combination:<ul id="ul0004" list-style="none"><li>a) The application program executed by a computer itself gives the command to send out a data packet and thus initiate a round of transmission. Since, in multicomputer systems, all computers usually process the same application program in parallel, it can happen that all computers receive the command simultaneously or approximately simultaneously, ie at very short time intervals, to initiate a round of transmission. The scenario then arises, which has been described above for the exemplary embodiment according to claim 2.</li><li>b) Any computer in the multi-computer system regularly initiates a transmission round at predetermined intervals or in the course of operation. In this way it is ensured in particular that status information is exchanged regularly between the computers, so that any malfunctions of computers or transmission paths that may have occurred are quickly discovered. If, however, only one computer is responsible for triggering, no transmission rounds are triggered in the event of the failure of precisely this computer. Therefore, it will usually be more appropriate to have all computers initiate transmission rounds at regular intervals. If there are no messages to be exchanged on a computer, the computer in question sends out a data packet whose useful data area is empty.</li><li>c) A computer triggers a transmission round by sending out a data packet at least whenever the amount of messages stored in its message buffer (see below for more information) exceeds a predetermined amount. This reduces the exchange of empty data packets. The number of messages can also take the place of the amount of messages that can be measured in bytes. In this way it can be determined that a computer initiates a transmission round as soon as, for example, 5 messages are stored in its message buffer.</li><li>d) In security-critical multicomputer systems, it makes sense to provide messages with a priority value that is a measure of the importance of the user data. In a variant of the invention, it is provided to transmit the priority value in the data packets together with the user data and possibly further data. In this variant, a transmission round is initiated if the priority value of a message lies above a predetermined threshold. This ensures that important messages are exchanged immediately.</li></ul>
The criteria a) to d) are preferably combined in such a way that a portable compromise is achieved between the high security requirements on the one hand and the requirement for a low transmission volume on the communication connections on the other hand. As already mentioned, when the criteria are combined, it can happen that several computers initiate a transmission round simultaneously or approximately simultaneously.
Example 5:
8 schematically shows in a highly simplified model how messages and data packets are processed within a computer and exchanged in the multi-computer system according to another exemplary embodiment of the method according to the invention. The modules belonging to a computer R are enclosed by a dashed line. An application process AP carried out by any computer R in the multicomputer system sends messages NA to the computer R. These messages NA are buffered in an input message buffer ENP. When a message arrives, the incoming message buffer ENP communicates this to an exchange regulations module ATVM via messages MA. The ATVM exchange regulations module stores the basic exchange regulations according to which the data exchange between the computers within the multi-computer system is to be carried out. The exchange regulations define, for example, the criteria used to initiate transmission rounds. Furthermore, the exchange regulations module ATVM has the task of controlling the message stream in the input message buffer ENP (and also in the output message buffer ANP, see below).
In this exemplary embodiment, the messages MA from the input message buffer ENP to the exchange regulations module ATVM contain information about the priority value of the last received message. The exchange regulations module ATVM compares this priority value with a predefined threshold value and, if the threshold value is exceeded via commands B, causes the input message buffer ENP to create a data packet DP from the buffered messages NA and to deliver this data packet to the exchange module AM to initiate a transmission round. The data packet DP has a message header, which among other things contains the sender and recipient address, and a core in which the user data to be transmitted are stored.
In this exemplary embodiment, the exchange module AM fulfills, inter alia, the functions of the network and transport layers in the OSI layer model. The physical layers of the communication connections, the transmission methods used and the number of computers in the multi-computer system are thus hidden from the layers above. The exchange module AM also has the task of adding the following information to the header of the data packets:<ul id="ul0005" list-style="bullet"><li>a number identifying the round of transmission;</li><li>the newly determined local view of the system status;</li><li>the local time;</li><li>an identification of the computer on which the exchange module AM operates.</li></ul>
The data packet DP reaches the driver layer T from the exchange module, which corresponds to the connection layer in the OSI layer model and which maps the data packets to a physical bit stream. The driver layer ensures that the data packet is transmitted to other computers, not shown in FIG. 8, within a limited period of time. The driver layer should ensure that data packets are not rearranged during transmission, since otherwise u. U. it is not guaranteed that all computers in the multi-computer system can process a uniform data stream. In order to ensure a clear error assignment, no data packets should be falsified or newly generated during the transmission. The driver layer is transparent for the layers above.
In the receiving direction, data packets DP reach the exchange module AM via the driver layer T, in which in particular the information contained in the header of the data packets (local view of the system status, local time, number of the transmission round, etc.) is evaluated. The exchange module creates error messages FM and forwards them together with the data packets DP to an output message buffer ANP. The output message buffer ANP extracts the messages NA from the data packets and makes them available to an application process AP '.
FIG. 9 shows a summary of the structure of a data packet DP, which is used in a preferred exemplary embodiment of the invention when data is exchanged between the computers. The core of the data packet DP consists of the user data ND. The header H of the data packet contains, among other things<ul id="ul0006" list-style="bullet"><li>Data W for identifying which route the data packet has taken through the multi-computer system,</li><li>Data (LS) that characterize the local view of the sending computer from the system status in the previous transmission round,</li><li>Data LZ, from which the local time of the computer sending the data packet can be determined,</li><li>RN data to identify the current round of transmission,</li><li>Control data K for checking whether all or part of the above data have been received correctly.</li></ul>
The exemplary embodiments of the invention explained so far all relate to a 2-out-of-3 computer system, since the method according to the invention can be used particularly advantageously there. However, the method can also be used advantageously for multi-computer systems with other degrees of redundancy without modifications. The exemplary embodiments described above can also be transferred, for example, to a 2-out-of-2 computer system. Likewise, in a 3-out-of-5 computer system, the communication between the computers can still be carried out using the method according to the invention. Instead of individually addressing the computers, however, it can then be advantageous to carry out the communication in the "broadcasting" mode, ie data packets are either transmitted to all or to no computer. 10 shows an example of such a 3-out-of-5 computer system. Five computers R1 ... R5 communicate with one another in broadcasting mode via a first bus B1 and a second bus B2. The second bus B2 is optional and is used to increase the availability of the multi-computer system.
In this exemplary embodiment, a transmission round comprises the following steps:<ul id="ul0007" list-style="bullet"><li>a computer, for example computer R1, sends a data packet A to all computers R2 ... R5;</li><li>computers R2 ... R5 then send their own data packets B, C, D and E, which are received by all other computers.</li></ul>
An additional redundancy can be created with this arrangement if, after the first round of transmission via the first bus B1, a second round of transmission is carried out via the bus B1 and / or via the bus B2. The possibilities of error localization in broadcast mode are not as varied as in the exemplary embodiments described above.
8 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8
Every citation, both waysCites: the store holds 2 of 3
| Document | Relation | Office |
|---|---|---|
| EP0246218A | Cites | European Patent Office (EPO) |
| US5506962A | Cites | United States of America |
7 members in 3 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 19742918 | Germany | A | |
| 19742918 | Germany | A | |
| 19742918 | Germany | – | |
| 19742918 | – | – | – |
| DE1997142918 | – | – | – |
Members7
| Document | Office | Kind | |
|---|---|---|---|
| EP0905623A2 | European Patent Office (EPO) | A2 | |
| DE19742918A1 | Germany | A1 | |
| EP0905623A3 | European Patent Office (EPO) | A3 | |
| EP0905623B1This record | European Patent Office (EPO) | B1 | |
| AT313828T | Austria | T | |
| ATE313828T1 | Austria | T1 | |
| DE59813290D1 | Germany | D1 |
46 legal events, as 6 offices reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | Office | |
|---|---|---|---|
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| ExpiryMK07 | MK07 | AT | |
| Patent expired after termination of 20 yearsExpiredPE20 | PE20 | GB | |
| Patent ceasedCeasedPL | PL | CH | |
| Expiry of rightR071 | R071 | DE | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Fee paymentPLFP | PLFP | FR | |
| Fee paymentPLFP | PLFP | FR | |
| Be: lapsedLapsedBERE | BERE | EP | |
| No opposition filedOpposition26N | 26N | EP | |
| No opposition filed within time limitOppositionORIGINAL CODE: 0009261PLBE | PLBE | EP | |
| Information on the status of an ep patent application or granted ep patentGrantedSTATUS: NO OPPOSITION FILED WITHIN TIME LIMITSTAA | STAA | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Fr: translation filedET | ET | EP | |
| Nl: lapsed or annulled due to failure to fulfill the requirements of art. 29p and 29m of the patents actLapsedNLV1 | NLV1 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Gb: translation of ep patent filed (gb section 77(6)(a)/1977)GBT | GBT | EP | |
| Corresponds to:REF | REF | EP | |
| New agentNV | NV | CH | |
| European patent takes effect as a national patent in ch/liEP | EP | CH | |
| Designated contracting statesAK | AK | EP | |
| European patent grantedGrantedNOT ENGLISHFG4D | FG4D | GB | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| (expected) grantORIGINAL CODE: 0009210GRAA | GRAA | EP | |
| Grant fee paidORIGINAL CODE: EPIDOSNIGR3GRAS | GRAS | EP | |
| Despatch of communication of intention to grant a patentORIGINAL CODE: EPIDOSNIGR1GRAP | GRAP | EP | |
| Title (correction)METHOD FOR EXCHANGING DATA PACKETS IN A SAFE MULTICOMPUTER SYSTEM AND MULTICOMPUTER SYSTEM FOR CARRYING OUT THE SAMERTI1 | RTI1 | EP | |
| First examination report despatched17Q | 17Q | EP | |
| Designation fees paidAT BE CH DE DK ES FI FR GB IT LI NL PT SEAKX | AKX | EP | |
| Request for examination filed17P | 17P | EP | |
| Information provided on ipc code assigned before grant6G 06F 11/16 A, 6G 06F 11/18 B, 6G 06F 11/00 BRIC1 | RIC1 | EP | |
| Designated contracting statesAK | AK | EP | |
| Request for extension of the european patentAL;LT;LV;MK;RO;SIAX | AX | EP | |
| Search report despatchedORIGINAL CODE: 0009013PUAL | PUAL | EP | |
| Designated contracting statesAK | AK | EP | |
| Request for extension of the european patentAL;LT;LV;MK;RO;SIAX | AX | EP | |
| Public reference made under article 153(3) epc to a published international application that has entered the european phaseORIGINAL CODE: 0009012PUAI | PUAI | EP |
Numbers
- Publication
- 0905623
- Publication, DOCDB
- 0905623
- Publication, EPODOC
- EP0905623
- Application
- 98440198
- Application, DOCDB
- 98440198
- Application, EPODOC
- EP19980440198
Titles3
- German
- Verfahren zum Austausch von Datenpaketen innerhalb eines sicheren Mehrrechnersystems und Mehrrechnersystem zur Ausführung des Verfahrens
- English
- Method for exchanging data packets in a safe multicomputer system and multicomputer system for carrying out the same
- French
- Méthode pour l'échange de packets de données dans un système sécurisé de multiordinateur et multiordinateur mettant la méthode en oeuvre
Classification
- CPC, 1
- G06F11/18
- IPC, 5
- G06F9 46
- G06F11 00
- G06F11 16
- G06F11 18
- G06F15 163
Designated states14
- Contracting states, 14
- Austria
- Belgium
- Switzerland
- Germany
- Denmark
- Spain
- Finland
- France
- United Kingdom
- Italy
- Liechtenstein
- Netherlands (Kingdom of the)
- Portugal
- Sweden
