Network node failover using failover or multicast address
Abstract
Public network node failure repair. The first node joins the multicast group (102). Join by performing one of three operations (104). First, the fault repair address is associated with the first node, and the first node effectively joins the group that uses the fault repair address as the multicast address. Second, associate the multicast address with the first node. Third, the multicast port of the switch is mapped to the port of the first node. When the first node fails (106), one of three operations is performed. If the joining involves a fault repair address, the fault repair address is associated with the second node, and the second node effectively joins the group (114). If the joining involves a multicast address, the second node joins the group, which address is related to the second node (110). If it joins the multicast port of the mapped switch, the port is remapped to the second node port (112).

Term
Term ended
Projected expiry passed 26 July 2022, 4.2 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
10 claims: 1 independent, 9 dependent
- 1一种方法,包括:网络的第一节点加入具有多播地址的多播组(102),这里所述加入选自实质上包括下述之一的组:使故障修复地址与第一节点相关联,从而第一节点有效加入把故障修复地址作为多播地址的多播组,给故障修复地址的通信通过网络被引向第一节点;使多播地址与第一节点相关联,从而给多播地址的通信通过网络被引向第一节点;和把网络的交换机上的多播端口映射到第一节点上的端口,从而给多播地址的通信从交换机上的多播端口被引向第一节点上的端口(104);和当第一节点发生故障时(108),如果加入使故障修复地址与第一节点相关联,则使故障修复地址与第二节点相关联,从而第二节点有效加入多播组,并且给故障修复地址的通信由第二节点处理(114);如果加入使多播地址与第一节点相关联,网络的第二节点加入多播组,从而多播地址与第二节点相关联,并且给多播地址的通信由第二节点处理(110);和如果加入把交换机上的多播端口映射到第一节点上的端口,则把交换机上的多播端口重映射到第二节点上的端口,从而给多播地址的通信被引向第二节点上的端口(112)。
- 2按照权利要求1所述的方法,其中网络是Infiniband网络。
- 3按照权利要求1所述的方法,其中故障修复地址选自实质上包括下述之一的组:值小于故障修复LID阈值的故障修复位置标识符(LID),网络包括Infiniband网络;有效LID范围内的故障修复位置标识符(LID),网络包括Infiniband网络;和作为源LID对于通过网络的任意传送方法来说不被检查、并且作为多播目的地LID(DLID)对于通过网络的任意传送方法被接受的故障修复位置标识符(LID),其中网络包括Infiniband网络。
- 4按照权利要求1、2或3所述的方法,还包括如果加入使多播地址或故障修复地址与第一节点相关联,在网络的第二节点加入多播组之前,通过第二节点代表第一节点向子网管理器(SM)发送离开请求,第一节点离开多播组(706)。
- 5按照权利要求1、2、3或4所述的方法,还包括如果加入使多播地址或故障修复地址与第一节点相关联,当第一节点消除故障(failback)时,使故障修复地址与第一节点相关联,从而给故障修复地址的通信重新由第一节点处理(712)。
- 6按照权利要求1、2、3、4或5所述的方法,其中如果加入使多播地址或故障修复地址与第一节点相关联,则网络的第一节点加入多播组包括第一节点向子网管理器(SM)请求加入多播组。
- 7按照权利要求1或2所述的方法,其中如果加入把交换机上的多播端口映射到第一节点上的端口,通过第二节点向子网管理器(SM)请求把交换机上的多播端口重映射到第二节点,SM把交换机上的多播端口重映射到第二节点上的端口,交换机上的多播端口被重映射到第二节点上的端口。
- 8按照权利要求1、2或7所述的方法,还包括如果加入把交换机上的多播端口映射到第一节点上的端口,当第一节点故障消除时,把交换机上的多播端口重映射到第一节点的端口上,从而给多播地址的通信再次被引向第一节点上的端口(912)。
- 9一种用于权利要求1、2、3、4、5、6、7或8所述方法的故障修复节点。
- 10一种产品,包括:计算机可读介质;和所述介质中用于实现按照权利要求1、2、3、4、5、6、7或8所述方法的装置。
Independent claims10
75 paragraphs, as filed
Using fault repair or multicast address network node fault repair
Technical field
The present invention relates to a network, such as an Infiniband network, and particularly relates to the failover of nodes in such a network.
Background technique
An input/output (I/O) network, such as a system bus, can be used in a computer's processor to communicate with peripheral devices such as network adapters. However, the structure of common I/O networks, such as Peripheral Component Interface (PCI) bus constraints, limit the overall performance of the computer. Therefore, a new type of I/O network was proposed.
A known new type of I/O network is called an Infiniband network. The Infiniband network replaces the PCI or other buses in the current computer with a packet-switched network with one or more routers. The host channel adapter (HCA) is coupled to the processor and the subnet, and the target channel adapter (TCA) is coupled to the peripheral and the subnet. The subnet includes at least one switch, and links that connect the HCA and TCA to the switch. For example, a simple Infiniband network may have a switch, and HCA and TCA are connected to it through a link. More complex layouts are also possible and predictable.
Each end node of the Infiniband network includes one or more channel adapters (CA), and each CA includes one or more ports. Each port has a local identifier (LID) assigned by the local subnet manager (SM). Within the subnet, the LID is unique. The switch uses LID to route packets within the subnet. Each packet of data contains a source LID (SLID) and a destination LID (DLID). The source LID identifies the port where the packet is injected into the subnet, and the destination LID identifies the port where the Infiniband structure or network will transmit the packet.
The Infiniband network method provides multiple virtual ports within a physical port by defining a LID mask count (LMC). LMC specifies the number of least significant bits of the LID that the physical port masks or ignores when it is confirmed that the packet DLID matches its assigned LID. But the switch does not ignore these bits. Thus, SM can program different paths through the Infiniband structure according to the least significant bit. Thus, the ports can be considered as 2LMC ports used for routing purposes within the Infiniband structure.
For critical applications that require non-faulty continuous availability, fault repair of a single application is usually required, thereby requiring fault repair of communication endpoints or end nodes. The communication endpoint in the Infiniband network environment is related to the CA port. Applications use endpoints to communicate within the Infiniband network, such as communicating with other applications. The transparent fault repair of the endpoint means that the other endpoint takes over the responsibility of the failed endpoint in a way that does not interfere with the communication within the network itself.
However, due to the way the endpoints are addressed, it is more difficult to repair transparent failures of the endpoints or other nodes in the Infiniband network. Failure repair requires that the LID be reassigned to a new port that takes over the failed port. However, the new port usually already has a LID assigned to it. Therefore, the only way to allocate additional LIDs is to extend the LMC range on the port to ensure that the new LID falls within that range.
However, it is actually difficult to expand the LMC range on the port, and sometimes considerable overhead is required to ensure that the takeover ports can have the LID assigned to their failed ports. Therefore, LID fault repair is considered to be a problem and obstacle to the successful rollout of the Infiniband network that requires transparent fault repair. For the above reasons, the present invention is needed.
Summary of the invention
The present invention relates to fault repair of nodes in a network using fault repair or multicast addresses. In one method of the present invention, the first node of the network joins a multicast group with a multicast address. The joining is achieved by performing one of three operations. First, the fault repair address can be associated with the first node, so that the first node effectively joins the multicast group that uses the fault repair address as the multicast address. The communication to the fault repair address is directed to the first node through the network. Second, the multicast address can be associated with the first node, so that communications to the multicast address are directed to the first node through the network. Third, the multicast port on the switch of the network can be mapped to the port on the first node. Communications given to the multicast address are directed from the multicast port on the switch to the port on the first node.
When the first node fails, one of the three operations is performed corresponding to the method for the first node to join the network. If the join associates the fault repair address with the first node, the fault repair address is related to the second node, so that the second node effectively joins the multicast group, and the communication to the fault repair address is processed by the second node. If joining makes the multicast address associated with the first node, the second node joins the multicast group, so that the multicast address is associated with the second node, and the communication to the multicast address is handled by the second node. If you join to map the multicast port on the switch to the port on the first node, the multicast port on the switch is remapped to the port on the second node. Thus the communication to the multicast address is directed to the port on the second node.
The invention also includes fault repair nodes and manufactured products. The fault repair node is a node that implements the method of the present invention, and the manufactured product has a computer-readable medium and a device for implementing the method of the present invention in the medium. With reference to the accompanying drawings, other features and advantages of the present invention will be apparent according to the following detailed description of the preferred embodiments of the present invention.
Description of the drawings
Figure 1 is a flowchart of a method according to a preferred embodiment of the present invention, and is suggested to be printed on the first page of a patent issued.
Figure 2 is a diagram of an Inifiniband network that can be implemented in conjunction with an embodiment of the present invention.
FIG. 3 is a diagram of an exemplary Inifiniband system area network (SAN) that an embodiment of the present invention can be implemented in conjunction with.
Figure 4 is a diagram of the communication interface of an exemplary end node of an Inifiniband network.
Figures 5 and 6 are diagrams of Inifiniband networks showing how Inifiniband addressing is performed.
FIG. 7 is a flowchart showing how an embodiment of the present invention can associate a fault repair address of a multicast group and/or a multicast address of a multicast group with another node to implement a method for repairing a network node fault.
Fig. 8 is a diagram showing the performance of the embodiment of Fig. 7.
FIG. 9 is a flowchart showing how an embodiment of the present invention can remap a switch multicast port to a port on another node to implement a method for repairing a network node failure.
Fig. 10 is a diagram showing the performance of the embodiment of Fig. 9.
detailed description
Overview Figure 1 shows a method 100 according to a preferred embodiment of the invention. The first node of the network effectively joins the multicast group initially (102). The multicast group has a multicast address or a fault repair address. Perform at least one of the three operations (104). In the first mode, the multicast address is assigned to the first node. The communication to the multicast address can then be automatically directed to the first node, where the network may have been previously established manually or automatically in order to achieve such communication. In the second mode, the multicast port on the switch of the network is mapped to or associated with the port on the first node. Communication to the multicast address can then be directed from the multicast port on the switch to the port on the first node, where the switch does not support multicast. In the third mode, the fault repair address is assigned to the node. The communication to the fault repair address is then automatically directed to the first node, where the network has been previously established manually or automatically in order to realize this communication. The network is preferably an Infiniband network. The first and second nodes may be hosts with channel adapters (CA) and ports on such a network.
The first node subsequently fails (108), so that the transparent fault repair of the first node is preferably achieved by the second node of the network. This can involve performing one of three operations. First, the second node can join the multicast group so that the multicast address is also assigned to the second node (110). The communication to the multicast address is thus directed to the second node and directed to the first node (the failed node), so that the second node takes over the processing of this communication from the first node. Second, the multicast port on the switch can be remapped to the port on the second node (112). The communication to the multicast address is thus directed to the port on the second node, so that the second node takes over the processing of this communication. Third, the second node is associated with the fault repair address, so that the second node effectively joins the multicast group (114). The communication to the fault repair address is thereby directed to the second node and directed to the first node (the failed node), so that the second node takes over the processing of this communication from the first node.
A management component such as the subnet manager (SM) of the Infiniband subnet can assign the multicast address of the multicast group initially assigned to the first node to the second node. The management component can also remap the multicast port of the switch originally mapped to the port on the first node to the port on the second node. The device in the computer-readable medium of the article of manufacture can also realize this function. The device may be a recordable data storage medium, a modulated carrier signal, or another type of medium or signal.
So in the first mode, the multicast address is used for unicast communication. Multicast addresses allow local identifier (LID) failure repairs, because only multicast LIDs can be shared by more than one port. In the second mode, the node in question is connected to the main multicast port of the switch. When the fault is repaired, the switch configuration is modified by redistributing the main port, so that the packet is propagated to the fault repair node. In the third mode, the fault repair LID is allowed to be associated with any multicast group. In addition, the fault-recovery LID does not include the multicast group address.
Technical Background FIG. 2 shows an exemplary Infiniband network structure 200 with which embodiments of the present invention can be implemented. Infiniband network is a kind of network. The present invention can also be implemented with other types of networks. The processor 202 is coupled to the host interconnect 204, and the memory controller 206 is also coupled to the host interconnect 204. The memory controller 206 manages the system memory 208. The memory controller 206 is also connected to a host channel adapter (HCA) 210. The HCA 210 allows the processor and the memory subsystem to communicate through the Infiniband network. The processor and the memory subsystem include the processor 202, the host interconnect 204, the memory controller 206, and the system memory 208.
The Infiniband network in FIG. 2 is called a subnet 236, and the subnet 236 includes Infiniband links 212, 216, 224, and 230, and an Infiniband switch 214. There may be more than one Infiniband switch, but only switch 214 is shown in FIG. 2. Links 212, 216, 224, and 230 enable the HCA and target channel adapters (TCA) 218 and 226 to communicate with each other, and also enable the Infiniband network to communicate with other Infiniband networks through the router 232. Specifically, the link 212 connects the HCA 210 and the switch 214. Links 216 and 224 connect TCA 218 and 226 to switch 224, respectively. The link 230 connects the router 232 and the switch 214.
The TCA 218 is a specific peripheral, in this case, the target channel adapter of the Ethernet adapter 220. TCA can accommodate multiple peripherals, such as multiple network adapters, SCSI adapters, etc. TCA 218 enables the network adapter 220 to send and receive data through the Infiniband network. The adapter 220 itself allows communication via a communication network, especially Ethernet, as shown by line 222. Other communication networks are also suitable for the present invention. The TCA 226 is another peripheral, the target channel adapter of the target peripheral 228. The target peripheral 228 is not described in detail in FIG. 2. The router 232 allows the Infiniband network of FIG. 2 to be connected to other Infiniband networks, and the line 234 represents the connection.
The Infiniband network is a packet-switched input/output (I/O) network. Thus, by interconnecting 204 and the memory controller 206, the processor 202 sends and receives data packets via the HCA 210. Similarly, the target peripheral 228 and the network adapter 220 send and receive data packets through TCA 226 and 218, respectively. Data packets can also be sent and received through the router 232, which connects the switch 214 and other Infiniband networks. The links 212, 216, 224, and 230 may have varying capacities, depending on the bandwidth required by the particular HCA, TCA, etc. they are connected to the switch 214.
The Infiniband network provides communication between TCA and HCA in the different ways described briefly here. Similar to other types of networks, Infiniband networks have a physical layer, link layer, network layer, transport layer and advanced protocols. As in other types of packet-switched networks, in the Infiniband network, specific transactions are divided into messages, and the messages themselves are divided into packets for transmission through the Infiniband network. When received by the intended recipient, the packet is recorded in the constituent message of the specified transaction. The Infiniband network provides queues and channels in which packets are received and sent.
In addition, the Infiniband network allows many different delivery services, including reliable and unreliable connections, reliable and unreliable datagrams, and raw packet support. In reliable connections and datagrams, a packet sequence number that confirms and guarantees packet ordering is generated. Duplicate packets are rejected, and lost packets are detected. In unreliable connections and datagrams, no acknowledgments are generated, and packet ordering is not guaranteed. Do not reject duplicate packets, and do not detect lost packets.
The Infiniband network can also be used to define a system area network (SAN) to connect multiple independent processor platforms, or main processor nodes, I/O platforms, and I/O devices. Figure 3 shows an exemplary SAN 300 in conjunction with which an embodiment of the invention can be implemented. SAN300 is a communication and management infrastructure that supports I/O and inter-processor communication of one or more computer systems. Infiniband systems range from small servers to large parallel supercomputer central stations. In addition, the Internet Protocol (IP)-friendly nature of the Infiniband network allows bridging to the Internet, corporate intranet, or connection to remote computer systems.
The SAN 300 has a switched communication structure 301, or subnet, and the communication structure 301 or subnet allows many devices to communicate simultaneously with high bandwidth and low latency in a protected remote management environment. End nodes can communicate through multiple Infiniband ports and can utilize multiple paths through structure 301. The diversity of ports and paths through the network 300 is used for fault tolerance and increased data transfer bandwidth. Infiniband hardware removes most of the processor and I/O communication operations. This allows multiple simultaneous communications without the traditional overhead associated with communication protocols.
The structure 301 specifically includes a number of switches 302, 304, 306, 310, and 312, allowing the structure 301 to be linked with other Infiniband subnets, wide area networks (WAN), local area networks (LAN) and host routers 308 (as shown by arrow 303). Structure 301 allows several hosts 318, 320, and 322 to communicate with each other, as well as with different subsystems, management consoles, drives, and I/O racks. In Figure 3, these various subsystems, management consoles, drives, and I/O racks are represented as redundant array of information disks (RAID) subsystem 324, management console 326, I/O racks 328 and 330 , Drive 332 and storage subsystem 334.
Figure 4 shows the communication interface of an exemplary end node 400 of an Infiniband network. The end node may be one of the hosts 318, 320, and 322 of FIG. 3. The end node 400 has processes 402 and 404 running on it. Each process has one or more queue pairs (QP) associated with it, and each QP communicates with the channel adapter (CA) 418 of the node 400 to link with the Infiniband structure, as shown by arrow 420. For example, process 402 has QPs 406 and 408, and process 404 has QP 410.
Define QP between HCA and TCA. Each end of the link has a message queue to be transmitted to the other end. QP includes paired sending work queues and receiving work queues. In general, the sending work queue holds instructions that cause data to be transferred between the client's memory and the memory of another process, and the receiving work queue holds instructions on where to place data received from another process.
QP represents a virtual communication interface with the Infiniband client process, and provides a virtual communication port for the client. CA can provide up to 224 QPs, and the operations on each QP are independent of each other. The client generates a virtual communication port by allocating QP. The client initiates any communication establishment required to connect this QP with another QP, and configures the QP environment with certain information, such as destination address, service level, protocol operating limits, etc.
Figures 5 and 6 show how to address in the Infiniband network. In FIG. 5, a simple Infiniband network 500 is shown, and the Infiniband network 500 includes an end node 502 and a switch 504. The end node 502 has a process 504 running on it, and the process 504 has related QPs 506, 508, and 510. The end node 502 also includes one or more CAs, such as CA 512. CA 512 includes one or more communication ports, such as ports 514 and 516. The QPs 506, 508, and 510 have a queue pair number (QPN) assigned by the CA, and the queue pair number uniquely identifies the QP in the CA 512. Data packets other than the original datagram contain the QPN of the destination work queue. When CA 512 receives a packet, it uses the environment of the destination QPN to process the packet appropriately.
The local subnet manager (SM) assigns a local identifier (LID) to each port. SM is a management component connected to the subnet and is responsible for configuring and managing switches, routers, and CAs. Other equipment such as CA or switch can be used to embed SM. For example, the SM can be embedded in the CA 512 of the end node 502. As another example, the SM may be embedded in the switch 504.
Within the Infiniband subnet, the LID is unique. Switches such as switch 504 use LID to send packets within the subnet. Each packet contains a source LID (SLID) and a destination LID (DLID). The source LID identifies the port where the packet is injected into the subnet, and the destination LID identifies the port that the structure will transmit the packet to. Switches such as switch 504 also each have many ports. Each port on the switch 504 can be associated with a port on the end node 502. For example, port 518 of switch 504 is associated with port 516 of end node 502, as indicated by arrow 520. The data packet destined for the port 516 of the node 502 received by the switch 504 is thus sent from the port 518 to the port 516. More specifically, when the switch 504 receives a packet with a DLID, the switch only checks whether the DLID is non-zero. Otherwise, the switch sends the packet according to the form designed by SM.
In addition to identifying the DLID of a specific port in the Infiniband subnet, a multicast DLID or multicast address can also be specified. Generally, a group of end nodes can join a multicast group, so that the SM allocates a multicast DLID of the multicast group to a port of each node. The data packet sent to the multicast DLID is sent to each node that joins the multicast group. Each switch, such as switch 504, has a default primary multicast port and a default non-primary multicast port. The primary/non-primary multicast port is used for all multicast packets and is not related to any specific DLID. One port of each node that joins the multicast group is either associated with the primary multicast port of the switch or associated with the non-primary multicast port of the switch.
When a data packet with a multicast DLID is received, the multicast DLID is detected, and the data packet is forwarded according to the table planned by the SM. If the multicast DLID is not in the table, or the switch does not save the table, the switch forwards the packet on the primary default multicast port and the non-primary default multicast port. If received on the main port, the packet goes out of the non-primary multicast port, and if it is received on any other port of the switch, the packet goes out of the main multicast port. The data packet specifying the multicast DLID received by the switch 504 is thus sent from one of these multicast ports to the relevant port of the multicast group node. The switch 504 can be configured with routing information about the multicast communication that specifies the port to which the packet should be sent.
In addition, although any Infiniband node can transmit packets to any multicast group, if a switch, such as switch 504, incorrectly forwards the packet, there is no guarantee that the data packet will be correctly received by the members of the multicast group. Therefore, the switch should be set up so that the multicast data packet is received by the group members. This can be achieved by ensuring that multicast data packets always file through one or more switches that are pre-programmed or specially programmed, thereby ensuring that the multicast data reaches their correct destination. On the other hand, if all switches fully support multicast, the end node joining the multicast group will cause the SM to program the switch so that the packet is correctly received by all members of the multicast group. Other methods can also be performed.
In FIG. 6, a more complex Infiniband network 600 is shown, and the Infiniband network 600 has two subnets 602 and 604. The subnet 602 has end nodes 604, 606, and 608 connected to the switches 610 and 612 differently. Similarly, subnet 604 has end nodes 614, 616, 618, and 620 that are connected to switches 622 and 624 differently. The switches 610 and 612 of the subnet 602 are differently connected to the switches 622 and 624 of the subnet 604 through routers 626 and 628, and the routers 626 and 628 can implement communication between the subnets. In this case, connecting differently means that one or more ports of one entity are associated with one or more ports of another entity. For example, the node 604 may have two ports, one associated with the switch 610 and the other associated with the switch 612.
In order to repair the fault of the first node, associate the fault repair (multicast) address with the second node. By associating the fault repair address of the multicast group with another node, the embodiment of the present invention can realize network node fault repair. Figure 7 shows a method 700 according to such an embodiment of the invention. In this embodiment, it is better to redefine the Inifiniband specification of the location identifier (LID) as follows:
ThLID is a threshold specified by the administrator, so it is best that only LIDs higher than ThLID can be fault repair LIDs, which are also effectively multicast LIDs. In addition, the Inifiniband specification is preferably enhanced to allow the fault repair LID to be associated with the multicast group identifier (GID). It is allowed to use this fault repair LID in the presence or absence of a GID. In the case that ThLID is equal to 0XC000 (the starting value of the multicast range in the current Inifiniband specification), this embodiment is consistent with the current specification.
In another embodiment of the present invention, in addition to the permitted LID, any valid LID can be associated with the multicast group, so that it can effectively function as a fault repair LID. The subnet manager (SM) is enhanced to allow any such valid LIDs other than the permitted LID to be associated with the multicast group. That is, the Infiniband specification is modified so that the SM can allow any valid LID other than the permitted LID to be associated with the multicast group in order to allow node failure repair. Finally, in an alternative embodiment of the present invention, no changes are made to the Infiniband specification, so that, in contrast to the fault repair LID, which is effectively also a multicast group LID, the following description of the method 700 of FIG. 7 is only relevant to the multicast group LID related.
Referring now to the method 700 of FIG. 7, the first node of the Infiniband network is related to the fault repair LID (or multicast LID), and the fault repair LID is effectively the multicast group LID, so that the first node effectively joins the multicast group (702) . The SM of the subnet of which the first node is a part associates the fault repair LID with the multicast group, for example, in response to the first node's request to join the multicast group. The first node may be the channel adapter (CA) of the host on the subnet of the Infiniband network. The first node subsequently fails (704), which is usually detected by another node of the subnet. The second node through the subnet sends a leave request to the SM on behalf of the first node, and the first node can optionally leave the multicast group (706).
The second node then joins the multicast group, thereby receiving packets sent to the fault repair LID (or sent to the multicast LID) (708). More specifically, in response to the join request from the second node, the SM programs the switch so that the packets sent to the multicast group will be received by the second node. The second node can also be the CA of the host on the subnet of the Infiniband network. The host of the second node may be the same host as the host of the first node. The communication planned for the fault repair LID is handled by the second node instead of the first node, so that the first node seamlessly fails over to the second node.
At a certain moment, the first node may eliminate the failure (failback) (710) and get back online. Subsequently, the fault repair LID (or multicast LID) is again associated with the first node (712), so that the first node can reprocess the communication planned for the fault repair LID. Before the first node rejoins the multicast group, the second node of the subnet can leave the multicast group. Therefore, before the first node sends a join request to the SM so that the SM associates the fault repair LID with the first node, the second node can send a leave request to the SM. Failure elimination (failback) may also include that the first node obtains a state dump from the second node, where the second node freezes all connections until the failure elimination is completed. On the other hand, the second node may not leave the multicast group until the existing connection with the second node expires.
Therefore, the third node that originally communicated with the first node will not know that the fault repair has been performed to the second node. That is, it will continue to communicate with the fault repair address without having to know whether the fault repair address is associated with the first node or the second node. Generally, even if the multicast fault repair address is used to implement fault repair, the communication between the third node and the first node or the second node is actually a unicast communication. The third node does not know that the fault repair address is in fact a multicast address, leading to the belief that the communication between it and the fault repair address is actually a unicast communication. That is, when the multicast address is actually being used to complete the communication, the third node may see that the communication is proceeding normally.
Figure 8 shows the first node failure repair relative to the second node. The multicast group is represented as a multicast group 802A in order to represent the pre-failure state of the first node 804. Then, the packet 806 with the fault repair LID is sent to the first node 804. The multicast group is represented as a multicast group 802B in order to represent the post-failure state of the first node 804, so that after the first node 804 fails, the group 802A becomes the group 802B, as indicated by the arrow 808. The first node 804 of the group 802A becomes the first node 804' of the group 802B in order to indicate the failure. The second node 810 joins the multicast group 802B. The first node 804' is represented as being in group 802B, but may have left group 802B. Thus, in addition to the first node 804', the packet 806 is now sent to the second node 810.
To repair the failure of the first node, remap the switch multicast port to the port on the second node. By remapping the switch multicast port to the port on another node, the embodiment of the present invention can also implement network node failure repair. Figure 9 shows a method 900 according to such an embodiment of the invention. The first node of the Infiniband network joins the multicast group. In this case, the primary multicast port on the switch is mapped to the port on the first node (902). In response to the first node's joining request, the first node and the subnet manager (SM) of the subnet of which the switch is a part implement this mapping. The first node may be the channel adapter (CA) of the host on the network subnet.
The first node subsequently fails (904), which is usually detected by another node of the subnet. The second node through the subnet sends a leave request to the SM on behalf of the first node, and the first node can optionally leave the multicast group (906). The primary multicast port on the switch is then remapped to the port on the second node (708). More specifically, in response to a corresponding request from the second node (optionally a dedicated request), the SM remaps the primary multicast port on the switch to the port on the second node. The second node may also be the CA of the host on the subnet of the Inifiniband network. The host of the second node may be the same host as the host of the first node. The communication given to the multicast address is directed to the port on the second node instead of the port on the first node, so that the first node seamlessly transfers to the second node.
At a certain moment, the first node may eliminate the fault (910) and thus return to the online. The primary multicast port on the switch is then remapped to the port on the first node (912), so that the first node can again process communications scheduled for the multicast address, which can be a multicast destination location identifier (DLID) . The second node of the subnet may have to leave the multicast group initially, so that a leave request can be sent to the SM before the main multicast port is remapped to the port on the first node. The fault elimination may also include that the first node obtains a state dump from the second node. In this case, the second node freezes all connections until the fault elimination is completed. In addition, the second node may not leave the multicast group until the existing connection with the second node expires.
FIG 10 shows a first node of the second node with respect to the failure to repair the complex. A part of the subnet is represented as part 1002A in order to represent the pre-failure state of the first node 1004. The first node 1004 has a port 1006. The switch 1008 has a main multicast port 1010. The primary multicast port 1010 of the switch 1008 is mapped to the port 1006 of the first node 1004, as shown by the line 1012. The multicast communication directed to the switch 1008 is thus sent to the port 1006. This part of the subnet is represented as part 1002B in order to represent the post-failure state of the first node 1004, so that after the failure of the first node 1004, part 1002A becomes part 1002B, as shown by arrow 1014. The second node 1016 has a port 1018. The multicast port 1030 of the switch 1008 now becomes the primary multicast port and is mapped to the port 1018 of the second node 1016, as shown by the line 1020. The multicast communication directed through switch 1008 is now sent to port 1018.
Switches, datagrams, and connection service types Infiniband networks use switches that usually only check that the destination location identifier (DLID) is not zero, and send data packets according to a table designed by the subnet manager (SM). Preferably, each switch is configured with routing information of the multicast communication, the routing information designating all ports through which the multicast data packet needs to pass. This ensures that multicast packets are sent to their correct destination.
In addition, Infiniband networks can use different types of datagrams and connection services. Datagrams are used when the order in which packets are received is not important compared to the order in which packets are sent. Datagram packets can be received out of order compared to the order in which they were sent. Datagrams can be primitive, which means they conform to non-Inifiniband specifications such as Ethertype, Internet Protocol (IP) version 6, etc. Conversely, use connection services when the order of receiving packets is more important than the order in which packets are sent. The connection service packets are received in the same order in which the packets were sent.
Both datagrams and connected services can be reliable or unreliable. Reliability generally involves whether to maintain the sequence number of the packet, whether to send an acknowledgment message on the received packet, and/or whether to perform other verification measures to ensure that the sent packet is received by their intended recipient. Unreliable datagrams and unreliable connection services do not perform such verification measures, while reliable datagrams and unreliable connection services perform such verification measures.
For unreliable original datagrams, the first node uses the multicast location identifier (LID) as its source LID (SLID). In the case where the second node is a fault repair node, the third node is able to receive such packets because they are sent to its unicast DLID and because the SLID of the packet has not been checked. The third node should reply to the multicast LID of the first node. To this end, the client can be sent a multicast LID association, which is recorded by the client. In the case of the SIDR protocol for unreliable datagrams, the third node can be sent a multicast LID, and/or the second node can pick up the multicast LID from the received packet. In the Inifiniband specification, there is no validity check specified for unreliable datagram mode.
If the third node determines the LID based on the path record maintained by the SM, the appropriate value of the LID can be placed in the path record before initiating communication with the third node. When the first node, or the failure repair second node, receives the reply packet from the client, the packet has a non-multicast queue pair (QP), but has a multicast DLID. In the case of reliable datagram and connection mode delivery, the connection manager exchanges the LID that will be used for communication. At this stage, multicast or fault repair LIDs can be exchanged. This LID can be recorded by the second node without any validity check and used as a unicast LID in all communications.
Both link layer and transport layer checks are also verified. The link layer check only verifies the client's LID, either a multicast LID or a unicast LID. In the transport layer check, the receiving QP is first verified as valid because the sender sets the QP. Finally, the QP is verified as not being 0xFFFFFFF (hexadecimal), so the data packet is not considered as a multicast packet, so the existence of the multicast global routing header (GRH) is not checked.
However, in one embodiment of the present invention, the Infiniband specification is redefined to provide node failure repair by not strictly performing these transport layer checks. In this embodiment, the source LID (SLID) is not checked for any transmission method, and the multicast destination LID (DLID) is accepted for any transmission method. As a result, the Inifiniband specification has been modified so that SLID and DLID checks are not as strictly enforced as before.
In another embodiment of the present invention, on the other hand, by setting the QP to a specific value (for example, 0xFFFFFFE), it represents multicast communication and provides node failure repair. This embodiment of the invention is only suitable for unreliable connection services. The specific QP value is configurable and can be maintained by the SM.
For reliable datagrams and reliable and unreliable connection services, multicast is not allowed because it is not defined. However, if the two end nodes work in a unicast manner, this limitation can be overcome. The server sends the packet to the client using the multicast LID. The remote client checks whether the SLID is a multicast LID. If so, you can modify the host channel adapter (HCA) of the client to receive the multicast SLID, otherwise you can modify the SM to associate the unicast LID with the multicast group.
That is, only when it is greater than 0xC000 (hexadecimal), the unmodified receiver concludes that the SLID is multicast. Therefore, the SM is modified so that it assigns a value lower than 0xC000 (hexadecimal) to the multicast group, so that the receiver constantly determines that the SLID is a multicast. The client responds to the server, and the server receives the packet with the specified DLID. The server checks whether the DLID is a multicast LID. If so, the HCA of the server can be modified to receive the multicast DLID, or the SM can be modified to associate the single-access communication LID with the multicast group.
Advantages over the prior art Embodiments of the present invention provide advantages over the prior art. By using the multicast address and port of the Infiniband network, node failure repair is realized. Even if the specified Infiniband structure does not allow multicast, in the case that the faulty node leaves the multicast group before another node joins the multicast group, the embodiment of the present invention can still be used, so that only one node exists in the multicast group each time . The fault repair of the faulty node does not require involving the remote node with which the faulty node has been communicating.
Thus, the fault repair node transparently assumes the responsibility of the faulty node, and usually does not notify the remote node. Preferably, any host can take over the responsibility of the failed host. The embodiments of the present invention are also applicable to all Infiniband transmission types. The embodiments of the present invention generally do not require non-proprietary extensions to the Infiniband specification, so the embodiments work under the support of the Infiniband specification.
In addition, in other embodiments of the present invention, node fault repair is achieved by using the fault repair address of the Infiniband network specified in the present invention. The fault repair of the faulty node does not need to involve the remote node with which the faulty node has been communicating. In contrast, the takeover node transparently assumes the responsibility of the failed node and usually does not notify the remote node. Preferably, any host can take over the responsibility of the failed host. The embodiments of the present invention are also applicable to all Infiniband transmission types.
Although the alternative embodiments have described specific embodiments of the present invention for illustrative purposes, various modifications can be made without departing from the spirit and scope of the present invention. For example, although the present invention is mainly described about Inifiniband networks, the present invention is also applicable to other types of networks. Therefore, the scope of the present invention is limited only by the following claims and their equivalents.
11 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| CN110097660A | Cited by | China | Search report |
| WO2010000172A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| CN104836755A | Cited by | China | Search report |
| CN114500375A | Cited by | China | Search report |
17 members in 9 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 09917464 | United States of America | – | |
| 91746401 | United States of America | A | |
| 91746401 | United States of America | A | |
| 09917464 | – | – | – |
| US20010917464 | – | – | – |
Members17
| Document | Office | Kind | |
|---|---|---|---|
| US2003023896A1 | United States of America | A1 | |
| CA2451101A1 | Canada | A1 | |
| WO03013059A1 | World Intellectual Property Organization (WIPO) | A1 | |
| KR20040012978A | Republic of Korea | A | |
| EP1419612A1 | European Patent Office (EPO) | A1 | |
| CN1528069AThis record | China | A | |
| JP2004537918A | Japan | A | |
| US6944786B2 | United States of America | B2 | |
| KR100537583B1 | Republic of Korea | B1 | |
| CA2451101C | Canada | C | |
| JP4038176B2 | Japan | B2 | |
| EP1419612A4 | European Patent Office (EPO) | A4 | |
| CN100486159C | China | C | |
| EP1419612B1 | European Patent Office (EPO) | B1 | |
| AT445943T | Austria | T | |
| ATE445943T1 | Austria | T1 | |
| DE60234037D1 | Germany | D1 |
5 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Termination of patent right due to non-payment of annual feeCF01 | CF01 | |
| Transfer of patent application or patent right or utility modelC41 | C41 | |
| Grant of patent or utility modelGrantedC14 | C14 | |
| Entry into substantive examinationC10 | C10 | |
| PublicationC06 | C06 |
Numbers
- Publication
- 1528069
- Publication, DOCDB
- 1528069
- Publication, EPODOC
- CN1528069
- Application
- 28140427
- Application, DOCDB
- 02814042
- Application, EPODOC
- CN2002814042
Titles3
- Chinese
- 使用故障修复或多播地址的网络节点故障修复
- English
- Using fault repair or multicast address network node fault repair
- Chinese
- 使用故障修复或 多播地址的网络节点故障修复
Classification
- CPC, 4
- H04L12/185
- G06F11/20
- H04L12/1877
- H04L41/0663
- IPC, 5
- G06F13 00
- G06F11 20
- H04L12 18
- H04L12 24
- H04L12 56