Virtual network fault detection and location method
Abstract
The present invention relates to a method for detecting and locating virtual network faults, which includes the following steps: 1. Establish a virtual network fault management system in which the fault management center periodically sends status query requests to the physical nodes managed by it; 2: physical nodes Perform health checks on nodes and related link resources through its own detection mechanism, and send resource status information to the fault management center; 3: The fault management center determines whether an abnormality occurs in the virtual network based on the received information, and if an abnormality occurs, execute it 4. Otherwise, end; 4: The fault management center sends an abnormal query request to the associated nodes at both ends of the abnormal link in the switch; 5: The associated node sends an abnormal query response message to the fault management center according to the query content, and the fault management center confirms according to the message The exact location and type of the fault. The invention can automatically, quickly and accurately detect and locate the fault of the virtual network.
Term
No projected expiry on record.
- Priority and filed
- Published
- Today
6 claims: 1 independent, 5 dependent
- 1一种虚拟网故障探测和定位方法,其特征是:含有下列步骤: 步骤I :建立一个虚拟网故障管理系统,该虚拟网故障管理系统含有故障管理中心和 基础支撑环境中的物理节点,故障管理中心用于进行网络资源监控和故障发现与定位;基 础支撑环境中的物理节点用于为虚拟网构建提供基础资源;故障是指用于构建虚拟网的资 源发生异常,无法为虚拟网构建及其上运行的各类业务提供网络服务; 步骤2 :故障管理中心周期性向其管理的物理节点发送状态查询请求; 步骤3 :收到查询请求的物理节点通过自身的检测机制对节点和相关链路资源进行健 康检查; 步骤4:各物理节点通过故障管理接口向故障管理中心发送资源状态信息,根据查询 内容及健康检查的结果进行通告; 步骤5 :故障管理中心根据收到的资源状态信息判定虚拟网中是否发生异常,如发生 异常,则执行步骤6,否则,执行步骤9 ; 步骤6 :当虚拟网中发生异常时,物理节点会探测到某交换中的异常链路,故障管理中 心通过向该异常链路两端的关联节点发送异常查询请求消息来确认故障的准确位置和类 型;由于异常链路连接两个关联节点,因此至少需要向这两个关联节点都发送异常查询请 求消息;请求查询的内容含有:分享该异常链路资源的所有虚拟网中该异常链路的状态; 步骤7 :收到异常查询请求消息的关联节点根据查询内容向故障管理中心发送异常查 询应答消息; 步骤8 :故障管理中心根据异常查询应答消息确认故障的准确位置和类型; 步骤9 :结束。
- 2根据权利要求1所述的虚拟网故障探测和定位方法,其特征是:所述故障含有物理 节点故障、物理链路故障、虚拟节点故障和虚拟链路故障。
- 3根据权利要求1所述的虚拟网故障探测和定位方法,其特征是:所述故障管理中心 进行下列工作:监控各虚拟网中的资源运行状态;及时更新基础支撑环境中可用资源分布 情况;对虚拟网中发生的故障进行迅速精确定位;利用可用资源对故障进行修复处理。
- 4根据权利要求1所述的虚拟网故障探测和定位方法,其特征是:所述步骤2中,故障 管理中心向物理节点发送的状态查询内容含有:物理节点的资源总量、物理节点资源被分 配和映射到的虚拟网、各个虚拟网中分配的资源数量、剩余未分配的物理节点资源数量 、分配和映射给各个虚拟网的资源运行状态是否正常。
- 5根据权利要求1所述的虚拟网故障探测和定位方法,其特征是:所述步骤5中,判定 虚拟网中发生异常的方法如下: 方法1 :如果故障管理中心在限定时间内未收到针对某物理节点的状态查询请求的回 复信息,或物理节点的回复信息与故障管理中心预测的资源状态信息不符,则判定虚拟网 中发生异常; 方法2 :如果故障管理中心收到物理节点发出的未请求的异常状态通告消息,则判定 虚拟网中发生异常。
- 6根据权利要求5所述的虚拟网故障探测和定位方法,其特征是:所述方法2中,物理 节点在发现异常情况时主动向故障管理中心通告异常状态通告消息,异常情况含有以下类 型: 类型1 :与物理节点连接的物理链路故障:发生该物理链路故障时,分享该物理链路资 源的各虚拟网都会探测到同样的异常情况,但是由于探测到异常情况的时间有先后,因此, 异常状态通告消息中只通告其中一个虚拟网的链路异常; 类型2 :虚拟链路故障:由于上层软件漏洞造成某虚拟网中的虚拟链路故障,则只有该 虚拟网会探测到链路异常; 类型3 :物理节点故障:如果是物理节点故障,则分享该物理节点资源的虚拟网都会探 测到异常,但是由于探测时间的关系,所以故障管理中心收到的异常通告只是其中的一个 虚拟网异常; 类型4 :虚拟节点故障:如果是上层软件漏洞的问题引起某虚拟网中的虚拟节点故障, 则只有该虚拟网会通告异常。 7.根据权利要求1所述的虚拟网故障探测和定位方法,其特征是:所述步骤8中,确认 故障的准确位置时,有以下情况: 情况1 :对于两个关联节点中的任一个关联节点,如果指定时间内故障管理中心未收 到该关联节点的任何应答消息,则判断该关联节点发生物理故障,与该关联节点物理相邻 的节点都会探测到与该关联节点相连的链路故障,向该关联节点的所有物理邻接节点都发 送异常查询请求消息,从而更准确地定位故障; 情况2:如果两个关联节点都有应答消息,且其中第一个关联节点的应答消息中通告 了处于该虚拟网中的所述异常链路,而第二个关联节点由于配置故障已经释放了分配给该 虚拟网中相应的资源,因此,第二个关联节点的应答消息中没有所述异常链路在虚拟网中 的资源状态,该资源状态在第一个关联节点的异常通告消息中通告过,这说明第二个关联 节点映射到该虚拟网的虚拟节点发生故障; 情况3 :如果在两个关联节点的应答消息中,分享所述异常链路资源的所有虚拟网都 通告该异常链路异常,则说明:为该异常链路提供基础资源的物理链路故障,该物理链路资 源都变为不可用状态; 情况4:如果在两个关联节点的应答消息中,都只是通告某虚拟网中的所述异常链路 异常,则说明:只是该异常链路映射到该虚拟网中的虚拟链路故障。
Independent claims6
29 paragraphs, as filed
Virtual network fault detection and location method
[0001] (1) Technical Field: The present invention relates to a method for detecting and locating network faults, and in particular, to a method for detecting and locating virtual network faults.
[0002] (2) Background technology: network failures affect the normal operation of the network system, and the causes of network failures are complicated and inevitable, such as configuration errors, fiber breaks, unstable switching equipment, malicious attacks, misoperations, and accidental disconnections. Electricity etc. Virtual network technology is a new network technology, and it is also impossible to avoid various network failures. Moreover, due to the complexity of virtual network technology, more complex network failures may be introduced. Therefore, in order to make the virtual network system run stably, it is necessary to quickly detect and accurately locate the virtual network fault, so as to provide support for fault repair.
[0003] At present, the detection of network failures is mainly accomplished through upper-layer routing protocols. If unreachable between routers is detected, a rerouting mechanism is used to avoid the point of failure. The fault location mainly relies on manual methods and largely relies on the experience of the network administrator. Therefore, how to quickly locate the network fault point has become an important indicator for evaluating the ability of the network administrator. For network management, it is not an easy task to locate the fault point, mainly including ping the target address, checking the router indicator on the spot, and identifying the appearance of the routing device. The above methods have higher requirements for technical personnel, and in extreme cases, they are likely to cause a large area of network paralysis and cause serious consequences. Moreover, for virtual networks, in addition to the types of failures that may occur in traditional networks, there may also be virtual failures generated in the process of virtual network construction. If the original traditional network troubleshooting mechanism is still used, it will be difficult for network administrators. Putting forward higher requirements may also have a more serious impact on the network. This urgently needs to propose a brand-new troubleshooting method based on the characteristics and technical conditions of the virtual network.
[0004] (3). Summary of the invention: The technical problem to be solved by the present invention is to provide a method for detecting and locating virtual network faults, which can automatically, quickly and accurately perform fault detection and locating on the virtual network, thereby providing Provide support for fault repair.
[0005] The technical scheme of the present invention: A method for detecting and locating virtual network faults, including the following steps: Step 1: Establish a virtual network fault management system, the virtual network fault management system contains a fault management center (Fault Management Center, FMC ) And physical nodes in the basic support environment , the fault management center is used for network resource monitoring and fault discovery and location; the physical nodes in the basic support environment are used to provide basic resources for virtual network construction; faults are used to build virtual networks The resources of the network are abnormal and cannot provide network services for the construction of the virtual network and various services running on it; Step 2: The fault management center periodically sends status query requests to the physical nodes it manages; Step 3: The physical node that receives the query request Perform health checks on nodes and related link resources through its own detection mechanism; Step 4: Each physical node sends resource status information to the fault management center through the fault management interface, and announces according to the query content and the results of the health check; Step 5: Fault The management center determines whether an abnormality occurs in the virtual network according to the received resource status information. If an abnormality occurs, perform step 6, otherwise, perform step 9; due to the existence of the above-mentioned multiple failure possibilities, for the failure management center, if Receiving an abnormal notification message from a physical node cannot infer the exact location and type of the failure point. Therefore, in order to locate the failure point and the type of failure,
The fault management center must initiate the fault location process; Step 6: When an abnormality occurs in the virtual network, the physical node will detect an abnormal link in the RSCN of a certain switch, and the fault management center sends an abnormal query to the associated nodes at both ends of the abnormal link Request a message to confirm the exact location and type of the fault; because the abnormal link connects two associated nodes, it is necessary to send an abnormal query request message to both associated nodes at least; the content of the requested query includes: sharing the abnormal link resource The status of the abnormal link in all virtual networks; Step 7: The associated node that receives the abnormal query request message sends an abnormal query response message to the fault management center according to the query content; Step 8: The fault management center confirms the fault according to the abnormal query response message Accurate location and type; Step 9: End.
[0006] Failures include physical node failures, physical link failures, virtual node failures, and virtual link failures.
[0007] The fault management center performs the following tasks: monitors the operating status of resources in each virtual network; timely updates the distribution of available resources in the basic support environment; quickly and accurately locates faults in the virtual network; uses available resources to repair faults deal with.
[0008] Since the link cannot exist alone, it is always associated with the connected node. Therefore, the physical link connecting the physical node is also described as a node resource.
[0009] In step 2, the content of the status query sent by the fault management center to the physical node includes: the total amount of resources of the physical node, the virtual network to which the physical node resources are allocated and mapped, the number of resources allocated in each virtual network, and the remaining outstanding resources. Whether the number of allocated physical node resources and the operating status of the resources allocated and mapped to each virtual network are normal.
[0010] In step 5, the method for determining an abnormality in the virtual network is as follows: Method 1: If the fault management center does not receive the response information for the status query request of a certain physical node within a limited time, or the response information of the physical node and If the resource status information predicted by the fault management center does not match, it is determined that an abnormality has occurred in the virtual network; Method 2: If the fault management center receives an unrequested abnormal status notification message from the physical node, it is determined that an exception has occurred in the virtual network.
[0011] In method 2, the physical node proactively announces the abnormal status notification message to the fault management center when an abnormal situation is found, and the abnormal situation includes the following types: Type 1: physical link failure connected to the physical node: the physical link failure occurs At this time, each virtual network sharing the physical link resource will detect the same abnormal situation, but because the abnormal situation is detected in a sequential order, the abnormal state notification message only announces that the link of one of the virtual networks is abnormal; 2: Virtual link failure: a virtual link failure in a virtual network caused by a bug in the upper layer software, only the virtual network will detect the link abnormality; Type 3: physical node failure: if a physical node advertises the link Anomaly, it may also be caused by the failure of another physical node connected to the link to respond to the link failure; if it is a physical node failure, the virtual network sharing the resources of the physical node will detect the abnormality, but due to detection Because of the relationship of time, the exception notification received by the fault management center is only one of the virtual network exceptions; Type 4: Virtual node failure: If a virtual node failure in a virtual network is caused by a bug in the upper-layer software, only the virtual node The network will notify the exception.
[0012] In step 8, when confirming the exact location of the fault, there are the following situations:
Case 1: For any one of the two associated nodes, if the fault management center does not receive any response message from the associated node within a specified time, it is judged that the associated node has a physical failure, and all resources of the associated node are changed. In the unavailable state, the nodes physically adjacent to the associated node will detect the link failure to the associated node. Therefore, the associated range is expanded, and abnormal query request messages are sent to all physical adjacent nodes of the associated node, thereby Locate the fault more accurately; Case 2: If both associated nodes have response messages, and the response message of the first associated node announces the abnormal link in the virtual network, and the second associated node Because the configuration failure has released the corresponding resources allocated to the virtual network, there is no resource status of the abnormal link in the virtual network in the response message of the second associated node, and the resource status is in the first associated node Announced in the abnormal notification message, which indicates that the virtual node mapped to the virtual network by the second associated node is faulty; Case 3: If in the response message of the two associated nodes, all virtual nodes that share the abnormal link resource If the network announces that the abnormal link is abnormal, it means: the physical link that provides the basic resources for the abnormal link fails, and the physical link resources become unavailable; Case 4: If the response message of the two associated nodes In, they are just notifying the abnormal link in a certain virtual network Abnormal, it means: only the abnormal link is mapped to the virtual link failure in the virtual network.
[0013] The beneficial effects of the present invention:
1. The present invention can not only detect the faults in the network in time, but also actively notify them instead of passively waiting for the network management to discover them, thereby improving the efficiency of fault handling.
[0014] 2. After detecting possible network failures, the present invention can perform the next step of automatic fault location instead of simply relying on the experience of network administrators to locate the fault point, which can effectively improve the accuracy and speed of fault location. One-step troubleshooting provides strong support.
[0015] (4) Description of the drawings: Figure 1 is a schematic structural diagram of a virtual network fault management system; Figure 2 is a schematic diagram of an information interaction process between a fault management center and a physical node; Figure 3 is a schematic diagram of a fault management center and a physical node A schematic diagram of the content of the interactive messages between the two; Figure 4 is a schematic flowchart of the fault location process.
[0016] (5) Specific implementation mode: The virtual network fault detection and location method includes the following steps: Step 1: Establish a virtual network fault management system (as shown in Figure 1), the virtual network fault management system contains a fault management center (Fault Management Center, FMC) and physical nodes in the basic support environment. The fault management center is used for network resource monitoring and fault detection and location; the physical nodes in the basic support environment are used to provide basic resources for virtual network construction; the fault is Refers to the abnormality of the resources used to construct the virtual network, which cannot provide network services for the construction of the virtual network and the various services running on it; Step 2: The fault management center periodically sends status query requests to the physical nodes managed by it (Figure 2) Figure 3); Step 3: The physical node receiving the query request performs a health check on the node and related link resources through its own detection mechanism; the physical node first checks its own resource status, and then reports to the fault management center based on the current status Send regular query response messages; the current status includes: the total number of resources of the node, which virtual networks have participated in the construction, how many resources have been allocated to the virtual distribution networks participating in the construction, and whether the resources allocated to each virtual network are operating normally;
Step 4: Each physical node sends resource status information to the fault management center through the fault management interface, and announces according to the query content and the results of the health check; Step 5: The fault management center determines whether an abnormality occurs in the virtual network based on the received resource status information , If an exception occurs, proceed to step 6, otherwise, proceed to step 9; due to the existence of the above-mentioned multiple failure possibilities, for the fault management center, if an exception notification message is received from a physical node, the failure point cannot be inferred Therefore, in order to locate the fault point and fault type, the fault management center must initiate the fault location process (as shown in Figure 4); Step 6: When an abnormality occurs in the virtual network, the physical node will detect the RSCN of a certain switch The fault management center confirms the exact location and type of the fault by sending an exception query request message to the associated nodes at both ends of the abnormal link; because the abnormal link connects two associated nodes, it needs to send at least these two All associated nodes send an abnormal query request message; the content of the request for query includes: the status of the abnormal link in all virtual networks that share the abnormal link resource; Step 7: the associated node that receives the abnormal query request message reports the failure according to the query content The management center sends an abnormal query response message; Step 8: The fault management center confirms the exact location and type of the fault according to the abnormal query response message; Step 9: End.
[0017] Failures include physical node failures, physical link failures, virtual node failures, and virtual link failures.
[0018] The fault management center performs the following tasks: monitors the operating status of resources in each virtual network; timely updates the distribution of available resources in the basic support environment; quickly and accurately locates faults that occur in the virtual network; uses available resources to repair faults deal with.
[0019] Since the link cannot exist alone, it is always associated with the connected node. Therefore, the physical link connecting the physical node is also described as a node resource.
[0020] In step 2, the content of the status query sent by the fault management center to the physical node includes: the total amount of resources of the physical node, the virtual network to which the physical node resources are allocated and mapped, the number of resources allocated in each virtual network, and the remaining outstanding resources. Whether the number of allocated physical node resources and the operating status of the resources allocated and mapped to each virtual network are normal. If a virtual network is constructed or cancelled during the query cycle, and the queried node participates in the construction or withdrawal of the virtual network, that is, the node allocates or releases resources during this cycle, the fault management center can use the Query the message to obtain this information.
[0021] In step 5, the method for determining that an abnormality has occurred in the virtual network is as follows: Method 1: If the fault management center does not receive the response information for the status query request of a certain physical node within a limited time, or the response information of the physical node is If the resource status information predicted by the fault management center does not match, it is determined that an abnormality has occurred in the virtual network; Method 2: If the fault management center receives an unrequested abnormal status notification message from the physical node, it is determined that an exception has occurred in the virtual network.
[0022] In method 2, the physical node proactively announces the abnormal status notification message to the fault management center when an abnormal condition is found, and the abnormal condition includes the following types: Type 1: physical link failure connected to the physical node: the physical link failure occurs When the physical link resource is shared, the virtual networks that share the physical link resource will detect the same abnormal situation, but because the time of detecting the abnormal situation is sequential, the abnormal state notification message only announces that the link of one of the virtual networks is abnormal;
Type 2: Virtual link failure: A virtual link failure in a virtual network is caused by a bug in the upper layer software, and only the virtual network will detect the link abnormality; Type 3: Physical node failure: If a physical node advertises the link The path abnormality may also be caused by the failure of another physical node connected to the link. If the physical node fails, the virtual network sharing the resources of the physical node will detect the abnormality. The relationship between detection time, so the abnormal notification received by the fault management center is only one of the virtual network abnormalities; Type 4: Virtual node failure: If the problem of the upper-layer software vulnerability bug causes the virtual node failure in a virtual network, only this The virtual network will notify the exception.
[0023] In step 8, when confirming the exact location of the fault, there are the following situations: Case 1: For any one of the two associated nodes, if the fault management center does not receive any response message from the associated node within a specified time , It is judged that the associated node has a physical failure, all the resources of the associated node become unavailable, and the nodes physically adjacent to the associated node will detect the failure of the link connected to the associated node. Therefore, expand the scope of association , To send abnormal query request messages to all physical adjacent nodes of the associated node, so as to locate the fault more accurately; Case 2: If two associated nodes have response messages, and the response message of the first associated node announces The abnormal link in the virtual network, and the second associated node has released the corresponding resources allocated to the virtual network due to a configuration failure, therefore, the second associated nodes response message does not contain the abnormal chain The resource status of the path in the virtual network. The resource status has been notified in the abnormal notification message of the first associated node, which means that the virtual node mapped to the virtual network by the second associated node has failed; Case 3: In the response message of two associated nodes, all virtual networks sharing the abnormal link resource announce that the abnormal link is abnormal, indicating that the physical link that provides the basic resources for the abnormal link is faulty, and the physical link resources are all Becomes unavailable; Case 4: If both the response messages of the two associated nodes only notify the abnormal link in a certain virtual network that the abnormal link is abnormal, it means that only the abnormal link is mapped to the virtual link failure in the virtual network.
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US10771323B2 | Cited by | United States of America | Applicant |
| CN106664214A | Cited by | China | Search report |
| CN105933176A | Cited by | China | Search report |
| WO2016127482A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| CN106170947A | Cited by | China | Search report |
| US10892943B2 | Cited by | United States of America | Applicant |
| CN106130761A | Cited by | China | Search report |
| CN101710869A | Cites | China | Search report |
| CN101917460A | Cites | China | Search report |
| US2010165876A1 | Cites | United States of America | Search report |
2 priority claims, no other members on record
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 201410311441 | China | A | |
| CN20141311441 | – | – | – |
3 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Patent grantGrantedGR01 | GR01 | |
| Entry into substantive examinationC10 | C10 | |
| PublicationC06 | C06 |
Numbers
- Publication
- 104243232
- Publication, DOCDB
- 104243232
- Publication, EPODOC
- CN104243232
- Application
- 103114413
- Application, DOCDB
- 201410311441
- Application, EPODOC
- CN20141311441
Titles2
- Chinese
- 虚拟网故障探测和定位方法
- English
- Virtual network fault detection and location method
Classification
- IPC, 2
- H04L12 26
- H04L12 24