Distributed fault detection for data storage networks
Claim Score by NHIP
Abstract
A distributed fault detection system and method for diagnosing a storage network fault in a data storage network having plural network access nodes connected to plural logical storage units. When a fault is detected, the node that detects it (designated the primary detecting node) issues a fault information broadcast advising one or more other access nodes (peer nodes) of the fault. The primary detecting node also sends a fault report pertaining to the fault to a fault diagnosis node. When the peer nodes receive the fault information broadcast, they attempt to recreate the fault. Each peer node that successfully recreates the fault (designated a secondary detecting node) sends its own fault report pertaining to said fault to the fault diagnosis node. The fault diagnosis node performs fault diagnosis based on all of the fault reports.

Term
Term ended
Projected expiry passed 11 October 2023, 3 years ago.
- Priority and filed
- Published
- Projected expiry
- Today
29 claims: 4 independent, 25 dependent
- 1In a data storage network having plural network access nodes connected to plural logical storage units, a distributed fault detection method for diagnosing a storage network fault, comprising:broadcasting fault information pertaining to a fault (fault information broadcast) from one of said access nodes that detects said fault (primary detecting node) to at least one other of said access nodes that are peers of said primary detecting node (peer nodes);attempting to recreate said fault at said peer nodes;providing fault reports pertaining to said fault to a fault diagnosis node from said primary detecting node and any of said peer nodes that are able to recreate said fault (secondary detecting nodes);and said fault diagnosis node performing fault diagnosis based on said fault reports.
- 11In a system adapted for use as a network access node of a data storage network having plural network access nodes connected to plural logical storage units, a fault detection system enabling said access node to participate in distributed diagnosis of a storage network fault, comprising:means for detecting a fault P 1 when said access node acts as a primary detecting node;means for broadcasting fault information pertaining to said fault P 1 (first fault information broadcast) to one or more other access nodes that are peers of said access node;means for receiving a second fault information broadcast pertaining to a fault P 2 detected at one of said other access nodes when said access node acts as a secondary detecting node;means responsive to receiving said second fault information broadcast for attempting to recreate said fault P 2 ;and means for providing fault reports pertaining to said faults P 1 and P 2 to a fault diagnosis node.
- 20Broadest claimClaim Score 65, broad(NHIP)In a system adapted for use as a fault diagnostic node that communicates with plural network access nodes of a data storage network in which said access nodes are connected to plural logical storage units, a fault diagnosis system for performing distributed diagnosis of a storage network fault, comprising:means for receiving fault reports from one or more of said access nodes;means for evaluating fault information contained in said fault reports;and means for generating a distributed diagnosis of said storage network fault based on said fault information evaluation.
- 21A computer program product for use in a data storage network having plural network access nodes communicating with plural data storage logical units, comprising:one or more data storage media;means recorded on said data storage media for controlling one of said access nodes to detect a fault P 1 when said access node acts as a primary detecting node;means recorded on said data storage media for controlling one of said access nodes to broadcast fault information pertaining to said fault P 1 (first fault information broadcast) to at least one other of said access nodes that are peers of said access node;means recorded on said data storage media for controlling one of said access nodes to receive a second fault information broadcast pertaining to a fault P 2 detected at another of said access nodes when said access node acts as a secondary detecting node;means recorded on said data storage media for controlling one of said access nodes to respond to receipt of said second fault information broadcast by attempting to recreate said fault P 2 ;and means recorded on said data storage media for controlling one of said access nodes to send fault reports pertaining to said faults P 1 and P 2 to a fault diagnosis node.
Independent claims4
56 paragraphs in 4 sections, as filed
BACKGROUND OF THE INVENTION
P-0001[0001] 1. Field of the Invention
P-0002[0002] The present invention relates to data storage networks, and especially switched networks implementing SAN (Storage Area Network) functionality, NAS (Network Attached Storage) functionality, or the like. More particularly, the invention concerns a distributed fault detection system and method for diagnosing storage network faults.
P-0003[0003] 2. Description of the Prior Art
P-0004[0004] By way of background, data storage networks, such as SAN systems, provide an environment in which data storage peripherals are managed within a high speed network that is dedicated to data storage. Access to the storage network is provided through one or more access nodes that typically (but not necessarily) function as file or application servers (e.g., SAN application servers, NAS file server gateways, etc.) on a conventional LAN (Local Area Network) or WAN (Wide Area Network). Within the data storage network, the access nodes generally have access to all devices within the pool of peripheral storage, which may include any number of magnetic disk drive arrays, optical disk drive arrays, magnetic tape libraries, etc. In all but the smallest storage networks, the required connectivity is provided by way of arbitrated loop arrangements or switching fabrics, with the latter being more common.
P-0005[0005]FIG. 1 is illustrative. It shows a typical data storage network <b>2</b> in which a plurality of access nodes <b>4</b><i>a</i>-<b>4</b><i>e, </i>all peers of each other, are connected via switches <b>6</b><i>a</i>-<b>6</b><i>b </i>and controllers <b>8</b><i>a</i>-<b>8</b><i>c </i>to plural LUNs <b>10</b><i>a</i>-<b>10</b><i>f, </i>representing virtualized physical storage resources. Note that each access node <b>4</b><i>a</i>-<b>4</b><i>e </i>has several pathways to each LUN <b>10</b><i>a</i>-<b>10</b><i>f, </i>and that portions of each pathway are shared with other access nodes.
P-0006[0006] One of the problems with this kind of topology is that network faults due to controller failures, LUN failures, link failures and other problems are often difficult to isolate. Often, an access node will detect a problem, but will not know whether that problem is isolated to itself or has a larger scope. With this incomplete knowledge, it is difficult for the storage network to react optimally.
P-0007[0007] Each access node <b>4</b><i>a</i>-<b>4</b><i>e </i>can only evaluate pathways between itself and the switches, controllers and LUNs to which it is connected. This permits the detection of limited information, which may or may not allow complete fault isolation. Examples of the kind of path information that can be determined by a single access node include the following:
P-0008[0008] 1. A path from the node to a LUN is down;
P-0009[0009] 2. All paths from the node to a LUN through a controller (path group) are down;
P-0010[0010] 3. All paths from the node to a LUN through any controller are down;
P-0011[0011] 4. All paths from the node to all LUNs through a controller are down; and
P-0012[0012] 5. All paths from the node to all LUNs through all controllers are down.
P-0013[0013] Depending on which of the foregoing conditions is satisfied, a single access node can at least partially isolate the source of a storage network fault. However, to complete the diagnosis, the following information, requiring distributed fault detection, is needed:
P-0014[0014] 1. Whether a path to a LUN from all nodes is down;
P-0015[0015] 2. Whether all paths from all nodes to a LUN through a controller (path group) are down;
P-0016[0016] 3. Whether all paths from all nodes to a LUN through any controller are down;
P-0017[0017] 4. Whether all paths from all nodes to all LUNs through a controller are down; and
P-0018[0018] 5. Whether all paths from all nodes to all LUNs through all controllers are down.
P-0019[0019] If such additional information can be determined, the following likely diagnoses can be made:
P-0020[0020] 1. A switch's connection to a controller is likely defective;
P-0021[0021] 2. A controller's connection to a LUN is likely defective;
P-0022[0022] 3. A LUN is likely defective;
P-0023[0023] 4. A controller is likely defective; and
P-0024[0024] 5. A total system failure has occurred.
P-0025[0025] One solution to the problem of isolating faults in a data storage network is proposed in commonly assigned European Patent Application No. EP 1115225A2 (published Nov. 7, 2001). This application discloses a method and system for end-to-end problem determination and fault isolation for storage area networks in which a communications architecture manager (CAM) uses a SAN topology map and a SAN PD (Problem Determination) information table (SPDIT) to create a SAN diagnostic table (SDT). A failing component in a particular device may generate errors that cause devices along the same network connection path to generate errors. As the CAM receives error packets or error messages, the errors are stored in the SDT, and each error is analyzed by temporally and spatially comparing the error with other errors in the SDT. This allows the CAM to identify the faulty component.
P-0026[0026] It is to storage network fault analysis systems and methods of the foregoing type that the present invention is directed. In particular, the invention provides an alternative fault detection system and method in which access nodes act as distributed fault detection agents that assist in isolating storage network faults.
SUMMARY OF THE INVENTION
P-0027[0027] The foregoing problems are solved and an advance in the art is obtained by a distributed fault detection system and method for diagnosing a storage network fault in a data storage network having plural network access nodes connected to plural logical storage units. When a fault is detected, the node that detects it (designated the primary detecting node) issues a fault information broadcast advising the other access nodes (peer nodes) of the fault. The primary detecting node also sends a fault report pertaining to the fault to a fault diagnosis node. When the peer nodes receive the fault information broadcast, they attempt to recreate the fault. Each peer node that successfully recreates the fault (designated a secondary detecting node) sends its own fault report pertaining to the fault to the fault diagnosis node. Peer nodes that cannot recreate the fault may also report to the fault diagnosis node. The fault diagnosis node can then perform fault diagnosis based on reports from all of the nodes.
P-0028[0028] The secondary detecting nodes also preferably issue their own fault information broadcasts pertaining to the fault if they are able to recreate the fault. This redundancy, which can result in multiple broadcasts of the same fault information, helps ensure that all nodes will receive notification of the fault. In order to prevent multiple fault recreation attempts by the nodes as they receive the fault information broadcasts, the primary detecting node preferably assigns a unique identifier to the fault and includes the unique identifier in the initial fault information broadcast. By storing a record of the fault with the unique identifier, the peer nodes can test when they receive subsequent fault information broadcasts whether they have previously seen the fault. This will prevent repeated attempts to recreate the same fault event because a node will ignore the fault information broadcast if it has previously seen the fault event.
P-0029[0029] The fault reports sent to the fault diagnosis node preferably includes localized fault diagnosis information determined by the primary and secondary detecting nodes as a result of performing one or more diagnostic operations to ascertain fault information about a cause of the fault. The same localized fault information could also be provided between nodes as part of the fault information broadcasts. This sharing of information between nodes would be useful to each node as it performs its own localized diagnosis of the fault. The fault diagnosis performed by the fault diagnosis node may include determining one or more of a switch-controller connection being defective, a controller-storage device connection being defective, a storage device being defective, or a total system failure.
P-0030[0030] The invention further contemplates a computer program product that allows individual instances of the above-described fault detection functionality to execute on one or more access nodes of a data storage network.
BRIEF DESCRIPTION OF THE DRAWINGS
P-0031[0031] The foregoing and other features and advantages of the invention will be apparent from the following more particular description of preferred embodiments of the invention, as illustrated in the accompanying Drawings, in which:
P-0032[0032]FIG. 1 is a functional block diagram showing a prior art data storage network;
P-0033[0033]FIG. 2 is a functional block diagram showing a data storage network having access nodes and a fault diagnosis node constructed according to the principals of the invention;
P-0034[0034]FIG. 3 is another functional block diagram of the data storage network of FIG. 2;
P-0035[0035]FIG. 4 is a diagrammatic illustration of a computer program product medium storing computer program information that facilitates distributed fault detection in accordance with the invention;
P-0036[0036]FIGS. 5A and 5B constitute a two-part flow diagram showing a fault detection method implemented in the data storage network of FIGS. 2 and 3;
P-0037[0037]FIG. 6 is a block diagram showing functional elements of fault processing software for controlling an access node of the data storage network of FIGS. 2 and 3; and
P-0038[0038]FIG. 7 is a block diagram showing functional elements of fault diagnosis software for controlling a fault diagnosis node of the data storage network of FIGS. 2 and 3.
DETAILED DESCRIPTION OF PREFERRED EMBODIMENTS
P-0039[0039] Turning now to the figures, wherein like reference numerals represent like elements in all of the several views, FIG. 2 illustrates a data storage network <b>20</b> that is adapted to perform fault detection in accordance with the invention. The storage network <b>20</b> includes a plurality of access nodes, five of which are shown at <b>22</b>, <b>24</b>, <b>26</b>, <b>28</b> and <b>30</b>. The storage network <b>20</b> further includes a node that performs fault diagnosis functions. Although one of the access nodes <b>22</b>-<b>30</b> could be used for this purpose, a separate fault diagnosis node <b>38</b> is preferably provided. Each access node <b>22</b>-<b>30</b> of the storage network <b>20</b> is assumed to connect via a switching fabric <b>32</b> to a plurality of controllers <b>34</b>. The controllers <b>34</b> are in turn connected to, and control the flow of data to and from, a plurality of LUNs <b>36</b>.
P-0040[0040]FIG. 3 illustrates an exemplary topology that can be used to configure the storage network <b>20</b> shown in FIG. 2. In this configuration, the switching fabric <b>32</b>, the controllers <b>34</b> and the LUNs <b>36</b> are arranged to match the topology of FIG. 1 for convenience and ease of description. In the topology of FIG. 3, the switching fabric <b>32</b> includes a pair of switching elements <b>32</b><i>a </i>and <b>32</b><i>b. </i>The controllers <b>34</b> include three controllers <b>34</b><i>a, </i><b>34</b><i>b </i>and <b>34</b><i>c. </i>The LUNs <b>36</b> include six LUNs <b>36</b><i>a, </i><b>36</b><i>b, </i><b>36</b><i>c, </i><b>36</b><i>d, </i><b>36</b><i>e </i>and <b>36</b><i>f. </i>Each of the access nodes <b>22</b>-<b>30</b> in FIG. 3 connects to the two switching elements <b>32</b><i>a </i>and <b>32</b><i>b. </i>Each switching element <b>32</b><i>a </i>and <b>32</b><i>b </i>connects to at least two of the controllers <b>34</b><i>a, </i><b>34</b><i>b </i>and <b>34</b><i>c. </i>Each controller <b>34</b><i>a, </i><b>34</b><i>b </i>and <b>34</b><i>c </i>connects to two or more of the LUNs <b>36</b><i>a</i>-<b>36</b><i>f. </i>
P-0041[0041] There are a variety of system components that can be used to implement the various elements that make up the storage network <b>20</b>, depending on design preferences. Underlying the network design will be the selection of a suitable communication and media technology, such as Fibre Channel, Ethernet or SCSI. Selection of one of these core technologies will dictate the choice of devices that will be used to implement the switching elements <b>32</b><i>a </i>and <b>32</b><i>b, </i>as well as the network interfaces that reside in the access nodes <b>22</b>-<b>30</b> and the controllers <b>34</b><i>a</i>-<b>34</b><i>c. </i>Selection of the controllers <b>34</b><i>a</i>-<b>34</b><i>c </i>will be dictated by the physical storage devices that provide the LUNs <b>36</b><i>a</i>-<b>36</b><i>f. </i>The latter could be implemented as RAID arrays, JBOD arrays, intelligent disk subsystems, tape libraries, etc., or any combination thereof. As persons skilled in the art will appreciate, each LUN represents a logical virtualization of some defined subset of the total pool of available physical storage space in the storage network <b>20</b>. For example, a LUN could represent a set of disks in a RAID or JBOD array, or it could represent the entire array.
P-0042[0042] The access nodes <b>22</b>-<b>30</b> can be configured according to the data storage services they are intended to provide. Typically, they will function as file or application servers on a LAN or WAN. For example, one or more of the access nodes could be implemented as SAN application servers offering block I/O interfaces to client devices. Examples of such application servers include database servers, web servers, to name but a few. Other access nodes could be implemented as NAS file server gateways offering file I/O interfaces based on a network file system service such as NFS (Network File System), CIFS (Common Internet File System), or the like. Still other access nodes could be implemented as hosts that connect to both the data storage network <b>20</b> and a LAN or WAN, and which are programmed with client software such as the SANergy™ software product from International Business Machines Corporation (“IBM”) and Tivloi Systems, Inc. The SANergy™ product allows such hosts to handle NFS or CIFS file requests from other LAN hosts.
P-0043[0043] Regardless of the foregoing implementation choices, each access node <b>22</b>-<b>30</b> will preferably be built from a conventional programmable computer platform that is configured with the hardware and software resources needed to implement the required data storage network functions. Exemplary computer platforms include mainframe computers such as an IBM S/390 system running IBM's OS/390 operating system, mid-range computers such as an IBM AS/400 system running IBM's OS/400 operating system, workstation computers such as an IBM RISC/System 6000 system running IBM's Advanced Interactive Executive (AIX) operating system, or any number of microprocessor-based personal computers running a Unix-based operating system or operating system kernel, or a Unix-like operating system or operating system kernel, such as Linux, FreeBSD, etc.
P-0044[0044] Each access node <b>22</b>-<b>30</b> further includes an appropriate network interface, such as a Fibre Channel Host Bus Adapter (HBA), that allows it to communicate over the data storage network <b>20</b>. An additional network interface, such as an Ethernet card, will typically be present in each access node <b>22</b>-<b>30</b> to allow data communication with other hosts on a LAN or WAN. This data communication pathway is shown by the double-headed arrows <b>40</b> extending from each of the access nodes <b>22</b>-<b>30</b> in FIG. 2. It will also be seen in FIG. 2 that the access nodes <b>22</b>-<b>30</b> maintain communication with each other, and with the fault diagnosis node <b>38</b>, over an inter-node communication pathway shown by the dashed line <b>42</b>. Any suitable signaling or messaging communication protocol, such as SNMP (Simple Network Message Protocol), can be used to exchange information over the inter-node communication pathway <b>42</b>. Various physical resources may be used to carry this information, including the LAN or WAN to which the access nodes are connected via the communication pathways <b>40</b>, one or more pathways that are dedicated to inter-node communication, or otherwise.
P-0045[0045] As additionally shown in FIG. 2, each access node <b>22</b>-<b>30</b> is programmed with fault processing software <b>44</b> that allows the access node to detect and process faults on network pathways to which the access node is connected. The fault diagnosis node <b>38</b> is programmed with fault diagnosis software <b>46</b> that allows it to perform distributed fault diagnosis by evaluating fault processing information developed by each access node's fault processing software <b>44</b>. The functionality provided by the software <b>44</b> and <b>46</b> is described below. Note that one aspect of the present invention contemplates a computer program product in which the foregoing software is stored in object or source code form on a data storage medium, such as one or more portable (or non-portable) magnetic or optical disks. FIG. 4 illustrates an exemplary computer program product <b>50</b> in which the storage medium is an optical disk. The computer program product <b>50</b> can be used by storage network administrators to add distributed fault detection functionality in accordance with the invention to conventional data storage networks. In a typical scenario, one copy of the fault processing software <b>44</b> will be installed onto each of the access nodes <b>22</b>-<b>30</b>, and one copy of the fault diagnosis software <b>46</b> will be installed onto the fault diagnosis <b>38</b>. The software installation can be performed in a variety of ways, including local installation at each host machine, or over a network via a remote installation procedure from a suitable server host that maintains a copy of the computer program product <b>50</b> on a storage medium associated with the server host.
P-0046[0046] Persons skilled in the art will appreciate that data storage networks conventionally utilize management software that provides tools for managing the storage devices of the network. An example of such software is the Tivoli® Storage Network Manager product from IBM and Tivoli Systems, Inc. This software product includes agent software that is installed on the access nodes of a data storage network (such as the access nodes <b>22</b>-<b>30</b>). The product further includes management software that is installed on a managing node of a data storage network (such as the fault diagnosis node <b>38</b>) that communicates with the access nodes as well as other devices within and connected to the network switching fabric. Using a combination of inband and out-of-band SNMP communication, the management software performs functions such as discovering and managing data network topology, assigning/unnassigning network storage resources to network access nodes, and monitoring and extending file systems on the access nodes. It will be appreciated that the fault processing software <b>44</b> and the fault diagnosis software <b>46</b> of the present invention could be incorporated into a product such as the Tivoli® Storage Network Manager as a functional addition thereto.
P-0047[0047] Turning now to FIGS. 5A and 5B, an exemplary distributed fault detection process according to the invention will now be described from a network-wide perspective. Thereafter, a description of the processing performed at an individual one of the access nodes <b>22</b>-<b>30</b> will be set forth with reference to FIG. 6, and a description of the processing performed at the fault diagnosis node <b>38</b> will be set forth with reference to FIG. 7.
P-0048[0048] In discussing the flow diagrams of FIGS. 5A and 5B, it will be assumed that the access node <b>22</b> cannot communicate with the LUN <b>36</b><i>b </i>(LUN <b>1</b>), but can communicate with the LUN <b>36</b><i>a</i>(LUN <b>0</b>) and the LUN <b>36</b><i>c </i>(LUN <b>2</b>). Assume further that the cause of the problem is a defective link <b>60</b> between the switching element <b>32</b><i>a </i>and the controller <b>34</b><i>a, </i>as shown in FIG. 3. It will be seen that the node <b>22</b> cannot determine the source of the problem by itself For all it knows, the controller <b>34</b><i>a </i>could be defective, LUN <b>1</b> could be defective, or a connection problem could exist on any of the links between the switching element <b>32</b><i>a </i>and LUN <b>1</b>. Note that the access node <b>22</b> does know that the path between itself and the switching element <b>32</b><i>a </i>is not causing the problem insofar as it is able to reach LUN <b>0</b> and LUN <b>2</b> through the controller <b>34</b><i>b. </i>
P-0049[0049] As a result of being unable to communicate with LUN <b>1</b>, the access node <b>22</b> will detect this condition as a fault event in step <b>100</b> of FIG. 5A. The access node <b>22</b> will thereby become a primary fault detecting node. After the fault event is detected, the primary detecting node <b>22</b> assigns it a unique identifier, let us say P<b>1</b> (meaning “Problem 1”), in step <b>102</b>. The primary detecting node <b>22</b> then logs the fault P<b>1</b> in step <b>104</b> by storing a record of the event in a suitable memory or storage resource, such as a memory space or a disk drive that is local to the primary detecting node. Remote logging, e.g., at the fault diagnosis node <b>38</b>, would also be possible. In step <b>106</b>, the primary detecting node <b>22</b> issues a fault information broadcast advising the other access nodes <b>24</b>-<b>30</b> (peer nodes) of the fault P<b>1</b>. The primary detecting node <b>22</b> also provides a fault report pertaining to the fault P<b>1</b> to the fault diagnosis node <b>38</b>. This is shown in step <b>108</b>. Reporting a fault can be done in several ways. In one scenario, if the primary detecting node <b>22</b> logs the fault P<b>1</b> remotely at the fault diagnosis node <b>38</b> in step <b>104</b>, this would also serve to report the fault for purposes of step <b>108</b>. Alternatively, if the primary detecting node <b>22</b> logs the fault P<b>1</b> locally in step <b>104</b>, then the reporting of fault P<b>1</b> would constitute a separate action. Again, this action could be performed in several ways. In one scenario, the primary detecting node <b>22</b> could transfer a data block containing all of the fault information required by the fault diagnosis node <b>38</b>. Alternatively, the primary detecting node <b>22</b> could send a message to the fault diagnosis node <b>38</b> to advise it of the fault, and provide a handle or tag that would allow the fault diagnosis node to obtain the fault information from the primary detecting node's fault log.
P-0050[0050] The quantum of fault information provided to the peer nodes in step <b>106</b>, and to the fault diagnosis node in step <b>108</b>, can also vary according to design preferences. For example, it may be helpful to provide the peer nodes <b>24</b>-<b>30</b> and/or the fault diagnostic node <b>38</b> with all diagnostic information the primary detecting node <b>22</b> is able to determine about the fault P<b>1</b> by performing its own local diagnostic operations. In the fault example given above, this might include the primary detecting node <b>22</b> advising that it is unable to communicate with LUN <b>1</b>, and further advising that it is able to communicate with LUN <b>0</b> and LUN <b>2</b> through the controller <b>34</b><i>b. </i>This additional information might help shorten overall fault diagnosis time by eliminating unnecessary fault detection steps at the peer nodes <b>24</b>-<b>30</b> and/or the fault diagnosis node <b>38</b>.
P-0051[0051] When the peer nodes <b>24</b>-<b>30</b> receive the fault information broadcast from the primary detecting node <b>22</b> in step <b>106</b>, they test the fault identifier in step <b>110</b> to determine whether they have previously seen the fault P<b>1</b>. If so, no further action is taken. If the peer nodes <b>24</b>-<b>30</b> have not seen the fault P<b>1</b>, they store the fault identifier P<b>1</b> in step <b>112</b>. All peer nodes that are seeing the fault P<b>1</b> for the first time preferably attempt to recreate the fault in step <b>114</b>. In step <b>116</b>, a test is made to determine whether the recreation attempt was successful. For each peer node <b>24</b>-<b>30</b> where the fault P<b>1</b> cannot be recreated, no further action needs to be taken. However, such nodes preferably report their lack of success in recreating the fault P<b>1</b> to the fault diagnosis node <b>38</b>. This ensures that the fault diagnosis will be based on data from all nodes. In the present example, the fault P<b>1</b> would not be recreated at the peer nodes <b>26</b>-<b>30</b> because those nodes would all be able to communicate with LUN <b>1</b> through the switching element <b>32</b><i>b </i>and the controller <b>34</b><i>a </i>(see FIG. 3). On the other hand, the peer node <b>24</b> would be able to recreate the fault P<b>1</b> because it can only communicate with LUN <b>1</b> via the link <b>60</b>, which is defective. Having successfully recreated the fault P<b>1</b>, the peer node <b>24</b> would become a secondary detecting node. The test in step <b>116</b> would produce a positive result. In step <b>118</b>, the secondary detecting node <b>24</b> logs the fault P<b>1</b> using the same unique identifier. It then provides its own fault report pertaining to the fault P<b>1</b> to the fault diagnosis node <b>38</b> in step <b>120</b>. In step <b>122</b>, the secondary detecting node <b>24</b> preferably issues its own fault information broadcast pertaining to the fault P<b>1</b>. This redundancy, which can result in multiple broadcasts of the same fault information, helps ensure that all nodes will receive notification of the fault P<b>1</b>.
P-0052[0052] In step <b>124</b>, the fault diagnosis node <b>38</b> performs fault diagnosis based on the distributed fault detection information received from the access nodes as part of the fault reports. The fault diagnosis may include determining one or more of a switch-controller connection being defective, a controller-storage device connection being defective, a storage device being defective, a total system failure, or otherwise. In the current example, the fault diagnosis node <b>38</b> receives fault reports from access nodes <b>22</b> and <b>24</b> advising that they cannot reach LUN <b>1</b>. The fault diagnosis node <b>38</b> will also preferably receive reports from the access nodes <b>26</b>, <b>28</b> and <b>30</b> advising that they were unable to recreate the fault P<b>1</b>. If the fault diagnosis node <b>38</b> receives no fault reports from the access nodes <b>26</b>, <b>28</b> and <b>30</b>, it will assume that these nodes could not recreate the fault P<b>1</b>. Because the fault diagnosis node <b>38</b> is preferably aware of the topology of the data storage network <b>20</b>, it can determine that the controller <b>34</b><i>a, </i>LUN <b>1</b> and the link extending between the controller <b>34</b><i>a </i>and LUN <b>1</b> are all functional. It can further determine that the switching element <b>32</b><i>a </i>is functional, and that the link between the node <b>22</b> and the switching element <b>32</b><i>a </i>is functional. This leaves the link <b>60</b> as the sole remaining cause of the fault P<b>1</b>.
P-0053[0053] Turning now to FIG. 6, the various software functions provided by each copy of the fault processing software <b>44</b> that is resident on the access nodes <b>22</b>-<b>30</b> is shown in block diagrammatic form. For reference purposes in the following discussion, the access node whose fault processing software functions are being described shall be referred to as the programmed access node. The fault processing software <b>44</b> includes a functional block <b>200</b> that is responsible for the programmed access node detecting a fault P<b>1</b> when the programmed access node acts as a primary detecting node. A functional block <b>202</b> of the fault processing software <b>44</b> is responsible for assigning the fault P<b>1</b> its unique identifier. A functional block <b>204</b> of the fault processing software <b>44</b> is responsible for logging the fault P<b>1</b> using the unique identifier. A functional block <b>206</b> of the fault processing software <b>44</b> is responsible for providing a fault report pertaining to the fault P<b>1</b> to the fault diagnosis node <b>38</b>. A functional block <b>208</b> of the fault processing software <b>44</b> is responsible for broadcasting fault information pertaining to the fault P<b>1</b> (fault information broadcast) to other access nodes that are peers of the programmed access node when it acts as a primary detecting node.
P-0054[0054] A functional block <b>210</b> of the fault processing software <b>44</b> is responsible for receiving a fault information broadcast pertaining to a fault P<b>2</b> detected at another access node when the programmed access node acts as a secondary detecting node. A functional block <b>212</b> of the fault processing software <b>44</b> is responsible for responding to receipt of the second fault information by testing whether the programmed access node has previously seen the fault P<b>2</b>, and if not, for storing the unique identifier of the fault P<b>2</b>. A functional block <b>214</b> of the fault processing software <b>44</b> is responsible for attempting to recreate the fault P<b>2</b>. A functional block <b>216</b> of the fault processing software <b>44</b> is responsible for detecting the fault P<b>2</b> as a result of attempting to recreate it. A functional block <b>218</b> of the fault processing software <b>44</b> is responsible for logging the fault P<b>2</b>. A functional block <b>220</b> of the fault processing software <b>44</b> is responsible for providing a fault report pertaining to the fault P<b>2</b> to the fault diagnosis node <b>38</b>. A functional block <b>222</b> of the fault processing software <b>44</b> is responsible for broadcasting fault information pertaining to the fault P<b>2</b> (fault information broadcast) to other access nodes that are peers of the programmed access node when it acts as a secondary detecting node.
P-0055[0055] Turning now to FIG. 7, the various software functions provided by the fault diagnosis software <b>46</b> that is resident on the fault diagnosis node <b>38</b> are shown in block diagrammatic form. The fault diagnosis software <b>46</b> includes a functional block <b>300</b> that is responsible for receiving fault reports from the access nodes <b>22</b>-<b>30</b> concerning a storage network fault. A functional block <b>302</b> of the fault diagnosis software <b>46</b> is responsible for evaluating fault information contained in the fault reports. A functional block <b>304</b> of the fault diagnosis software <b>46</b> is responsible for generating a distributed diagnosis of the storage network fault based on the fault information evaluation. As mentioned above, the fault diagnosis performed by the fault diagnosis node <b>38</b> may include determining one or more of a switch-controller connection being defective, a controller-storage device connection being defective, a storage device being defective, a total system failure, or otherwise.
P-0056[0056] Accordingly, a distributed fault detection system and method for a data storage network has been disclosed, together with a computer program product for implementing distributed fault detection functionality. While various embodiments of the invention have been described, it should be apparent that many variations and alternative embodiments could be implemented in accordance with the invention. It is understood, therefore, that the invention is not to be in any way limited except in accordance with the spirit of the appended claims and their equivalents.
Contents4
7 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| EP1746489A1 | Cited by | European Patent Office (EPO) | Search report |
| US2006184823A1 | Cited by | United States of America | Pre-grant |
| US2007286087A1 | Cited by | United States of America | Pre-grant |
| US7206912B2 | Cited by | United States of America | Applicant |
| US7516352B2 | Cited by | United States of America | Search report |
| US10015084B2 | Cited by | United States of America | Applicant |
| US9753797B1 | Cited by | United States of America | Search report |
| CN118740590A | Cited by | China | Search report |
| CN104569785A | Cited by | China | Search report |
| US8161013B2 | Cited by | United States of America | Applicant |
| US2005262233A1 | Cited by | United States of America | Pre-grant |
| US8826032B1 | Cited by | United States of America | Applicant |
| US2008270684A1 | Cited by | United States of America | Pre-grant |
| US2007088763A1 | Cited by | United States of America | Pre-grant |
| US8332860B1 | Cited by | United States of America | Applicant |
| US9042263B1 | Cited by | United States of America | Applicant |
| US12562971B2 | Cited by | United States of America | Search report |
| US2007094447A1 | Cited by | United States of America | Pre-grant |
| US7961594B2 | Cited by | United States of America | Search report |
| US8775387B2 | Cited by | United States of America | Applicant |
| US6810462B2 | Cited by | United States of America | Search report |
| US7231491B2 | Cited by | United States of America | Search report |
| US2005033915A1 | Cited by | United States of America | Pre-grant |
| US10833914B2 | Cited by | United States of America | Search report |
| US7617320B2 | Cited by | United States of America | Applicant |
| US2007011371A1 | Cited by | United States of America | Pre-grant |
| US7415629B2 | Cited by | United States of America | Applicant |
| US2003204671A1 | Cited by | United States of America | Pre-grant |
| US2007226537A1 | Cited by | United States of America | Pre-grant |
| US8566650B2 | Cited by | United States of America | Search report |
| US2020050524A1 | Cited by | United States of America | Search report |
| US7856525B2 | Cited by | United States of America | Applicant |
| US2007162717A1 | Cited by | United States of America | Pre-grant |
| US7478267B2 | Cited by | United States of America | Search report |
| US7702667B2 | Cited by | United States of America | Applicant |
| US2011035620A1 | Cited by | United States of America | Pre-grant |
| US2012311391A1 | Cited by | United States of America | Pre-grant |
| US8812916B2 | Cited by | United States of America | Search report |
| US11237936B2 | Cited by | United States of America | Search report |
| US7444468B2 | Cited by | United States of America | Applicant |
| CN105119765A | Cited by | China | Search report |
| US2003031126A1 | Cites | United States of America | Pre-grant |
| US3873819A | Cites | United States of America | Pre-grant |
| US5390326A | Cites | United States of America | Pre-grant |
| US5537653A | Cites | United States of America | Pre-grant |
| US5684807A | Cites | United States of America | Pre-grant |
| US5712968A | Cites | United States of America | Pre-grant |
| US5724341A | Cites | United States of America | Pre-grant |
| US5774640A | Cites | United States of America | Pre-grant |
| US5784547A | Cites | United States of America | Pre-grant |
| US5966730A | Cites | United States of America | Pre-grant |
| US5987629A | Cites | United States of America | Pre-grant |
| US6278690B1 | Cites | United States of America | Pre-grant |
| US6282112B1 | Cites | United States of America | Pre-grant |
| US6308282B1 | Cites | United States of America | Pre-grant |
2 members in 1 office; this record represents the family
Members2
| Document | Office | Kind | |
|---|---|---|---|
| US2003191992A1 | United States of America | A1 | |
| US6973595B2 | United States of America | B2 |
31 transactions on the USPTO file
Allowed after 1 non-final rejection.
- Non-final rejections
- 1
- Final rejections
- 0
- RCEs
- 0
- Appeals
- 0
Over time
Point at a mark for the transactionTransactions
| Event | Code | |
|---|---|---|
| Expire PatentEXP. | EXP. | |
| Recordation of Patent Grant MailedPGM/ | PGM/ | |
| Patent Issue Date Used in PTA CalculationAllowedPTAC | PTAC | |
| Issue Notification MailedAllowedWPIR | WPIR | |
| Dispatch to FDCD1935 | D1935 | |
| Application Is Considered Ready for IssuePILS | PILS | |
| Issue Fee Payment VerifiedN084 | N084 | |
| Workflow - Drawings FinishedDRWF | DRWF | |
| Issue Fee Payment ReceivedIFEE | IFEE | |
| Mail Notice of AllowanceAllowedMN/=. | MN/=. | |
| Notice of Allowance Data Verification CompletedAllowedN/=. | N/=. | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Response after Non-Final ActionA... | A... | |
| Mail Examiner Interview Summary (PTOL - 413)MEXIN | MEXIN | |
| Interview Summary RecordEXIN | EXIN | |
| Correspondence Address ChangeC.ADB | C.ADB | |
| Mail Non-Final RejectionNon-final rejectionMCTNF | MCTNF | |
| Non-Final RejectionNon-final rejectionCTNF | CTNF | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| IFW TSS Processing by Tech Center CompleteTSSCOMP | TSSCOMP | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Application Dispatched from OIPEOIPE | OIPE | |
| Application Is Now CompleteCOMP | COMP | |
| IFW Scan & PACR Auto Security Review | – | |
| IFW Scan & PACR Auto Security Review | – | |
| Reference capture on IDSRCAP | RCAP | |
| Information Disclosure Statement (IDS) Filed | – | |
| Information Disclosure Statement (IDS) Filed | – | |
| Initial Exam Team nnIEXX | IEXX |
7 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Lapsed due to failure to pay maintenance feeLapsedFP | FP | |
| Information on status: patent discontinuationPATENT EXPIRED DUE TO NONPAYMENT OF MAINTENANCE FEES UNDER 37 CFR 1.362STCH | STCH | |
| Information on status: patent discontinuationPATENT EXPIRED DUE TO NONPAYMENT OF MAINTENANCE FEES UNDER 37 CFR 1.362STCH | STCH | |
| Lapse for failure to pay maintenance feesLapsedLAPS | LAPS | |
| Maintenance fee reminder mailedREMI | REMI | |
| Fee paymentFPAY | FPAY | |
| AssignmentAS | AS |
Numbers
- Application
- 11747802
Titles
- English
- Distributed fault detection for data storage networks
Patent term adjustment
- A delay
- +615 daysthe office missed an examination deadline
- Applicant delay
- −61 days
- Net adjustment
- 554 days
Classification
- CPC, 6
- G06F11/0784
- G06F11/0727
- G06F11/079
- H04L67/1097
- H04L69/40
- H04L69/329
- IPC, 2
- G06F11 07
- H04L69 40