Isolating a drive from disk array for diagnostic operations
Summary by NHIP
RAID Drive Isolation System
The system isolates a suspected faulty hard disk drive within a redundant storage array to enable diagnostics without disrupting the rest of the network. Distinctive elements include redundant disk array switches that bypass the faulty drive port and establish a private zone containing the drive and a SCSI Enclosure Services sub-processor for testing.
Claim Score by NHIP
Abstract
A storage system includes a RAID adapter, disk array switches, sub-processors, and hard disk drives (HDDs). The system permits the isolation of a suspected faulty HDD to allow diagnostics to be performed without impacting operation of the rest of the system. Upon detection of a possible fault in a target HDD, a private zone is established including the target HDD and one of the sub-processors, thereby isolating the target HDD. The sub-processor performs diagnostic operations, then transmits its results to the adapter. A faulty HDD can then be fully isolated and the private zone is disassembled, allowing the sub-processor to rejoin the network.

Term
Projected expiry 1 June 2027.
- Priority and filed
- Granted
- Today
- Projected expiry
19 claims: 4 independent, 15 dependent
- 1Broadest claimClaim Score 53, average(NHIP)A redundant storage system, comprising:first and second redundant disk array switches;a group of hard disk drives (HDDs), each coupled separately to the first and second switches through a pair of independent ports;first and second redundant sub-processors coupled to the first and second switches, respectively;an adapter separately interconnected with the first and second switches through a network;means for detecting a suspected faulty HDD;the first switch comprising means for bypassing the port through which the suspected faulty HDD is coupled;the second switch comprising means for establishing a private zone, comprising the second sub-processor and the suspected faulty HDD whereby the private zone is isolated from the network;the second sub-processor comprising: means for performing diagnostic operations on the suspected faulty HDD;and means for transmitting results of the diagnostic operations to the adapter through the first sub-processor.
- 5A method for isolating and performing diagnostics on a hard disk drive in a redundant storage system, the system having first and second redundant disk array switches, a group of hard disk drives (HDDs), each coupled separately to the first and second switches through a pair of independent ports, first and second redundant sub-processors coupled to the first and second switches, respectively, and an adapter separately interconnected with the first and second switches through an interconnecting network, the method comprising:detecting possible faults in a target HDD;using the first switch to bypass the port through which the suspected faulty HDD is coupled;using the second switch to establish a private zone, isolated from the network, comprising the target HDD and the second sub-processor;using the second sub-processor to perform diagnostics on the target HDD;and transmitting the results of the diagnostics to the adapter through the first sub-processor.
- 10A computer program product of a computer readable recordable-type medium usable with a programmable computer, the computer program product having computer-readable code embodied therein for isolating and performing diagnostics on a suspected faulty hard disk drive in a redundant storage system, the system having first and second redundant disk array switches, a group of hard disk drives (HDDs), each coupled separately to the first and second switches through a pair of independent ports, first and second redundant sub-processors coupled to the first and second switches, respectively, and an adapter separately interconnected with the first and second switches through an interconnecting network, the computer-readable code comprising instructions for:detecting possible faults in a target HDD;using the first switch to bypass the port through which the suspected faulty HOD is coupled;using the second switch to establish a private zone, isolated from the network, comprising the target HDD and the second sub-processor;using the second sub-processor to perform diagnostics on the target HDD;and transmitting the results of the diagnostics to the adapter through the first sub-processor.
- 15A method for deploying computing infrastructure, comprising integrating computer readable code into a computing system, the system having first and second redundant disk array switches, a group of hard disk drives (HDDs), each coupled separately to the first and second switches through a pair of independent ports, first and second redundant sub-processors coupled to the first and second switches, respectively, and an adapter separately interconnected with the first and second switches through an interconnecting network, wherein the code, in combination with the computing system, is capable of performing the following:detecting possible faults in a target HDD;using the first switch to bypass the port through which the suspected faulty HDD is coupled;using the second switch to establish a private zone, isolated from the network, comprising the target HDD and the second sub-processor;using the second sub-processor to perform diagnostics on the target HDD;and transmitting the results of the diagnostics to the adapter through the first sub-processor.
Independent claims4
16 paragraphs in 6 sections, as filed
RELATED APPLICATION DATA
p-0002The present application is related to commonly-assigned and co-pending U.S. application Ser. No. 11/386,066, entitled ENCLOSURE-BASED RAID PARITY ASSIST, and Ser. No. 11/386,025, entitled OFFLOADING DISK-RELATED TASKS FROM RAID ADAPTER TO DISTRIBUTED SERVICE PROCESSORS IN SWITCHED DRIVE CONNECTION NETWORK ENCLOSURE filed on the filing date hereof, which applications are incorporated herein by reference in their entireties.
TECHNICAL FIELD
p-0003The present invention relates generally to RAID storage systems and, in particular, to isolating and diagnosing a target drive with minimal impact on the balance of the system.
BACKGROUND ART
p-0004Many computer-related systems now include redundant components for high reliability and availability. Nonetheless, the failure or impending failure of a component may still affect the performance of other components or of the system as a whole. For example, in a RAID storage system, an enclosure includes an array of hard disk drives (HDDs) which are each coupled through independent ports to both of a pair of redundant disk array switches. One of a pair of redundant sub-processors is coupled to one of the switches while the other of the pair is coupled to the other switch. Alternatively, a single sub-processor is coupled to both switches and logically partitioned into two images, each logically coupled to one of the switches. Each switch is also coupled through a fabric or network to both of a pair of redundant RAID adapters external to the enclosure. The system may include additional enclosures, each coupled in daisy-chain fashion in the network to the disk array switches of the previous enclosure.
p-0005If the system is fibre channel-arbitrated loop (FC-AL) architecture, when the system is initialized, either or both RAID adapters (collectively referred to as “adapter”) performs a discovery operation using a “pseudo-loop” through the switches. During discovery, the addresses of all of the devices on the network are determined. The system then enters its normal switched mode. However, if a drive becomes faulty during normal system operations, it may repeatedly enter and exit the network, each time causing the adapter to enter the discovery mode again, resulting in system-wide disruption.
p-0006If diagnostics are performed on the suspected faulty drive, the system is further disrupted. While it is possible to isolate the suspected faulty drive by by-passing the ports through which it is coupled to the switches, effectively removing the drive from the network, the drive is then inaccessible for diagnostic operations to be performed on it.
p-0007Consequently, a need remains to be able to perform diagnostic operations on a drive without disrupting access to the rest of the disk array or to the network.
SUMMARY OF THE INVENTION
p-0008The present invention includes a storage system, a RAID adapter, disk array switches, sub-processors, and hard disk drives (HDDs). The system permits the isolation of a target HDD to allow diagnostics to be performed without impacting operation of the rest of the system. The status of HDDs is monitored for a variety of factors, such as unstable network behaviors, slow response or some other trigger event or process. Upon detection of such an event or process (also referred to herein as a “possible fault”), a private zone is established including the target HDD and one of the sub-processors, thereby isolating the target HDD. The sub-processor performs diagnostic operations, then transmits its results to the adapter. The target HDD is then fully isolated and the private zone is disassembled, allowing the sub-processor to rejoin the network.
BRIEF DESCRIPTION OF THE DRAWINGS
p-0009<figref idrefs="DRAWINGS">FIG. 1</figref> is a block diagram of a RAID storage system in which the present invention may be implemented;
p-0010<figref idrefs="DRAWINGS">FIG. 2</figref> is a block diagram of the RAID storage system illustrating the process of isolating and diagnosing a target drive; and
p-0011<figref idrefs="DRAWINGS">FIG. 3</figref> is flowchart of a method of the present invention.
DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENT
p-0012<figref idrefs="DRAWINGS">FIG. 1</figref> is a block diagram of a RAID storage system <b>100</b> in which the present invention may be implemented. The system <b>100</b> includes a redundant pair of RAID adapters or controllers <b>110</b>A, <b>110</b>B (collectively referred to as <b>110</b>) which are coupled to one or more servers. The system <b>100</b> further includes an enclosure <b>115</b> housing a pairs of redundant disk array switches <b>120</b>A and <b>120</b>B (collectively referred to as <b>120</b>). The enclosure <b>115</b> also houses a group of hard disk drives (HDDs) <b>130</b>A, <b>130</b>B, <b>130</b>C, <b>130</b>D, <b>130</b>E, <b>130</b>F (collectively referred to as <b>130</b>). Each HDD is coupled through ports <b>122</b> with both switches. The system <b>100</b> also includes a pair of redundant sub-processors or service processors <b>140</b>A, <b>140</b>B (and collectively referred to as <b>140</b>), such as SCSI Enclosure Services (SES) processors, each coupled through a fabric or network <b>142</b> with one of the switches <b>120</b>A, <b>120</b>B. The sub-processors <b>140</b>A, <b>140</b>B are coupled to each other with a processor-to-processor link <b>144</b>. In the system <b>100</b> illustrated, the service processors <b>140</b>A, <b>140</b>B are SCSI Enclosure Service (SES) processors which manage switch functions and the enclosure environment. The adapters <b>110</b> are coupled to the switches <b>120</b> through fabric or network links <b>112</b>. The system <b>100</b> may include additional enclosures coupled in daisy-chain fashion to ports of the upstream enclosure. Thus, any communications between an adapter <b>110</b> and a switch or HDD in an enclosure passes through the switches of upstream enclosures.
p-0013The system <b>100</b> may be based on a fibre channel-arbitrated loop (FC-AL) architecture, a serial attached SCSI (SAS) architecture, or other architecture which includes dual-ported access to the HDD.
p-0014Referring to <figref idrefs="DRAWINGS">FIG. 2</figref> and to the flowchart of <figref idrefs="DRAWINGS">FIG. 3</figref>, a possible fault has been detected in one of the HDDs <b>130</b>F (step <b>300</b>). Rather than perform the diagnostics in the adapter <b>110</b>, the task is offloaded to a sub-processor <b>140</b>. A “private zone” <b>200</b> is established by one of the switches (step <b>302</b>), switch <b>120</b>A in <figref idrefs="DRAWINGS">FIG. 2</figref>, including one of the sub-processors (<b>140</b>A), a target drive <b>130</b>F and the port <b>122</b>F<sub>1 </sub>through which the target drive <b>130</b>F is coupled to the switch <b>120</b>A. The other port <b>122</b>F<sub>2 </sub>through which the target drive <b>130</b>F is coupled to the other switch <b>120</b>B is disabled or by-passed by the other switch <b>120</b>B. The components within the private zone are thus isolated from the balance of the system <b>100</b>. The sub-processor <b>140</b>A is then able to perform diagnostics on the target drive <b>130</b>F (step <b>304</b>) without impacting the rest of the system <b>100</b>.
p-0015Upon completion of the diagnostic operations, the sub-processor <b>140</b>A communicates the results to the other sub-processor <b>140</b>B over the processor-to-processor link <b>144</b> (step <b>306</b>). The other sub-processor <b>140</b>B then communicates the results through the switch <b>120</b>B to the adapter <b>110</b> over the network <b>112</b> (step <b>308</b>). Subsequently, if the target drive <b>130</b>F is determined to be faulty, both ports <b>122</b>F<sub>1 </sub>and <b>122</b>F<sub>2 </sub>through which the drive <b>130</b>F is coupled to the switches <b>120</b>A, <b>120</b>B, respectively, are disabled or by-passed to fully isolate the drive <b>130</b>F (step <b>310</b>) and the private zone is disassembled (step <b>312</b>), allowing the sub-processor <b>140</b>A to rejoin the full network.
p-0016It is important to note that while the present invention has been described in the context of a fully functioning data processing system, those of ordinary skill in the art will appreciate that the processes of the present invention are capable of being distributed in the form of a computer readable medium of instructions and a variety of forms and that the present invention applies regardless of the particular type of signal bearing media actually used to carry out the distribution. Examples of computer readable media include recordable-type media such as a floppy disk, a hard disk drive, a RAM, and CD-ROMs and transmission-type media such as digital and analog communication links.
p-0017The description of the present invention has been presented for purposes of illustration and description, but is not intended to be exhaustive or limited to the invention in the form disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art. The embodiment was chosen and described in order to best explain the principles of the invention, the practical application, and to enable others of ordinary skill in the art to understand the invention for various embodiments with various modifications as are suited to the particular use contemplated. Moreover, although described above with respect to methods and systems, the need in the art may also be met with a computer program product containing instructions for isolating and performing diagnostics on a hard disk drive in a redundant storage system or a method for deploying computing infrastructure comprising integrating computer readable code into a computing system for isolating and performing diagnostics on a hard disk drive in a redundant storage system.
Contents6
4 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US8645652B2 | Cited by | United States of America | Applicant |
| US2008052566A1 | Cited by | United States of America | Pre-grant |
| US9513996B2 | Cited by | United States of America | Applicant |
| US9542273B2 | Cited by | United States of America | Search report |
| US2015100821A1 | Cited by | United States of America | Pre-grant |
| US2009292932A1 | Cited by | United States of America | Pre-grant |
| JP2002037427A | Cites | Japan | Applicant |
| US2002191537A1 | Cites | United States of America | Search report |
| US2003191992A1 | Cites | United States of America | Search report |
| JP2004199551A | Cites | Japan | Applicant |
| US2005010843A1 | Cites | United States of America | Applicant |
| US2005228943A1 | Cites | United States of America | Search report |
| US5285451A | Cites | United States of America | Applicant |
| US5699510A | Cites | United States of America | Applicant |
| US5975738A | Cites | United States of America | Applicant |
| US6654831B1 | Cites | United States of America | Applicant |
| US6678839B2 | Cites | United States of America | Search report |
| US6751136B2 | Cites | United States of America | Applicant |
| US6766466B1 | Cites | United States of America | Search report |
| US6813112B2 | Cites | United States of America | Applicant |
| US6826778B2 | Cites | United States of America | Applicant |
| US7047450B2 | Cites | United States of America | Search report |
| US7222259B2 | Cites | United States of America | Search report |
2 priority claims, no other members on record
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 38538806 | United States of America | A | |
| US20060385388 | – | – | – |
46 transactions on the USPTO file
Allowed after 1 non-final rejection.
- Non-final rejections
- 1
- Final rejections
- 0
- RCEs
- 0
- Appeals
- 0
Over time
Point at a mark for the transactionTransactions
| Event | Code | |
|---|---|---|
| Recordation of Patent Grant MailedPGM/ | PGM/ | |
| Patent Issue Date Used in PTA CalculationAllowedPTAC | PTAC | |
| Issue Notification MailedAllowedWPIR | WPIR | |
| Dispatch to FDCD1935 | D1935 | |
| Dispatch to FDCD1935 | D1935 | |
| Correspondence Address ChangeC.AD | C.AD | |
| Application Is Considered Ready for IssuePILS | PILS | |
| Issue Fee Payment VerifiedN084 | N084 | |
| Issue Fee Payment ReceivedIFEE | IFEE | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Email NotificationEML_NTR | EML_NTR | |
| Email NotificationEML_NTF | EML_NTF | |
| Mail Examiner's AmendmentMEX.A | MEX.A | |
| Mail Notice of AllowanceAllowedMN/=. | MN/=. | |
| Notice of Allowance Data Verification CompletedAllowedN/=. | N/=. | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Examiner's Amendment CommunicationEX.A | EX.A | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Response after Non-Final ActionA... | A... | |
| Mail Post CardPST_CRD | PST_CRD | |
| Email NotificationEML_NTF | EML_NTF | |
| Mail Non-Final RejectionNon-final rejectionMCTNF | MCTNF | |
| Non-Final RejectionNon-final rejectionCTNF | CTNF | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| PG-Pub Issue NotificationPG-ISSUE | PG-ISSUE | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| IFW TSS Processing by Tech Center CompleteTSSCOMP | TSSCOMP | |
| Transfer Inquiry to GAUTI1050 | TI1050 | |
| Application Return from OIPEWROIPE | WROIPE | |
| Application Is Now CompleteCOMP | COMP | |
| Application Return TO OIPEROIPE | ROIPE | |
| Application Is Now CompleteCOMP | COMP | |
| Application Is Now CompleteCOMP | COMP | |
| Application Return from OIPEWROIPE | WROIPE | |
| Application Return TO OIPEROIPE | ROIPE | |
| Application Return from OIPEWROIPE | WROIPE | |
| Application Return TO OIPEROIPE | ROIPE | |
| Application Dispatched from OIPEOIPE | OIPE | |
| Application Is Now CompleteCOMP | COMP | |
| Cleared by OIPE CSRL194 | L194 | |
| IFW Scan & PACR Auto Security ReviewSCAN | SCAN | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Reference capture on IDSRCAP | RCAP | |
| Information Disclosure Statement (IDS) FiledM844 | M844 | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Initial Exam Team nnIEXX | IEXX |
11 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Lapsed due to failure to pay maintenance feeLapsedFP | FP | |
| Lapse for failure to pay maintenance feesLapsedPATENT EXPIRED FOR FAILURE TO PAY MAINTENANCE FEES (ORIGINAL EVENT CODE: EXP.); ENTITY STATUS OF PATENT OWNER: LARGE ENTITYLAPS | LAPS | |
| Information on status: patent discontinuationPATENT EXPIRED DUE TO NONPAYMENT OF MAINTENANCE FEES UNDER 37 CFR 1.362STCH | STCH | |
| Fee payment procedureMAINTENANCE FEE REMINDER MAILED (ORIGINAL EVENT CODE: REM.); ENTITY STATUS OF PATENT OWNER: LARGE ENTITYFEPP | FEPP | |
| Fee paymentFPAY | FPAY | |
| Fee paymentFPAY | FPAY | |
| Surcharge for late paymentSULP | SULP | |
| Maintenance fee reminder mailedREMI | REMI | |
| Information on status: patent grantGrantedPATENTED CASESTCF | STCF | |
| Fee payment procedurePAYOR NUMBER ASSIGNED (ORIGINAL EVENT CODE: ASPN); ENTITY STATUS OF PATENT OWNER: LARGE ENTITYFEPP | FEPP | |
| AssignmentAS | AS |
Numbers
- Publication, DOCDB
- 7516352
- Publication, EPODOC
- US7516352
- Application
- 11385388
- Application, DOCDB
- 38538806
- Application, EPODOC
- US20060385388
Titles
- English
- Isolating a drive from disk array for diagnostic operations
Patent term adjustment
- A delay
- +437 daysthe office missed an examination deadline
- Net adjustment
- 437 days
Classification
- CPC, 1
- G06F11/2221
- IPC, 1
- G06F11 00
- USPC, 2
- 714003000
- 714006130