US8448013B2

Failure-specific data collection and recovery for enterprise storage controllers

Summary by NHIP

Storage controller failure recovery

The method detects storage controller failures linked to specific host-device relationships and executes targeted recovery processes using failure IDs. It suspends I/O for the affected pair while maintaining connectivity for unrelated hosts and devices through the controller.

Claim Score by NHIP

Read claim 1, the broadest

Abstract

A method, apparatus, and computer program product for handling a failure condition in a storage controller is disclosed. In certain embodiments, a method may include initially detecting a failure condition in a storage controller. The failure condition may be associated with a specific host and a specific storage device connected to the storage controller. The method may further include determining a failure ID associated with the failure condition. Using the failure ID, en entry may be located in a data collection and recovery table. This entry may indicate one or more data collection and/or recovery processes to execute in response to the failure condition. The method may then execute the data collection and/or recovery processes indicated in the entry. While executing the data collection and/or recovery processes, connectivity may be maintained between hosts and storage devices not associated with the failure condition.

US8448013B2, drawing sheet 1
Sheet 1 of 7

Term

Projected expiry 3 November 2029.

  1. Priority and filed
  2. Granted
  3. Today
  4. Projected expiry

20 claims: 3 independent, 17 dependent

  1. 1
    Broadest claimClaim Score 47, average(NHIP)A method to handle a failure condition in a storage controller, the storage controller enabling one or more host devices to access data in one or more storage devices, the method comprising:detecting a failure condition in the storage controller;identifying a specific host-device/storage-device relationship associated with the failure condition;determining a failure ID associated with the failure condition;locating, in a table using the failure ID, an entry identifying a set of data structures and traces to collect and a recovery process to execute on the storage controller in response to the failure condition;collecting the set of data structures and traces identified in the entry for the specific host-device/storage-device relationship;executing the recovery process on the storage controller for the specific host-device/storage-device relationship;and while collecting the set of data structures and traces and executing the recovery process, suspending I/O between the specific host device and storage device associated with the failure condition, while maintaining, through the storage controller, I/O between host devices and storage devices not associated with the failure condition.
  2. 11
    An apparatus to handle a failure condition in a storage controller, the storage controller enabling one or more host devices to access data in one or more storage devices, the apparatus comprising:a storage controller storing modules for execution thereon, the modules comprising: a detection module to detect a failure condition in the storage controller;the detection module further configured to identify a specific host-device/storage-device relationship associated with the failure condition;a determination module to determine a failure ID associated with the failure condition;a location module to locate, in a table using the failure ID, an entry identifying a set of data structures and traces to collect and a recovery process to execute on the storage controller in response to the failure condition;a data collection module to collect the set of data structures and traces identified in the entry for the specific host-device/storage-device relationship;a recovery module to execute the recovery process on the storage controller for the specific host-device/storage-device relationship;and a maintenance module to, while collecting the set of data structures and traces and executing the recovery process, suspend I/O between the specific host device and storage device associated with the failure condition, while maintaining I/O between host devices and storage devices not associated with the failure condition.
  3. 16
    A computer program product to handle a failure condition in a storage controller, the computer program product comprising a non-transitory computer-usable storage medium having computer-usable program code embodied therein, the computer-usable program code comprising:computer-usable program code to detect a failure condition in a storage controller;computer-usable program code to identify a specific host-device/storage-device relationship associated with the failure condition;computer-usable program code to determine a failure ID associated with the failure condition;computer-usable program code to locate, in a table using the failure ID, an entry identifying a set of data structures and traces to collect and a recovery process to execute on the storage controller in response to the failure condition;computer-usable program code to collect the set of data structures and traces identified in the entry for the specific host-device/storage-device relationship;computer-usable program code to execute the recovery process on the storage controller for the specific host-device/storage-device relationship;and computer-usable program code to, while collecting the set of data structures and traces and executing the recovery process, suspend I/O between the specific host device and storage device associated with the failure condition, while maintaining I/O between host devices and storage devices not associated with the failure condition.