EP0709779A2

Virtual shared disks with application-transparent recovery

Abstract

A system and method for recovering from failures in the disk access path of a clustered computing system. Each node of the clustered computing system is provided with proxy software for handling physical disk access requests from applications executing on the node and for directing the disk access requests to an appropriate server to which the disk is physically attached. The proxy software on each node maintains state information for all pending requests originating from that node. In response to detection of a failure along the disk access path, the proxy software on all of the nodes directs all further requests for disk access to a secondary node physically attached to the same disk.

EP0709779A2, drawing sheet 1
Sheet 1 of 7

Term

Term ended

Projected expiry passed 6 October 2015, 11 years ago.

  1. Priority
  2. Filed
  3. Published
  4. Projected expiry
  5. Today

7 claims: 3 independent, 4 dependent

  1. 1
    A method for recovering from failures in a disk access path of a clustered computing system, comprising the steps of:a) providing each given node of the clustered computing system with proxy logic for handling physical disk access requests from applications executing on the given node and for directing the disk access requests to a primary node to which the disk is physically attached;the proxy logic on each given node maintaining state information for all pending requests originating from that given node;b) detecting a failure along the access path to a disk;c) in response to detection of the failure, notifying the proxy software on all of the nodes to direct all further requests for access to the disk to a secondary node which is also physically attached to the the disk.
  2. 4
    A clustered multi-processing system comprising a plurality of N nodes;a multiported disk having plurality of ports connected to M of the nodes, where N is greater than M;a failure detection mechanism, coupled to the nodes, for detecting failures along a disk access path between the disk and the nodes which are not physically connected to the disk;and, proxy logic on each of the nodes, coupled to the failure detection mechanism, for redirecting access requests to the multiported disk to another disk access path between the disk and the nodes not physically connected to the disk, when a failure is detected.
  3. 7
    A method for recovering from failures along a disk access path in a clustered computing system, comprising the steps of:detecting a failure in the disk access path in the clustered computing system;upon detection of the failure broadcasting a message to all nodes of the system having access to the disk;in response to the message, suspending virtual shared disks on each node, saving pending requests that had been sent to the disk along the failed access path and saving requests that arrive while a virtual shared disk is suspended;broadcasting a second message to the nodes to resume the affected virtual shared disks;upon receipt of the second message, resuming the virtual shared disks at each node by recording the node holding a new primary tail in a destination map for that virtual device;and, re-issuing all of the requests to the new primary tail.