EP1908261B1

Client failure fencing mechanism for fencing network file system data in a host-cluster environment

Abstract

A method and system performs a fencing technique in a host cluster storage environment. The fence program executes on each cluster member in the cluster, and the cluster is coupled to a storage system by a network. When a cluster member fails or cluster membership changes, the fence program is invoked and a host fencing API message is sent via the network to the storage system. The storage system in turn modifies export lists to restrict further access by the failed cluster node to otherwise fence the failed cluster node off from that storage system or from certain directories within that storage system.

EP1908261B1, drawing sheet 1
Sheet 1 of 9

Term

Term ended

Expired 20 July 2026, 0.2 years ago.

  1. Priority and filed
  2. Granted
  3. Expired
  4. Today

22 claims: 9 independent, 13 dependent

  1. 1
    A method for performing fencing in a clustered storage environment (100), the method comprising the steps of:providing a plurality of nodes configured in a cluster (120) for sharing data, each node being a cluster member (130a;130b);providing a storage system (200) that supports a plurality of data containers, said storage system (200) supporting a protocol that configures export lists (406;408;410) that assign each cluster member (130a;130b) certain access permission rights including read-write access permission or read only access permission as to each respective data container associated with the storage system (200);coupling the cluster (120) to the storage system (200) by a network (160);and providing a fencing program in each cluster member (130a;130b) such that when a change in cluster membership is detected, a surviving member transmits an application program interface message to said storage system (200) commanding said storage system to modify one or more of said export lists (406;408;410) such that the access permission rights of one or more identified cluster members (130a;130b) are modified.
  2. 4
    The method as defined in any of claims 1, 2 or 3, wherein said cluster member (130a;130b) transmits said application program interface messages encapsulated in HyperText Transfer Protocol messages.
  3. 5
    The method as defined in any of the preceding claims, further comprising the step of initially configuring the storage system (200) such that each cluster member (130a;130b) has a set of access permission rights set forth in a configuration file.
  4. 7
    The method as defined in any of the preceding claims, including the further step of validating the storage system (200) and the exports configuration.
  5. 8
    The method as defined in any of the preceding claims, further comprising checking the initial configuration for the presence of any wildcard entries in said export lists (406;408;410), and if said wildcard entries are found, issuing an error message.
  6. 9
    The method as defined in any of the preceding claims, wherein a change in cluster membership is determined by the associated cluster infrastructure.
  7. 10
    A system adapted to perform fencing operations in a clustered storage environment (100), the system comprising:a cluster (120) of interconnected cluster members (130a;130b) that provide storage services for one or more clients (150);a networked storage system (200) coupled to said cluster (120) by way of a network, said storage system (200) having a plurality of storage devices (230), said storage system (200) including export lists (406;408;410) that contain rules regarding access permission rights for specified cluster members (130a;130b) as to specified associated storage devices, and/or files and/or directories stored therein;and a fencing program (136a;136b) executing within each cluster member (130a;130b) that includes program instructions for sending an application program interface message from one of said cluster members (130a;130b) to said storage system (200) and that contains instructions for modifying one or more of said export lists (406;408;410) to change the access permission rights of one or more of said cluster members (130a;130b).
  8. 13
    The system as defined in any of claims 10, 11 or 12, where said storage system (200) is initially configured with a set of export lists that sets forth read-only and/or read-write access permissions for each cluster member (130a;130b) of the cluster (120) that is interfaced with the storage system (200) with respect to one of an entire storage system, a file of a storage system, and a directory of a storage system.
  9. 14
    The system as defined in any of claims 10 to 13, wherein said fencing program (136a;136b) is initiated upon a change in cluster membership as determined by the associated cluster infrastructure.
  10. 15
    The system as defined in any of claims 10 to 14, wherein said cluster members (130a;130b) are interfaced with a client requesting data from the storage system (200) via one or more networks (160).
  11. 16
    The system as defined in any of claims 10 to 15, further comprising a quorum device (172) that is directly coupled to each cluster member (130a;130b) and which quorum device (172) is configured in such manner that the cluster member (130a;130b) that asserts a claim to the quorum device (172) first is thereby granted access to the storage system (200).
  12. 18
    A computer readable medium for performing failure fencing in a clustered environment with networked storage (200), including program instructions for performing the steps of:providing a plurality of nodes configured in a cluster (120) for sharing data, each node being a cluster member (130a;130b);providing a storage system (200) that has access to multiple storage devices including disks (230), said storage system (200) supporting a protocol that configures export lists that assign each cluster member (130a;130b) certain access permission rights including read-write access permission or read only access permission as to each respective storage device, and files and directories of said storage device, associated with the storage system (200);coupling the cluster (120) to the storage system (200) by a network (160);and providing a fencing program (136a;136b) in each cluster member (130a;130b) such that when a change in cluster membership is detected, a surviving member transmits an application program interface message via said protocol over said network to said storage system (200) commanding said storage system to modify one or more of said export lists such that the access permission rights of one or more identified cluster members (130a;130b) are modified.
  13. 21
    The computer readable medium for performing failure fencing as defined in any of claims 18 to 20, further comprising a simple and complete user interface that can be plugged into a host cluster framework which can accommodate different types of shared data containers.
  14. 22
    The computer readable medium as defined in any of claims 18 to 21, including means for supporting NFS as a shared data source in a high-availability environment that includes one or more storage system clusters and one or more host clusters having end-to-end availability in mission-critical deployments having substantially continuous availability.