EP3528124A2

Partition tolerance in cluster membership management

Abstract

Techniques are disclosed for managing a cluster of computing nodes following a division of the cluster into at least a first and second partition, where the cluster aggregates local storage resources of the nodes to provide an object store, and objects stored in the object store are divided into data components stored across the nodes. In accordance with one method, it is determined that a majority of data components comprising a first object are stored within nodes in the first partition. It is determined that a majority of data components comprising a second object are stored within nodes in the second partition. In one aspect, configuration objects are permitted to be performed on the first object in the first partition while denying access to the first object from the second partition, and on the second object in the second partition while denying access to the second object from the first partition. In another aspect, I/O operations are permitted upon determining that that a quorum number of the component objects contain a full copy of the data of the first virtual disk object.

EP3528124A2, drawing sheet 1
Sheet 1 of 8

Term

7.7 yearsto projected expiry

Projected expiry 5 June 2034, counted from filing; an application has no term until it is granted.

  1. Priority and filed
  2. Published
  3. Today
  4. Projected expiry

15 claims: 1 independent, 14 dependent

  1. 1
    A computer-implemented method, comprising:presenting an object store (116) to virtual machines (112) running on host computing nodes (111) in a cluster (110), wherein objects in the object store (116) include virtual disk objects (200) which each represent a virtual disk associated with a corresponding virtual machine (112), wherein each virtual disk object (200) is a composite object associated with a set of component objects (220) in the object store (116), wherein each component object (220) represents a data component storing data of the virtual disk, and wherein the data components of the component objects (220) are separately stored across local storage resources (117, 118) of the respective computing nodes (111) according to a specified storage policy (215) of the virtual disk object (220);and managing the cluster (110) of computing nodes (111) following a failure or management event that divides the cluster (110) into at least a first partition and a second partition, including determining, by a coordinator node (111 1 ) in the first partition as owner of a first virtual disk object (200), that a quorum number of the component objects (220) that comprise the first virtual disk object (200) are stored as data components within the computing nodes (111) in the first partition, and, in response, permitting configuration operations which alter the specified storage policy of the first virtual disk object (200) to be performed by the computing nodes (111) in the first partition while denying any access to the first virtual disk object (200) from the computing nodes (111) in the second partition.
  2. 14
    A non-transitory computer readable storage medium storing instructions, which, when executed on a processor, perform the method of any of claims 1 to 13.
  3. 15
    A system, comprising:a processor;and a memory hosting an application, which, when executed on the processor, performs the method of any of claims 1 to 13.