Providing executing programs with reliable access to non-local block data storage
Abstract
A computerized method for managing, through running programs, access to data storage functionality at the block level, said method comprising: receiving one or more indications of a first group of access data requests initiated by a first copy of a first program (375) running on a first computer system (180, 300, 370, 390) to access block level data stored in a data storage volume at the non-local block level (155), the volume of data storage at the block level (155) being provided by a second different data storage system (165, 360) by means of one or more networks (170, 185, 385) and being connected to a first computer system (180, 300, 370, 390) so that the first copy of the running program (375) initiates the requests for data access for the volume of data storage at the block level (155) through interactions with a first storage device at the local logical block level to the first computer system (180, 300, 370, 390) representing the volume of data storage at block level (155); automatically respond to the indications received from the first group of data access requests by interacting with the second data storage system (165, 360) on behalf of the first copy of the program being executed (375) to start the realization of the request for data access of the first group in the volume of data storage at block level (155) provided by the second data storage system (165, 360); After determining that the first copy of the program (375) is no longer available, identify a third computer system (180, 300, 370, 390) in which a second copy of the first program (375) is running, the third computer system other than the first computer system and the second data storage system, the second copy of the first program already being executed before the first copy of the program becomes unavailable, and connecting the data storage volume at the block level (155) to the third computer system (180, 300, 370, 390) so that the second program copy (375) has access to the second storage device at the block level local logic to the third computer system (180, 300, 370, 390) representing the volume of data storage at block level (155), in which to identify the third computer system (180, 300, 370, 390) in which the second copy of the first program (375) is running and the connection of the data storage volume at block level (155) to the third computer system (180, 300, 370, 390) is automatically made to maintain the access of the first program (375) to the volume of data storage at block level (155), automatic maintenance of the access being carried out in response to the automatic definition that the first program copy (375) is no longer available; receive one or more indications from a second group of other data access requests initiated by the second copy of the program being executed (375) for the volume of data storage at block level (155) by means of interactions with the second device storage storage at the local and logical block level in the third computer system (180, 300, 370, 390); and automatically respond to the indications received from the second group of data access requests by interacting with the second data storage system (165, 360) on behalf of the second copy of the running program (375) to initiate the realization of the data access requests of the second group in the block-level data storage volume (155) in the second data storage system (165, 360).

Term
2.9 yearsto projected expiry
Projected expiry 7 August 2029, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
15 claims: 4 independent, 11 dependent
- 1ES 2 575 155 T3 REIVINDICACIONES 1. Un método computarizado para gestionar mediante programas en ejecución el acceso a la funcionalidad de almacenamiento de datos a nivel de bloque, dicho método comprendiendo:recibir una o más indicaciones de un primer grupo de peticiones de datos de acceso iniciadas por una primera copia de un primer programa (375) ejecutándose en un primer sistema informático (180, 300, 370, 390) para acceder a datos a nivel de bloque almacenados en un volumen de almacenamiento de datos a nivel de bloque no local (155), el volumen de almacenamiento de datos a nivel de bloque (155) siendo provisto por un segundo sistema de almacenamiento de datos distinto (165, 360) por medio de una o más redes (170, 185, 385) y estando conectado a un primer sistema informático (180, 300, 370, 390) de manera que la primera copia de programa en ejecución (375) inicie las peticiones de acceso a datos para el volumen de almacenamiento de datos a nivel de bloque (155) por medio de interacciones con un primer dispositivo de almacenamiento a nivel de bloque lógico local al primer sistema informático (180, 300, 370, 390) que representa el volumen de almacenamiento de datos a nivel de bloque (155);responder automáticamente a las indicaciones recibidas del primer grupo de peticiones de acceso a datos por medio de la interacción con el segundo sistema de almacenamiento de datos (165, 360) en nombre de la primera copia de programa en ejecución (375) como para iniciar la realización de la petición de acceso a datos del primer grupo en el volumen de almacenamiento de datos a nivel de bloque (155) provisto por el segundo sistema de almacenamiento de datos (165, 360);después de determinar que la primera copia de programa (375) ha dejado de estar disponible, identificar un tercer sistema informático (180, 300, 370, 390) en el que se está ejecutando una segunda copia del primer programa (375), siendo el tercer sistema informático distinto del primer sistema informático y del segundo sistema de almacenamiento de datos, la segunda copia del primer programa ya siendo ejecutada antes de que la primera copia del programa deje de estar disponible, y conectar el volumen de almacenamiento de datos a nivel de bloque (155) al tercer sistema informático (180, 300, 370, 390) de manera que la segunda copia de programa (375) tenga acceso al segundo dispositivo de almacenamiento a nivel de bloque lógico local al tercer sistema informático (180, 300, 370, 390) que representa el volumen de almacenamiento de datos a nivel de bloque (155), en el que identificar el tercer sistema informático (180, 300, 370, 390) en el que la segunda copia del primer programa (375) se está ejecutando y la conexión del volumen de almacenamiento de datos a nivel de bloque (155) al tercer sistema informático (180, 300, 370, 390) se realiza automáticamente para mantener el acceso del primer programa (375) al volumen de almacenamiento de datos a nivel de bloque (155), el mantenimiento automático del acceso siendo realizado en respuesta a la definición automática de que la primera copia de programa (375) ha dejado de estar disponible;recibir una o más indicaciones de un segundo grupo de otras peticiones de acceso a datos iniciadas por la segunda copia de programa en ejecución (375) para el volumen de almacenamiento de datos a nivel de bloque (155) por medio de interacciones con el segundo dispositivo de almacenamiento de almacenamiento a nivel de bloque local y lógico en el tercer sistema informático (180, 300, 370, 390);y responder automáticamente a las indicaciones recibidas del segundo grupo de peticiones de acceso a datos mediante la interacción con el segundo sistema de almacenamiento de datos (165, 360) en nombre de la segunda copia de programa en ejecución (375) para iniciar la realización de las peticiones de acceso a datos del segundo grupo en el volumen de almacenamiento de datos a nivel de bloque (155) en el segundo sistema de almacenamiento de datos (165, 360).
- 2El método de la reivindicación 1 que además comprende definir automáticamente que la primera copia de programa (375) ha dejado de estar disponible, estando la no disponibilidad de la primera copia de programa (375) basada en al menos un fallo del primer sistema informático (180, 300, 370, 390), en un fallo en la conectividad con el primer sistema informático (180, 300, 370, 390), y en una incapacidad del primer sistema informático (180, 300, 370, 390) de continuar con la ejecución de la primera copia de programa (375).
- 3El método de la reivindicación 1 en el que al menos una de las acciones de determinar que la primera copia de programa (375) ha dejado de estar disponible, identificar el tercer sistema informático (180, 300, 370, 390) en el cual se está ejecutando la segunda copia del primer programa (375), y conectar el volumen de almacenamiento de datos a nivel de bloque (155) al tercer sistema informático (180, 300, 370, 390) se lleva a cabo en respuesta a una o más indicaciones recibidas de un usuario asociado al primer programa (375) y al volumen de almacenamiento de datos a nivel de bloque (155).
- 4El método de la cláusula 1 en el que la primera copia del primer programa (375) es una de las múltiples copias del primer programa (375) que se ejecutan en múltiples sistemas informáticos distintos (180, 300, 370, 390), en el que al menos una de las múltiples copias (375) es una copia alternativa para una o más de las otras múltiples copias (375), y en el que la identificación del tercer sistema informático en el que se está ejecutando la segunda copia del primer programa incluye seleccionar la segunda copia del primer programa basada en la segunda copia siendo al menos una de las copias alternativas.
- 5El método de la reivindicación 1 que además comprende, antes de conectar el volumen de almacenamiento de ES 2 575 155 T3 datos a nivel de bloque (155) al primer sistema informático (180, 300, 370, 390), crear el volumen de almacenamiento de datos a nivel de bloque (155) en el segundo sistema de almacenamiento de datos (165, 360) y crear una copia espejo del volumen de almacenamiento de datos a nivel de bloque (155) en un cuarto sistema de almacenamiento de datos a nivel de bloque (165, 360) distinto.
- 6El método de la reivindicación 1 que además comprende, antes de recibir las indicaciones del primer grupo de una o más peticiones de acceso a datos, conectar (420) el volumen de almacenamiento de datos a nivel de bloque (155) al primer sistema informático (180, 300, 370, 390) para que lo utilice la primera copia de programa en ejecución (375), incluyendo la conexión del volumen de almacenamiento de datos a nivel de bloque (155) la asociación del primer dispositivo de almacenamiento a nivel de bloque lógico al volumen de almacenamiento de datos a nivel de bloque (155) provisto por el segundo sistema de almacenamiento de datos (165, 360).
- 7El método de la reivindicación 1 en el que el primer y tercer sistema informático (180, 300, 370, 390) y el segundo sistema de almacenamiento de datos (165, 360) son un subconjunto de una pluralidad de sistemas informáticos (180, 300, 370, 390) co-localizados en una primera ubicación geográfica, en donde la pluralidad de sistemas informáticos (180, 300, 370, 390) incluye múltiples sistemas de almacenamiento de datos a nivel de bloque (165, 360) que están provistos por un servicio de almacenamiento de datos a nivel de bloque, en el que el segundo y tercer sistema de almacenamiento de datos (165, 360) son cada uno distintos de los múltiples sistemas de almacenamientos de datos a nivel de bloque (165, 360), y la primera ubicación geográfica es un centro de datos.
- 8Un medio legible por ordenador cuyos contenidos permiten a uno o más sistemas informáticos gestionar el acceso a la funcionalidad de almacenamiento de datos a nivel de bloque mediante programas en ejecución, al llevar a cabo un método que comprende:recibir una o más indicaciones de un primer grupo de peticiones de acceso a datos iniciadas por una primera copia de un primer programa (375) ejecutándose en un primer sistema informático (180, 300, 370, 390) para acceder a datos a nivel de bloque almacenados en un volumen de almacenamiento de datos a nivel de bloque no local (155), el volumen de almacenamiento de datos a nivel de bloque (155) estando provisto por un segundo sistema de almacenamiento de datos distinto (165, 360) por medio de una o más redes (170, 185, 385) y estando conectado al primer sistema informático (180, 300, 370, 390) de manera que la primera copia de programa en ejecución (375) inicie las peticiones de acceso a datos para el volumen de almacenamiento de datos (155) por medio de interacciones con un primer dispositivo de almacenamiento a nivel de bloque lógico local al primer sistema informático (180, 300, 370, 390) que representa el volumen de almacenamiento de datos a nivel de bloque (155);responder automáticamente a una o más de las indicaciones recibidas del primer grupo de peticiones de acceso a datos por medio de la interacción con el segundo sistema de almacenamiento de datos (165, 360) en nombre de la primera copia de programa en ejecución (375), incluyendo la respuesta iniciar la realización de las peticiones de acceso a datos en el volumen de almacenamiento de datos a nivel de bloque (155) en el segundo sistema de almacenamiento de datos a nivel de bloque (165, 360);después de determinar que la primera copia de programa (155) ha dejado de estar disponible, identificar un tercer sistema informático (180, 300, 370, 390) en el que se está ejecutando una segunda copia del primer programa (375), siendo el tercer sistema informático distinto del primer sistema informático y del segundo sistema de almacenamiento de datos, la segunda copia del primer programa ya siendo ejecutada antes de que la primera copia del programa deje de estar disponible, y conectar el volumen de almacenamiento de datos a nivel de bloque (155) al tercer sistema informático (180, 300, 370, 390) de manera que la segunda copia de programa (375) tenga acceso al segundo dispositivo de almacenamiento a nivel de bloque lógico local al tercer sistema informático (180, 300, 370, 390) que representa el volumen de almacenamiento de datos a nivel de bloque (165, 360), en el que la identificación del tercer sistema informático (180, 300, 370, 390) en el que la segunda copia del primer programa (375) se está ejecutando y en el que la conexión del volumen de almacenamiento de datos a nivel de bloque (155) al tercer sistema informático (180, 300, 370, 390) se realiza automáticamente para mantener el acceso del primer programa (375) al volumen de almacenamiento de datos a nivel de bloque (165, 360), llevándose a cabo el mantenimiento automático del acceso en respuesta a la determinación automática de que la primera copia de programa (155) ha dejado de estar disponible;recibir una o más indicaciones de un segundo grupo de otras peticiones de acceso a datos iniciadas por la segunda copia de programa en ejecución (155) para el volumen de almacenamiento de datos a nivel de bloque (155) por medio de interacciones con el segundo dispositivo de almacenamiento a nivel de bloque local y lógico en el tercer sistema informático (180, 300, 370, 390);y responder automáticamente a las indicaciones recibidas del segundo grupo de peticiones de acceso a datos mediante la interacción con el segundo sistema de almacenamiento de datos (165, 360) en nombre de la segunda copia de programa en ejecución (375) para iniciar la realización de las peticiones de acceso a datos del segundo grupo en el volumen de almacenamiento de datos a nivel de bloque (155) en el segundo sistema de almacenamiento de datos (165, 360).
- 9El medio legible por ordenador de la reivindicación 8 en el que la primera copia de programa (375) y la segunda ES 2 575 155 T3 copia de programa (375) son copias en ejecución de un único programa (375), las peticiones de acceso a datos para el volumen de datos a nivel de bloque (155) mediante la primera y segunda copia de programa (375) son para acceder a los datos a nivel de bloque almacenados en el volumen de almacenamiento de datos a nivel de bloque (155), el primer y tercer sistema informático (180, 300, 370, 390) y el segundo sistema de almacenamiento de datos a nivel de bloque (165, 360) están ubicados en una única ubicación geográfica, el primer y segundo sistema informático (180, 300, 370, 390) y el segundo sistema de almacenamiento de datos a nivel de bloque (165, 360) están separados por una o más redes (170, 185, 385), la realización de las peticiones de acceso a datos y las otras peticiones de acceso a datos en el volumen de almacenamiento de datos a nivel de bloque (155) en el segundo sistema de almacenamiento de datos a nivel de bloque (165, 360) se realizan automáticamente mediante un módulo de administrador de sistema (340) de un servicio de almacenamiento de datos a nivel de bloque, un módulo de administrador de nodo (380) del servicio de almacenamiento de datos a nivel de bloque está asociado al primer sistema informático (180, 300, 370, 390) y gestiona el acceso de la primera copia de programa en ejecución (375) a una o más redes (170, 185, 385), y la respuesta automática a las indicaciones recibidas del primer grupo de peticiones de acceso a datos se realiza en parte bajo el control del módulo de administrador de nodo (380) al iniciar el envío de las peticiones de acceso a datos recibidas por medio de una o más redes (170, 185, 385) hacia el segundo sistema de almacenamiento de datos a nivel de bloque (165, 360).
- 10El medio legible por ordenador de la reivindicación 9 en el que el módulo de administrador de nodo (380) además gestiona el acceso de la segunda copia de programa en ejecución (375) a una o más de las redes (170, 185, 385), el tercer y primer sistema informático (180, 300, 370, 390) son parte de un único sistema informático físico, y la respuesta automática a las indicaciones recibidas del segundo grupo de otras peticiones de acceso a datos se realiza en parte bajo el control del módulo de administrador de nodo (380) al iniciar el envío de las otras peticiones de acceso a datos por medio de una o más redes (170, 185, 385) hacia el sistema de almacenamiento de datos a nivel de bloque (165, 360).
- 11El medio legible por ordenador de la reivindicación 8 en el que el medio legible por ordenador es al menos una de las memorias (330, 354, 374) de un sistema informático (180, 300, 370, 390) que almacena los contenidos y un medio de transmisión de datos que incluye una señal de datos almacenada y generada que contiene los contenidos, o en el que los contenidos son instrucciones que al ejecutarse hacen que uno o más sistemas informáticos (180, 300, 370, 390) lleven a cabo el método.
- 12Un sistema configurado para gestionar el acceso a la funcionalidad de almacenamiento de datos a nivel de bloque mediante programas en ejecución, que comprende:una o más memorias (330, 354, 374);y un sistema de almacenamiento de datos a nivel de bloque (340) configurado para proveer un servicio de almacenamiento de datos a nivel de bloque que utiliza múltiples sistemas de almacenamiento de datos a nivel de bloque (165, 360) para almacenar volúmenes de almacenamiento de datos a nivel de bloque (155) creados por usuarios del servicio de almacenamiento de datos a nivel de bloque y a los que se accede en nombre de uno o más programas en ejecución (375) asociados a los usuarios, el suministro del servicio de almacenamiento de datos a nivel de bloque incluyendo: crear uno o más volúmenes de almacenamiento de datos a nivel de bloque (155) para ser utilizados por uno o más programas en ejecución (375), estando cada uno de los volúmenes de almacenamiento de datos a nivel de bloque (155) almacenados en uno de los múltiples sistemas de almacenamiento de datos a nivel de bloque (165, 360);recibir una o más indicaciones de un primer grupo de peticiones de acceso a datos iniciadas por una primera copia de un primer programa (375) ejecutándose en un primer sistema informático (180, 300, 370, 390) para acceder a datos a nivel de bloque almacenados en un volumen de almacenamiento de datos a nivel de bloque no local (155) creado de uno o más de los volúmenes de almacenamiento de datos a nivel de bloque, estando el primer volumen de almacenamiento de datos a nivel de bloque (155) creado provisto por un segundo sistema de almacenamiento de datos distinto (165, 360) por medio de una o más redes (170, 185, 385) y estando conectado a un primer sistema informático (180, 300, 370, 390) de manera que la primera copia de programa en ejecución (375) inicie las peticiones de acceso a datos para el primer volumen de almacenamiento de datos a nivel de bloque (155) creado por medio de interacciones con un primer dispositivo de almacenamiento a nivel de bloque lógico local al primer sistema informático (180, 300, 370, 390) que representa al primer volumen de almacenamiento de datos a nivel de bloque (155) creado;responder automáticamente a las indicaciones recibidas del primer grupo de peticiones de acceso a datos mediante la interacción con el segundo sistema de almacenamiento de datos (165, 360) en nombre de la segunda copia de programa en ejecución (375) para iniciar la realización de las peticiones de acceso a datos del primer grupo en el primer volumen de almacenamiento de datos a nivel de bloque (155) creado provisto por el segundo sistema de almacenamiento de datos (165, 360);después de determinar que la primera copia de programa (375) ha dejado de estar disponible, identificar un tercer sistema informático (180, 300, 370, 390) en el que se está ejecutando una segunda copia del primer programa (375), siendo el tercer sistema informático distinto del primer sistema ES 2 575 155 T3 informático y del segundo sistema de almacenamiento de datos, la segunda copia del primer programa ya siendo ejecutada antes de que la primera copia del programa deje de estar disponible, y conectar el primer volumen de almacenamiento de datos a nivel de bloque creado (155) al tercer sistema informático (180, 300, 370, 390) de manera que la segunda copia de programa (375) tenga acceso a un segundo dispositivo de almacenamiento a nivel de bloque lógico local al tercer sistema informático (180, 300, 370, 390) que representa el primer volumen de almacenamiento de datos a nivel de bloque (155) creado, en el la identificación del tercer sistema informático (180, 300, 370, 390) en el que la segunda copia del primer programa (375) se está ejecutando y la conexión del primer volumen de almacenamiento de datos a nivel de bloque (155) creado para el tercer sistema informático (180, 300, 370, 390) se realizan automáticamente para mantener el acceso del primer programa (375) al primer volumen de almacenamiento de datos a nivel de bloque creado (155), llevándose a cabo el mantenimiento automático del acceso en respuesta a la determinación automática de que la primera copia de programa (375) ha dejado de estar disponible;recibir una o más indicaciones de un segundo grupo de otras peticiones de acceso a datos iniciadas por la segunda copia de programa en ejecución (375) para el primer volumen de almacenamiento de datos a nivel de bloque (155) creado por medio de interacciones con el segundo dispositivo de almacenamiento de almacenamiento a nivel de bloque local y lógico en el tercer sistema informático (180, 300, 370, 390);y responder automáticamente a las indicaciones recibidas del segundo grupo de peticiones de acceso a datos mediante la interacción con el segundo sistema de almacenamiento de datos (165, 360) en nombre de la segunda copia de programa en ejecución (375) para iniciar la realización de las peticiones de acceso a datos del segundo grupo en el primer volumen de almacenamiento de datos a nivel de bloque (155) creado en un segundo sistema de almacenamiento de datos (165, 360).
- 13El sistema de la reivindicación 12 en el que la primera copia de programa (375) y la segunda copia de programa (375) son copias en ejecución de un único programa (375), las peticiones de acceso a datos para el primer volumen de almacenamiento de datos a nivel de bloque (155) creado por la primera y segunda copia de programa (375) son para acceder a los datos a nivel de bloque almacenados en el primer volumen de almacenamiento de datos a nivel de bloque (155) creado, el primer volumen de almacenamiento de datos a nivel de bloque (155) creado está almacenado en un primero de los sistemas de almacenamiento de datos (165, 360), el primer y tercer sistema informático (180, 300, 370, 390) y el primer sistema de almacenamiento de datos a nivel de bloque (165, 360) están ubicados en una única ubicación geográfica y separados por una o más redes (170, 185, 385), y el sistema además comprende uno o más módulos de administradores de nodo (380) del servicio de almacenamiento de datos a nivel de bloque asociado a la primera y segunda copia de programa (375), estando uno o más de los módulos de administrador de nodo (380) configurados para gestionar el acceso de la primera y segunda copia de programa (375) a una o más redes (170, 185, 385), de manera que la respuesta a las peticiones de acceso a datos y a otras peticiones de acceso a datos se lleven a cabo en parte bajo el control de uno o más de los módulos de administrador de nodo (380) al iniciar el envío de las peticiones de acceso a datos y de las otras peticiones de acceso a datos por medio de una o más redes (170, 185, 385) hacia el primer sistema de almacenamiento de datos a nivel de bloque (165, 360).
- 14El sistema de la reivindicación 12 en el que una primera copia del primer volumen de almacenamiento de datos a nivel de bloque (155) creado está almacenado en un primer sistema de almacenamiento de datos a nivel de bloque (165, 360), y el primer sistema informático (180, 300, 370, 390) está co-localizado con el primer sistema de almacenamiento de datos a nivel de bloque (165, 360) en una primera ubicación geográfica.
- 15El sistema de la reivindicación 12 en el que el primer sistema informático (180, 300, 370, 390) incluye al menos una de una o más memorias (330, 354, 374), y el módulo de administrador de sistema de almacenamiento de datos a nivel de bloque (340) incluye instrucciones de software para su ejecución mediante el primer sistema informático (180, 300, 370, 390) que utiliza al menos una de una o más de las memorias (330, 354, 374).
Independent claims15
242 paragraphs in 12 sections, as filed
ES 2 575 155 T3
DESCRIPTION
Providing Execution Programs with Reliable Access to Non-Local Block Level Data Storage BACKGROUND
Many companies and other organizations operate computer networks that interconnect numerous computer systems to support their operations, such as with computer systems that are co-located (for example, as part of a local network) or that are instead located in multiple different geographic locations (For example, connected by one or more private or public intermediate networks). For example, data centers that host significant numbers of co-located interconnected computer systems have become commonplace and ordinary, such as private data centers that are operated by and on behalf of a single organization, and private data centers that are operated on behalf of a single organization. public data that is operated by entities such as companies. Some public data center operators provide secure network access, power, and installation services for hardware owned by multiple customers, while other public data center operators provide comprehensive services that also include hardware resources available for use. by your customers. However, as the scale and scope of typical data centers and computer networks have increased, the task of providing, administering, and managing the associated physical computing resources has become increasingly complicated.
The advent of virtualization technologies for basic hardware has provided some benefits regarding large-scale computing resource management for many customers with different needs, allowing multiple computing resources to be shared efficiently and securely among multiple customers. For example, virtualization technologies such as those provided by XEN, VMWare, or User-Mode Linux can allow a single physical computer system to be shared among multiple users by providing each user with one or more virtual machines hosted by the single physical computer system. , and each such virtual machine is a software simulation that acts as a different logical computing system that provides users with the illusion that they are the sole operators and administrators of certain hardware computing resource, while also providing application isolation and security between different virtual machines. Additionally, some virtualization technologies provide virtual resources that span one or more physical resources, such as a single virtual machine with multiple virtual processors that actually span multiple different physical computer systems.
The Red Hat Cluster Suite Summary: Red Hat Cluster Suite for Red Hat Enterprise Linux 5, dated January 1, 2007, pages 1-67, Raleigh, NC 27606-2072 USA, refers to a system and method based on storage clusters to provide a consistent filesystem image across all servers in a cluster, and allow servers to simultaneously read and write to a single shared filesystem. A storage cluster simplifies storage management by limiting application installation and patching to one file system. Also, with a cluster-level file system, a storage cluster eliminates the need for redundant copies of application data and simplifies backup and disaster recovery. The Red Hat Cluster Suite provides storage clustering through Red Hat GFS. In particular, this prior art document describes highly available service management that provides the ability to create and manage highly available cluster services in a Red Hat cluster.
This document describes a failover to the failure of a service running on a primary node and accessing a shared block-level storage system remote to a secondary node where a second service instance is started afterwards. to detect the failure of the first service instance.
US 2008/189468 A1 refers to a system that includes: (a) plural virtualization systems configured in a cluster; (b) accessible storage for each cluster virtualization system, where for each operational virtual machine in a cluster virtualization system, the storage maintains a representation of virtual machine state that includes at least one description of a virtualized hardware system and an image of the virtualized memory state for the virtual machine; and (c) a failover system that, in response to an outage of, or in, a particular system of virtualization systems, converts at least one affected virtual machine to another virtualization system in the cluster and continues with calculations. of the converted virtual machine according to the state encoded by a corresponding state of the virtual machine states represented in the storage.
BRIEF DESCRIPTION OF THE DRAWINGS
Figure 1 is a network diagram illustrating an exemplary embodiment in which multiple computer systems execute programs and access reliable non-local block-level data storage.
Figures 2A-2F illustrate examples of providing reliable non-local block-level data storage functionality to clients.
Figure 3 is a block diagram illustrating exemplary computer systems suitable for
ES 2 575 155 T3 manage the provision to and use by customers of reliable non-local block level data storage functionality.
Figure 4 illustrates a flow chart of an exemplary embodiment of a Block-level Data Storage System Manager routine.
Figure 5 illustrates a flow chart of an exemplary embodiment of a Node Manager routine.
Figure 6 illustrates a flow chart of an exemplary embodiment of a Block-level Data Storage Server routine.
Figures 7A-7B illustrate a flow chart of an exemplary embodiment of a Program Execution Services System Administrator routine.
Figure 8 illustrates a flow chart of an exemplary embodiment of a Block-level Data Storage Archive Manager routine.
DETAILED DESCRIPTION
The object of the present invention is to provide techniques for managing the access of running programs to data storage at the non-local block level.
The present object is solved by the invention as claimed in the independent claims.
Preferred embodiments are defined by the dependent claims.
Techniques for managing access by running programs to data storage at the non-local block level are described. In at least some embodiments, the techniques include providing a block-level data storage service that uses multiple server storage systems to reliably store block-level data that is accessible and usable in one or more networks by programs running on other physical computer systems. Users of the block-level data warehouse service can each create one or more block-level data warehouse volumes that each have a specified amount of block-level data warehouse space, and can initiate the use of said block-level data storage volume (also referred to as a volume herein) by one or more running programs, having at least some of said volumes copies stored by two or more of the multiple server storage systems to improve the reliability and availability of the volume for running programs. By way of example, the multiple server block level data storage systems that store data at the block level may, in some embodiments, be organized into one or more sets or other groups each having multiple physical systems. server storage co-located in a geographic location, such as in each of the single or more geographically distributed data centers, and the program (s) that use a volume stored in a server block-level data storage system in a data center may run on one or more different physical computer systems in that center of data. The following are additional details regarding the implementations of a block-level data storage service, and at least some of the techniques described for providing a block-level data storage service may be performed automatically by the implementations of a Block Level Data Storage System (BDS) Administrator module.
Furthermore, in at least some embodiments, running programs that access and use one or more of said non-local block-level data storage volumes on one or more networks may each have an associated node manager that manages access to such non-local volumes by the program, for example, a node manager module that is provided by the block-level data storage service and / or that operates in conjunction with one or more BDS System Manager modules. For example, a first user who is a customer of the block-level data warehouse service can create a first block-level data warehouse volume, and run one or more program copies on one or more computing nodes to which it is located. that they are commanded to access and use the first volume (eg, serially, simultaneously, or otherwise overlapping, etc.). When a program running on a compute node initiates use of a non-local volume, the program can be mounted or otherwise provided with a logical block-level data storage device that is local to the compute node. y representing the non-local volume such as to allow the running program to interact with the local logical block level data storage device in the same way as any other local hard disk or other physical block level data storage device that is connected to the compute node ( For example, to perform data read and write access requests, to implement a file system or database or other higher-level data structure on the volume, etc.). For example, in at least some embodiments, a representative local, logical block level data storage device may be available to a running program through the use of GNBD (Global Network Block Device) technology. In addition, as discussed in greater detail below, when the running program interacts with the representative, local, logical block level data storage device, the associated node manager can manage those interactions.
ES 2 575 155 T3 communicating on one or more networks with at least one of the server block-level data storage systems that stores a copy of the associated non-local volume (For example, in a transparent way for the running program and / or the computing node) to perform the interactions on said volume copy stored on behalf of the running program. Furthermore, in at least some embodiments, at least some of the techniques described for managing access of running programs to non-local block-level data storage volumes are performed automatically by embodiments of a Node Manager module.
Furthermore, in at least some embodiments, at least some block-level data storage volumes (or portions of such volumes) may further be stored in one or more remote archive storage systems that are other than data storage systems. server-side block-level used to store volume copies. In various embodiments, one or more remote archive storage systems may be provided by the block-level data storage service (for example, at a remote location from a data center or other geographic location that has a set of co-located server block-level data storage systems) or, instead, can be provided by a remote long-term storage service and used by block-level data storage and, in at least some embodiments, the archive storage system may store data in a different format than block-level data. block (For example, you can store one or more chunks or portions of a volume as separate objects.) Such archive storage systems can be used in various ways in various embodiments to provide various benefits, as discussed in greater detail below. In some embodiments where a remote long-term storage service provides the archiving storage systems, the users of the block-level data storage service (For example, customers of the block-level data storage service that pay fees to use the block-level data storage service) who are also users of the remote long-term storage service (For example, remote long-term storage service customers who pay fees to use the remote long-term storage service) may have at least portions of their block-level data storage volumes stored by archival storage systems, such as example in response to instructions from those customers. In other embodiments, a single organization may provide at least some of the capabilities of the block-level data storage service and the capabilities of the remote long-term storage service (for example, in an integrated manner such as, for example, part of a single service), while in still other embodiments the block-level data storage service can be provided in environments that do not include the use of archival data storage systems. Furthermore, in at least some embodiments, the use of the archive storage systems is performed automatically under the control of one or more archive manager modules, such as an archive manager module provided by the archive storage service. block-level data or otherwise provided to operate in conjunction with block-level data storage service modules (for example, provided by the remote long-term storage service to interact with the block-level data storage service).
In some embodiments, at least some of the disclosed techniques are performed on behalf of a program execution service that manages the execution of multiple programs on behalf of multiple program execution service users. In some embodiments, the program execution service may have groups of multiple physical host computer systems co-located in one or more geographic locations, such as in one or more geographically distributed data centers, and may execute user programs in such physical host computer systems, such as under the control of a program execution service system (PES) administrator, as mentioned in more detail below. In such embodiments, program execution service users (e.g., program execution service customers who pay fees to use the program execution service) who are also block-level data storage service users may run programs that access and use non-local block-level data storage volumes provided by the block-level data warehouse service. In other embodiments, a single organization may provide at least some of the capabilities of the program execution service and the capabilities of the block-level data storage service (e.g., in an integrated manner, such as part of a only service), while in still other embodiments the block-level data storage service can be provided in environments that do not include a program execution service (For example, internally to a business or other organization to support the organization's operations) .
In addition, the host computer systems on which the programs run can have various forms in various embodiments. Such multiple host computer systems can, for example, be co-located in one physical location (e.g., a data center), and can be managed by multiple node manager modules that are each associated with a subset of one or more of the host computer systems. At least some of the host computer systems may each include sufficient computing resources (For example, non-permanent memory, CPU cycles, or other measure of CPU usage, network bandwidth, swap space, etc.) to run multiple programs simultaneously and, in at least some
In embodiments, some or all computer systems may each have one or more physically connected local block-level data storage devices (e.g., hard drives, magnetic tape drives, etc.) that they can be used to store local copies of programs to be executed and / or data used by those programs. Furthermore, at least some of the host computer systems in some of said embodiments may each host multiple virtual machine computing nodes that may each run one or more programs on behalf of a different user, each computer system having a host a running hypervisor or other virtual machine monitor that manages the virtual machines for that host computer system. For host computer systems running multiple virtual machines, the associated node manager module for the host computer system may, in some embodiments, run on at least one of multiple hosted virtual machines (For example, as part of or in conjunction with the virtual machine monitor for the host computer system), while in other situations a node manager may run on a different physical computer system than one or more different host computer systems being managed.
The server block level data storage systems in which the volumes are stored may also take various forms in various embodiments. As previously mentioned, multiple server block-level data storage systems can, for example, be co-located in one physical location (e.g. a data center), and can be managed by one or more modules of BDS System Administrator. In at least some embodiments, some or all of the server block-level data storage systems may be physical computer systems similar to the host computer systems that run programs, and in some of those embodiments, they may each run software from the host. server storage system to assist in provisioning and maintenance of volumes on these server storage systems. For example, in at least some embodiments, one or more of said server block level data storage computer systems may run at least part of the BDS System Manager, as if, for example, one or more Manager modules of BDS System were provided in a distributed way between pairs (peer-to-peer) by multiple computerized data storage systems at the block level in server that interact. In other embodiments, at least some of the server block-level data storage systems may be network storage devices that may lack some I / O components and / or other physical computer system components, as if, for For example, at least part of the provisioning and maintenance of volumes on such server storage systems was performed by other remote physical computer systems (For example, by a BDS System Administrator module running on one or more different computer systems). Furthermore, in some embodiments, at least some server block level data storage systems each maintain multiple local hard drives, and strip at least some volumes along a portion of each of some or all local hard drives. Additionally, various types of techniques can be used to create and use volumes, including, in some embodiments, the use of LVM (Logical Volume Manager) technology.
As previously mentioned, in at least some embodiments, some or all of the block-level data storage volumes each have copies stored on two or more different server-side block-level data storage systems, such as, for example, to improve the reliability and availability of volumes. Through such action, the failure of a single server block-level data storage system may not cause access by running programs to a volume to be lost, since the use of said volume by such running programs may change. to another available server block-level data storage system that has a copy of that volume. In such embodiments, consistency between multiple copies of a volume across multiple server-side block-level data storage systems can be maintained in a number of ways. For example, in some embodiments, one of the server-side block-level data storage systems is designated as a system that stores the primary copy of the volume, and the other block-level data storage system or systems at servers are designated as systems that store mirror copies of the volume - in such embodiments, the server block-level data storage system that has the primary copy of the volume (referred to as the primary server block-level data storage system for the volume) can receive and handle access requests to data for the volume, and in some of these embodiments it can further act to maintain consistency of the other volume mirrors (For example, by sending update messages to the other server block level data stores that provide volume mirror copies when the data on the primary volume copy is modified, such as in the form of a master-slave computing relationship). Various types of volume consistency techniques can be used, with additional details included below.
In at least some embodiments, the disclosed techniques include providing reliable and available access by a program running on a computing node to a block-level data storage volume by managing the use of the volume's primary and mirror copies. For example, the node manager for the running program can, in some embodiments, interact only with the primary copy of the volume through the block-level data storage system on the primary server, such as if the primary copy of volume was responsible for maintaining the volume mirror copies or if another
ES 2 575 155 T3 replication mechanism. In such embodiments, if the primary server block-level data storage system fails to respond to a request sent by the node administrator (for example, a request for data access initiated by the running program, a ping message or other request initiated by the node administrator to periodically verify that the block-level data storage system on the primary server is available, etc.) within a predetermined period of time, or if the node administrator is otherwise alerted that the primary volume copy is not available (For example, by a message from the BDS System Administrator), the node administrator node can automatically switch its interactions to one of the volume mirrors on a mirror server block-level data storage system (For example, the running program ignores such a change, apart from waiting possibly a slightly longer time to obtain a response to a request for access to data made by the running program if it was said request for access to data that had expired and initiated the change to the volume mirror copy). The volume mirror copy can be selected in several ways, such as, for example, if it were the only one, if an order in which to access multiple volume mirror copies was previously indicated, interacting with the BDS System Administrator to request a indication of which volume mirror is promoted to act as the primary volume copy, etc. In other embodiments, at least some volumes may have multiple primary copies, as if, for example, one volume is available for simultaneous read access by multiple running programs and the resulting data access load is spread across multiple primary volume copies - in such embodiments, a node administrator can select one of multiple primary volume copies to interact with in various ways (for example, randomly, according to an instruction from a BDS System Administrator module, etc.).
In addition, the BDS System Administrator can take various actions in various embodiments to maintain reliable and available access of a program running on a computing node to a block-level data storage volume. In particular, if the BDS System Administrator becomes aware that a particular server-side block-level data storage system (or a particular volume on a particular server-side block-level data storage system) stops working. be available, the BDS System Administrator may take various actions for some or all volumes stored by that server-side block-level data storage system (or for the particular unavailable volume) to maintain their availability. For example, for each volume primary copy stored in the unavailable server block-level data storage system, the BDS System Administrator can promote one of the existing volume mirrors to be the new primary copy of volume, and optionally notify one or more node administrators of the change (for example, the node administrators for any running programs that are currently using the volume). Additionally, for each stored volume copy, the BDS System Administrator can initiate the creation of at least one new mirror copy of the volume on a different server-side, block-level data storage system, such as through replication. of an existing copy of the volume on another available server block-level data storage system that has an existing copy (for example, by replicating the primary volume copy). In addition, in at least some embodiments, other benefits can be achieved in at least some situations by using at least portions of a volume that are stored on remote archive storage systems to assist in the replication of a new mirror copy of the volume (e.g. , higher data reliability, ability to minimize an amount of storage used for volume mirrors and / or continuous processing power to maintain full volume mirrors, etc.), as discussed in more detail below.
The BDS System Administrator may become aware of the unavailability of a server-side block-level data storage system in various ways such as, for example, by a message from a node administrator that he cannot contact the server-side storage system. server block level data, according to a message from the server block level data storage system (For example, to indicate that it has experienced an error condition, has initiated a shutdown or failure mode operation, etc.), based on an inability to contact the server block-level data storage system (For example, based on periodic or constant monitoring of some or all data storage systems block-level data on server), etc. Furthermore, the unavailability of a server block-level data storage system can be caused by several episodes in various embodiments, such as the failure of one or more hard drives or other storage media in which the storage system block-level data server stores at least a portion of one or more volumes, the failure of one or more different components of the server block-level data storage system (for example, CPU memory, a fan, etc.), a power failure in the data storage system to server block level (For example, a power failure in a single server block level data storage system, in a rack of multiple server block level data storage systems, across a data center, etc.), a network failure, or other communication that prevents the server block-level data storage system from communicating with a node administrator and / or the BDS System Administrator , etc. In some embodiments, failure of or problems with any component of a server-side block-level data storage system may be considered an unavailable condition for the entire server-side block-level data storage system. For example, in embodiments where a server block-level data storage system maintains multiple local hard drives, the failure of (s)
ES 2 575 155 T3 problems with any of the local hard drives can be considered an unavailability condition for the entire server block level data storage system), while in other embodiments a server level data storage system Block on server will not be considered unavailable as long as it can respond to data access requests.
Additionally, apart from moving one or more volumes from an existing server-side block-level data storage system when that server-side block-level data storage system becomes unavailable, the BDS System Administrator can, in some realizations, decide to move one or more volumes from an existing server-side block-level data warehouse to a different server-side block-level data warehouse and / or decide to create a new copy of one or more volumes at various times different and for different reasons. Such movement or creation of a new copy of a volume can be done in a manner similar to that mentioned in greater detail elsewhere in this document (For example, by replicating the primary copy of the volume to create a new copy, and by optional deletion of the previous volume copy in at least some situations (such as when the volume copy is moving). Situations that can cause a volume move or the creation of a new volume copy include, for example, the following non-exclusive list: a particular server block-level data storage system can become overused (e.g. based on CPU usage, network bandwidth, I / O access, storage capacity, etc.) such as, for example, to trigger the movement of one or more volumes from said server block level data storage system; A particular server block level data storage system may lack sufficient resources for a desired modification of an existing volume (For example, it may lack sufficient available storage space if the size of an existing volume is required to be expanded ) such as, for example, to trigger the movement of one or more volumes from said data storage system at the server block level; a particular server block-level data storage system may require maintenance or updates that will make it unavailable for a period of time, such as to trigger the temporary or permanent movement of one or more volumes from that system data storage at the block level on the server; based on the recognition that usage patterns for a particular volume or other characteristics of a volume may be better accommodated in other server-side block-level data storage systems, such as another server-level data storage system. server block with additional capabilities (For example, for volumes that have frequent data modifications, use a block-level data storage system on the primary server with disk write capacities above average capacities, and / or for volumes that are very large in size, use a block-level data storage system on primary server with a storage capacity above average capacity); in response to a request from a user who has created or is otherwise associated with a volume (For example, in response to the user's purchase of premium access to a server-side block-level data storage system that has enhanced capabilities); to provide at least one fresh copy of a volume in a different geographic location (for example, another data center) in which programs run, such as to trigger movement and / or copying of the volume from a system block-level data storage on the server in a first geographic location when the use of a volume is requested by a program running in another geographic location; etc.
Additionally, after a volume has been moved or a new copy has been created, the BDS System Administrator may, in some embodiments and situations, update one or more node administrators, as appropriate (For example, only node administrators for running programs that are currently using the volume, all node managers, etc.). In other embodiments, different information about volumes can be maintained in another way, such as by using one or more copies of a volume information database that is accessible on the network to node administrators and / or the Administrator of the volume. BDS system. A non-exclusive list of types of information about volumes that can be maintained includes the following: an identifier for a volume, such as an identifier that is unique to server-side block-level data stores that store copies of the volume or that it is universally unique to the block-level data storage service; access-restricted information for a volume, such as passwords or encryption keys, or lists or other indications of authorized users for the volume; information about the primary server block-level data storage system for the volume, such as a network address and / or other access information; information about one or more mirror-server block-level data storage systems for the volume, such as information about an order indicating which mirror-server block-level data storage system will be promoted to be the system primary if the existing primary server storage system becomes unavailable, a network address and / or other login information, etc .; information about any volume snapshot that was created for the volume, as described in more detail below; information on whether or not the volume will be available to users other than the creator of the volume, and, if so, under what circumstances (for example, for read-only access, for other users to make their own volumes that are copies of it volume, pricing information for other users to receive various types of volume access); etc.
ES 2 575 155 T3
In addition to maintaining reliable and available access by running programs to block-level data storage volumes by moving or replicating volume copies when server-side block-level data storage systems are out of service. available, the block-level data warehouse service can take other actions in other situations to maintain access by running programs to the block-level data warehouse volumes. For example, if a first running program becomes unavailable unexpectedly, in some embodiments the block-level data warehouse service and / or the program execution service can take actions to cause a different, second running program (For example, a second copy of the same program that is running on a different host computer system) connects to some or all of the block-level data storage volumes that were in use by the first unavailable program, so that the second program can quickly take over at least some operations from the first unavailable program. The second program may, in some situations, be a new program whose execution starts due to the unavailability of the first existing program, while in other situations the second program may already be running (For example, if multiple copies of the program are run simultaneously to share a total workload, such as multiple Web server programs receiving different incoming requests from clients through a load balancer, with one of the multiple program copies selected to be the second program; if the second program is a standby copy of the running program to allow a hot swap from the first existing program in case of unavailability, such as without the standby copy of the program being actively used until the non-availability of the first existing program occurs; etc.). Furthermore, in some embodiments, a second program in which the connection and continued use of an existing volume are switched may be on another host physical computer system in the same geographic location (for example, the same data center) as the first program. , while in other embodiments the second program may be in a different geographic location (For example, a different data center such as, for example, in conjunction with a copy of the volume that has previously or simultaneously been moved to said other data center and that will be used by said second program). Furthermore, in some embodiments, other related actions may be performed to further facilitate the switch to the second program, such as redirecting some communications directed to the first unavailable program to the second program.
In addition, in at least some embodiments, other techniques can be used to provide reliable and available access to block-level data storage volumes, as well as other benefits, such as allowing a copy of a specified volume to be saved. for one or more remote archive storage systems (For example, in a second geographic location that is remote from a first geographic location where server block-level data storage systems store the active primary and mirror copies of the volume and / or that is remote from physical computer systems hosts running the programs that use the volume) such as for long-term backup and / or other purposes. For example, in some embodiments, the archive storage systems may be provided by a network accessible remote storage service. Also, the copies of a volume that are saved for archive storage systems may, in at least some situations, be snapshots of the volume at a particular point in time, but not automatically updated as continued use of the volume makes that their stored block-level data contents change, and / or that they are not available to be connected to and used by running programs in the same way as volumes. Thus, by way of example, a long-term snapshot of a volume may be used, for example, as a backup of a volume, and may further, in some embodiments, serve as the basis for one or more new volumes that are created from the snapshot (for example, so that the new volumes start with the same block-level data storage contents as the snapshot).
Additionally, shadow copies of a volume on archive storage systems can be stored in various ways, such as to represent smaller chunks of a volume (For example, if archive storage systems store data as larger objects). rather than a large, linear, sequential data block). For example, a volume may be represented as a series of multiple smaller chunks (with a volume that is, for example, one gigabyte or terabyte in size, and with a chunk that is, for example, 1⁄4 '' in size). a few megabytes), and information about some or all of the fragments (for example, each fragment that is modified) can be stored separately in archival storage systems such as, for example, treating each chunk as a separate stored object. In addition, in at least some embodiments, a second subsequent snapshot of a particular volume can be created such that only incremental changes from a previous snapshot of the volume are stored, such as by including stored copies. of new storage chunks that have been created or modified since the previous snapshot, but sharing stored copies of some previously existing fragments with the previous snapshot if those fragments have not changed. In such embodiments, if a previous snapshot is subsequently deleted, previously existing fragments stored by said earlier snapshot that are shared by subsequent snapshots can be retained for use by said subsequent snapshots, while the fragments
ES 2 575 155 T3 previously existing unshared stored by said previous snapshot can be deleted.
Also, in at least some embodiments, when a snapshot of a volume is created at one point in time, access to the primary copy of the volume by running programs can be allowed to continue, including permission to modify the data. stored on the primary volume copy, but without such ongoing data modifications being reflected in the snapshot, such as if the snapshot was based on volume chunks stored in the archive storage systems that are not updated once the snapshot creation had started until the snapshot creation was complete. For example, in at least some embodiments, copy-on-write techniques are used when the creation of a snapshot of a volume is initiated and a chunk of the volume is later modified, such as to initially keep stored copies of the unused chunk. Modified and Modified Chunk in the primary server block-level data storage system that stores the primary copy of the volume (and optionally, also in mirror server block-level data storage systems that store one or more mirrored copies of the volume). When confirmation is received that the archive storage systems have successfully stored the snapshot of the volume (including a copy of the unmodified shard), the copy of the unmodified shard on the server-side block-level data storage system primary (and optionally in mirror server block-level data storage systems) can be erased.
Furthermore, such volume chunks or other volume data stored in the archive storage systems can be used in other ways in at least some embodiments, such as, for example, to use the archive storage systems as a backing store for the archives. primary and / or mirror copies of some or all volumes. For example, volume data stored on archive storage systems can be used to help maintain consistency between multiple copies of a volume on multiple server-side block-level data stores in at least some situations. By way of example, one or more mirrors of a volume can be created or updated based on, at least partially, volume fragments stored in archive storage systems, such as to minimize or eliminate a need to access the file. primary volume copy to get at least some of the volume chunks. For example, if the primary volume copy is updated faster or more reliably than the modified chunks on the archive storage systems, a new volume mirror copy can be created using at least some volume chunks stored on the storage systems. archive files that are known to be accurate (for example, from a recent volume snapshot), and accessing the primary volume copy only to obtain portions of the volume that correspond to fragments that may have been modified after the creation of the volume snapshot. Similarly, if modified shards on archive storage systems reliably reflect a current state of a volume primary copy, a volume mirror can be updated using those modified shards rather than through interactions by part of or with the primary volume copy.
Also, in some embodiments, the amount of data that is stored on a volume mirror (and the resulting size of the volume mirror) may be much smaller than that of the primary copy of the volume, such as if the volume information in archive storage systems was used in place of at least some data that would otherwise be stored in said minimum volume mirror. By way of example, once a snapshot of a volume is created on one or more archive storage systems, a minimal mirror of a volume does not need, in such embodiments, to store the volume data that is present in the instant volume copy. While modifications are made to the primary volume copy after the snapshot is created, some or all of the data modifications can also be made to the minimum volume mirror (For example, all data modifications, only data modifications). data modifications that are not reflected in modified volume chunks stored in archive storage systems, etc.); so if access to the minimum volume mirror is then necessary as, for example, if the minimum volume mirror is promoted to be the primary volume copy, the other data missing from the minimum volume mirror ( For example, unmodified portions of the volume) can be stored by retrieving them from archive storage systems (for example, from the previous volume snapshot). In this way, volume reliability can be improved, while also minimizing the amount of storage space used in server-side block-level data storage systems for volume mirroring.
In still other embodiments, the final or master copy of a volume can be maintained in the archive storage systems, and the primary and mirror copies of the volume can reflect a cache or other subset of the volume (for example, a subset that recently accessed and / or expected to be accessed soon); In such embodiments, the non-local block-level data storage volume of the block-level data storage service can be used to provide a closer volume data source for access by running programs that remote archive storage systems. Furthermore, in at least some of such embodiments, a volume can be described in
ES 2 575 155 T3 is viewed by users as having a particular size that corresponds to the master copy kept in the archive storage systems, but with the primary and mirror copies of a smaller size. Furthermore, in at least some of these embodiments, lazy update techniques can be used to immediately update a copy of data in a first data store (for example, a primary volume copy in a data storage system to block level at server) but to update the copy of the same data in a second different data store (for example, archive storage systems) later, such as to maintain strict data consistency in the second data store by ensuring that write operations or other data modifications to a portion of a volume are updated in the second data store before any subsequent read or other data access is performed. that portion of the volume from the second datastore (for example, using lazy-write cache update techniques). Such lazy update techniques can be used, for example, when updating changed chunks of a volume on archive storage systems, or when updating a volume mirror from changed chunks of the volume that are stored on storage systems. of archiving. In other embodiments, other techniques can be used when updating modified chunks of a volume on archive storage systems, such as using write-away cache techniques to immediately update the copy of data in the second data store (For example, For example, in archive storage systems) when the copy of the data in the first data store is modified.
Such volume snapshots stored on archive storage systems also offer several benefits. For example, if all primary and mirror copies of a volume are stored in multiple server-side block-level data storage systems in a single geographic location (for example, a data center), and the computer and storage systems in that geographic location they are no longer available (For example, electricity is cut in an entire data center), the existence of a recent snapshot of the volume in a different remote storage location can ensure that a recent version of the volume is available when the computer and storage systems in the geographic location become available again later (for example, when power is turned off). reset), such as if data is lost from one or more server storage systems in the geographic location. In addition, in such a situation, one or more new copies of the volume may be created in one or more new geographic locations based on a recent long-term snapshot of the volume from remote archive storage systems, such as to allow that one or more copies of the program running outside of an unavailable geographic location access and use those new volume copies. Additional details regarding archive storage systems and their use are included below.
As previously mentioned, in at least some embodiments, some or all of the block-level data storage volumes each have copies stored in two or more different server-side block-level data storage systems in a single server. geographic location, such as within the same data center where running programs will access the volume; By locating all volume copies and running programs in the same data center or other geographic location, various desired data access characteristics can be maintained (For example, based on one or more internal networks in that data center or other geographic location ) such as latency and throughput. For example, in at least some embodiments, the disclosed techniques can provide non-local block-level data storage access that has access characteristics similar to, or better than, the access characteristics of data storage devices at the block level. local physical block, but with much more reliability similar to, or exceeding, the reliability characteristics of dedicated RAID (Redundant Set of Independent / Inexpensive Disks) and / or SANs (Storage Area Networks) systems at a much lower cost. In other embodiments, the primary and mirror copies for at least some volumes may instead be stored in another way, such as in different geographic locations (For example, different data centers), such as to further maintain the availability of a volume even when an entire data center becomes unavailable. In embodiments where volume copies can be stored in different geographic locations, a user may, in some situations, request that a particular program be run near a particular volume (for example, in the same data center where it is locates the primary copy of the volume), or that a particular volume is located next to a particular running program, such as to provide relatively high network bandwidth and low latency for communications between the running program and the primary volume copy.
In addition, for at least some users, access to some or all of the described techniques may be provided, in some embodiments, on a fee or other form of payment. For example, users can pay one-time fees, recurring fees (For example, monthly), and / or one or more types of usage-based fees to use the block-level data storage service with the in order to store and access volumes, use the program execution service to run programs, and / or use archival storage systems (for example, provided by a remote long-term storage service) to store long-term backups or other snapshots of volumes. Rates can be determined according to one or more factors and activities, such as, for example, as indicated in the following non-exclusive list: depending on the size of a volume, such as creating the volume (For example,
ES 2 575 155 T3 a one-time fee), to have storage and / or continued use of the volume (For example, a monthly fee), etc .; according to characteristics other than the size of a volume, such as a number of mirrors, characteristics of data storage systems at the server block level (for example, data access rates, storage sizes, etc.) in which primary and / or volume mirror copies are stored; and / or a way in which the volume is created (For example, a new volume that is empty, a new volume that is a copy of an existing volume, a new volume that is a copy of a snapshot volume, etc. .); based on the size of a volume snapshot, such as to create the volume snapshot (For example, as a one-time fee) and / or to have regular volume storage (For example, a monthly fee); based on characteristics other than the size of one or more volume snapshots, such as a number of snapshots of a single volume, whether or not a snapshot is incremental relative to one or more previous snapshots, etc .; based on the usage of a volume, such as the amount of data transferred to and / or from a volume (for example, to reflect an amount of network bandwidth used), a number of data access requests sent to a volume, a number of running programs connecting to and using a volume (either sequentially or concurrently), etc .; based on the amount of data transferred to and / or from a snapshot, such as similar to that for volumes; etc. Furthermore, the access provided may take various forms in various embodiments, such as a one-time purchase fee, a regular rental fee, and / or under another regular subscription. In addition, in at least some embodiments and situations, a first group of one or more users may provide data to other users at a fee, such as charging the other users to receive access to current volumes and / or historical snapshots of volumes created by one or more users in the first group (For example, allowing them to make new volumes that are volume copies and / or volume snapshots; allowing them to use one or more created volumes; etc.), either on a one-time purchase rate, a regular rental rate, or another regular subscription.
In some embodiments, the block-level data storage service, the program execution service, and / or the remote long-term storage service may provide one or more APIs (Application Programming Interfaces) such as, to allow other programs to programmatically start various types of operations to be performed (For example, as directed by users of the other programs). Such operations may allow some or all of the types of functionality described above to be invoked, including, but not limited to, the following types of operations: create, delete, connect, disconnect, or describe volumes; create, delete, copy or describe snapshots; specify access rights or other metadata for volumes and / or snapshots; manage the execution of programs; make payment to obtain other types of functionality; obtain reports and other information about the use of the capabilities of one or more of the services and / or about the fees paid or due for such use; etc. The operations provided by the API may be invoked by, for example, programs running on host computer systems of the program execution service and / or by computer systems of clients or other users that are external to the one or more geographic locations used by the block-level data storage service and / or the program execution service.
For illustrative purposes, some embodiments are described below in which specific types of block-level data storage are provided in specific ways for specific types of programs that run on specific types of computer systems. These examples are provided for illustrative purposes and are simplified for brevity, and the inventive techniques can be used in a wide variety of different situations, some of which are described below, and the techniques are not limited to use with machines. virtual spaces, data centers or other specific types of data storage systems, computer systems or computer system arrangements. Furthermore, while some embodiments are described as providing and using reliable non-local block-level data storage, in other embodiments, types of data storage other than block-level data storage may be similarly provided.
Figure 1 is a network diagram illustrating an exemplary embodiment in which multiple computer systems run programs and access reliable non-local block-level data storage, such as under the control of a storage service. block-level data and / or a program execution service. In particular, in the present example, a program execution service manages the execution of programs on several host computer systems located within a data center 100 and a block-level data storage service uses multiple data storage systems. block-level servers on different servers in the data center to provide reliable non-local block-level data storage to such running programs. Multiple remote archive storage systems external to the data center can also be used to store additional copies of at least some portions of at least some volumes of block-level data storage.
In the present example, the data center 100 includes a number of racks 105, and each rack includes a number of host computer systems, as well as an optional rack support computer system 122 in the present exemplary embodiment. The host computer systems 110a-c in the illustrated rack 105 each house one or more virtual machines 120 in the present example, as well as an Administrator module.
ES 2 575 155 T3 of Different Node 115 associated with the virtual machines in said host computer system to manage said virtual machines. One or more different host computer systems 135 each also host one or more virtual machines 120 in the present example. Each virtual machine 120 can act as a separate computing node to run one or more program copies (not shown) for a user (not shown) such as a client of the program execution service. In addition, the present exemplary data center 100 further includes additional host computer systems 130a-b that do not include distinct virtual machines, but can nonetheless each act as a computing node for one or more programs (no show) that are being executed for a user. In the present example, a Node Manager module 125 that is running on a computer system (not shown) other than the host computer systems 130a-b and 135 is associated with said host computer systems to manage the computer nodes provided by said host computers. host computer systems such as, for example, in a manner similar to Node Manager modules 115 for host computer systems 110. The rack support computing system 122 can provide various utility services for other computer systems local to its rack 105 (For example, long-term program storage, measurement and other monitoring of program execution and / or storage access of non-local block-level data performed by other computer systems local to the rack, etc.), as well as possibly other computer systems located in the data center. Each computer system 110, 130 and 135 may also have one or more local storage devices attached (not shown) such as for storing local copies of programs and / or data created by or otherwise used by the programs in execution, as well as several different components.
In the present example, an optional computer system 140 is also illustrated, which runs a PES System Administrator module for the program execution service to assist in managing program execution on the computer nodes provided by the systems. computer hosts located within the data center (or, optionally, on computer systems located in one or more different data centers 160, or other remote computer systems 180 external to the data center). As described in greater detail elsewhere in this document, a PES System Administrator module can provide a variety of services in addition to managing program execution, including user account management (e.g. creation, deletion, billing, etc.); the registration, storage, and distribution of programs to be executed; the collection and processing of performance and audit data relating to the execution of programs; obtaining payment from clients or other users for the execution of programs; etc. In some embodiments, the PES System Administrator module may coordinate with the Node Administrator modules 115 and 125 to manage the execution of programs in computing nodes associated with the Node Administrator modules, while in other embodiments the Node Administrator modules Node Administrator 115 and 125 may not collaborate in the administration of said program execution.
The present exemplary data center 100 also includes a computer system 175 running a Block Level Data Storage System Manager (BDS) module. for block-level data storage service to help manage the availability of non-local block-level data storage to programs running on computing nodes provided by host computing systems located within the data center (or, optionally, in computer systems located in one or more different data centers 160, or other remote computer systems 180 external to the data center). In particular, in the present example, data center 100 includes a group of multiple server block-level data storage systems 165, each having local block-level storage for use in storing one or more copies of volume 155. Access to copies of volume 155 is provided on the internal network (s) 185 to programs running on computing nodes 120 and 130. As described in greater detail elsewhere in this document, a BDS System Administrator module can provide a variety of services related to providing non-local block-level data storage functionality, including management of account accounts. user (For example, creation, deletion, billing, etc.); the creation, use, and disposal of block-level data storage volumes and snapshots of those volumes; the collection and processing of performance and auditing data relating to the use of block-level data storage volumes and snapshots of such volumes; obtaining payments from customers or other users for the use of block-level data storage volumes and snapshots of those volumes; etc. In some embodiments, the BDS System Manager module can coordinate with Node Manager modules 115 and 125 to manage volume usage by programs running on associated computing nodes, while in other embodiments the System Manager modules Node 115 and 125 may not be used to manage such volume usage. Additionally, in other embodiments, one or more BDS System Administrator modules may be structured differently, such as to have multiple instances of the BDS System Administrator running in a single data center (For example, to share the management of data storage at the non-local block level by programs running on the computing nodes provided by the host computing systems located within the data center), and / or to have at least part of the functionality of a BDS System Administrator module provided in a distributed manner by software running on some or all server-side block-level data storage systems 165 (For example , peer
ES 2 575 155 T3 to-peer), without separate centralized BDS System Manager modules in a computer system 175).
In the present example, the different host computer systems 110, 130 and 135, the server block level data storage systems 165, and the computer systems 125, 140 and 175 are interconnected by one or more internal networks 185 of the center. data, which may include various network devices (eg routers, switches, gateways, etc.) that are not shown. Furthermore, the internal networks 185 are connected to an external network 170 (eg, the Internet or other public network) in the present example, and the data center 100 may further include one or more optional devices (not shown) in the interconnection. between data center 100 and an external network 170 (eg, network proxies, load balancers, network address translation devices, etc.). In the present example, data center 100 is connected via external network 170 to one or more different data centers 160 which may each include some or all of the computer systems and storage systems illustrated with respect to the data center. 100, as well as other remote computer systems 180 external to the data center. The other computer systems 180 may be operated by various parties for various purposes such as, for example, by the operator of the data center 100 or third parties (For example, customers of the program execution service and / or data storage service to block level). In addition, one or more of the other computer systems 180 may be archival storage systems (e.g., as part of a remote network accessible storage service) with which the block-level data storage service can interact as for example, under the control of one or more archive manager modules (not shown) running on the one or more 180 different computer systems, or instead on one or more computer systems in data center 100, as described in greater detail elsewhere in this document. Furthermore, while not illustrated herein, in at least some embodiments, at least some of the server block level data storage systems 165 may further be interconnected with one or more different networks or other connection means such as, For example, a high-bandwidth connection in which server storage systems 165 can share volume data (For example, in order to replicate volume copies and / or maintain consistency between the primary and mirror copies of volumes), said connection with high bandwidth is not available for the different host computer systems 110, 130 and 135 in at least some of said accomplishments.
It will be appreciated that the example in Figure 1 has been simplified for the purpose of explanation, and that the number and organization of host computer systems, server block-level data storage systems, and other devices can be very large. larger than illustrated in Figure 1. For example, according to an illustrative embodiment, there may be approximately 4,000 computer systems per data center, at least some of said computer systems are host computer systems that can each host 15 virtual machines, and / or some of said computer systems are server-side block-level data storage systems that can each store multiple volume copies. If each hosted virtual machine runs a program, then that data center can run up to 60,000 program copies at the same time. In addition, hundreds or thousands (or more) of volumes can be stored on server block-level data storage systems, depending on the number of server storage systems, the size of the volumes, and the number of mirrors per server. volume. It will be appreciated that other numbers of computer systems, programs, and volumes may be used in other embodiments.
Figures 2A-2F illustrate examples of providing reliable non-local block level data storage functionality to clients. In particular, Figures 2A and 2B illustrate examples of server-side block-level data storage computer systems that can be used to provide reliable non-local block-level data storage functionality to clients (e.g., programs in execution), such as on behalf of a block-level data warehouse service, and Figures 2C-2F illustrate examples of the use of archive storage systems to store at least some portions of some block-level data storage volumes. In the present example, Figure 2A illustrates several server block-level data storage systems 165 each storing one or more copies of volume 155, such as each volume with a primary copy and at least one copy. mirror copy. In other embodiments, other arrangements may be used, as mentioned in greater detail elsewhere in this document, such as by multiple primary volume copies (For example, all primary volume copies being available for simultaneous read access by one or more programs) and / or by multiple volume mirrors. Exemplary server block level data storage systems 165 and volume copies 155 may, for example, correspond to a subset of server block level data storage systems 165 and volume copies Figure 155
1.
In the present example, the server storage system 165a stores at least three volume copies, including the primary copy 155A-a for volume A, a mirror copy 155B-a for volume B, and a mirror copy 155C-a. for volume C. One or more volume copies not illustrated in the present example may further be stored by the server storage system 165a, as well as by the other server storage systems 165. Another example server block level data storage 165b stores the primary copy 155B-b for volume B in the present example, as well as a copy
ES 2 575 155 T3 mirror 155D-b for volume D. In addition, the example server block-level data storage system 165n includes a mirror copy 155A-n of volume A and a primary copy 155D-n of volume D. Thus, if a running program (not shown) is connected to and using volume A, the node manager for that running program will interact with the server block level data storage system 165a to access the primary copy 155A-a for volume A, such as by server storage system software (not shown) running on server block level data storage system 165a. Similarly, for one or more running programs (not shown) connected to and using volumes B and D, the node manager (s) for the running program (s) will interact with server block level data storage systems 165b and 165n, respectively, to access primary copies 155B-b for volume B and 155D-n for volume D, respectively. In addition, other server block level data storage systems may also be present (For example, server block level data storage systems 165c-165m and / or 165o and others), and can store the primary copy volume for volume C and / or other primary copies and volume mirror, but are not shown in the present example. Thus, in the present example, each server block level data storage system can store more than one volume copy, and can store a combination of primary and volume mirror copies, although in other embodiments the volumes may be stored in other ways.
Figure 2B illustrates server block level data storage systems 165 similar to those in Figure 2A, but at a later point in time after the server storage system 165b of Figure 2A fails or stops. be available otherwise. In response to the unavailability of the server storage system 165b, and its stored primary copies of volume B and mirror copies of volume D, the stored volume copies of the server storage systems 165 in Figure 2B have been modified to maintain availability of volumes B and D. In particular, due to the unavailability of primary copy 155B-b of volume B, the previous mirror copy 155B-a of volume B on server storage system 165a has been promoted to be the new primary copy for volume B . Thus, if one or more programs have previously been connected to or otherwise interacting with the previous primary 155B-b copy of volume B when it became unavailable, those programs may have been automatically switched (For example, by node administrators associated with those programs) to continue ongoing interactions with the server block-level data storage system 165a to access the new primary copy 155B-a for volume B. In addition, a new mirror copy 155B-c for volume B has been created on server storage 165c.
Although the 155A-n mirror copy of volume A of the 165n server storage system in Figure 2A is not illustrated in Figure 2B for brevity, it continues to be available on the 165n server storage system along with the primary copy 155D-n from volume D, and thus any program that has previously been connected to or otherwise interacting with the primary copy 155D-n of volume D when the server storage system 165b became unavailable will continue to interact with said same primary copy 155D-n of the volume D on the server storage system on the server storage system 165n without modification. However, due to the unavailability of mirror copy 155D-b of volume D on the unavailable server storage system 165b, at least one additional mirror copy of volume D has been created in Figure 2B as, for example, the mirror copy 155D-o of volume D of the server storage system 165o. In addition, Figure 2B illustrates that at least some volumes can have multiple mirrors such as volume D which also includes a 155D-c mirror copy of the previously existing volume D (but not shown in Figure 2A) in server storage 165c.
Figures 2C-2F illustrate examples of the use of archive storage systems to store at least some portions of some block-level data storage volumes. In the present example, Figure 2C illustrates multiple server block level data storage systems 165 each storing one or more volume copies 155 as, for example, to correspond to data storage systems 165 The exemplary server block level data storage system illustrated in Figure 2A before the server block level data storage system 165b becomes unavailable. Figure 2C further illustrates multiple archive storage systems 180, which may, for example, correspond to a subset of the computer systems 180 of Figure 1. In particular, in the present example, Figure 2C illustrates server block-level data storage systems 165a and 165b of Figure 2A, although in the present example only the primary and mirror copies of volume B are illustrated for such storage systems. server block level data storage. As described with respect to Figure 2A, the server storage system 165b stores the primary copy 155B-b of volume B, and the server storage system 165a stores the mirror copy 155B-a of volume B.
In the example of Figure 2C, a user associated with volume B has requested that a new initial snapshot of volume B be stored on remote archive storage systems, for example to be a long-term backup. Consequently, volume B has been separated into multiple chunks of fragments that will each be stored separately by the storage systems.
ES 2 575 155 T3 archive storage as, for example, to correspond to a typical or maximum storage size for archive storage systems, or instead otherwise as determined by the archive storage service block-level data. In the present example, the primary 155B-b copy of volume B has been separated into N fragments 155B-b1 through 155B-bN, and the mirror copy 155B-a of volume B similarly stores the same data using the fragments 155B-a1 up to 155BaN. Each of the N chunks of volume B is stored as a separate data object in one of two exemplary archive storage systems 180a and 180b, and thus said corresponding multiple stored data objects in total form the initial instant volume copy for volume B. In particular, fragment 1 155B-b1 of the primary copy of volume B is stored as data object 180B1 in archive storage system 180a, fragment 2 155B-b2 is stored as data object 180B2 in storage system archive 180b, chunk 3 155B-b3 is stored as data object 180B3 in archive storage 180a, and fragment N 155B-bN is stored as data object 180BN in archive storage system 180a. In the present example, the separation of volume B into multiple chunks is performed by the data warehouse service at the block level, so that individual chunks of volume B can be transferred individually to archive storage systems, although in other realizations all of volume B may instead be sent to archive storage systems, which can then separate the volume into multiple chunks or otherwise process the volume data if desired.
Furthermore, in the present example, the archive storage system 180b is an archive storage computer system running an Archive Manager module 190 to manage operations of the archive storage systems, such as to manage storage. and data object retrieval, to track which stored data objects correspond to which volumes, to separate transferred volume data into multiple data objects, to measure and otherwise track usage of archival storage systems, etc. The Archive Manager module 190 may, for example, maintain a variety of information about the different data objects that correspond to a particular volume, such as for each snapshot of the volume, as mentioned in greater detail with reference to Figure 2F, while in other embodiments such volume snapshot information may instead be maintained in other ways (For example, by server block-level data storage systems or other block-level data storage service modules). In other embodiments, only a single archive storage system can be used, or data objects corresponding to chunks of volume B can be stored across many more archive storage systems (not shown). In addition, in other embodiments, each archive storage system may run at least part of an archive manager module, such as for each archive storage system to have a different archive manager module, or for all Archival storage systems provide the functionality of the archive manager module in a distributed peer-to-peer manner. In other embodiments one or more archive manager modules may instead run on one or more computer systems that are local to the other block-level data storage service modules (e.g., on the same computer system or on a computer system close to one running a BDS System Administrator module), or the operations of the archive storage systems may instead be managed directly by one or more different modules of the block-level data warehouse service without using an archive manager module (for example, by a BDS System Administrator module).
Furthermore, in at least some embodiments, the archive storage systems can perform various operations to improve the reliability of the stored data objects, such as replicating some or all of the data objects across multiple archive storage systems. Thus, for example, the other data objects 182b of the archive storage system 180b may include mirror copies of one or more of the data objects 180B1, 180B3, and 180BN of the archive storage system 180a, and the others Data objects 182a from archive storage system 180a may similarly store a mirror copy of data object 180B2 from archive storage system 180b. Furthermore, as described in greater detail elsewhere in this document, in some embodiments at least some fragments of volume B may already be stored in the archive storage systems prior to receipt of the request to create the initial snapshot of the volume B such as if the data objects stored in the archive storage systems to represent the chunks of volume B were used as a backing store or other remote long-term backup for volume B. If this is the case, the snapshot on the archive storage systems can instead be created without transferring additional volume data at that time, such as if the data objects on the archive storage systems represented a current state of the fragments in volume B, while in other embodiments additional measures can be taken to ensure that the already stored data objects are up-to-date with respect to the fragments of volume B.
Figure 2D continues the example of Figure 2C, and reflects modifications to volume B that are made after the initial snapshot is stored relative to Figure 2C. In particular, in the present
For example, after the initial volume snapshot has been created, volume B is modified, such as by one or more programs (not shown) that are connected to the volume. In the present example, data is modified in at least two portions of volume B corresponding to fragment 3 155B-b3 and fragment N 155B-bN of the primary copy of volume B, with the modified fragment data illustrated as data 3a and Na, respectively. In the present example, after the primary copy 155B-b of volume B is modified, the server storage system 165b initiates the corresponding updates to the mirror copy 155B-a of volume B on the server storage system 165a, so that the 3 155B-a3 fragment of the mirror copy is modified to include the modified data 3a, and the N 155B-aN fragment of the mirror copy is modified to include the Na modified data. Thus, the mirror copy of volume B is kept in the same state as the primary copy of volume B in the present example.
Furthermore, in some embodiments, the data in the archive storage systems may further be modified to reflect the changes on volume B, although such new data modifications on volume B are not currently part of a volume snapshot for the volume. B. In particular, since the previous version of the data in chunk 3 and chunk N is part of the initial volume snapshot stored in the archive storage systems, the corresponding 180B3 and 180BN data objects are not modified to reflect the data changes on volume B that occur after the creation of the initial volume snapshot. On the other hand, if the copies are optionally made from the data of the modified volume B, they are stored in the present example as additional data objects such as, for example, the optional data object 180B3a to correspond to the modified data 3a of the fragment 3 155B-b3, and as, for example, optional data object 180BNa to correspond to the modified Na data of fragment N 155B-bN. In this way, the data for the initial volume snapshot is preserved even while making changes to the primary and mirror copies of volume B. If the optional data objects 180B3a and 180BNa are created, the creation can be initiated in several ways. as, for example, through the server block level data storage system 165b in a similar way to the updates that are initiated for the mirror copy 155B-a of volume B.
Figure 2E illustrates an alternative embodiment with respect to the embodiment previously described with reference to Figure 2D. In particular, in the example of Figure 2E, volume B is modified again after the initial snapshot of the volume is stored in the archive storage systems in a manner similar to that described with respect to Figure 2D and, consequently, the primary copy 155B-b of volume B on server storage 165b is updated so that fragment 3 155B-b3 and fragment N 155B-bN are updated to include the modified data 3a and Na, respectively. However, in the present embodiment, the mirror copy 155B-a of volume B on the server storage system 165a is not maintained as a full copy of volume B. Instead, the volume snapshot copy of volume B on systems archive storage is used in conjunction with mirror copy 155B-a to maintain a copy of volume B. Thus, in the present example, while modifications are made to the primary copy 155Bb of volume B after the creation of the initial volume snapshot, such modifications are also made to the mirror copy 155B-a on the system server storage 165a, so that the mirror copy stores modified data 3a for fragment 3 155B-a3 and modified data Na for fragment N 155B-aN. However, the mirror copy of volume B does not initially store copies of the other fragments of volume B that have not been modified since the initial volume snapshot was created, since the initial volume snapshot copy of volume B on systems Archival storage includes copies of such data. Consequently, if server storage 165b subsequently becomes unavailable, as described above with reference to Figure 2B, mirror copy 155B-a of volume B on server storage 165a may be promoted to be the new primary copy of volume B. For the purpose of achieving this promotion in the present exemplary embodiment, the remaining portions of mirror 155B-a of volume B are restored using the initial volume snapshot of volume B on archive storage systems such as, for example, in order to use stored data object 180B1 to restore fragment 155B-a1, in order to use stored data object 180B2 to restore fragment 155B-a2, and so on. Furthermore, in the present example, the data objects 180B3a and 180BNa may be similarly and optionally stored in the archive storage systems to represent the modified data 3a and Na. If this is the case, in some embodiments, the modified data 3a and Na will not be initially stored in the server block level data storage system 165a for the mirror copy 155B-a, and instead the fragments of mirror copy 155B-a3 and 155B-aN can be restored, similarly, from the archive storage system data objects 180B3a and 180BNa in a similar way to that described above for the other fragments of the mirror copy.
well the volume snapshot of volume B is used in the above example to restore the mirror copy of volume B when the mirror is promoted to be the new primary volume copy, the volume snapshot on archive storage systems it can be used in other ways in other embodiments. For example, a new copy of volume B can be created that matches the initial volume snapshot by using the volume snapshot on archive storage systems in a similar way as described above to restore the volume mirror as, for
ES 2 575 155 T3 example, to create a new mirror copy of volume B from the moment the volume snapshot occurs, to create a totally new volume based on the volume snapshot of volume B, in order to to help move volume B from one server-side block-level storage system to another, and so on. In addition, when the block-level data storage systems in the block-level data service server are available in multiple different data centers or other geographic locations, the remote archive storage systems may be available to all server block level data storage systems, and thus can be used to create a new volume copy based on a volume snapshot in any of these geographic locations.
Figure 2F continues the examples in Figures 2C and 2D, starting from a later point in time after additional modifications are made to volume B. In particular, after modifications are made to fragment 3 and fragment N as described with reference to Figures 2C or 2D, a second volume snapshot copy of volume B is created on the archive storage systems. Subsequently, additional modifications are made to the data on volume B that are stored in fragments 2 and 3. Consequently, the primary copy of volume B 155B-b as illustrated in Figure 2F includes original data 1 in fragment 1 155B-b1, data 2a in fragment 2 155B-b2 that are modified after the creation of the second volume snapshot, data 3b in chunk 3 155B-b3 which is also modified after the second volume snapshot is created, and Na data in fragment N 155B-bN that was modified after the creation of the first initial volume snapshot but after the creation of the second volume snapshot. Consequently, after a third volume snapshot of volume B is instructed to create, additional data objects are created on the archive storage systems to correspond to the two shards changed since the second snapshot of volume was created. volume, data object 180B2a corresponding to fragment 155B-b2 and including modified data 2a, and fragment 180B3b corresponding to fragment 155B-b3 and including modified data 3b.
In addition, the present example does not show the server block-level data storage system 165a, but a copy of the information 250 maintained by the Archive Manager module 190 (for example, stored in the storage system) is shown. archive storage 180b or elsewhere) to provide information about volume snapshots stored in archive storage systems. In particular, in the present example, the information 250 includes multiple rows 250a-250d, each corresponding to a different volume snapshot. Each of the rows of information in the present example includes a unique identifier for the volume copy, an indication of the volume to which the snapshot volume corresponds, and an indication of an ordered listing of the data objects stored on the systems. archive storage that comprise the volume snapshot. Thus, for example, row 250a corresponds to the initial volume snapshot of volume B described with reference to Figure 2C, and indicates that the initial volume snapshot includes the stored data objects 180B1, 180B2, 180B3, and so on up to 180BN. Row 250b corresponds to an example volume snapshot for a different volume A that includes several stored data objects not shown in the present example. Row 250c corresponds to the second volume snapshot copy of volume B, and row 250d corresponds to the third volume snapshot copy of volume B. In the present example, the second and third volume copies for volume B are copies incrementals and not global copies, so that chunks on volume B that do not change from a previous volume snapshot will continue to be represented using the same stored data objects. Thus, for example, the second snapshot of volume B in row 250c indicates that the second volume snapshot shares data objects 180B1 and 180B2 with those of the initial volume snapshot of volume B (and possibly , some or all of the data objects for fragments 4 through fragments N-1, not illustrated). Similarly, the third snapshot of volume B in row 250d also continues to use the same data object 180B1 as the initial and second snapshots of volume.
By sharing common data objects across multiple Volume Shadow Copies, the amount of storage on archive storage systems is minimized, since a new copy of an unchanging volume chunk such as Chunk 1 does not have separate copies on the archive storage systems for each volume snapshot. In other embodiments, however, some or all of the volume shadow copies may not be incremental and instead each may include a separate copy of each volume chunk regardless of whether the data in the chunk has changed or not. Additionally, when using incremental volume snapshots that can share one or more overlapping data objects with one or more different volume snapshots, the overlapping data objects are managed when additional types of operations are performed with respect to the volume snapshots. For example, if a request is subsequently received to erase the initial volume snapshot for volume B listed in row 250a and to consequently free up storage space that is no longer needed on archive storage systems , you can delete only some of the data objects listed for that initial volume snapshot on the archive storage systems. For example, fragment 3 and fragment N were modified after the initial snapshot was created
ES 2 575 155 T3 volume, and thus the corresponding stored data objects 180B3 and 180BN for the initial volume snapshot are used only by said initial volume snapshot. Thus, these two data objects can be permanently erased from the archive storage 180a if the initial volume snapshot of volume B is erased. However, data objects 180B1 and 180B2 will remain even after that initial volume snapshot of volume B is erased, since they are still part of at least the second volume snapshot of volume B.
While not illustrated in the present example, the information 250 can include a variety of other types of information about volume shadow copies, including information about which archive storage system stores each of the data objects, information about who has permission to access the volume snapshot information and under what circumstances, etc. By way of example, in some embodiments, some users may create Volume Shadow Copies and make access to those Volume Shadow Copies available to at least some users in at least some circumstances, such as charging a fee to allow have other users create copies of one or more particular volume snapshots. In this case, said access-related information can be stored in information 250 or elsewhere, and the archive manager module 190 can use said information to determine whether or not to satisfy requests for information made corresponding to snapshot copies of particular volume. Alternatively, in other embodiments, access to volume shadow copies may instead be managed by other block-level data warehouse service modules (for example, a BDS System Administrator module) such as , for example, to prevent requests from being sent to archive storage systems unless such requests are authorized.
It will be appreciated that the examples in Figures 2A-2F have been simplified for ease of explanation, and that the number and organization of server block level data storage systems, archive storage systems, and other devices they can be of a size much larger than the one represented. Similarly, in other embodiments, volume primary copies, volume mirrors, and / or volume snapshots may be stored and managed in other ways.
Figure 3 is a block diagram illustrating exemplary computing systems suitable for managing provisioning and use of reliable non-local block-level data storage functionality to clients. In the present example, a server computing system 300 executes one embodiment of a BDS System Administrator module 340 to manage the provision of non-local block level data storage functionality to programs running on host computer systems. 370 and / or in at least some different computer systems 390 such as, for example, to lock block-level data storage volumes (not illustrated) provided by server-wide block-level data storage systems 360. Each of the host computer systems 370 in the present example also executes an embodiment of a Node Manager module 380 to manage access of programs 375 running on the host computer system to at least some of the storage volumes of the host computer system. non-local block-level data, such as in coordination with the BDS 340 System Administrator module on a 385 network (For example, an internal network of a data center, not shown, including computer systems 300, 360, 370, and optionally at least some of the other computer systems 390). In other embodiments, some or all of the Node Manager modules 380 may instead manage one or more different computer systems (eg, other computer systems 390).
In addition, multiple server-side block-level data storage systems 360 are illustrated and each store at least some of the non-local block-level data storage volumes (not illustrated) used by running programs 375 with access to said volumes also provided on network 385 in the present example. One or more of the server-side block-level data storage systems 360 may also each store a server-side software component (not illustrated) that manages the operation of one or more server-level data storage systems. block on server 360 as well as different information (not shown) about the data that is stored by data storage systems at the block level on server 360. Thus, in at least some embodiments, the server computing system 300 of Figure 3 may correspond to the computing system 175 of Figure 1, one or more Node Manager modules 115 and 125 of Figure 1 may correspond to the Node Manager 380 modules from Figure 3, and / or one or more of the server block level data storage computer systems 360 of Figure 3 may correspond to server block level data storage systems 165 of Figure 1. Furthermore, in the present exemplary embodiment, multiple archive storage systems 350 are illustrated, which can store snapshots and / or other copies of at least portions of at least some stored block-level data storage volumes. in block-level data storage systems in server 360. The archive storage systems 350 may also interface with some or all of the computer systems 300, 360, and 370, and in some embodiments may be remote archive storage systems (e.g., from a remote storage service, not illustrated) that interact with computer systems 300, 360 and
ES 2 575 155 T3
370 on one or more different external networks (not illustrated).
The other computer systems 390 may also include other nearby or remote computer systems of various types in at least some embodiments, including computer systems whereby clients or other users of the block-level data storage service interact with the computer systems. 300 and / or 370. Also, one or more of the other computer systems 390 may further execute a PES System Administrator module to coordinate the execution of programs on the host computer systems 370 and / or other host computer systems 390, or the computer system 300 or one. of the other illustrated computer systems may instead run said PES System Administrator module, although a PES System Administrator module is not illustrated in the present example.
In the present exemplary embodiment, the computer system 300 includes a CPU (central processing unit) 305, a local storage 320, a memory 330, and various I / O (input / output) components 310, the I / O components S illustrated in the present example including a display 311, a network connection 312, a computer-readable media drive 313, and other I / O devices 315 (eg, a keyboard, mouse, speakers, microphone, etc.) . In the illustrated embodiment, the BDS System Administrator module 340 is running in memory 330 and one or more different programs (not shown) may optionally also run in memory 330.
Each computer system 370 similarly includes a CPU 371, a local storage 377, a memory 374, and various I / O components 372 (For example, I / O components similar to the I / O components 310 of the computer system in server 300). In the illustrated embodiment, a Node Manager module 380 is running in memory 374 to manage one or more different programs 375 that run in memory 374 in the computer system, such as on behalf of clients of the service. execution of programs and / or the data storage service at the block level. In some embodiments, some or all of the computer systems 370 may host multiple virtual machines, and in this case, each of the running programs 375 may be a full virtual machine image (e.g., with one operating system and one or more more application programs) running on a different hosted virtual machine compute node. The Node Manager module 380 may similarly be running on another hosted virtual machine, such as a privileged virtual machine monitor managing the other hosted virtual machines. In other embodiments, the copies of the running program 375 and the Node Manager module 380 may run as separate processes on a single operating system (not shown) running on the computer system 370.
Each archive storage system 350 in the present example is a computer system that includes a CPU 351, a local storage 357, a memory 354, and various I / O components 352 (For example, I / O components similar to E components). / S 310 of the computer system on server 300). In the illustrated embodiment, an Archive Manager module 355 is running in memory 354 to manage the operation of one or more archive storage systems 350 such as, for example, on behalf of clients of the data warehouse service at the data level. block and / or a separate storage service provided by the archive storage systems. In other embodiments, the Archive Manager module 355 may instead be running on another computer system, such as one of the other computer systems 390 or on the computer system 300 in conjunction with the System Manager module. of BDS 340. In addition, while not illustrated herein, in some embodiments, different information about the data that is stored by archival storage systems 350 may be maintained in storage 357 or elsewhere, as described above with reference to Figure 2F. Furthermore, while not illustrated herein, each of the server block level data storage systems 360 and / or other computer systems 390 may similarly include some or all of the types of components illustrated with with respect to archival storage systems 350 such as a CPU, local storage, memory, and various I / O components.
The BDS 340 System Manager module and Node Manager 380 modules can perform various actions to manage provisioning and use of reliable non-local block-level data storage functionality to clients (For example , running programs), as described in greater detail elsewhere in this document. In the present example, the BDS System Administrator module 340 may maintain a database 325 in storage 320 that includes information about volumes stored in server-side block-level data storage systems 360 and / or on servers. 350 archive storage systems (For example, for use in volume management), and may also store other information (not shown) about users or other aspects of the block-level data storage service. In other embodiments, the information about the volumes may be stored in other ways, such as in a distributed manner by Node Manager modules 380 in your computer systems and / or by other computer systems. Furthermore, in the present example, each Node Manager module 380 in a host computer system 370 can store information 378 in local storage 377 about the current volumes connected to the host computer system and used by running programs 375 on the computer system. host, such as to coordinate interactions with storage systems
ES 2 575 155 T3 block-level data on server 360 that provide the primary copies of the volume, and to determine how to switch to a mirror copy of a volume if the primary copy of the volume becomes unavailable. While not illustrated herein, each host computer system may further include a different, logical, local block-level data storage device interface for each volume connected to the host computer system and used by a program running on the computer system, which also, for running programs, may not be distinguished from one or more physically attached local storage devices that provide local storage 377.
As noted, computer systems 300, 350, 360, 370, and 390 are merely illustrative and are not intended to limit the scope of the present invention. For example, computer systems 300, 350, 360, 370 and / or 390 may be connected to other devices not illustrated, even through network 385 and / or one or more different networks such as, for example, the Internet or via the Internet. World Wide Web (Web). More generally, a computing node or other computing system or data storage system may comprise any combination of hardware or software that can interact and perform the types of functionality described, including, but not limited to, desktop computers or other, database servers, network storage devices and other network devices, PDAs (personal digital assistants), mobile phones, cordless phones, personal pagers, electronic organizers, Internet devices, television-based systems (for example, through the use of encoders and / or personal / digital video recorders), and other consumer products that include appropriate communication capabilities. Furthermore, the functionality provided by the illustrated modules can, in some embodiments, be combined into fewer modules or distributed into additional modules. Similarly, in some embodiments, the functionality of some of the illustrated modules may not be provided and / or additional functionality may be available.
It will also be appreciated that, while various items are illustrated as being stored in memory or in storage while in use, such items or portions thereof may be transferred from memory and other storage devices for the purposes of managing the memory. memory and data integrity. Alternatively, in other embodiments, some or all of the software modules and / or systems may run in memory on another device and communicate with the illustrated computer systems via computer-to-computer communication. Furthermore, in some embodiments, some or all of the systems and / or modules may be implemented or provided in other ways such as, for example, at least partially in firmware and / or hardware, including, but not limited to, one or more integrated circuits. for specific applications (ASICs), standard integrated circuits, controllers (For example, executing appropriate instructions, and including microcontrollers and / or built-in controllers), Programmable Gate Arrays (FPGAs), Complex Programmable Logic Devices (CPLDs), etc. Some or all of the modules, systems, and data structures may also be stored (for example, as software instructions or structured data) on a computer-readable medium such as a hard drive, memory, network, or article. portable media that will be read by an appropriate drive or through an appropriate connection. Data systems, modules, and structures can also be transmitted as generated data signals (e.g., as part of a carrier wave or other analog or digital propagated signal) on a variety of computer-readable transmission media, including wireless-based media and cable / wired-based media, and can take different forms (e.g. as part of a single signal or analog multiplex, or as multiple discrete digital packages or frames). Such computer program products may also take other forms in other embodiments. Accordingly, it is possible to practice the present invention with other computer system configurations.
Figure 4 is a flow chart of an exemplary embodiment of a Block level Data Storage System Manager routine 400. The routine may be provided, for example, by executing the Block-level Data Storage System Manager module 175 of Figure 1 and / or the BDS System Manager module 340 of Figure 3 as, for For example, to provide a block-level data storage service for use by running programs. In the illustrated embodiment, the routine can interact with multiple server block level data storage systems in a single data center or other geographic location (For example, if each data center or other geographic location has a different realization of the routine running at the geographic location), although in other embodiments a single routine 400 may support multiple different data centers or other geographic locations.
The illustrated embodiment of the routine begins at block 405, where a request or other information is received. The routine continues to block 410 to determine if the request received was to create a new block-level data storage volume, such as from a user of the block-level data storage service and / or a running program that wants to access the new volume and, if this is the case, continues to block 415 to carry out volume creation. In the illustrated embodiment, the routine at block 415 selects one or more server-side block-level data storage systems in which copies of the volume will be stored (e.g., depending on, at least in part, location and / or the capabilities of the selected server storage systems), initializes the volume copies on those selected server storage systems, and updates the stored information about volumes to reflect the
ES 2 575 155 T3 new volume. For example, in some embodiments, creating a new volume may include initializing a specified size of linear storage on each of the selected servers in a specific way, such as to be empty, to include a copy of another volume. indicated (For example, another volume in the same data center or other geographic location, or instead a volume stored in a remote location), to include a copy of a specified volume snapshot (For example, a volume snapshot stored by one or more archive storage systems, such as interacting with the archive storage systems to obtain the volume snapshot ), etc. In other embodiments, a linear storage logical block of a specified size can be created for a volume in one or more server-side block-level data storage systems, such as by using multiple non-contiguous storage areas. which are presented as a single logical block and / or by striping a linear storage logical block on multiple local physical hard drives. For the purpose of creating a copy of a volume that already exists in another data center or other geographic location, the routine may, for example, coordinate with another instance of routine 400 that supports data warehouse service operations at block level at that location. Furthermore, in some embodiments, at least some volumes will each have multiple copies including at least one primary volume copy and one or more mirror copies on multiple different server storage systems, and if this is the case, they will be they can select and initialize multiple server storage systems.
If instead it is determined in block 410 that the received request is not to create a volume, the routine continues to block 420 to determine if the received request is to connect an existing volume to a copy of the running program such as, for example, a request received from the copy of the running program or another computer system operated on behalf of a user associated with the copy of the running program and / or the indicated volume. If this is the case, the routine continues to block 425 to identify at least one of the server-side block-level data storage systems that stores a copy of the volume, and to associate at least one of the storage systems in server identified to the running program (For example, associate the primary server storage system for the volume with the compute node on which the program is running, for example, causing a logical local block level storage device to be mounted on the compute node representing the primary copy of the volume). The volume to be attached can be identified in various ways, such as by a unique identifier for the volume and / or an identifier for a user who has created or is otherwise associated with the volume. After connecting the volume to the copy of the running program, the routine can also update stored information about the volume to indicate the connection of the running program, such as whether only a single program could connect to the volume at a time, or whether only a single program could have write access or other modify access to the volume at a time. Furthermore, in the indicated embodiment, information about at least one of the identified server storage systems can be provided to a node manager associated with the running program, such as to facilitate the actual connection of the volume to the running program. , although in other embodiments the node administrator may have other access to such information.
If it is determined instead at block 420 that the received request is not to connect a volume to a running program, the routine continues to block 430 to determine if the received request is to create a snapshot for a specified volume. such as a request received from a running program that is attached to the volume or instead another computer system (For example, a computer system operated by a user associated with the volume and / or a user who has gained access to create a snapshot of another user's volume). In some embodiments, a volume snapshot can be created from a volume regardless of whether or not the volume is attached or in use by any running program, and / or regardless of whether or not the volume is stored on it. data center or other geographic location where routine 400 is running. If determined to be so, the routine continues to block 435 to initiate the creation of a volume snapshot of the indicated volume, such as interacting with one or more archive manager modules that coordinate operations of one or more systems. archival storage (for example, archival storage systems at a remote storage location such as, for example, in conjunction with a remote long-term storage service that is accessible on one or more networks). In some embodiments, the creation of the volume snapshot can be performed through a third-party remote storage service in response to an instruction from routine 400, such as if the remote storage service already stored the minus some fragments of the volume. Also, other different parameters can be specified in at least some embodiments, such as whether the volume snapshot will be incremental relative to one or more different volume snapshots, and so on.
If instead it is determined at block 430 that the received request is not to create a volume snapshot, the routine instead continues to block 440 to determine whether the information received at block 405 is an indication. failure or other situation of unavailability of one or more server block-level data storage systems (or of one or more volumes, in other embodiments). For example, as described below with respect to block 485, the routine may, in some embodiments, monitor the status of some or all of the server block-level data storage systems and determine whether or not to do so.
ES 2 575 155 T3 availability on this basis, for example, periodically or constantly sending ping messages or other messages to data storage systems at the server block level to determine if a response is received, or obtaining information about the status of server storage systems. If it is determined at block 440 that the information received indicates the possible failure of one or more server storage systems, the routine continues to block 445 to take actions to maintain the availability of the single or more stored volumes. on the indicated server storage system (s). In particular, the routine at block 445 determines whether any of said volumes stored on the indicated server storage system (s) are primary volume copies, and for each primary volume copy, promotes one of the mirrors for said volume to another server storage system to be the new primary copy for that volume. At block 450, the routine causes at least one new copy of each volume to replicate to one or more different server storage systems, such as by using an existing copy of the volume that is available on a server storage system. server storage other than those indicated as unavailable. In other embodiments, the promotion of mirrors to primary copies and / or the creation of new mirrors may instead be carried out in other ways such as, for example, in a distributed manner by the data storage systems at the level of the database. block on server (For example, using a protocol of choice between mirroring a volume). Also, in some embodiments, volume mirrors can be minimal copies that include only portions of a primary copy of a volume (for example, only portions that have been modified since a snapshot of the volume was created), and promotion of A mirror to a primary copy may further include gathering information for the new primary copy in order to complete it (for example, from the most recent snapshot).
At block 455, the routine then optionally initiates connections from one or more running programs to new volume primary copies that were promoted from mirror copies, such as for running programs that have previously been connected to Primary volume copies on the server storage system (s) are not available, although in other embodiments said re-connection to new primary volume copies will be made, instead, in other ways (For example, through a node manager associated with the running program through which the re-connection will occur). At block 458, the routine updates information about volumes on unavailable server storage systems, such as to indicate new volume copies created in block 450 and new primary volume copies promoted in block 445. In other embodiments, new primary volume copies can be created in other ways, such as by creating a new volume copy as a primary volume copy and not by promoting an existing volume mirror copy, although this is not the case. it may take longer than promoting an existing volume mirror. Also, if there are no volume copies available from which to replicate new volume copies in block 450, such as if multiple server storage systems that store the primary and mirror copies for a volume all fail substantially simultaneously, The routine may, in some embodiments, attempt to obtain information for the volume for use in that replication in other ways, such as from one or more recent volume snapshots for the volume that are available on archive storage systems, from a copy of the volume in another data center or other geographic location, and so on.
If, instead, it is determined at block 440 that the received information is not an indication of failure or other unavailability of one or more server-side block-level data storage systems, the routine continues, instead. , to block 460 to determine whether or not the information received at block 405 indicates moving one or more volumes to one or more new server block level data storage systems. Such volume movement can be carried out for various reasons, as mentioned in greater detail elsewhere in this document, including to other server-side block-level data storage systems in the same geographic location (For example, moving existing volumes to storage systems that are better equipped to support the volumes) and / or to one or more server data storage systems in one or more different data centers or other geographic locations. In addition, the movement of a volume can be initiated in various ways, such as, for example, following a request from a user of the block-level data storage system that is associated with the volume, a request from a human operator of the block-level data storage, based on an automatic detection of a better server storage system for a volume than the currently used server storage system (For example, due to overuse of the server storage system at that time and / or underuse of the new server storage system), etc. If it is determined at block 460 that the information received is for moving one or more volume copies, the routine continues to block 465 and creates a copy of each indicated volume in one or more block-level data storage systems in new servers such as, for example, similar to that previously described with respect to block 415 (For example, using an existing volume copy in a server block-level data storage system, using a snapshot or other copy of the volume on one or more archive storage systems, etc.), and further updates the information stored for the volume in block 465. In addition, in some embodiments, the routine may perform actions to support the move, such as deleting the old volume copy from a server-side block-level data storage system after the new volume copy is created. Also, in situations where one or more
ES 2 575 155 T3 running programs have been connected to the previous volume copy that is being moved, the routine may initiate disconnection of the previous volume copy that is being moved for a running program and / or it may initiate a re -connection of said running program to the new volume copy that is being created, for example, by sending associated instructions to a node manager for the running program, although in other embodiments the node manager may instead carry out such actions.
If, instead, it is determined at block 460 that the received information is not an instruction to move one or more volumes, the routine instead continues to block 485 to perform one or more different indicated operations as appropriate. Other operations can have several forms in various implementations, such as, for example, one or more of the following non-exclusive list: monitoring some or all data storage systems at the server block level (For example, sending ping messages or other status messages to data storage systems at the server block level and waiting for a response); initiate the creation of a primary volume copy and / or a replacement volume mirror in response to the determination that a primary or mirror copy of a volume is not available, such as based on monitoring being performed , in a message received from a primary server block-level data storage system that stores a primary copy of a volume but is unable to update one or more mirror copies of that volume, in a message received from a node manager module, etc .; disconnect, delete, and / or describe one or more volumes; delete, describe and / or copy one or more volume shadow copies; track usage of volumes and / or volume snapshots by users, for example to measure such usage for payment purposes; etc. After blocks 415, 425, 435, 458, 465, or 485, the routine continues to block 495 to determine whether or not to continue as, for example, until an explicit end instruction is received. If this is the case, the routine returns to block 405, and otherwise the routine continues to block 499 and ends.
Furthermore, for at least some types of requests, the routine may, in some embodiments, further verify that the requestor is authorized to make the request, such as according to access rights specified for the requestor and / or an associated destination of the request. request (For example, a specified volume). In some of these embodiments, the authorization verification may further include obtaining payment from the requester for the requested functionality (or verifying that such payment has already been made), such as not making the request if the payment is not made. Has made. For example, request types that may have associated payment in at least some realizations and situations include requests to create a volume, attach a volume, create a snapshot, move a specified volume (For example, to a server storage system premium), and other types of operations indicated. In addition, some or all types of actions carried out on behalf of users may be monitored and measured, for example, for further use by determining the corresponding usage-based fees for at least some of those actions.
Figure 5 is a flow diagram of an exemplary embodiment of a Node Manager routine 500. The routine may be provided, for example, by execution of the Node Manager module 115 and / or 125 of Figure 1, and / or execution of a Node Manager module 380 of Figure 3, such as to manage the use of one or more running programs of data storage at the non-local block level. In the illustrated embodiment, the block-level data warehouse service offers functionality through a combination of one or more ADB System Manager modules and multiple Node Manager modules and optionally one or more Archive Manager modules, although in other embodiments other configurations may be used (For example a single ADB System Manager module without any Node Manager module and / or Archive Manager modules, multiple Node Manager modules running together in a coordinated manager without a ADB System Administrator module, etc.).
The illustrated embodiment of the routine begins at block 505 where a request related to program execution is received at an associated computing node. The routine continues to block 510 to determine if the request is related to the execution of one or more of the indicated programs at a indicated associated computing node, such as, for example, a request from a program execution service and / or a user associated with such programs. If so, the routine continues to block 515 to obtain a copy of the indicated program / s and to initiate execution of the program / s on an associated computing node. In some embodiments, one or more of the indicated programs may be obtained at block 515 based on the indicated programs sent to routine 500 as part of the request received at block 505, while in other embodiments the indicated programs may be retrieved from a local or non-local storage (for example, from a remote storage service). In other embodiments, routine 500 may instead not perform operations related to running programs, such as if another routine that supports the program execution service would instead perform those operations on behalf of associated computing nodes.
If, instead, it is determined in block 510 that the received request is not to execute one or more of the indicated programs, the routine continues instead to block 520 to define if a request is received to connect a indicated volume to an indicated running program, such as, for example, the running program, routine 400 of Figure 4, and / or a user associated with the indicated volume and / or the indicated running program. If so, the routine continues to block 525 to obtain an indication of a primary copy of the
ES 2 575 155 T3 volume, and to associate that primary volume copy with a data storage device at the local, logical and representative block level for the computing node. In some embodiments, the representative, logical, block-level data storage device may be indicated for the running program and / or computing node by routine 500, while in other embodiments, the running program may instead initiate the creation of the data storage device at the local logical block level. For example, in some embodiments, routine 500 may use GNBD (global network block device) technologies to make available to a virtual machine computing node a logical and local block-level data storage device by importing a block-level device in a specific virtual machine and the mounting of that data storage device at the local and logical block level. In some embodiments, the routine may perform additional actions at block 525 such as obtaining and storing indications of one or more mirrored volume copies for the volume, such as allowing the routine to dynamically connect to a mirrored volume copy if it copies it. Primary volume later becomes unavailable.
If, instead, it is determined in block 520 that the request received from block 505 is not to connect an indicated volume, the routine instead continues to block 530 to define whether the received request is a request for access to data from a running program for an attached volume, such as a read request or a write request. If so, the routine continues to block 535, where the routine identifies the associated primary volume copy that corresponds to the data access request (for example, based on the data storage device at the logical block level, local and representative used by the program run for the data access request) and initiates the requested data access to the primary volume copy. As discussed in greater detail elsewhere in this document, a deferred writing scheme may be used in some embodiments, such as to immediately modify the actual and / or mirror primary volume copies to reflect a write data access request (For example, to always update the mirror volume copy, to update a mirror volume copy only if the mirror volume copy is promoted to primary volume copy, etc.), but not to immediately modify a corresponding segment stored in one or more archive storage systems to reflect the access request to write data (For example, to eventually update the copy stored in the archive storage systems when sufficient modifications have been made and / or when read access to corresponding information is requested, etc.). In the illustrated embodiment, maintenance of mirror volume copies is performed by a routine other than routine 500 (for example, by the primary server of the block-level data storage system that stores the primary volume copy) although in other embodiments routine 500 may further collaborate at block 535 to maintain one or more mirror volume copies by sending requests for access to data similar or identical to those mirror volume copies. Additionally, in some embodiments, a volume may not be stored in archive storage systems until expressly requested by a corresponding user (for example, as part of a request to create a snapshot of the volume), while in other embodiments a copy can be kept on the archive storage systems of at least some portions of at least some volumes (For example if the copy of the archive storage systems is used as a backing store for the copies primary volume and / or mirror).
After block 535, the routine continues to block 540 to define whether a response has been received from the primary block level data warehouse server for the request sent at block 535 within a predefined time limit, such as for indicate the success of the operation. If not, the routine defines that the primary server of the block-level data storage system is unavailable, and continues to block 545 to initiate a modification to connect one of the mirror volume copies as the new copy of primary volume, and to associate the server block level data storage system for said mirror volume copy as the new primary server block level data storage system for the volume. Additionally, similarly, the routine sends the data access request to the new primary volume copy in a manner similar to that noted above for block 535, and may further, in some embodiments, control whether a request is received. appropriate response and proceed again to block 545 if not (eg to promote another mirror volume copy and repeat the process). In some embodiments, initiating the change of a mirror volume copy as a new primary volume copy can be done in coordination with routine 400, for example by initiating contact with routine 400 to define which mirror volume copy should become the one. new copy of primary volume, upon receiving instructions from routine 400 when a volume copy is promoted to become a primary volume copy by routine 500 (For example as shown in the indication sent by routine 500 at block 545 that the primary volume copy not available), etc.
If instead it is determined at block 530 that the received request is not a data access request for a connected volume, the routine instead continues to block 585 to perform one or more different indicated operations, as appropriate. The other operations can take various forms in various embodiments, such as new volume information routine 400 instructions (For example, a new primary volume copy promoted for a volume to which one or more running programs that are connected are connected. being handled), to detach a volume from a running program on an associated compute node
ES 2 575 155 T3 with routine 500, etc. Furthermore, in at least some embodiments, routine 500 may further perform one or more different actions of a virtual machine monitor, such as whether routine 500 operates as part of or otherwise in conjunction with a machine monitor. virtual machine that manages one or more associated virtual machine compute nodes.
After blocks 515, 525, 545, or 585, or if instead it is determined in block 540 that a response is received within a predefined time limit, the routine continues to block 595 to define whether to continue, for example , until an explicit completion instruction is received. If so, the routine returns to block 505, otherwise it continues to block 599 and ends.
Furthermore, in at least some types of requests, the routine may further verify in some embodiments that the requestor is authorized to make the request, based, for example, on specific access rights for the requestor and / or for an associated destination of the request. request (for example a specified volume). In some of these embodiments, the authorization verification may also include obtaining payment from the applicant for the requested functionality (or verifying that said payment has already been made), so as not to carry out the request if the payment has not been made. done. For example, types of requests that may have an associated payment in at least some embodiments and situations include requests to run indicated programs, connect a volume, perform one or any type of data access request, and other types of indicated operations. In addition, some or all types of actions carried out on behalf of users may be monitored and measured, for example, for further use by determining the corresponding usage-based fees for at least some of those actions.
FIG. 6 is a flow diagram of an exemplary embodiment of a Server Block-level Data Storage System 600 routine. This routine may be provided, for example, by executing a software component in a server block level data storage system, such as to manage block level data storage in one or more data storage volumes to block level in said server storage system (For example for server block level data storage systems 165 of Figure 1 and / or Figure 2). In other embodiments, some or all of the functionality of the routine may be provided in other ways such as, for example, by running software on one or more computer systems to manage one or more block-level data storage systems in server.
The illustrated embodiment of the routine begins at block 605, where a request is received. The routine continues to block 610 to determine if the request received is related to the creation of a new volume, such as associating a block of available storage space for the server storage system (For example, storage space in one or more local hard drives) with a new volume indicated. The request can, for example, be from routine 400 and / or from a user associated with the new volume being created. If so, the routine continues to block 615 to store information about the new volume, and at block 620 it starts a storage space for the new volume (eg, logical linear block of storage space of a specified size). As discussed in more detail elsewhere in this document, in some embodiments, new volumes can be created based on another existing volume or snapshot volume copy, and if so, the routine at block 620 can start the storage space. storage for the new volume, copying appropriate data to the storage space, while in other embodiments you can start a new volume storage space in other ways (For example to initialize the storage space to a default value, for example to all zeros).
If, instead, it is determined in block 610 that the received request is not to create a new volume, the routine continues instead to block 625 to define whether a data access request has been received for an existing volume stored in the server storage system, such as from a node manager associated with a running program that initiated the data access request. If so, the routine continues to block 630 to perform the data access request on the indicated volume. The routine then continues to block 635 to, in the illustrated embodiment, optionally initiate the corresponding updates for one or more mirrored copies of the volume, such as if the indicated volume on the current server storage system is the copy primary volume for the volume. In other embodiments, consistency between a primary volume copy and mirror volume copies can be maintained in other ways. As described in greater detail elsewhere in this document, in some embodiments, at least some modifications to the stored data contents of at least some volumes may also be made to one or more archival storage systems (e.g. to a remote storage service), such as to maintain a backup or other copy of those volumes and, if so, the routine may further initiate updates to the archive storage systems to initiate corresponding updates for one or more copies of the volume on the archive storage systems. Also, if the routine at block 635 or elsewhere defines that there is no mirror copy of the volume available (for example, based on a lack of response to a data access request sent at block 635 within a predefined time) or to a ping message or to another status message initiated by the routine
ES 2 575 155 T3
600 to periodically verify that the mirror volume copy and its mirror server block-level data storage system are available; according to a message from the mirror server block-level data storage system that it has experienced an error condition or that shutdown or failure mode operation has begun; etc.), the routine can initiate actions to create a new mirror copy of a volume, for example, by sending a message corresponding to routine 400 in Figure 4 or, instead, by directly initiating the creation of a mirror volume copy .
If instead it is determined at block 625 that the received request is not a data access request for a volume, the routine continues to block 685 to perform one or more of the different indicated operations, as appropriate. Said other operations may take various forms in various embodiments such as, for example, one or more of the following non-restrictive listing: to delete a volume (For example to have the associated storage space available for other uses); to copy a volume to a specified destination (For example, to another new volume on another server block-level data storage system, one or more archive storage systems for use as instant volume copy, etc.); to provide information about volume usage (For example to measure volume usage, for customers who have a rate based on volume usage; to perform ongoing maintenance or diagnostics on the server block-level data storage system (For example, defragment local hard drives); etc. After blocks 620, 635 or 685, the routine continues to block 695 to define whether to continue, for example until it receives an explicit end instruction. If so, the routine returns to block 605, otherwise it continues to block 699 and ends.
In addition, for at least some types of requests, the routine may in some embodiments further verify that the requestor is authorized to make said request, based for example on access rights specified by the requestor and / or an associated destination of the request (For example a specified volume), while in other embodiments the routine may assume that the requests have been previously authorized by a routine from which it receives requests (for example a Node Manager routine and / or an ADB System Manager routine). Additionally, some or all types of actions taken on behalf of users may be monitored and measured, for later use in determining usage-based fees for at least some of those actions.
Figures 7A and 7B are a flow chart of an exemplary embodiment of a PES System Administrator routine 700. The routine may be provided, for example, by executing a PES System Administrator module 140 of Figure 1. In other embodiments, one or all of the functionality of routine 700 may instead be provided in other ways, such as by routine 400 as part of the block-level data storage service.
In the illustrated embodiment, the routine begins at block 705, where a status message or other request related to the execution of a program is received. The routine continues to block 710 to determine the type of message or request received. If at block 710 it is determined that the type is a request to execute a program, such as a user or running program, the routine continues to block 720 to select one or more host computer systems where to execute the indicated program. as, for example, a group of possible host computer systems for the execution of programs. In some embodiments, one or more host computer systems may be selected based on user instructions or other stated criteria of interest. The routine then continues to block 725 to initiate program execution by each of the selected host computer systems, for example, through interaction with a Node Manager associated with the selected host computer system. At block 730, the routine then optionally performs one or more cleaning tasks (eg, execution of monitoring programs by users, such as for metering and / or other billing purposes).
If, instead, it is determined in block 710 that the request received is to register a new program as available for later execution, the routine instead continues to block 740 to store an indication of the program and associated administrative information for its use (e.g. access control information related to users who are authorized to use the program and / or types of authorized uses) and you can also store at least one centralized copy of the program in some situations. The routine then continues to block 745 to optionally initiate the distribution of copies of the indicated program to one or more host computer systems for later use, so as to allow a quick start of the program by these host computer systems upon retrieval of the copy. stored from the local storage of these host computer systems. In other embodiments, one or more copies of the indicated program may be stored in other ways, such as in one or more remote archive storage systems.
If instead it is determined at block 710 that a status message has been received at block 705 related to one or more host computer systems, the routine instead continues to block 750 to update information related to said host computer systems such as to track the use of running programs and / or other status information about the host computer systems (For example use of storage volumes of non-local block-level data). In some embodiments, the modules of
ES 2 575 155 T3 node manager will periodically send status messages, while in other embodiments, status messages can be sent at other times (for example when there are significant modifications). In still other embodiments, routine 700 may instead request information from node manager modules and / or host computer systems as desired. Status messages can include a variety of types of information such as the number and identity of programs currently running on a specific computer system, the number and identity of copies of programs currently stored in the local program repository at a specific computer system, attachments and / or other use of non-local block-level data storage volumes, performance and resource related information (for example CPU, network, disk, memory utilization, etc.) for a computer system, configuration information for a computer system, and error reports or failure conditions related to hardware or software on a specific computer system.
If the routine at block 705, instead, defines that another type of request or message is received, the routine instead continues to block 785 to perform one or more of the other indicated operations, as appropriate. Such other operations may include, for example, suspending or terminating the execution of currently running programs, and otherwise managing the administrative aspects of the program execution service (registering new users, determining and obtaining payment for the use of the service. program execution, etc.). After blocks 745, 750, or 785, the routine continues to block 730 to optionally perform one or more cleaning tasks. The routine then continues to block 795 to determine whether or not to continue, for example, until it receives an explicit completion instruction. If so, the routine returns to block 705, otherwise it continues to block 799 and ends.
While not illustrated herein, in at least some embodiments, a variety of additional types of functionality may be provided for executing programs by a program execution service, such as in conjunction with a data storage service to block level. In at least some embodiments, the execution of one or more copies or instances of a program on one or more computer systems may be initiated in response to a current execution request for immediate execution of those instances of the program. Alternatively, the initiation may be based on a previously received program execution request that anticipates or otherwise reserves the future execution of those program instances for the current time. Program execution requests can be received in different ways, directly from a user (For example, through an interactive console or another GUI (graphical user interface) provided by the program execution service), or of a running program of a user that automatically starts the execution of one or more instances of other programs or of itself (For example, through an API (application programming interface) provided by the program execution service, such and as an API that uses web services). Program execution requests may include different information to be used in the initiation of execution of one or more instances of a program, such as an indication of a program previously registered or otherwise provided for future execution, and a number of instances of the program that will run simultaneously (For example expressed as a single number of desired instances, as a minimum and maximum number of desired instances, etc.). In addition, in some embodiments, the program execution requests may include several different types of information, such as the following: an indication of a user account or other indication of a previously registered user (For example, to use in identifying a previously stored program and / or to determine if the execution of the requested program instance is authorized); an indication of a source of payment with which to make the payment of the program execution service for the execution of program instances; an indication of a prepayment or other authorization for the execution of program instances (For example, a subscription previously purchased valid for a certain period, for a number of program execution instances, for a certain number of uses of a resource , etc.); and / or an executable or other copy of a program that will be executed immediately and / or stored for later execution. Furthermore, in some embodiments, the program execution requests may further include a variety of other types of preferences and / or requirements for the execution of one or more program instances. Said preferences and / or requirements may include indications for some or all of the program instances to be executed in an indicated geographical and / or logical location, such as in one of the multiple data centers that host computer systems available for use, in multiple computer systems that are close to each other and / or in one or more computer systems that are close to computer systems that have other indicated characteristics (for example, that provide a copy of a specified block-level data storage volume).
Figure 8 is a flow chart of an exemplary embodiment of an Archive Manager routine 800. The routine may be provided, for example, by execution of one of the Archive Manager modules 355 of Figure 4, from the Archive Manager module. Archive manager 190 of Figures 2C-2F and / or one or more archive manager modules (not shown) in the computer systems 180 of Figure 1. In other embodiments, some or all of the functionality of routine 800 may instead be provided in other ways, such as by routine 400 as part of a block-level data storage service. In the illustrated embodiment, archive storage systems store data in chunks, each of which corresponds to a portion of a block-level data storage volume, but in other embodiments the data may be stored in other ways.
ES 2 575 155 T3
The illustrated embodiment of routine 800 begins at block 805, where information or a request is received. The routine then continues to block 810 to define whether the request or information is authorized, such as whether the requester has paid to gain access based on a fee, or whether he instead has access rights to indicate that a request is performed. petition. If in block 815 it is defined that the request or information is authorized, the routine continues to block 820 and, if not, returns to block 805. In block 820, the routine defines if the received request is to store a new snapshot copy for a specified volume. If so, the routine continues to block 825 to obtain multiple volume chunks for the volume, store each chunk as an archive storage system data object, and then store information about the data objects for the chunks that they are associated with instant volume copy. As discussed in greater detail elsewhere in this document, volume chunks can be obtained in various ways, for example by being received at block 805 as multiple distinctive blocks, received at block 805 as a single large group of block-level data that is chunked in block 825, retrieved in block 825 as individual chunks or as a single large group of block-level data that will be chunked, previously stored in archival storage systems, etc.
If instead it is determined at block 820 that the received request is not to store a new snapshot volume copy, the routine continues instead to block 830 to define whether the received request is to store an incremental snapshot. of a volume that reflects changes from a previous instant volume copy. If so, the routine continues to block 835 to identify snapshots that have changed since a previous snapshot of the volume, and to obtain copies of the modified snapshots in a manner similar to that discussed above with reference to block 825. The routine then continues to block 840 to store copies of the modified shards, and to store information about the new modified shards and the other previous unmodified shards, whose corresponding data objects are associated with the new snapshot volume copy. Fragments that have been modified since a previous snapshot volume copy can be identified in a number of ways, for example by server-side block-level data storage systems that store primary and / or mirror copies of the volume (For example, by tracking any requests access to write data or other modification requests for the volume).
If instead it is determined in block 830 that the received requests are not to store an incremental snapshot volume copy, the routine continues instead to block 845 to determine if the received request is to provide one or more chunks. from an instant volume copy, for example from corresponding stored data objects. If so, the routine continues to block 850 to retrieve the data from the indicated instant volume copy chunk (s), and send the retrieved data to the requester. Such requests can be, for example, part of creating a new volume based on an existing instant volume copy by retrieving all fragments for instant volume copy, part of recovering a subset of a volume copy fragments. snapshot to restore a minimal mirrored volume copy, etc.
If, instead, it is determined at block 845 that the received request is not to provide one or more snapshot volume copy chunks, the routine continues to block 855 to define whether the received request is to perform one or more backup requests. data access for one or more volume fragments that are not part of an instant volume copy, such as to perform read data access requests and / or write data access requests for one or more data objects that represent specific volume fragments (For example, if those stored data objects serve as a backup store for said fragments of volume). If so, the routine continues to block 860 to carry out the data access request (s) for the corresponding stored data object (s) for the indicated volume chunk (s). As described in greater detail elsewhere in this document, in at least some embodiments, lazy update techniques may be used when modifying stored data objects, such that a request for access to write data may not be carried out. out immediately. If so, before a post-read data access request for the same object is completed, one or more pre-write data access requests can be made to ensure strict data consistency.
If, instead, it is determined at block 855 that the received request is not to carry out data access requests for one or more volume chunks, the routine instead continues to block 885 to perform one or more other indicated operations, as appropriate. Such operations may include, for example, repeatedly receiving information corresponding to modifications that are made to a volume to update the corresponding stored data objects that represent the volume (for example as a backing store or for other purposes) and take appropriate corresponding actions, respond to requests to delete or otherwise modify stored shadow volume copies, respond to requests from a user to manage an account with a storage service that provides archiving storage systems, etc. After blocks 825, 840, 850, 860, or 885, the routine continues to block 895 to define whether to continue, for example until it receives an explicit end instruction. If so, the routine returns to block 805, otherwise it continues to block 899 and ends.
ES 2 575 155 T3
As described above, for at least some types of requests, the routine may in some embodiments further verify that the requestor is authorized to make said request, based on, for example, access rights specified by the requestor and / or a destination. associated with the request (for example an indicated volume or snapshot volume copy), while in other embodiments the routine may assume that the requests have been previously authorized by a routine from which it receives requests (for example a Node Administrator routine and / or an ADB System Administrator routine). In addition, some or all types of actions taken on behalf of users can be monitored and measured in at least some embodiments, for example for later use in determining the corresponding usage-based fees for at least some of those actions.
Also, as noted above, some embodiments may use virtual machines, and if so, the programs to be run by the program execution service may include full virtual machine images. In some embodiments, a program to be executed may comprise a complete operating system, a file system, and / or other data, and possibly one or more user-level processes. In other embodiments, a program to be executed may comprise one or more types of executables that interoperate to offer some functionality. In still other embodiments, a program to be executed may comprise a physical or logical collection of instructions and data that can be executed natively on the provided computer system or indirectly through interpreters or other software-implemented hardware abstractions. More generally, in some embodiments, a program to be executed may include one or more application programs, application frameworks, libraries, archives, class files, scripts, configuration files, data files, and the like.
Additionally, as noted above, in at least some embodiments and situations, volumes can be migrated or otherwise moved from one server storage system to another. Various techniques can be used to translate volumes, and such movement can be initiated in a number of ways. In some situations, the move may reflect problems related to the server storage systems on which the volumes are stored (For example, failure of the server storage systems and / or network access to the server storage systems ). In other situations, the move can be performed to accommodate other volume copies that will be stored on existing server storage systems, such as higher priority volumes, or to consolidate volume copy storage on a limited number of storage systems. server storage, such as to enable the original server storage systems that store the volume copies to be shut down for reasons such as maintenance, energy conservation, etc. In a specific example, if one or more volume copies stored on a server storage system require more resources than are available on that server storage system, one or more volume copies can be migrated to one or more storage systems on server with additional resources. Excessive use of available resources can occur for various reasons, for example when one or more server storage systems have fewer resources than expected, one or more of the server storage systems uses more resources than expected (or allowed ) or, in some embodiments where the available resources of one or more server storage systems are intentionally committed in excess of the possible resource needs of one or more reserved or stored volume copies. For example, if the expected resource requirements of the volume copies are within the available resources, the maximum resource requirements may exceed the available resources. Excessive use of available resources can also occur if the actual resources required for storage or use of the volume exceed the available resources.
It will be appreciated that in some embodiments the functionality provided by the aforementioned routines can be provided in alternative ways, for example, by dividing between more routines or consolidating into fewer routines. Similarly, in some embodiments, the illustrated routines may provide more or less functionality than described, for example when other illustrated routines instead exclude or include such functionality respectively, or when the amount of functionality that is provided is altered. Furthermore, while various operations may be illustrated as being performed in a certain way (eg in series or parallel) and / or in a particular order, in other embodiments the operations may be performed in different orders or in different ways. Similarly, the data structures listed above can be structured in different ways in other embodiments, such as by presenting a single data structure divided into multiple data structures consolidated into a single data structure, and can store more or less information than that described (For example when other illustrated data structures instead exclude or include such information respectively, or when the amount or types of information that is stored is altered).
From the foregoing, it will be appreciated that, although specific embodiments have been described herein for illustrative purposes, various modifications can be made without departing from the scope of the invention. Accordingly, the invention is not limited except by the appended claims and the elements described therein. Furthermore, while certain aspects of the invention are presented below in certain forms of
In claims, the inventors contemplated the various aspects of the invention in any available claim form. For example, while only some aspects of the invention can currently be listed as claimed on a computer-readable medium, other aspects can be done as well.
The following clauses describe technical concepts related to the invention, but do not constitute embodiments within the scope of the claims.
Clause 1. A method for a block-level data storage service computer system to provide remote block-level data storage capabilities for running programs, said method comprising:
receiving a request to initiate access of a first running program to non-local block-level data storage provided by a block-level data storage service, the first program executing on the first of a plurality of computer systems that are co-located in a geographic location and that share one or more internal networks, using the block-level data storage service a first group of another multiple plurality of computer systems as block-level data storage systems that provide block-level data storage to multiple running programs, the first not being computer system part of the first group; and being under the control of a first block-level data storage service node manager module that manages the access of the first computer system to internal networks, receive one or more data access requests initiated by the first running program to a logical local storage device of the first computer system representing the non-local block-level data storage provided by the data storage service at the level of block;
automatically respond to received data access requests by interacting over internal networks with a second computer system on behalf of the first running program to carry out data access requests received in a first block-level data storage volume having a primary copy stored in the second computer system and having a mirror copy stored in a third computer system, the second and third computer systems being each part of the first group of block-level data storage systems, the interaction over the internal networks being carried out in a manner that is transparent to the first running program;
after performing the data access requests received on the first block-level data storage volume, automatically define that the primary copy of the first block-level data storage volume on the second computer system is not available ; and after receiving one or more different requests for access to data initiated by the first running program to the first storage device at the local and logical block level of the first computer system, automatically respond to the different requests for access to data received through the interaction through the internal networks with the third internal computer system on behalf of the first running program to carry out the other requests for access to data in the first mirror copy of the block-level data storage volume in the third computer system.
Clause 2. The method of clause 1, which further comprises, under the control of a system administrator module for the block-level data storage service, administering the supply, by the data storage service to block level of the block level data storage to the multiple running programs, including such management:
create multiple block-level data storage volumes to be each used by one or more of the multiple running programs, creating each of the block-level data storage volumes including storing a primary copy of the block-level data storage volume in one of multiple block-level data storage systems and including the storing a mirror copy of the block-level data storage volume in another of multiple block-level data storage systems, the multiple programs running on multiple computer systems featuring multiple associated node manager modules; maintain mirror copies of block-level data storage volumes by performing such modification to block-level data stored in the mirror copy of the block-level data storage volume, when a block-level data storage system that stores the primary copy of one of the created block-level data storage volumes performs a modification of the block-level data stored in said primary copy according to an access request to write data received from a program in
ES 2 575 155 T3 execution; and maintain access to block-level data storage volumes by promoting the mirror copy of a block-level data storage volume to become a new primary copy of a block-level data storage volume, when the primary copy of one of the created block-level data storage volumes is not available.
Clause 3. The method of clause 2 in which the multiple programs are executed by a program execution service in a second group of multiple computer systems of the plurality of computer systems, wherein the computer systems of the second group are different from the systems computer scientists of the first group, in which the first host computer system houses multiple virtual machines each of which is capable of running at least one program, in which the first running program is one of the multiple programs and is a virtual machine image being executed by at least one of the multiple virtual machines hosted by the first computer system, and in which the first administrator module is executed node as part of a virtual machine monitor for the first computer system.
Clause 4. A computer-implemented method for providing block-level data storage functionality to a running program, the method comprising:
receiving one or more indications of a first group of one or more data access requests initiated by a first running program to a local block-level storage device to a first computer system on which the first program is running, the local block level storage device being a logical device representing a non-local block level data storage volume provided by a second distinct data storage system over one or more networks;
automatically respond to the indications received from the first group of data access requests by interacting with the second data storage system over one or more networks on behalf of the first running program to cause the data access requests of the first group on block-level data stored by the block-level data storage volume provided by the second data storage system;
determine that the block-level data storage volume provided by the second data storage system is no longer available and automatically identify a third storage system that contains a mirror copy of the provided block-level data storage volume by the second data storage system, the third data storage system being distinct from the first computer system and the second data storage system;
receiving one or more indications from a second group of one or more data access requests initiated by the first running program to the block-level storage device in the first computer system; and automatically respond to the indications received from the second group of data access requests by interacting with the third data storage system through one or more networks on behalf of the first running program to cause the data access requests to be carried out of the second group on the mirror copy of the block-level data storage volume in the third identified data storage system.
Clause 5. The method of clause 4 further comprising, before receiving the indications of the first group of one or more data access requests, connecting the block-level data storage volume to the first computer system so that the first program in execution use it, connecting the block-level data storage volume to the first computer system, including associating the logical and local block-level storage device for the first computer system to the block-level data storage volume provided by the second data storage system.
Clause 6. The method of clause 5 further comprising, before connecting the block-level data storage volume to the first computer system, create the block-level data storage volume by storing a primary copy of the block-level data storage volume created on the second data storage system and storing the mirror copy of the data storage volume at the block level in the third data storage system, so that performing the data access requests of the first group on the block-level data storage volume provided by the second data storage system are performed on the primary stored copy of the data storage volume at the block level.
Clause 7. The method in clause 6 where the block-level data storage volume is provided by a block-level data storage service, where the volume creation
ES 2 575 155 T3 block-level data storage is performed by a system administrator module of the block-level data storage service in response to a request from a user who is associated with the first running program, and wherein the connection of the block-level data storage volume to the first computer system is performed by a node manager module of the block-level data storage service that manages the access of the first computer system to one or more more of the networks.
Clause 8. The method in clause 4 further comprising, after determining that the block-level data storage volume provided by the second data storage system is no longer available, automatically initiating the creation of another copy of the block-level data storage volume in a fourth data storage system that is different from the first computer system and the second and third data storage systems.
Clause 9. The method of clause 8 in which the block-level data storage volume provided by a second data storage system has become unavailable according to at least one phallus of the second data storage system, in a failure in connectivity to the second data storage system, and in an inability of the second data storage system to reliably access the stored block-level data storage volume.
Clause 10. The method of clause 4 wherein performing at least one of the first group's data access requests includes modifying the block-level data stored in the block-level data storage volume in the second system of data storage, and further includes maintaining the mirror copy of the block-level data storage volume in the third data storage system by modifying the block-level data stored in the mirror copy of the block-level data storage volume in the third data storage system.
Clause 11. The method of clause 4 wherein the first computer system and the second data storage system are a subset of a plurality of computer systems co-located in a single geographic location, wherein the plurality of computer systems includes multiple systems block-level data storage that are provided by a block-level data storage service, and wherein each of the second and third data storage systems is distinct from the multiple block-level data storage systems.
Clause 12. The method of clause 11 in which the only geographic location in which the plurality of computer systems are co-located is a data center, and in which the method further comprises creating at least one copy of the volume of block-level data storage in one or more archival data storage systems of a remote storage service that are located external to the data center.
Clause 13. The method of clause 11 wherein the first program is executed by a program execution service running multiple programs for multiple users on multiple computer systems of the plurality of computer systems, wherein the first host computer system hosts multiple machines virtual machines, each of which is capable of running at least one program, wherein the first running program is one of the multiple programs and is a virtual machine image that is being executed by at least one of the multiple virtual machines hosted by the first computer system, wherein the reception of the indications of the first group of data access requests and the reception of the indications of the second group of the other data access requests and the automatic response to the indications received of the first and second group of requests for data access is done as part of running a virtual machine monitor for the first computer system, and wherein the first computer system includes one or more other actual local storage devices that are available for use by the first running program.
Clause 14. A computer-readable medium whose contents allow one or more computer systems to offer block-level data storage functionality to a running program, by performing a method that comprises:
receiving one or more indications of one or more data access requests initiated by a first program running on a first computer system from a block-level data storage volume, the block-level data storage volume being provided by a second block-level data storage system that is separated from the first computer system by one or more networks;
automatically responding to the received indications of data access requests by initiating performing data access requests in the block-level data storage volume in the second block-level data storage system;
ES 2 575 155 T3 automatically create a mirror copy of the block-level data storage volume in each one or more of the third block-level data storage systems that are different from the first computer system and the second storage system block-level data; and after the block-level data storage volume provided by the second data storage system has become unavailable, automatically respond to one or more prompts received from one or more different data access requests initiated by the first program to the block-level data storage volume by performing other data access requests on the mirror copy of the block-level data volume. block-level data storage created in at least one of the third block-level data storage systems.
Clause 15. The computer-readable medium of clause 14, in which the realization of requests for access to data of the first group, the automatic creation of the mirror copy of the data storage volume at the block level, and the realization of other requests for access to data in the second group are made automatically by a system administrator module of a block-level data warehouse service, wherein the first program initiates data access requests to the block-level data storage volume by interacting with a logical block-level data storage device that is local to the first computer system and represents the volume of data storage provided by the block-level data storage system, wherein receiving data access requests initiated by a first block-level data storage service node manager module that is associated with the first computer system after first program interactions with the device, and wherein the automatic response to the received indications of the data access requests is performed in part under the control of the node manager module when initiating the sending of the received data access requests over one or more networks.
Clause 16. The computer-readable medium of clause 14 in which the second block-level data storage system stores a primary copy of the block-level data storage volume at a time of automatic response to the received prompts from the requests for access to data, and in which the method further comprises, before automatically responding to the indications received from other requests for access to data, automatically define that the primary copy of the block-level data storage volume is no longer available, select the mirror copy of the block-level data storage volume on one of the third block-level data storage systems , and promoting the selected mirror copy of the block-level data storage volume to the third block-level data storage system to be a current primary copy of the block-level data storage volume.
Clause 17. The computer-readable medium of clause 16 in which the automatic creation of the mirror copy of the block-level data storage volume in each one or more of the third block-level data storage systems includes creating a first Mirror copy of the block-level data storage volume to one of the third-party block-level data storage systems before automatically responding to prompts received from the data access requests, where the automatic response to the received indications of the data access requests also includes carrying out the data access requests in the first mirror copy of the block-level data storage volume, in where the first mirror copy of the block-level data storage volume is the selected mirror, and wherein the automatic creation of the mirror copy of the block-level data storage volume in each one or more of the third block-level data storage systems further includes, after defining that the primary copy of the block-level data storage volume block-level data storage is no longer available, create a second mirror copy of the block-level data storage volume on another of the third block-level data storage systems.
Clause 18. The computer-readable medium of clause 14, wherein the second block-level data storage system stores a primary copy of the block-level data storage volume at a time of automatic response to received prompts from data access requests, where the indicated data access requests initiated by the first program are a second group of data access requests initiated by the first program, and in which the method also comprises, before receiving the indications of the requests for access to data from the second group:
create an initial primary copy of the block-level data storage volume in a fourth block-level data storage system other than the first computer system and the second block-level data storage system and one or more of the third
ES 2 575 155 T3 block level data storage systems;
automatically respond to one or more of the indications received from a first group of one or more data access requests initiated by the first program for the block-level data storage volume by performing data access requests from the first group on the primary copy created on the fourth block-level data storage system, the first group of data access requests being initiated by the first program before the second group of data access requests; and after responding to prompts from the first group of data access requests, determining to move the primary copy of the block-level data storage volume from the fourth block-level data storage system to the second storage system block-level data, so that requests for data access initiated by the first program to the block-level data storage volume after moving the primary copy of the data storage volume to the second block-level data storage system are made in the primary copy moved from the block-level data storage volume onto the second block-level data storage system.
Clause 19. The computer-readable medium of clause 18, wherein the definition of moving the primary copy of the block-level data storage volume from the fourth block-level data storage system to the second block-level data storage system is based on at least one automatic definition that the second block-level data storage system is better able than the fourth block-level data storage system to provide the primary copy of the volume of block-level data storage and a request received from a user of the block-level data storage service that is associated with the block-level data storage volume.
Clause 20. The computer-readable medium of clause 14, wherein the first program is one of multiple programs executed by a program execution service on multiple host computer systems on behalf of users of the program execution service in exchange for fees paid by users, where the first computer system is one of multiple host computer systems, and wherein the use of the block-level data storage volume by the first program is done in exchange for a fee paid by the user of the program execution service on whose behalf the first program is executed.
Clause 21. The computer-readable medium of clause 14 wherein the computer-readable medium is at least one of a memory of a computer system that stores the contents and a data transmission medium that includes a generated data signal and stored containing the contents.
Clause 22. The computer-readable medium of clause 14 in which the contents are instructions that, when executed, cause one or more computer systems to carry out the method.
Clause 23. A system configured to provide block-level data storage functionality for running programs, comprising:
one or more memories; and a block-level data storage system manager module configured to provide a block-level data storage service that uses multiple block-level data storage systems to store volumes of block-level data storage created by users of the block-level data warehouse service and accessed by one or more networks on behalf of one or more running programs associated with users, including the provision of block-level data storage service:
creating one or more block-level data storage volumes for use by one or more running programs, creating each of the block-level data storage volumes including creating a primary copy of the block-level data storage volume block-level data storage that is stored in one of multiple block-level data storage systems;
respond to data access requests, each of which is initiated by one of the running programs for one of the block-level data storage volumes created by, for each of the data access requests, if the primary copy of one of the block-level data storage volume for the data access request is available, the start of performing the data access request on the available primary copy; and in response to the received prompt, create a new copy of the first of the block-level data storage volumes, the primary copy of the block-level data storage system being stored in one of the first of the block-level data storage systems. block-level data storage, the newly created copy being stored in one or more of the others
ES 2 575 155 T3 storage systems of the first block-level data storage system.
Clause 24. The system of clause 23 wherein creating the first block-level data storage volume further includes creating a mirror copy of the first block-level data storage volume that is stored in one second of the multiple storage systems. block-level data storage other than the first block-level data storage system and one or more of the other data storage systems, and wherein the response to data access requests further includes, for each data access request that is initiated by one of the running programs for the first block-level data storage volume, if the primary copy the first bulk data storage volume is not available, initiate performing the data access request on the mirror copy of the first block-level data storage volume stored in the second block-level data storage system.
Clause 25. The system of clause 24 in which the response to data access requests further includes, for each data access request that is initiated by one of the running programs for the first block-level data storage volume , if the primary copy of the volume is available, initiate performing the data access request on the mirror copy of the first block-level data storage volume stored in the second block-level data storage system.
Clause 26. The system of clause 24 wherein the response to data access requests further includes, for an initial data access request initiated by one of the running programs for the first block-level data storage volume for which is not available the primary copy of the first block-level data storage volume in the first block-level data storage system, promote the mirror copy of the first block-level data storage volume stored on the second block-level data storage system to become the new primary copy of the first data storage volume on the first storage system block-level data, so that the initiation of the execution of said initial data access request in the mirror copy of the first data storage volume stored in the second block-level data storage system is carried out in the new promoted primary copy of the first block-level data storage volume, and so that the data access requests initiated by one of the running programs for the first block-level data storage volume subsequent to said initial data access request are carried out in the new promoted primary copy of the first block-level data storage volume in the second block-level data storage system.
Clause 27. The clause 26 system in which promoting the mirror copy of the first block-level data storage volume stored in the second block-level data storage system that will become the new primary copy of the first volume Block-level data storage also includes creating a second mirror copy of the first block-level data storage volume that is stored in a third of the multiple storage systems. block-level data storage other than the first and second block-level data storage system, and wherein the response to each of the data access requests subsequent to the initial data access request that are initiated by one of the running programs for the first block-level data storage volume further includes initiating performing the request for the second mirror copy of the first of the first block-level data storage volume stored in the third block-level data storage system.
Clause 28. The system of clause 23 in which the indication received in response to which the new copy of the first data storage volume is created at the block level is based on a definition of moving the primary copy of the first data storage volume block-level from the first block-level data storage system to a second distinct storage system, wherein one or more of the other block-level data storage systems is the second block-level data storage system, wherein creating the new copy of the first block-level data store volume on the second block-level data store includes designating the new copy created on the second block-level data store as a new primary copy of the first block-level data storage volume, and wherein the provision of the block-level data storage service carried out by the data storage system administrator module further includes, after creating the new copy of the first block-level data storage volume , respond to one or more additional data access requests each of which is initiated by one of the programs running for the first block-level data storage volume by initiating the performance of the additional data access requests in the new primary copy of the first block-level data storage volume stored in the second block-level data storage system.
ES 2 575 155 T3
Clause 29. The clause 28 system, wherein the decision to move the primary copy of the block-level data storage volume from the first block-level data storage system to the second block-level data storage system is based on at least one automatic decision that the second block-level data storage system is better able than the first block-level data storage system to provide one or more running programs with access to the primary copy of the first block-level data storage volume and of a request received from a user of the block-level data storage service that is associated with the first block-level data storage volume.
Clause 30. The system of clause 23 in which the indication received in response to which the new copy of the first block-level data storage volume is created is a request to create a snapshot of the first block-level data storage volume block, using a remote storage service of the first block-level data storage volume, wherein one or more other data storage systems are archival data storage systems that are part of the remote storage service, and wherein creating the new block-level data storage volume copy includes interacting with the remote storage service to create the new block-level data storage volume copy as a snapshot of the first block-level data storage volume.
Clause 31. The system of clause 23 further comprising the multiple block-level data storage systems, the multiple block-level data storage systems being co-located in a single geographic location, and further comprising a storage service remote using multiple different data storage systems in a location other than the single geographic location, storing the different data storage systems in a format other than block-level data, and wherein the provision of the block-level data storage service by the data storage service system administrator module to Block-level also includes, for each one or more of the block-level data storage volumes created, the storage of multiple block-level data storage volume portions in one or more of the multiple different data storage systems of the remote storage service, each of the multiple stored portions representing a portion of the data storage volume at the block level that has been modified by one or more data access requests.
Clause 32. The clause 23 system further comprising multiple node manager modules each associated with one or more of the running programs, each of the node manager modules managing the data access requests initiated by one or more associated running programs for one or more of the block-level data storage volumes created by forwarding said data access requests to the block-level data storage service over one or more networks.
Clause 33. The system of clause 23 further comprising a program execution service system administrator module configured to provide a program execution service that uses multiple host computer systems to execute multiple programs for users of the program execution service. , the multiple programs executed by the program execution service including one or more executing programs utilizing one or more created block-level data storage volumes.
Clause 34. The system of clause 23 in which the system includes a first computer system that includes at least one of the memories, and in which the block-level data storage system administrator includes software instructions for its execution by the first computer system that uses the at least one memory.
Clause 35. The system of clause 23 wherein the block-level data storage system manager module consists of one or more means for providing a block-level data storage service using multiple data storage systems to block-level to store block-level data storage volumes that are created by users of the block-level data warehouse service and accessed through one or more more networks on behalf of one or more running programs associated with users, including provision of block-level data storage service:
create one or more block-level data storage volumes to be used by one or more running programs, create each of the block-level data storage volumes including creating a primary copy of the data storage volume to block level that is stored in one of multiple block level data storage systems;
respond to data access requests each initiated by one of the running programs
ES 2 575 155 T3 for one of the block-level data storage volumes created, for each of the data access requests, if the first copy of the block-level data storage volume for the access request to data is available, initiating the realization of the data access request on the available primary copy; and in response to a received indication, creating a new copy of a first of the block-level data storage volumes, the primary copy of the first block-level data storage system being stored in a first block-level data storage system. block-level data, the newly created copy being stored in one or more of the other data storage systems other than the first block-level data storage system.
Clause 36. A method for a computer system of a block-level data storage service to manage access by running programs to remotely stored block-level data, the method comprising:
receive a request to initiate access of a first copy of an indicated program to a non-local block-level data store provided by a block-level data store, the first program copy running in the first of a plurality of computer systems that are co-located in a first geographic location and that share one or more internal networks, the block-level data storage service using a first group of multiple other computer systems of the plurality of computer systems as block-level data storage systems that provide block-level data storage to multiple running programs, the first computer system not being part of the first group;
in response to the received request, connect a first block-level data storage volume to the first computer system to be used by the first running program copy, the first block-level data storage volume presenting a primary copy stored in a second computer system and presenting a mirror copy stored in a third computer system, each first and second computer system being part of the first group of block-level data storage systems, the connection of the first block-level data storage volume including the association of a first logical and local block storage device of the first computer system to the first block-level data storage volume; and under the control of a system administrator module of the block-level data storage service, managing the provision of block-level data storage to the multiple running programs, through:
after receiving indications of one or more data access requests initiated by the first running program copy to the first logical and local block storage device, automatic response by performing data access requests on the primary and mirror copies of the first block-level data storage volume, the data access requests causing one or more modifications to the block-level data stored in the first block-level data storage volume in such a way that the same block-level data stored is maintained in each one of the primary and mirror copies of the first block-level data storage volume;
after carrying out the data access requests, automatically determining that the execution of the first copy of the program has ended, and in response to said determination, automatically maintaining access of the indicated program to the first data storage volume to block level, including maintaining access initiating the execution of a second copy of the indicated program on a fourth different computer system and connecting the first block-level data storage volume to the fourth computer system for use by the second copy of the running program, the connection including associating a second logical local block storage device of the fourth computer system to the first block level data storage volume; and after receiving indications of another request or requests for access to data to the second local and logical block storage device initiated by the second copy of the program in execution, automatic response carrying out the requests for access to data in the primary copy and the mirror copy of the first block-level data storage volume, the data access requests causing one or more additional modifications to the block-level data stored in the first block-level data storage volume such that each of the primary and mirror copies of the first volume block-level data storage stores the same block-level data.
Clause 37. The method of clause 36 in which the connection of the first block-level data storage volume to the first computer system is carried out by a system administrator module of the block-level data storage service that manages the access of the first computer system to the internal networks, and in which the method also comprises, under the control of the
ES 2 575 155 T3 node manager:
receiving requests for access to data from the first local and logical storage device initiated by the first copy of the running program; and facilitate the realization of data access requests by interacting over internal networks with the system administrator module to provide indications of data access requests, the interaction over internal networks being carried out in a way that is transparent for the first running copy of the program.
Clause 38. The method of clause 37 in which the multiple programs are executed by a program execution service in a second group of multiple computer systems of the plurality of computer systems, wherein the computer systems of the second group are different from the systems computer scientists of the first group, in which the first computer system houses multiple virtual machines each of which is capable of executing at least one program, in which the indicated program is one of the multiple programs and is a virtual machine image that is being executed by at least one of the multiple virtual machines hosted by the first computer system, and in which the administrator module of node as part of a virtual machine monitor for the first computer system.
Clause 39. A computer-implemented method for managing access to block-level data storage functionality through running programs, the method comprising:
receiving one or more indications of a first group of one or more access data requests initiated by a first copy of a first program running on a first computer system to access block-level data stored in a data storage volume a non-local block level, the block-level data storage volume being provided by a second distinct data storage system via one or more networks and being connected to a first computer system such that the first copy of the running program initiates requests for access to data to the block-level data storage volume through interactions with a first logical block-level storage device local to the first computing system that represents the volume of data storage at the block level; automatically respond to the prompts received from the first group of data access requests by interacting with the second data storage system on behalf of the first running program copy to initiate the execution of the data access requests of the first group in the block-level data storage volume provided by a second data storage system;
After determining that the first copy of the program is no longer available, identify a third computer system in which a second copy of the first program is running, the third computer system being distinct from the first computer system and the second computer storage system. data, and connecting the block-level data storage volume to the third computer system such that the second program copy has access to a second logical block storage device local to the third computer system representing the data-level storage volume block;
receive one or more indications from a second group of one or more distinct data access requests initiated by the second running program copy for the block-level data storage volume through interactions with the second storage device in local and logic block in the third computer system; and automatically respond to the indications received from the second group of data access requests by interacting with the second data storage system on behalf of the second copy of the running program so as to initiate the execution of access requests to data from the second group in the block-level data storage volume in the second data storage system.
Clause 40. The method of clause 39 further comprises automatically defining that the first copy of the program is no longer available, the unavailability of the first copy of the program being based on at least one failure of the first computer system, on a failure in connectivity with the first computer system, and in an inability of the first computer system to continue with the execution of the first copy of the program.
Clause 41. The method of clause 40 in which the identification of the third computer system on which the second copy of the first program is running and the connection of the block-level data storage volume to the third computer system are carried out automatically to keep the first program's access to the block-level data storage volume, the automatic maintenance of access being performed in response to the automatic determination that the first copy of the program is no longer available.
ES 2 575 155 T3
Clause 42. The method of clause 39 in which at least one of the actions of determining that the first copy of the program is no longer available, identifying the third computer system on which the second copy of the first program is running, and connecting the block-level data storage volume to the third computer system is performed in response to one or more indications received from a user associated with the first program and the block-level data storage volume.
Clause 43. The method of clause 39 in which the identification of the third computer system on which the second copy of the first program is running includes, after determining that the first copy of the program is no longer available, initiating execution of the second copy of the program in the third computer system.
Clause 44. The method of clause 39 wherein the first copy of the first program is one of multiple copies of the first program running on multiple different computer systems, wherein at least one of the multiple copies is an alternate copy of one or more than the other multiple copies, and wherein identifying the third computer system on which the second copy of the first program is running includes selecting the second copy of the first program based on the fact that the second copy is one of the alternate copies.
Clause 45. The method of clause 39 further comprising, before connecting the block-level data storage volume to the first computer system, creating the block-level data storage volume in the second data storage system and creating a mirror copy of the block-level data storage volume to a fourth different block-level data storage system, and wherein responding to first group and second group data access requests further includes initiating first group and second group data access requests on the created mirror.
Clause 46. The method of clause 39 further comprising, prior to receiving indications of one or more data access requests, connecting the block-level data storage volume to the first computer system for use by the first copy of the program in action, connecting the block-level data storage including associating the first logical block-level storage device to the block-level data storage volume provided by the second data storage system and being performed by an administrator module of node that manages the access of the first computer system to one or more networks, the node manager module being part of a block-level data storage service that previously creates the block-level data storage volume in response to a request from a user who is associated with the first running program.
Clause 47. The method of clause 39 wherein the first and third computer systems and the second data storage system are a subset of a plurality of computer systems co-located in a first geographic location, wherein the plurality of computer systems includes multiple block-level data storage systems that are provided by a block-level data storage service, wherein the second and third data storage systems are each distinct from the multiple block-level data storage systems.
Clause 48. The method of clause 47 in which the first geographic location is a data center, in which the first program is run by a program execution service that runs multiple programs for multiple users on multiple pluralities of computer systems at the center data, in which the first computer system hosts multiple virtual machines, each of which is capable of running at least one program, wherein the first running program is one of the multiple programs and is a virtual machine image run by at least one of the multiple virtual machines hosted by the first computer system, wherein the reception of the indications of the first group of data access requests and the reception of the indications of the second group of the other data access requests and the automatic response to the received indications of the first and second group of access requests data are performed as part of the execution of a virtual machine monitor for the first computer system, and wherein the first computer system includes one or more actual local storage devices that are available for use by the first running program.
Clause 49. A computer-readable medium whose contents allow one or more computer systems to manage access to block-level data storage functionality through running programs, by carrying out a method that comprises:
provide access to a non-local block-level data storage volume for a first program running on a first computer system, allowing the granted access the first program to initiate data access requests for the data storage volume
ES 2 575 155 T3 at the block level, the data storage volume being provided by a second data storage system that is separated from the first computer system by one or more networks;
automatically respond to one or more indications received from one or more data access requests initiated by the first program for the block-level data storage volume, the response including the initiation of data access requests in the volume of block-level data storage in the second block-level data storage system;
after the first program running on the first computer system is no longer available, granting access to the block-level data storage volume for a second program running on a third computer system in lieu of the access granted to the first program; and automatically respond to one or more indications received from one or more different data access requests initiated by the second program for the block-level data storage volume, the response including the initiation of the other access requests to data in the block-level data storage volume in the second block-level data storage system.
Clause 50. The computer-readable medium of clause 49 in which the first program and the second program are running copies of a single program, wherein granting access for the first program includes connecting the block-level data storage volume to the first computer system such that the first running program initiates data access requests for the data storage volume to block level by means of interactions with a first logical block level storage device local to the first computer system representing the volume of data storage at the level block, wherein granting access to the second program includes identifying the third computer system on which the second program is running and connecting the block-level data storage volume to the third computer system so that the second program has access to the second local logical block storage device to the third computer system representing the block level data storage volume, and wherein the block-level data storage volume data access requests by the first and second programs are for accessing block-level data stored in the block-level data storage volume.
Clause 51. The computer-readable medium of clause 50 in which the first and third computer systems and the second block-level data storage system are located in a single geographic location, in which the first computer system and the second system of block-level data storage are separated by one or more networks, wherein the realization of data access requests and other data access requests in the block-level data storage volume in the second block-level data storage system are performed automatically by a module as a system administrator of a block-level data warehousing service, wherein the first running program is granted access to the block-level data storage volume by a node manager module of the block-level data storage service that is associated with the first computer system and manages the access of the first running program to one or more networks, and wherein the automatic response to the received indications of the data access requests is carried out in part under the control of the node manager module when initiating the sending of the received data access requests via one or more networks towards the second block-level data storage system.
Clause 52. The computer-readable medium of clause 51 in which the second running program is granted access to the block-level data storage volume by a node manager module that further manages the second running program's access to a or more networks, in which the third and first computer systems are part of a single physical computer system, and wherein the automatic response to the received indications of the data access requests is carried out in part under the control of the node administrator module when initiating the sending of the received requests for access to other data by means of a or more networks to the second block-level data storage system.
Clause 53. The computer-readable medium of clause 49 in which the computer-readable medium is at least one of a memory of a computer system that stores the contents and a data transmission medium that includes a generated data signal and stored containing the contents.
Clause 54. The computer-readable medium of clause 49 in which the contents are instructions that, when executed, cause one or more computer systems to carry out the method.
Clause 55. A system configured to manage access to data storage functionality
ES 2 575 155 T3 at the block level through running programs, comprising:
one or more memories;
and a block-level data storage system configured to provide a block-level data storage service that utilizes multiple block-level data storage systems to store user-created block-level data storage volumes block-level data storage service and accessed through one or more networks on behalf of one or more running programs associated with users, including the provision of block-level data storage service:
create one or more block-level data storage volumes for use by one or more running programs, each of the block-level data storage volumes being stored in one of multiple data storage systems to block level;
After access to the first of the block-level data storage volumes is granted to the first of one or more running programs, respond to one or more data access requests initiated by the first program for the first volume of block-level data storage created by initiating the making of one or more data access requests in the first created block-level data storage volume; and after the first program becomes unavailable and access to the first bulk created data storage volume is granted to a second running program to replace the access granted to the first unavailable program, respond to one or more different data access requests initiated by the second program for the first block-level data storage volume created by initiating one or more different data access requests on the first block-level data storage volume. block-level data created.
Clause 56. The clause 55 system in which the first program and the second program are running copies of a single program, wherein granting access to the first program includes connecting the first block-level data storage volume to a first computer system running the first program such that the first running program initiates data access requests for the first data storage volume through interactions with a first logical block level storage device local to the first computer system representing the first volume block-level data storage, wherein granting access to the second program includes identifying a third computer system on which the second program is running and connecting the first block-level data storage volume to the third computer system such that the second program has access to the second device logical block-level storage local to the third computer system representing the first block-level data storage volume, and wherein the data access requests for the first block-level data storage volume by the first and second programs are for accessing block-level data stored in the block-level data storage volume.
Clause 57. The system of clause 56 wherein the first block-level data storage volume is stored in a first block-level data storage system, wherein the first and third computer systems and the first storage system block-level data is located in a single geographic location and separated by one or more networks, and wherein the system further comprises one or more node manager modules of the block-level data storage service associated with the first and second program to manage the access of the first and second program to one or more networks, so that the response to data access requests is carried out in part under the control of one or more system administrator modules when initiating the sending of data access requests through one or more networks to the first system block-level data storage.
Clause 58. The system of clause 55 in which a first copy of the first block-level data storage volume is stored in a first block-level data storage system, wherein performing the data access requests initiated by the first program are performed on the first copy of the first block-level data storage volume, wherein the first program runs on a first computer system that is co-located with the first block-level data storage system at a first geographic location, wherein the second program runs on a second computer system at a second different geographical location, and wherein granting access to the first block-level data storage volume created to the second running program includes initiating the connection of the second computer system to a second block-level data storage system at the second geographic location that stores a second copy of the first block-level data storage volume.
ES 2 575 155 T3
Clause 59. The system of clause 55 in which the system includes a first computer system that includes at least one of the memories, and in which the block-level data storage system administrator includes software instructions for its execution by the first computer system that uses at least one memory.
Clause 60. The system of clause 55 wherein the block-level data storage system manager module consists of one or more means for providing a block-level data storage service using multiple data storage systems to block-level to store block-level data warehouse volumes that are created by users of the block-level data warehouse service and accessed by one or more networks on behalf of one or more running programs associated with users, including the provision of block-level data storage service:
create one or more block-level data storage volumes for use by one or more running programs, each of the block-level data storage volumes being stored in one of multiple data storage systems to block level;
after access to the first of the block-level data storage volumes is granted to a first of one or more running programs, respond to one or more data access requests initiated by the first program for the first volume block-level data storage created by initiating one or more of the data access requests in the first block-level data storage volume created; and after the first program becomes unavailable and access to the first bulk created data storage volume is granted to a second, different running program to replace the access granted to the first unavailable program, respond to one or more different data access requests initiated by the second program to the first block-level data store volume created by initiating one or more different data access requests on the first data store volume at the created block level.
Contents12
13 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13
37 members in 6 offices
Priority claims4
| Document | Office | Kind | Date |
|---|---|---|---|
| 188943 | United States of America | – | |
| 18894308 | United States of America | A | |
| 188949 | United States of America | – | |
| 18894908 | United States of America | A |
Members37
| Document | Office | Kind | |
|---|---|---|---|
| US2010036850A1 | United States of America | A1 | |
| US2010036851A1 | United States of America | A1 | |
| WO2010017492A1 | World Intellectual Property Organization (WIPO) | A1 | |
| EP2324429A1 | European Patent Office (EPO) | A1 | |
| US8015343B2 | United States of America | B2 | |
| CN102177508A | China | A | |
| US8019732B2 | United States of America | B2 | |
| JP2011530748A | Japan | A | |
| US2012042142A1 | United States of America | A1 | |
| EP2426605A1 | European Patent Office (EPO) | A1 | |
| US2012060006A1 | United States of America | A1 | |
| JP2012053878A | Japan | A | |
| EP2324429A4 | European Patent Office (EPO) | A4 | |
| JP5015351B2 | Japan | B2 | |
| JP5118243B2 | Japan | B2 | |
| CN103645953A | China | A | |
| US8769186B2 | United States of America | B2 | |
| US8806105B2 | United States of America | B2 | |
| US2014281317A1 | United States of America | A1 | |
| US2014317370A1 | United States of America | A1 | |
| CN102177508B | China | B | |
| CN105138435A | China | A | |
| US9262273B2 | United States of America | B2 | |
| EP2426605B1 | European Patent Office (EPO) | B1 | |
| ES2575155T3This record | Spain | T3 | |
| EP3037970A1 | European Patent Office (EPO) | A1 | |
| US9529550B2 | United States of America | B2 | |
| CN103645953B | China | B | |
| US2017075606A1 | United States of America | A1 | |
| CN105138435B | China | B | |
| EP3699765A1 | European Patent Office (EPO) | A1 | |
| EP2324429B1 | European Patent Office (EPO) | B1 | |
| US10824343B2 | United States of America | B2 | |
| US2021064251A1 | United States of America | A1 | |
| US11768609B2 | United States of America | B2 | |
| US2023384948A1 | United States of America | A1 | |
| US12223182B2 | United States of America | B2 |
Numbers
- Publication
- 2575155
- Application
- 11009559
Titles2
- Spanish
- Suministro de programas de ejecución con acceso fiable al almacenamiento de datos a nivel de bloque no local
- English
- Provision of execution programs with reliable access to data storage at the non-local block level
Classification
- CPC, 6
- G06F11/2046
- G06F11/2071
- G06F11/2094
- G06F11/2033
- G06F11/1448
- G06F11/2056
- IPC, 4
- G06F11 20
- G06F21 53
- G06F21 60
- G06F21 62