Internally consistent file system image in distributed object-based data storage
Claim Score by NHIP
Abstract
A system and method to perform a system-wide file system image without time smear in a distributed object-based data storage system. A realm manager is elected as an image master using the Distributed Consensus Algorithm to execute image-taking. All pending write capabilities are invalidated prior to taking the system-wide file system image so as to quiesce the realm and prepare the storage system for the system-wide image. Once the system is quiesced, the image master instructs each storage manager in the system to clone each live object group contained therein without explicitly cloning any objects contained in such object group. In one embodiment, a file manager copies an object in the system before a write operation is performed on that object after the image is taken. Neither the cloning operation nor the copying operation update any directory objects in the system. At run time, a client application may use a mapping scheme to access objects contained in the system-wide image.

Term
Term ended
Projected expiry passed 31 March 2024, 2.5 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
69 claims: 18 independent, 51 dependent
- 1A method of initiating an internally-consistent system-wide file system image in an object-based distributed data storage system comprising:providing a plurality of management entities each of which maintains a record representing a configuration of a portion of said data storage system;and using a Distributed Consensus Algorithm (DCA) to elect one of said plurality of management entities to serve as a first image master to coordinate execution of said system-wide file system image.
- 7A computer-readable storage medium containing a program code, which, upon execution by a processor in an object-based distributed data storage system, causes said processor to perform the following:provide a plurality of management entities each of which maintains a record representing a configuration of a portion of said data storage system;and use a Distributed Consensus Algorithm (DCA) to elect one of said plurality of management entities to serve as an image master to coordinate execution of a system-wide file system image in said data storage system.
- 8An object-based data storage system comprising:means for maintaining a record representing a configuration of a portion of said data storage system;and means for electing one of said record maintaining means as an image master using a Distributed Consensus Algorithm (DCA), wherein said image master is configured to coordinate execution of an internally-consistent system-wide file system image.
- 9A method of achieving a quiescent state in a storage system having a plurality of object-based secure disks (OBDs) and a plurality of executable client applications, wherein each client application, upon execution, is configured to access one or more of said plurality of object-based secure disks, said method comprising:defining a plurality of capabilities required to perform a data write operation on corresponding one or more of said plurality of object-based secure disks;granting one or more of said plurality of capabilities to each of said plurality of client applications;and invalidating each of said plurality of capabilities so long as said quiescent state is to be maintained, thereby preventing each client application from accessing one or more of corresponding object-based secure disks to perform said data write operation thereon during said quiescent state.
- 16A computer-readable storage medium containing a program code, which, upon execution by a processor in an object-based distributed data storage system, causes said processor to perform the following:define a plurality of capabilities required to perform a data write operation on corresponding one or more of a plurality of object-based secure disks in said data storage system;grant one or more of said plurality of capabilities to each of a plurality of client applications configured to access one or more of said plurality of object-based secure disks;and invalidate each of said plurality of capabilities so long as a quiescent state is to be maintained, thereby preventing each client application from accessing one or more of corresponding object-based secure disks to perform said data write operation thereon during said quiescent state.
- 17An object-based data storage system comprising:a plurality of object-based secure disks (OBDs);a plurality of executable client applications, wherein each client application, upon execution, is configured to access one or more of said plurality of OBDs to perform a data write operation thereon;means for generating a plurality of capabilities required to perform said data write operation;means for receiving a corresponding request from each of said plurality of client applications for one or more of said plurality of capabilities;means for issuing one or more of said plurality of capabilities to each of said plurality of client applications in response to said request received therefrom;and means for preventing recognition of each of said plurality of capabilities so long as a quiescent state is to be maintained for said storage system, thereby preventing each said client application from performing said data write operation during said quiescent state.
- 18A method of performing an internally-consistent system-wide file system image in an object-based data storage system comprising:receiving a request for said system-wide file system image;quiescing said object-based data storage system in response to said request for said image;cloning each live object group stored in said data storage system after said data storage system is quiesced without cloning any object contained in each said object group during said system-wide file system image;and responding to said request for said system-wide file system image.
- 32An object-based data storage system comprising:means for receiving a request for an internally-consistent system-wide file system image;means for quiescing said object-based data storage system in response to said request for said image;means for cloning each object group stored in said data storage system after said data storage system is quiesced without primarily cloning any object contained in each said object group during said system-wide file system image;and means for responding to said request for said system-wide file system image.
- 33A computer-readable storage medium containing a program code, which, upon execution by a processor in an object-based distributed data storage system, causes said processor to perform the following:receive a request for a system-wide file system image in said data storage system;quiesce said object-based data storage system in response to said request for said image;clone each object group stored in said data storage system after said data storage system is quiesced without cloning any object contained in each said object group during said system-wide file system image;and respond to said request for said system-wide file system image.
- 34A method of performing an internally-consistent system-wide file system image in an object-based data storage system comprising:receiving a request for said system-wide file system image;placing a dummy image directory in a root directory object in said object-based data storage system;quiescing said object-based data storage system in response to said request for said image;informing each file manager in said data storage system about a timing of said system-wide file system image;indicating completion of said system-wide file system image;and configuring each said file manager to copy each corresponding object managed thereby prior to authorizing a write operation to said object after said completion of said system-wide file system image.
- 48An object-based data storage system comprising:means for receiving a request for an internally-consistent system-wide file system image;means for placing a dummy image directory in a root directory object in said object-based data storage system;means for quiescing said object-based data storage system in response to said request for said image;means for informing each file manager in said data storage system about a timing of said system-wide file system image;means for indicating completion of said system-wide file system image;and means for configuring a file manager to copy an object managed thereby prior to authorizing a write operation to said object after said completion of said system-wide file system image.
- 49A computer-readable storage medium containing a program code, which, upon execution by a processor in an object-based distributed data storage system, causes said processor to perform the following:receive a request for a system-wide file system image in said object-based data storage system;place a dummy image directory in a root directory object in said object-based data storage system;quiesce said object-based data storage system in response to said request for said image;inform each file manager in said data storage system about a timing of said system-wide file system image;indicate completion of said system-wide file system image;and configure each said file manager to copy each corresponding object managed thereby prior to authorizing a write operation to said object after said completion of said system-wide file system image.
- 50A method of performing an internally-consistent system-wide file system image in an object-based data storage system comprising:preparing said object-based data storage system for said system-wide file system image;and performing said system-wide file system image without updating any directory objects stored in said object-based data storage system during said system-wide file system image.
- 60A computer-readable storage medium containing a program code, which, upon execution by a processor in an object-based distributed data storage system, causes said processor to perform the following:prepare said object-based data storage system for an internally-consistent system-wide file system image;and perform said system-wide file system image without updating any directory objects stored in said object-based data storage system during said system-wide file system image.
- 61Broadest claimClaim Score 85, broad(NHIP)An object-based data storage system comprising:means for receiving a request for a system-wide file system image;means for quiescing said object-based data storage system in response to said request;and means for performing said system-wide file system image without updating any directory objects stored in said object-based data storage system during said system-wide file system image.
- 62In an object-based data storage system having a plurality of object-based secure disks (OBDs) and a storage manager facilitating data storage in one or more of said plurality of OBDs, a method of avoiding a need to rewrite metadata for an object, stored by said storage manager on one of said plurality of OBDs, when a system-wide file system image in said object-based data storage system is taken, said method comprising:obtaining information identifying said file system image;using said image identifying information, dynamically obtaining a mapping of a non-image identity of each object group appearing in a file path for said object into a corresponding identity of each said object group in said file system image;and for each said object group in said file path for said object, dynamically substituting said corresponding identity in said file system image in place of respective non-image identity therefor when accessing a version of said object in said file system image.
- 68An object-based data storage system comprising:a plurality of object-based secure disks (OBDs);a storage manager facilitating data storage in one or more of said plurality of OBDs;a realm manager configured to maintain a record of a file storage configuration for the files stored by said storage manager;and an executable client application, wherein said client application, upon execution, is configured to perform the following at run time: access a first OBD to obtain identity of a system-wide file system image from an image directory object stored thereon, and obtain a mapping from said realm manager of each object group corresponding to a respective file directory object in a file path for an object stored by said storage manager on a second OBD, wherein said realm manager maps a non-image identity of each said object group into a corresponding identity of each said object group in said file system image.
- 69A computer-readable storage medium containing a program code, which, upon execution by a processor in an object-based distributed data storage system, causes said processor to perform the following at run time to access an image version of an object stored in said data storage system:obtain information identifying a system-wide file system image in said data storage system;using said image identifying information, obtain a mapping of a non-image identity of each object group appearing in a file path for an object stored in said data storage system into a corresponding identity of each said object group in said file system image;and for each said object group in said file path for said object, substitute said corresponding identity in said file system image in place of respective non-image identity therefor when accessing said image version of said object in said file system image.
Independent claims18
110 paragraphs in 5 sections, as filed
REFERENCE TO RELATED APPLICATIONS
P-0001[0001] This application claims priority benefits of prior filed co-pending U.S. provisional patent applications Serial No. 60/368,785, filed on Mar. 29, 2002 and Serial No. 60/372,027, filed on Apr. 12, 2002, the disclosures of both of which are incorporated herein by reference in their entireties.
BACKGROUND
[0002] 1. Field of the Invention
[0003] The present invention generally relates to data storage systems and methods, and, more particularly, to methodologies for internally consistent system-wide file system image in a distributed object-based data storage system.
[0004] 2. Description of Related Art
[0005] With increasing reliance on electronic means of data communication, different models to efficiently and economically store a large amount of data have been proposed. A data storage mechanism requires not only a sufficient amount of physical disk space to store data, but various levels of fault tolerance or redundancy (depending on how critical the data is) to preserve data integrity in the event of one or more disk failures. One way of providing fault tolerance is to periodically take images or copies of various files stored in the data storage system to thereby store the file data for recovery purposes in the event that a disk failure occurs in the system. Thus, imaging is useful in facilitating system backups and related data integrity maintenance activity.
[0006] The term “image”, as used hereinbelow, refers to an immutable image or “copy” of some or all content of the file system at some point in time. Further, an image is said to be “internally consistent” or “crash consistent” if it logically occurs at a point in time at which no write activity is occurring anywhere in the data storage system. This guarantees that no files are left in an inconsistent state because of in-flight writes. On the other hand, an image is said to be “externally consistent” if the file system interacts with an external application program to assure that the external program is at a point from which it can be restarted, and flushed all of its buffers to storage, prior to taking the image. Both internal and external consistencies are desirable in order to guarantee that an image represents data from which an application can be reliably restarted.
[0007] The term “time smear” as used herein refers to an event where an image of the files in a distributed file processing system is not consistent with regard to the time that each piece of data was copied. In other words, “time smear” is the name for the effect of having two files in the image, wherein the contents of file A represent file A's state at some time T<sub>0</sub>, and the contents of file B represent file B's state at some other time T<sub>1</sub>≠T<sub>0</sub>. Such time smear in the images of different data files may occur when the data files are stored at different storage locations (or disks) and the server taking the image accesses, for example, separate data storage facilities consecutively over a period of time. For example, an image of a first data file in a first data storage facility may be taken at midnight, whereas an image of a second data file in a second data storage facility may be taken at one second after midnight, an image of a third data file in a third data storage facility may be taken at two seconds after midnight, and so on.
[0008] Another adverse result of time smear occurs when a single data file is saved on multiple machines or storage disks. In such a situation, if a save operation occurs nearly simultaneously with the image-taking operation, portions of the information contained in the image may correspond to different saved versions of the same file. If that file is then recovered from the image, the recovered file may not be usable because it contains data from different saves, resulting in inconsistent data and causing the file to be potentially corrupt.
[0009] Therefore, it is desirable to devise an image-taking methodology that substantially eliminates the time smear problem associated with prior art image mechanisms. To that end, it is desirable to obtain a time-wise consistent image of the entire file system in a distributed file processing environment. It is also desirable to simultaneously store multiple images online and to delete any image without affecting the content or availability of other images stored in the system.
SUMMARY
[0010] In one embodiment, the present invention contemplates a method of initiating a system-wide file system image in an object-based distributed data storage system. The method comprises providing a plurality of management entities (or realm managers) each of which maintains a record representing a configuration of a portion (which may include some part or all) of the data storage system; and using a Distributed Consensus Algorithm (DCA) to elect one of the plurality of management entities to serve as an image master to coordinate execution of the system-wide file system image.
[0011] In another embodiment, the present invention contemplates a method of achieving a quiescent state in a storage system having a plurality of object-based secure disks and a plurality of executable client applications, wherein each client application, upon execution, is configured to access one or more of the plurality of object-based secure disks. The method comprises defining a plurality of capabilities required to perform a data write operation on corresponding one or more of the plurality of object-based secure disks; granting one or more of the plurality of capabilities to each of the plurality of client applications; and invalidating each of the plurality of capabilities so long as the quiescent state is to be maintained, thereby preventing each client application from accessing one or more of corresponding object-based secure disks to perform the data write operation thereon during the quiescent state.
[0012] Upon receiving a request for a system-wide file system image, the image master establishes a quiescent state in the realm, thereby preparing the realm for the image. After the realm is quiesced, the system-wide image is performed, according to one embodiment of the present invention, by cloning each live object group stored in the data storage system without cloning any object contained in each such object group; and thereafter responding to the request for the system-wide file system image. In responding to the image request, the image master may create an image directory in the root object to record image-related information therein, thereby making the image accessible to other applications in the system.
[0013] In a still further embodiment, the present invention contemplates a method of performing a system-wide file system image wherein, in addition to quiescing the object-based data storage system in response to the request for the image, the method further includes placing a dummy image directory in the root directory object after receiving the request for the image. The method also includes informing each file manager in the data storage system about a timing of the system-wide file system image, indicating completion of the system-wide file system image; and configuring a file manager to copy an object managed thereby prior to authorizing a write operation to the object after the completion of the system-wide file system image. At the completion of the image, the image master converts the dummy image directory into a final image directory, indicating that the image directory now represents a valid image.
[0014] In one embodiment, the system-wide file system image is performed without updating any directory objects stored in the object-based data storage system during image-taking. The directory objects are also not updated after the completion of the image-taking. Neither the object group cloning operation nor the object copying operation update any directory object in the system. The correspondence between live directory objects and image objects is also established without updating the live directory objects to provide such correspondence.
[0015] In a still further embodiment, the present invention contemplates a method of avoiding a need to rewrite metadata for an object, stored by a storage manager on one of the plurality of object-based secure disks, when a file system image is taken. The method comprises obtaining information identifying the file system image; using the image identifying information, dynamically obtaining a mapping of a non-image identity of each object group appearing in a file path for the object into a corresponding identity of each object group in the file system image; and for each object group in the file path for the object, dynamically substituting corresponding identity in the file system image in place of respective non-image identity therefor when accessing a version of the object in the file system image. Thus, in order to traverse the image domain, the client obtains the mapping and performs the substitution step—both at run time—when doing the path name resolution in the image domain.
[0016] The image methodology according to the present invention allows taking system-wide file system images without time smear and without the need to pre-schedule the images (because of the implementation of capability invalidation). Further, there is no hard limit on the number of images that can be simultaneously kept on line. The images are performed without a significant overhead on system I/O operations. Because no directory objects are updated either during or after the image, the failure-handling procedures in the system are also simplified.
BRIEF DESCRIPTION OF THE DRAWINGS
P-0017[0017] The accompanying drawings, which are included to provide a further understanding of the invention and are incorporated in and constitute a part of this specification, illustrate embodiments of the invention that together with the description serve to explain the principles of the invention. In the drawings:
P-0018[0018]FIG. 1 illustrates an exemplary network-based file storage system designed around Object Based Secure Disks (OBSDs or OBDs);
P-0019[0019]FIG. 2 illustrates an implementation where various managers shown individually in FIG. 1 are combined in a single binary file;
P-0020[0020]FIG. 3 is a simplified diagram of the process when a client first establishes a contact with the data file storage system according to the present invention;
P-0021[0021]FIG. 4 illustrates an exemplary flowchart of steps involved in performing a system-wide image according to one embodiment of the present invention in the distributed object-based data storage systems shown in FIGS. 1 and 2;
P-0022[0022]FIG. 5 illustrates creation of a dummy image directory during the image request phase in the image process shown in FIG. 4;
P-0023[0023]FIG. 6 shows quiescent phase-related events in the image process in FIG. 4;
P-0024[0024]FIG. 7 depicts the reply phase-related events in the image process in FIG. 4;
P-0025[0025]FIG. 8 illustrates an exemplary tree structure with files and directories in the live tree and an image tree immediately after an image is taken according to the image process shown in FIG. 4;
P-0026[0026]FIG. 9 shows an image tree constructed in a distributed object-based data storage system using the leaf node approach;
P-0027[0027]FIG. 10 illustrates an exemplary flowchart of steps involved in performing a system-wide image according to another embodiment of the present invention;
P-0028[0028]FIG. 11 depicts the reply phase-related events in the image process in FIG. 10;
P-0029[0029]FIG. 12 shows a three-level storage configuration for objects stored in the object-based data storage systems in FIGS. 1 and 2;
P-0030[0030]FIG. 13 shows a simplified illustration of how the live and cloned object groups are contained in a storage manager;
P-0031[0031]FIG. 14 illustrates an exemplary set of steps carried out in an object group-based mapping scheme to access a specific live file object; and
P-0032[0032]FIG. 15 illustrates an exemplary mapping according to one embodiment of the present invention when a client application accesses an object in an image.
DETAILED DESCRIPTION
P-0033[0033] Reference will now be made in detail to the preferred embodiments of the present invention, examples of which are illustrated in the accompanying drawings. It is to be understood that the figures and descriptions of the present invention included herein illustrate and describe elements that are of particular relevance to the present invention, while eliminating, for purposes of clarity, other elements found in typical data storage systems or networks.
P-0034[0034] It is worthy to note that any reference in the specification to “one embodiment” or “an embodiment” means that a particular feature, structure or characteristic described in connection with the embodiment is included in at least one embodiment of the invention. The appearances of the phrase “in one embodiment” at various places in the specification do not necessarily all refer to the same embodiment.
P-0035[0035]FIG. 1 illustrates an exemplary network-based file storage system <b>10</b> designed around Object Based Secure Disks (OBSDs or OBDs) <b>12</b>. The file storage system <b>10</b> is implemented via a combination of hardware and software units and generally consists of managers <b>14</b>, <b>16</b>, <b>18</b>, and <b>22</b>, OBDs <b>12</b>, and clients <b>24</b>, <b>26</b>. It is noted that FIG. 1 illustrates multiple clients, OBDs, and managers—i.e., the network entities—operating in the network environment. However, for the ease of discussion, a single reference numeral is used to refer to such entity either individually or collectively depending on the context of reference. For example, the reference numeral “<b>12</b>” is used to refer to just one OBD or a group of OBDs depending on the context of discussion. Similarly, the reference numerals <b>14</b>-<b>22</b> for various managers are used interchangeably to also refer to respective servers for those managers. For example, the reference numeral “<b>14</b>” is used to interchangeably refer to the software file managers (FM) and also to their respective servers depending on the context. It is noted that each manager is an application program code or software running on a corresponding server. The server functionality may be implemented with a combination of hardware and operating software. For example, each server in FIG. 1 may be a Windows NT® server. Thus, the file system <b>10</b> in FIG. 1 is an object-based distributed data storage system implemented in a client-server configuration.
P-0036[0036] The network <b>28</b> may be a LAN (Local Area Network), WAN (Wide Area Network), MAN (Metropolitan Area Network), SAN (Storage Area Network), wireless LAN, or any other suitable data communication network including a TCP/IP (Transmission Control Protocol/Internet Protocol) based network (e.g., the Internet). A client <b>24</b>, <b>26</b> may be any computer (e.g., a personal computer or a workstation) electrically attached to the network <b>28</b> and running appropriate operating system software as well as client application software designed for the system <b>10</b>. FIG. 1 illustrates a group of clients or client computers <b>24</b> running on Microsoft Windows® operating system, whereas another group of clients <b>26</b> are running on the Linux® operating system. The clients <b>24</b>, <b>26</b> thus present an operating system-integrated file system interface. The semantics of the host operating system (e.g., Windows®, Linux®, etc.) may preferably be maintained by the file system clients.
P-0037[0037] The manager (or server) and client portions of the program code may be written in C, C<sup>++</sup>, or in any other compiled or interpreted language suitably selected. The client and manager software modules may be designed using standard software tools including, for example, compilers, linkers, assemblers, loaders, bug tracking systems, memory debugging systems, etc.
P-0038[0038] In one embodiment, the manager software and program codes running on the clients may be designed without knowledge of a specific network topology. In that case, the software routines may be executed in any given network environment, imparting software portability and flexibility in storage system designs. However, it is noted that a given network topology may be considered to optimize the performance of the software applications running on it. This may be achieved without necessarily designing the software exclusively tailored to a particular network configuration.
P-0039[0039]FIG. 1 shows a number of OBDs <b>12</b> attached to the network <b>28</b>. An OBSD or OBD <b>12</b> is a physical disk drive that stores data files in the network-based system <b>10</b> and may have the following properties: (1) it presents an object-oriented interface rather than a sector-based interface (wherein each “block” on a disk contains a number of data “sectors”) as is available with traditional magnetic or optical data storage disks (e.g., a typical computer hard drive); (2) it attaches to a network (e.g., the network <b>28</b>) rather than to a data bus or a backplane (i.e., the OBDs <b>12</b> may be considered as first-class network citizens); and (3) it enforces a security model to prevent unauthorized access to data stored thereon.
P-0040[0040] The fundamental abstraction exported by an OBD <b>12</b> is that of an “object,” which may be defined as a variably-sized ordered collection of bits. Contrary to the prior art block-based storage disks, OBDs do not export a sector interface (which guides the storage disk head to read or write a particular sector on the disk) at all during normal operation. Objects on an OBD can be created, removed, written, read, appended to, etc. OBDs do not make any information about particular disk geometry visible, and implement all layout optimizations internally, utilizing lower-level information than can be provided through an OBD's direct interface with the network <b>28</b>. In one embodiment, each data file and each file directory in the file system <b>10</b> are stored using one or more OBD objects.
P-0041[0041] In a traditional networked storage system, a data storage device, such as a hard disk, is associated with a particular server or a particular server having a particular backup server. Thus, access to the data storage device is available only through the server associated with that data storage device. A client processor desiring access to the data storage device would, therefore, access the associated server through the network and the server would access the data storage device as requested by the client.
P-0042[0042] On the other hand, in the system <b>10</b> illustrated in FIG. 1, each OBD <b>12</b> communicates directly with clients <b>24</b>, <b>26</b> on the network <b>28</b>, possibly through routers and/or bridges. The OBDs, clients, managers, etc., may be considered as “nodes” on the network <b>28</b>. In system <b>10</b>, no assumption needs to be made about the network topology (as noted hereinbefore) except that each node should be able to contact every other node in the system. The servers (e.g., servers <b>14</b>, <b>16</b>, <b>18</b>, etc.) in the network <b>28</b> merely enable and facilitate data transfers between clients and OBDs, but the servers do not normally implement such transfers.
P-0043[0043] In one embodiment, the OBDs <b>12</b> themselves support a security model that allows for privacy (i.e., assurance that data cannot be eavesdropped while in flight between a client and an OBD), authenticity (i.e., assurance of the identity of the sender of a command), and integrity (i.e., assurance that in-flight data cannot be tampered with). This security model may be capability-based. A manager grants a client the right to access the data storage (in one or more OBDs) by issuing to it a “capability.” Thus, a capability is a token that can be granted to a client by a manager and then presented to an OBD to authorize service. Clients may not create their own capabilities (this can be assured by using known cryptographic techniques), but rather receive them from managers and pass them along to the OBDs.
P-0044[0044] A capability is simply a description of allowed operations. A capability may be a set of bits (<b>1</b>'s and <b>0</b>'s) placed in a predetermined order. The bit configuration for a capability may specify the operations for which that capability is valid. Thus, there may be a “read capability,” a “write capability,” a “set-attribute capability,” etc. Every command sent to an OBD may need to be accompanied by a valid capability of the appropriate type. A manager may produce a capability and then digitally sign it using a cryptographic key that is known to both the manager and the appropriate OBD, but unknown to the client. The client will submit the capability with its command to the OBD, which can then verify the signature using its copy of the key, and thereby confirm that the capability came from an authorized manager (one who knows the key) and that it has not been tampered with in flight. An OBD may itself use cryptographic techniques to confirm the validity of a capability and reject all commands that fail security checks. Thus, capabilities may be cryptographically “sealed” using “keys” known only to one or more of the managers <b>14</b>-<b>22</b> and the OBDs <b>12</b>. In one embodiment, only the realm managers <b>18</b> may ultimately compute the keys for capabilities and issue the keys to other managers (e.g., the file managers <b>14</b>) requesting them. A client may return the capability to the manager issuing it or discard the capability when the task associated with that capability is over.
P-0045[0045] A capability may also contain a field called the Authorization Status (AS), which can be used to revoke or temporarily disable a capability that has been granted to a client. Every object stored on an OBD may have an associated set of attributes, where the AS is also stored. Some of the major attributes for an object include: (1) a device_ID identifying, for example, the OBD storing that object and the file and storage managers managing that object; (2) an object-group_ID identifying the object group containing the object in question; and (3) an object_ID containing a number randomly generated (e.g., by a storage manager) to identify the object in question. If the AS contained in a capability does not exactly match the AS stored with the object, then the OBD may reject the access associated with that capability. A capability may be a “single-range” capability that contains a byte range over which it is valid and an expiration time. The client may be typically allowed to use a capability as many times as it likes during the lifetime of the capability. Alternatively, there may be a “valid exactly once” capability.
P-0046[0046] It is noted that in order to construct a capability (read, write, or any other type), the FM or SM may need to know the value of the AS field (the Authorization Status field) as stored in the object's attributes. If the FM or SM does not have these attributes cached from a previous operation, it will issue a GetAttr (“Get Attributes”) command to the necessary OBD(s) to retrieve the attributes. The OBDs may, in response, send the attributes to the FM or SM requesting them. The FM or SM may then issue the appropriate capability.
P-0047[0047] Logically speaking, various system “agents” (i.e., the clients <b>24</b>, <b>26</b>, the managers <b>14</b>-<b>22</b> and the OBDs <b>12</b>) are independently-operating network entities. Day-to-day services related to individual files and directories are provided by file managers (FM) <b>14</b>. The file manager <b>14</b> is responsible for all file- and directory-specific states. The file manager <b>14</b> creates, deletes and sets attributes on entities (i.e., files or directories) on clients' behalf. When clients want to access other entities on the network <b>28</b>, the file manager performs the semantic portion of the security work—i.e., authenticating the requester and authorizing the access—and issuing capabilities to the clients. File managers <b>14</b> may be configured singly (i.e., having a single point of failure) or in failover configurations (e.g., machine B tracking machine A's state and if machine A fails, then taking over the administration of machine A's responsibilities until machine A is restored to service).
P-0048[0048] The primary responsibility of a storage manager (SM) <b>16</b> is the aggregation of OBDs for performance and fault tolerance. A system administrator (e.g., a human operator or software) may choose any layout or aggregation scheme for a particular object. The SM <b>16</b> may also serve capabilities allowing clients to perform their own I/O to aggregate objects (which allows a direct flow of data between an OBD and a client). The storage manager <b>16</b> may also determine exactly how each object will be laid out—i.e., on what OBD or OBDs that object will be stored, whether the object will be mirrored, striped, parity-protected, etc. This distinguishes a “virtual object” from a “physical object”. One virtual object (e.g., a file or a directory object) may be spanned over, for example, three physical objects (i.e., OBDs).
P-0049[0049] The storage access module (SAM) is a program code module that may be compiled into the managers as well as the clients. The SAM generates and sequences the OBD-level operations necessary to implement system-level I/O operations, for both simple and aggregate objects. A performance manager <b>22</b> may run on a server that is separate from the servers for other managers (as shown, for example, in FIG. 1) and may be responsible for monitoring the performance of the file system realm and for tuning the locations of objects in the system to improve performance. The program codes for managers typically communicate with one another via RPC (Remote Procedure Call) even if all the managers reside on the same node (as, for example, in the configuration in FIG. 2).
P-0050[0050] A further discussion of various managers shown in FIG. 1 (and FIG. 2) and the interaction among them is provided on pages <b>11</b>-<b>15</b> in the co-pending, commonly-owned U.S. patent application Ser. No. 10/109, 998, filed on Mar. 29, 2002, titled “Data File Migration from a Mirrored RAID to a Non-Mirrored XOR-Based RAID Without Rewriting the Data”, whose disclosure at pages <b>11</b>-<b>15</b> is incorporated by reference herein in its entirety.
P-0051[0051] The installation of the manager and client software to interact with OBDs <b>12</b> and perform object-based data storage in the file system <b>10</b> may be called a “realm.” The realm may vary in size, and the managers and client software may be designed to scale to the desired installation size (large or small). A realm manager <b>18</b> is responsible for all realm-global states. That is, all states that are global to a realm state are tracked by realm managers <b>18</b>. A realm manager <b>18</b> maintains global parameters, notions of what other managers are operating or have failed, and provides support for up/down state transitions for other managers. Realm managers <b>18</b> keep such information as realm-wide file system configuration, and the identity of the file manager <b>14</b> responsible for the root of the realm's file namespace. A state kept by a realm manager may be replicated across all realm managers in the system <b>10</b>, and may be retrieved by querying any one of those realm managers <b>18</b> at any time. Updates to such a state may only proceed when all realm managers that are currently functional agree. The replication of a realm manager's state across all realm managers allows making realm infrastructure services arbitrarily fault tolerant—i.e., any service can be replicated across multiple machines to avoid downtime due to machine crashes.
P-0052[0052] The realm manager <b>18</b> identifies which managers in a network contain the location information for any particular data set. The realm manager assigns a primary manager (from the group of other managers in the system <b>10</b>) which is responsible for identifying all such mapping needs for each data set. The realm manager also assigns one or more backup managers (also from the group of other managers in the system) that also track and retain the location information for each data set. Thus, upon failure of a primary manager, the realm manager <b>18</b> may instruct the client <b>24</b>, <b>26</b> to find the location data for a data set through a backup manager.
P-0053[0053]FIG. 2 illustrates one implementation <b>30</b> where various managers shown individually in FIG. 1 are combined in a single binary file <b>32</b>. FIG. 2 also shows the combined file available on a number of servers <b>32</b>. In the embodiment shown in FIG. 2, various managers shown individually in FIG. 1 are replaced by a single manager software or executable file that can perform all the functions of each individual file manager, storage manager, etc. It is noted that all the discussion given hereinabove and later hereinbelow with reference to the file storage system <b>10</b> in FIG. 1 equally applies to the file storage system embodiment <b>30</b> illustrated in FIG. 2. Therefore, additional reference to the configuration in FIG. 2 is omitted throughout the discussion, unless necessary.
P-0054[0054] Generally, the clients may directly read and write data, and may also directly read metadata. The managers, on the other hand, may directly read and write metadata. Metadata may include file object attributes as well as directory object contents. The managers may create other objects in which they can store additional metadata, but these manager-created objects may not be exposed directly to clients.
P-0055[0055] The fact that clients directly access OBDs, rather than going through a server, makes I/O operations in the object-based file systems <b>10</b>, <b>30</b> different from other file systems. In one embodiment, prior to accessing any data or metadata, a client must obtain (1) the identity of the OBD on which the data resides and the object number within that OBD, and (2) a capability valid on that OBD allowing the access. Clients learn of the location of objects by directly reading and parsing directory objects located on the OBD(s) identified. Clients obtain capabilities by sending explicit requests to file managers <b>14</b>. The client includes with each such request its authentication information as provided by the local authentication system. The file manager <b>14</b> may perform a number of checks (e.g., whether the client is permitted to access the OBD, whether the client has previously misbehaved or “abused” the system, etc.) prior to granting capabilities. If the checks are successful, the FM <b>14</b> may grant requested capabilities to the client, which can then directly access the OBD in question or a portion thereof.
P-0056[0056] Capabilities may have an expiration time, in which case clients are allowed to cache and re-use them as they see fit. Therefore, a client need not request a capability from the file manager for each and every I/O operation. Often, a client may explicitly release a set of capabilities to the file manager (for example, before the capabilities' expiration time) by issuing a Write Done command. There may be certain operations that clients may not be allowed to perform. In those cases, clients simply invoke the command for a restricted operation via an RPC (Remote Procedure Call) to the file manager <b>14</b>, or sometimes to the storage manager <b>16</b>, and the responsible manager then issues the requested command to the OBD in question.
P-0057[0057]FIG. 3 is a simplified diagram of the process when a client <b>24</b>, <b>26</b> first establishes a contact with the data file storage system <b>10</b>, <b>30</b> according to the present invention. At client setup time (i.e., when a client is first connected to the network <b>28</b>), a utility (or discovery) program may be used to configure the client with the address of at least one realm manager <b>18</b> associated with that client. The configuration software or utility program may use default software installation utilities for a given operating system (e.g., the Windows® installers, Linux® RPM files, etc.). A client wishing to access the file storage system <b>10</b>, <b>30</b> for the first time may send a message to the realm manager <b>18</b> (whose address is provided to the client) requesting the location of the root directory of the client's realm. A “Get Name Translation” command may be used by the client to request and obtain this information as shown by step-<b>1</b> in FIG. 3. The contacted RM may send the requested root directory information to the client as given under step-<b>2</b> in FIG. 3. In the example shown in FIG. 3, the root information identifies the triplet {device_ID, object-group_ID, object_ID}, which is {SM #3, object-group #29, object #6003}. The client may then contact the FM identified in the information received from the RM (as part of that RM's response for the request for root directory information) to begin resolving path names. The client may probably also acquire more information (e.g., the addresses of all realm managers, etc.) before it begins accessing files to/from OBDs.
P-0058[0058] After the client establishes the initial contact with the file storage system <b>10</b>, <b>30</b>—i.e., after the client is “recognized” by the system <b>10</b>, <b>30</b>—the client may initiate a data file write operation to one or more OBDs <b>12</b>.
P-0059[0059] In one embodiment, all files and directories in the storage system <b>10</b>, <b>30</b> may exist in an acyclic single-rooted file name space (i.e., a single file name space). In that embodiment, clients may be allowed to mount the entire tree or a single-rooted subtree. Clients may also be allowed to mount multiple subtrees through multiple mount commands. The software for the storage system <b>10</b>, <b>30</b> may be designed to provide support for a single, universal file name space. That is, a file name that works on one client should work on all clients in the system, even on those clients that are logically operating at disparate geographic locations. Under the single, global namespace environment, the file system abstraction that is exported to a client is one seamless tree of directories and files.
P-0060[0060] As discussed hereinbelow, a fully distributed file system-image solution according to the present invention allows managers to interact among themselves to produce a system-wide image without time smear. In the object-based distributed file system <b>10</b>, there is no single centralized point of control. Therefore, a system-wide image is handled by a “parliament” of controllers (i.e., realm managers <b>18</b> as discussed immediately below) because the image-taking affects the entire realm. However, because such system-wide control involves many steps and because the “voting” (discussed below) or communication among realm managers may consume system resources and time, it is desirable to configure realm managers to choose amongst themselves one realm manager or image master that is responsible to coordinate system-wide image activity in the realm. In one embodiment, if such image master fails, then the entire image is declared failed.
P-0061[0061]FIG. 4 illustrates an exemplary flowchart of steps involved in performing a system-wide image according to one embodiment of the present invention in the distributed object-based data storage systems <b>10</b>, <b>30</b> shown in FIGS. 1 and 2 respectively. At any time, the realm managers <b>18</b> in the system <b>10</b> (reference to the system <b>30</b> in FIG. 2 is omitted for the sake of brevity only) elect one of them to act as an image master to manage the image-taking operation (block <b>40</b>, FIG. 4). In one embodiment, the image master is elected at system boot time. The election of image master may be performed using the well-known Distributed Consensus Algorithm (DCA) (also known as Quorum/Consensus Algorithm). A more detailed discussion of DCA is provided in “The Part-Time Parliament” by Leslie Lamport, Digital Equipment Corporation, Sep. 1, 1989, which is incorporated herein by reference in its entirety. The DCA implements a Part Time Parliament (PTP) model in which one node (here, the realm manager acting as the image master) acts as a president, whereas other nodes (i.e., other realm managers) act as voters/witnesses. The nodes communicate with one another using this PTP model. The president issues “decrees” to voters who, in turn, determine whether the president's action is appropriate—this is known as the “voting phase.” The voters then send a message back to the president indicating whether or not they approve the president's action—this is known as the “Success” or “Failure” phase. If each node approves the president's decree, then the president informs all the nodes that the decree has passed.
P-0062[0062] The elected image master remains a master until it fails or is taken out of service or the system shuts down. The remaining RMs detect the absence of the image master via “heartbeats” (a type of internal messaging in the PTP model), which causes them to perform a quorum/consensus-based election of a new image master. Also, if an image master discovers that it is not in communication with a quorum of realm managers, it abdicates its position and informs other RMs of that action. The RMs may again elect another image master using DCA. Thus, at any time, there is at most one image master in the realm, and there is no master if there is no quorum of realm managers. When there is no quorum of realm managers, the entire realm is either in a failed state, or is transitioning to or from a failed state. For example, in a realm with five realm managers, the no-quorum condition arises when at least three of those managers fail simultaneously. It is noted here that the choice of the method of electing an image master (e.g., the PTP model or any other suitable method) is independent of the desired system fault tolerance.
P-0063[0063] The image taking process can be automated, in which case, the system-wide image may be taken automatically at a predetermined time or after a predetermined time period has elapsed. On the other hand, the image may be taken manually when desired. In order to take an image (block <b>42</b>, FIG. 4), a system administrator (software or human operator) running a system management console on a trusted client or manager first sends an authorized image request to the image master (step-<b>1</b> in FIG. 5). The image may be requested by sending an RPC command to the image master RM. After performing authentication and permission checks, and checking that the file system is not in a failure condition that would preclude taking the image, the image master places/writes a dummy image directory in the root object for the file system through appropriate file manager (block <b>44</b>, FIG. 4; step-<b>2</b>, FIG. 5), indicating that the image is under way. The master then sends image requests to all file managers <b>14</b> in the system <b>10</b>, informing them of the image (step-<b>3</b>, FIG. 5). These pre-image events are depicted through steps <b>1</b>-<b>3</b> (which are self-explanatory) in FIG. 5.
P-0064[0064] In response to the image request from the image master, each FM <b>14</b> stops issuing new write capabilities (i.e., queues all incoming requests for capabilities). This begins the phase to quiesce the realm or to achieve a quiescent state in the realm (block <b>46</b>, FIG. 4) to obtain a system-wide image with minimum time smear. It is desirable to quiesce write activity in the file system <b>10</b> to ensure that there is a consistent state to capture (i.e., to obtain an internally-consistent image). In the realm quiescing phase, each FM <b>14</b> scans its list of outstanding write capabilities and sends a message to each client <b>24</b>, <b>26</b> holding such capabilities. The message requests that the client finish any in-flight writes and then flush corresponding data (step-<b>4</b>, FIG. 6). The FM also stops issuing new write capabilities to clients. When a client has finished all in-flight writes, it flushes corresponding data and responds to the FM (step-<b>4</b>, FIG. 6). When the FM has received a response from all clients, or after a fixed delay regardless of client responses, the FM invalidates all outstanding write capabilities by issuing invalidation commands to the storage managers <b>16</b> upon which that FM has outstanding write capabilities (block <b>47</b>, FIG. 4; step-<b>5</b>, FIG. 6). This is to assure that the image is immutable, even when malicious or misbehaving clients re-use their capabilities after being requested to discard them. The FM can, however, make a note of any client that fails to respond in time. The FM may be configured to choose any technique for accomplishing the capability invalidations including, for example, changing the cryptographic key for a capability, or sending a bulk-invalidate command to all SM's <b>16</b> in the system <b>10</b>, or sending a series of single invalidation commands to each individual SM, etc. Upon completion of the invalidation task, each SM responds to the FM (step-<b>5</b>, FIG. 6) that it has completed the requested invalidations. These quiescent phase-related events are depicted through steps <b>4</b>-<b>5</b> (which are self-explanatory) in FIG. 6. It is noted that read capabilities need not be invalidated during image-taking because a read operation is generally not destructive. Further, the capability invalidation may be performed in the realm to prepare the system <b>10</b> for orderly shutdown or to repair a damaged volume on an OBD <b>12</b>.
P-0065[0065] Each FM then responds to the image master that it has completed its image work (step-<b>6</b>, FIG. 7). When the image master receives a response from all the file managers, it informs each file manager what the image time is (block <b>48</b>, FIG. 4). It is noted that when all outstanding write capabilities have been invalidated (i.e., when each FM receives a response from all necessary SM's), the FM makes a permanent record (by updating its internal state) that, from its own perspective, the image has happened. At this point the FM is free to start issuing write capabilities again (block <b>50</b>, FIG. 4). After receiving the response from all the file managers, the image master also converts or updates (through appropriate file manager) the dummy object in the root into a final image directory via which the image can be accessed (step-<b>7</b>, FIG. 7), indicating that the image directory now represents a valid image. In one embodiment, the image master also makes available in this image directory object the time at which the image was taken, the logical name for the image (which may be supplied, for example, by the client requesting the image), and the image version number. Thereafter, the image master informs the file managers <b>14</b> and the storage managers <b>16</b> in the system <b>10</b> that they may resume activity and may also include in this message the timestamp that it recorded for this new image (block <b>52</b>, FIG. <b>4</b>). Because each file manager and each storage manager in the system <b>10</b> receives the identical time for the image, the problem of time smear is avoided. The image master also responds to the image requesting console (step-<b>8</b>, FIG. 7) indicating completion of the system-wide image-taking operation (block <b>52</b>, FIG. 4). FIG. 7 graphically depicts steps <b>68</b> involved in the “reply phase” during an image in the system <b>10</b>. Based on the discussion above, the steps shown in FIG. 7 are self-explanatory.
P-0066[0066] In one embodiment, instead of quiescing the write activity in the file system <b>10</b> at an arbitrary time, the realm-wide images may be scheduled in advance so that the realm managers can inform corresponding file managers and storage managers that they must prepare to quiesce at the prearranged time. In that case, file managers <b>14</b> and storage managers <b>16</b> will track the time of the next upcoming image and will ensure that all capabilities they issue expire at or before the time of the image. Such an approach enables the system to minimize the time required to interfere with a large number of outstanding I/O operations in order to accomplish write capability invalidations for establishing the quiescent state. By synchronizing clocks of all file managers <b>14</b> in the system <b>10</b>, the file managers can be made to quiesce simultaneously or within a couple of seconds of clock skew.
P-0067[0067] Even though it is shown at block <b>50</b> in FIG. 4 that a file manager can start issuing write capabilities to clients at that point in the image process, it is noted that because of the requirement for recursive images (described in more detail hereinbelow with reference to block <b>54</b> in FIG. 4), that file manager may have to wait until all file managers in the system <b>10</b> have reported back (to the image master) that they are quiescent. A file manager attempting to issue write capabilities prior to the point at which all other FMs have quieted their write activity may encounter the dummy object at the file system root, rather than the final image directory (created at step-<b>7</b> in FIG. 7).
P-0068[0068] Once taken, an image becomes visible in the realm-wide file name space via a new directory that appears at the root of the file system. As noted before, the system administrator creating the image may be allowed to specify a name for the image, which appears in the root object as the image directory name. For example, the administrator may choose to name an image by its image version number, or by the time at which it was taken. Thus, the image of a file named “/xyz.com/a/b/c” may appear as “/xyz.com/image<sub>—</sub>1/a/b/c” or as “xyz.com/image_July<sub>—</sub>15<sub>—</sub>2002<sub>—</sub>04:00/a/b/c”.
P-0069[0069] The foregoing describes a process of taking a system-wide image in the distributed object-based data storage system <b>10</b> in FIG. 1 (or system <b>30</b> in FIG. 2, as noted before). Block <b>52</b> in FIG. 4 marks the end of actions taken at the time of an actual image. In order that the image be immutable, all subsequent file system write operations must cause the object being written to be duplicated into the image directory because the image directory at block <b>52</b> (FIG. 4) is just an image of the root object directory, with all directory entries in the image copy pointing to the entries in the “live” file tree. In other words, the “image tree” still does not contain images or copies of all system files as of the time the image was taken. Therefore, appropriate “leaf nodes” in the image tree representing duplicate versions of corresponding objects in the live tree must be “built” prior to allowing any write activity on a particular object. This recursive image approach is discussed below with reference to FIGS. 8 and 9.
P-0070[0070] As noted before, all files and directories in the storage system <b>10</b>, <b>30</b> may exist in an acyclic single-rooted file name space (i.e., a single file name space for the entire realm). FIG. 8 illustrates an exemplary tree structure with files and directories in the live tree and an image tree immediately after an image is taken. A “live tree” contains all files and directories active in the system <b>10</b>, <b>30</b>; whereas, an “image tree” contains just an image or copy of the files and directories in the system at a particular instant in time (e.g., the image time identified at block <b>52</b> in FIG. 4). In other words, the files in the live tree are “active” and their contents may constantly change in the normal course of data storage operations, whereas the contents of files in the image tree never change—i.e., the files in the image tree are “frozen” as of the time of image-taking. Each file object and directory object in the system may be represented as a “leaf node” in the appropriate tree. For example, in FIG. 8, the live tree is composed of leaf nodes <b>64</b>, <b>66</b>, <b>68</b>, <b>70</b>, <b>72</b> and <b>74</b>; whereas the image tree is composed of the leaf node <b>62</b>. The common root is represented by numeral <b>60</b>. As can be seen from the tree representation in FIG. 8, all directory entries in the image tree are pointing to entries in the live tree immediately after the image has been taken. Therefore, all post-image file system write operations must first cause the object being written to be duplicated into the file-system image to “construct” the image tree prior to proceeding with modifying the content of that object through a write operation. In other words, file and directory objects that have not been updated since the previous image will be visible in the image name space, but the links there will refer back to the live copy in the “live tree.”
P-0071[0071] The following describes the recursive image process in one embodiment of the present invention arising from post-image write activity (block <b>54</b>, FIG. 4). A file manager <b>14</b> detects a post-image write operation when it receives a request (e.g., from a client <b>24</b>, <b>26</b>) for a write capability. When an FM <b>14</b> receives a request for a write capability, the FM <b>14</b> checks the image version number attached to the “live” object to be written into against the latest image version number that the FM <b>14</b> holds in its own stable state. If the two image version numbers do not match, then prior to issuing the write capability, the FM <b>14</b> may issue a copy-on-write (FastCopy) command against that live object on the OBD <b>12</b> where the object is stored. As part of the FastCopy command, the FM <b>14</b> may supply the {object-group_ID, object_ID} values for the live object to be copied and also for the image of the copied object. It is noted that the image version number may be attached to a live object either at the time the object was created or at the time when that object was last updated after an image.
P-0072[0072] In the exemplary tree structure of FIG. 8, when a client requests a new write capability for the file file1, the FM <b>14</b> responsible for file1 object <b>74</b> will first perform the image version number comparison. Since an image has just occurred (as indicated by image1 node <b>62</b> in FIG. 8), the current image version number will be different from that attached to the live file1 object <b>74</b>, indicating that the file<b>1</b> object <b>74</b> requires a FastCopy operation before the FM <b>14</b> can issue any write capabilities against it. The FastCopy operation causes a storage manager <b>16</b> managing the OBD <b>12</b> containing file1 to instruct that OBD <b>12</b> to take a logical copy of the live file1 object <b>74</b> and create a new object (file1 object <b>75</b>) with a new object_ID (which may be randomly generated by the storage manager). It is noted that each object stored in the system <b>10</b>, <b>30</b> may have a unique object_ID assigned to it by the storage manager <b>16</b> storing it. For example, the object_ID for the live file1 object <b>74</b> may be “1011”, whereas the object_ID for the image of file1 object (object <b>75</b>) may be “2358.”
P-0073[0073] The FM <b>14</b> then records the new object_ID in a directory entry in the image of the directory that contains the original object (here, file1 object <b>74</b>). This implies that the FM <b>14</b> must first create the image of the parent directory (user1 directory <b>70</b>) for the file<b>1</b> object <b>74</b>, which in turn may imply a recursive traversal up several directory layers until the FM reaches a directory that (a) is owned by another file manager, (b) is the root of the file system (e.g., the root object <b>60</b> in FIG. 8), or (c) has already been modified since the last image was taken. In case of event (a), the FM initiating the recursive traversal must request the other file manager to take over and continue the recursive image of the directories along the path to the root, and wait for this other file manager to complete before continuing further. In cases (b) and (c), the FM initiating the recursive traversal has completed its path-image run and can now issue new write capabilities against the file in question. Thus, the traversal for recursive images of each directory upward in the live tree is depth-first (leaf node by leaf node), meaning that an FM recurses to the parent before taking the image of the child.
P-0074[0074] The following describes an exemplary set of steps that a file manager follows during a recursive image process to effect the change from the tree structure in FIG. 8 to that in FIG. 9, which illustrates an image tree constructed in a distributed object-based data storage system using the leaf node approach:
P-0075[0075] 1. The file manager (“FM<b>1</b>”) receiving a request for write capability on file1 object <b>74</b> first issues a GetAttr (“Get Attribute”) command on the OBD <b>12</b> storing file1 object <b>74</b> to obtain the image version number attached to it.
P-0076[0076] 2. Upon realizing that a FastCopy is needed for file1 object <b>74</b>, FM<b>1</b> recurses to the parent directory (user1 object <b>70</b>). This may cause FM<b>1</b> to fetch the attributes for user1 object <b>70</b> (using, for example, the GetAttr command).
P-0077[0077] 3. FM<b>1</b> then recurses to the parent of user1 <b>70</b>, which is usr <b>64</b>. FM<b>1</b> does a GetAttr on usr <b>64</b> and discovers that usr <b>64</b> is controlled by another file manager (“FM2”). FM<b>1</b> has now reached the point where it cannot recurse further, so FM<b>1</b> invokes an RPC on FM2 requesting that FM2 continue the directory image process farther up the live tree.
P-0078[0078] 4. FM2 wishes to image usr <b>64</b> and therefore fetches its attributes. In doing so, FM2 recurses to the parent, which is the root object “/” <b>60</b>.
P-0079[0079] 5. FM2, having reached the root object <b>60</b>, fetches the image object “/image1” <b>62</b> and confirms that it is a valid image (and not a dummy object created at the initiation of an image). The FM2 now begins the process of taking directory images back down along its recursion tree. The FM2 issues a FastCopy operation against usr <b>64</b>, and replaces the directory entry in image1 <b>62</b> to point to the image of usr <b>64</b> (i.e., usr <b>65</b>). This causes FM2 to reach the bottom of its recursion stack, and so it returns a “Success” indication to FM<b>1</b>.
P-0080[0080] 6. FM<b>1</b>, having received confirmation that the tree above it has been processed for image, is now able to image its own directories. FM<b>1</b> resolves the path name “/image1/usr” to identify the directory object in the image tree that it wishes to update. In that regard, FM<b>1</b> issues a FastCopy operation against user1 <b>70</b> and then issues an RPC to FM2 (distinct from the RPC in step <b>3</b> above), asking FM2 to replace the user1 directory <b>70</b> entry in “/image1/usr” with the new object_ID created by the FastCopy (or image) of user1 <b>70</b> (i.e., user1 <b>71</b>). Continuing down its recursion stack, FM<b>1</b> then images file1 <b>74</b>, and installs a new directory entry (i.e., file1 <b>75</b>) in “/image1/usr/user1.”
P-0081[0081] 7. This completes the image process for file1 <b>74</b> and FM<b>1</b> is now free to issue write capabilities against file1 <b>74</b>.
P-0082[0082] It is noted that if a file manager (FM<b>1</b>, FM<b>2</b> or any other file manager <b>14</b>) now wishes to create a new file (through a write operation) in the directory user2 <b>72</b>, the recursion chain will reach only as high as usr <b>64</b> because usr directory <b>64</b> has already been installed in the image tree (indicated as usr <b>65</b>), and, hence, there is no reason to recurse higher.
P-0083[0083] In one embodiment, at the time of the FastCopy operation, the FM initiating the FastCopy operation updates the image version number attached to the “live” object prior to issuing the write capability against that object. This assures that the next time that file manager receives a request for a write capability against that object, a FastCopy operation may not be needed because the object would have been marked as already having been updated in the most recent image.
P-0084[0084]FIG. 10 illustrates an exemplary flowchart of steps involved in performing a system-wide image according to another embodiment of the present invention. Blocks <b>40</b>-<b>47</b> shown in FIG. 10 are identical to those shown in FIG. 4 (with block <b>44</b> in FIG. 4 absent from the embodiment in FIG. 10) and, hence, a discussion of the steps described by those blocks is not repeated here for the sake of brevity. It is noted that the discussion given earlier with reference to FIGS. 5 and 6 equally applies to corresponding blocks in FIG. 10 (except for the discussion related to step-<b>2</b> in FIG. 5 because of the absence of creation of a dummy image object in the embodiment shown in FIG. 10), and, hence, that discussion is also not repeated here. However, the sequence of events shown in FIG. 7 does not completely apply to the embodiment described through FIG. 10. Therefore, a modified version of FIG. 7 is shown in FIG. 11, which depicts the reply phase-related events in the image process in FIG. 10. As shown at step-<b>1</b> in FIG. 11, when each FM has received a response from all necessary SM's that write capabilities are invalidated and the realm is quiesced (blocks <b>46</b>, <b>47</b> in FIGS. 4 and 10), that FM responds to the image master (IM) that the system is quiet and ready for the image.
P-0085[0085] When the IM receives a response from all file managers, it issues a Clone Object Group command to each SM in the system <b>10</b>, <b>30</b>, for each live object group that each SM contains or manages (block <b>80</b>, FIG. 10; step-<b>2</b>, FIG. 11). In one embodiment, there may be only one object group for each SM. In another embodiment, there may be more than one object group for one or more of the storage managers. The Clone Object Group command is similar to the FastCopy command, except that the Clone Object Group command uses only the {object-group_ID} values for the live and cloned object groups (instead of the pair of {object-group_ID, object_ID} values taken by the FastCopy command) and returns the results as a status indication (“success” or “fail”). It is noted that in the absence of an image, the object group(s) is called a “live” object group. An image of that live object group generates an image object group. In that case, there are two object groups for the corresponding SM—one live object group and one image object group. That storage manager may span (or store) a virtual object group over more than one physical storage disk or OBDs <b>12</b>.
P-0086[0086] In the system <b>10</b> in FIG. 1 (or the system <b>30</b> in FIG. 2), the storage configuration may be organized in three levels as illustrated in FIG. 12. Each higher level may be considered as a “container” for the next lower one. Thus, a device (e.g., a storage manager <b>16</b> or an OBD <b>12</b>) may contain one or more object groups, with each object group containing one or more objects, as illustrated in FIG. 12. However, it is noted that some devices may not contain any object group. Further, each entity in the storage hierarchy in FIG. 12 may have an ID assigned to it to easily identify that entity during data storage and image operations as discussed later. It is noted that, in one embodiment, each {device_ID, object-group_ID, object_ID} triplet must be unique in the realm. In other words, even if two objects have the same object_ID, they cannot have the same values for the corresponding {device_ID, object-group_ID, object_ID} triplet.
P-0087[0087] It is noted that the cloning operation is performed on a live object group and generates an image object group as a result. As illustrated in FIG. 11, upon receiving the Clone Object Group command from the IM, each storage manager duplicates object group header information (that includes the cloned object-group_ID, image_ID, and cloned object group attributes as discussed hereinbelow) for all of its local (live) object groups and marks the indicated object group(s) as copy-on-write. Then the storage manager issues corresponding Clone Object Group commands to each OBD that the storage manager uses (illustrated at step-<b>3</b> in FIG. 11). In response to the Clone Object Group command, each OBD takes a logical copy of an object group and creates a new (image) object group with a new object-group_ID (which may be randomly generated). Thereafter, each OBD reports the object-group_ID of each newly-created object group back to the corresponding storage manager, and that storage manager, in turn, records the new object-group_ID and identity of the respective OBD containing the new object group in its internal data structures. Each SM then reports the object-group_IDs of the newly-created SM-level object groups back to the image master, which, in turn, records the new object-group_IDs in an image-group_ID table maintained by the image master.
P-0088[0088]FIG. 13 shows a simplified illustration of how the live and cloned object groups are contained in a storage manager. As depicted in FIG. 13, two live object groups, A and B, are initially managed by the storage manager <b>16</b> before any image activity in the system. Thereafter, the storage manager <b>16</b> may receive a first Clone Object Group command and, in response, generate the first cloned object group A′. Next, the storage manager <b>16</b> receives a second Clone Object Group command to generate the second cloned object group B′. Thus, each object group is cloned one at a time, without the storage manager <b>16</b> knowing that a global image operation is taking place in the system. At the end of the first image operation, therefore, four object groups (A, A′, B, and B′) exist on the storage manager <b>16</b> in FIG. 13. During a second image operation in the system, the storage manager <b>16</b> does not clone the existing four object groups (i.e., object groups A, A′, B, and B′), but only clones the two live objects groups contained therein. That is, during a second, later occurring image, the storage manager <b>16</b> clones only object groups A and B to generate the corresponding cloned object groups A″ and B″ respectively as shown in FIG. 13.
P-0089[0089] Thus, at the storage manager level, according to the present invention, each object group in the system <b>10</b> is cloned—none after another—without primarily or explicitly focusing on cloning, one by one, each object in the system <b>10</b>. An individual object is therefore not the focus of the cloning operation by the storage manager; it is not addressed (by the storage manager) during the cloning. Individual objects are copied later as needed (copy-on-write), for example, when the first write operation is requested in the system after the image-taking operation is over. On the other hand, the leaf node approach discussed with reference to block <b>54</b> in FIG. 4 is an example of object-by-object copying scheme. It is noted that an object group, as discussed herein, is a flat name space; an object group has objects as members. There is no correlation between the file directory hierarchy (e.g., as shown in the file structure trees shown in FIGS. 8 and 9) in the file name space and the object groups. An image of an object group may contain an image of a random distribution of unrelated member objects. Whereas, an image of a file directory contains an image of those objects that are related to one another (and are nodes on the file tree).
P-0090[0090] Each object group in the system <b>10</b> may have properties or attributes associated therewith. Each attribute is a {parameter, value} pair, i.e., each attribute constitutes a parameter and its value. For example, one set of parameters includes image_time, image_name (a logical name given by the image requestor or user), and clone_source. These parameters allow a manager (e.g., a realm manager <b>18</b>) to identify which image an object is part of. With reference to FIG. 13, for example, the original (live) object group may have the following attributes: image_time=0 (i.e., t=0), image_name=blank (i.e., no name because of the live object status), and clone_source=0 (i.e., no clone source because of the live object status). On the other hand, the attributes for the cloned object group A′ may be: image_time=2:00 p.m., image_name=daily.2 pm.Jun. 12, 2002, and clone_source=A (because A′ is cloned from the live object group A). Similarly, the attributes for the cloned object group A″ include the same value for the clone_source parameter, but different values for the image time and image name parameters. In the event of a failure in the system <b>10</b>, each storage manager <b>16</b> may be asked (e.g., by a new image master elected after the system is restarted) to produce a list of its object groups and their attributes, and the system may be rebuilt from the time of failure by examining the attribute values.
P-0091[0091] The attributes of each object group stay with the corresponding storage manager <b>16</b>. However, each realm manager <b>18</b> may keep a copy of each object group attributes for indexing purpose. This is desirable to improve system performance when a user or a client <b>24</b>, <b>26</b> asks a realm manager <b>18</b> for a permission to access specific files in an image. Without such information pre-available to the realm manager <b>18</b>, the realm manager <b>18</b> may have to contact each relevant storage manager <b>16</b> to obtain a list of all relevant object groups and their attributes, thereby increasing network traffic and system overhead. This can be avoided when the realm manager is provided with an index of object groups and their attributes.
P-0092[0092] Each storage manager <b>16</b> may continue accumulating object group images (clones) until it runs out of the actual, physical storage space (on a corresponding OBD <b>12</b>) that is holding the system-wide file system image. During image-taking, because each object group is indicated as copy-on-write, only a small number of bytes may be needed to store the image-related information that includes values for the cloned object-group_ID, image_ID and cloned object group's attributes for each object group cloned by the respective storage manager. In one embodiment, each OBD <b>12</b> supports a 32-bit ID for an object group, which allows a creation and storage of a large number (theoretically 2<sup>32</sup>−2 (to exclude the ID for the object group to be cloned and the ID with all zeros)) of cloned object groups. All the object groups contained in a storage manager may have different, randomly-generated, object-group_ID. However, it may be possible that two object groups on different storage managers have the same object-group_ID.
P-0093[0093] The cloning of each object group results in creation of copies of object group header information without copying the objects contained in an object group. Because an object group in an image is made a copy-on-write entity, the constraints placed on the system architecture by the system-wide image are minimized because of the avoidance of copying the entire object group (with all its constituent objects) during image-taking.
P-0094[0094] Referring now to FIG. 10, upon successful completion of cloning of each object group (block <b>80</b>, FIG. 10), the image master informs the file managers <b>14</b> that they are free to start issuing write capabilities again (block <b>82</b>, FIG. 10). In one embodiment, the image master also stores an image-descriptor object or directory in the root (indicating a valid image) and makes available in this image-descriptor object the time at which the image was taken, the logical name for the image (which may be supplied, for example, by the client requesting the image), and the image version number (or image_ID). In any event, the image master informs the file managers <b>14</b> and the storage managers <b>16</b> in the system <b>10</b> that they may resume activity and also includes in this message the timestamp that it recorded for this new image (block <b>84</b>, FIG. 10). As noted before, because each file manager and each storage manager in the system <b>10</b> receives the identical time for the image, the problem of time smear is avoided. The image master also responds to the image requesting console (step-<b>4</b>, FIG. 11) indicating completion of the system-wide image-taking operation (block <b>84</b>, FIG. 10).
P-0095[0095] As part of re-enabling the write operations (block <b>82</b>, FIG. 10), the image master may generate a new value for each capability “key” and send that new value to respective file manager <b>14</b>. The file managers <b>14</b> may then start issuing new capabilities with these new keys. On the other hand, in an alternate embodiment, the image master may instead instruct the file managers <b>14</b> to start issuing new capabilities with the original “key” values, i.e., the key values prior to image-taking. In another embodiment, the image master may issue each key with embedded information about the state of the realm and for what operations the file manager <b>14</b> can use the key. It is noted that the realm managers <b>18</b> may employ the distributed consensus algorithm (DCA) to authorize the image master to change the key values because the activity of changing keys requires participation of all the realm managers who need to know the new key values in the event that there is a system failure and the realm managers need to communicate with one another to rebuild the system. Without the most recent key values available to them, the realm managers may not be able to communicate with their clients.
P-0096[0096] However, it may not be desirable to change the key values to something different from the original because of a possibility of creation of a system bottleneck when many clients having new capabilities different from the original ones start accessing their respective file managers to inform that their capabilities are no longer valid. But, if the original keys are re-enabled, then the original capabilities start working again and in the case of a client's request for an invalid capability, the corresponding file manager just has to respond by saying that “retry with your old capability.” Another advantage of keeping the original key values is that the image-taking activity becomes transparent to the clients. In other words, if a client's write operation or data access is briefly interrupted by the system-wide image-taking, the client does not need to be aware of image-taking. It may simply retry the interrupted operation with the original capability, which is already active after system-wide image-taking (which lasts for a very short time).
P-0097[0097] It is noted that the operation of cloning the object groups (block <b>80</b>, FIG. 10) does not involve the use of the distributed consensus algorithm (DCA). That operation is performed by the image master alone. However, when the image master fails before the image is completely over (block <b>84</b>, FIG. 10), then the realm managers may again use the DCA upon system re-start to elect another image master who deletes the partially created image by instructing all storage managers to remove the partial image information stored thereon. The new image master sends the image_ID of the failed image to each storage manager and instructs the storage manager to delete all object groups contained in the image having the failed image's image_ID.
P-0098[0098] After a system-wide file system image is taken, any future write requests to an object are entertained by first performing an object-copying operation (for the object to be written), if it has not already been done, at the target OBD (copy-on-write) prior to issuing any write capability on the object to be written. In other words, such copying operation is implicit in the copy-on-write requirement.
P-0099[0099] According to the present invention, when a client <b>24</b>, <b>26</b> wishes to access an object contained in a system-wide image (as opposed to a live object), the client <b>24</b>, <b>26</b> may perform a directory mapping (discussed below) to access the image version of the live object. Prior to discussing such mapping, it is apt to describe how a client may access a live object in an environment that is object group-based as opposed to a system that is file directory-based (e.g., the embodiment discussed hereinabove with reference to FIGS. 4 and 8-<b>9</b>). As discussed before, in the file directory-based system, a client may traverse the live tree (or the image tree) to access a specific node or file. In the present object group-based embodiment, the file path naming may still remain the same, but the way of accessing an object differs.
P-0100[0100]FIG. 14 illustrates an exemplary set of steps carried out in an object group-based mapping scheme to access a specific live file object. As shown in FIG. 14, the file path for the object “c.doc” may be given, for example, as “/a/b/c.doc”, where the first slash (/) points at the system root directory. In other words, in the file name space, there is a root directory and a number of other file directories emanating from the root in the live (or image) tree. In the system <b>10</b> in FIG. 1 (or the system <b>30</b> in FIG. 2), each directory object (including the root) on an OBD may contain a section for header information and another section for entries. The entries for a directory may include a field for names and a field for corresponding identifiers. An identifier may include the values for the entire {device_ID, object-group_ID, object_ID} triplet as illustrated and discussed with reference to FIG. 15 below. However, for the ease of discussion, only the object_ID values are shown for the identifier fields in FIG. 14. For example, as shown in FIG. 14, the root directory object may contain an entry having the name “a” and the identifier as number “37.” The number “37” may be the object_ID pointing to the next component in the path name (here, the component with name “b”). The client reaches the root directory by querying the appropriate realm manager <b>18</b> who, in turn, indicates that the root object has object_ID=3 as shown in FIG. 14. The initial name resolution to obtain root object identity is discussed hereinbefore with reference to FIG. 3.
P-0101[0101] The client <b>24</b>, <b>26</b> may identify the object “c.doc” as an object with the path name “/a/b/c.doc.” Therefore, after locating the root object, the client <b>24</b>, <b>26</b> may locate an entry with name “a” and read the associated identifier. Here, that identifier is “37”, which is pointing to the next component in the path name. Upon reading the object having object_ID=37, the client may find a directory entry having the name “b” (as that is the name the client wishes to reach during its file path name resolution) and associated identifier=53. Upon further accessing the object with object_ID=53, the client may locate an entry having the name “c.doc” and associated identifier <b>12</b>. The client finally accesses the object “c.doc” when it accesses the object having object_ID=12.
P-0102[0102] The above process of locating an object applies to a live object whose file path traverses a live tree. In the event that a client <b>24</b>, <b>26</b> wishes to access an object whose file path lies on an image tree, a modified version of the above name resolution process is performed according to the present invention. For example, a client application may name an object in an image using some specially defined path name (e.g., /image/a/b/c.doc), but the actual naming convention is not relevant to the mapping scheme described below with reference to FIG. 15. What is necessary is that the client be able to identify the triplet {device_ID, object-group_ID, object_ID} for each component in the file path.
P-0103[0103]FIG. 15 illustrates an exemplary mapping (block <b>86</b>, FIG. 10) when a client application accesses an object in an image. The mapping in FIG. 15 is illustrated in a table form for ease of discussion only. It is noted that the actual data corresponding to the entries in the table in FIG. 15 may not be stored in such a table form on respective OBDs. In the exemplary mapping shown in FIG. 15, the client <b>24</b>, <b>26</b> is assumed to attempt to access the object “/image/a/b/c.doc” in the image domain. The client <b>24</b>, <b>26</b> first obtains the object_ID for the root object as discussed above and also with reference to FIG. 3. The client also obtains the image_ID (for example, image_ID=S) for the image of interest from the image descriptor object or image pseudo-directory created in the root directory object in the system. Upon initially accessing the root object (with object_ID=3 in the present example), the client <b>24</b>, <b>26</b> encounters the triplet {α, A, 3} corresponding to the parameters {device_ID, object-group_ID, object_ID}. That is, the root object has device_ID=α, object-group_ID=A, and object_ID=3. Using the value for the image_ID (=S), the client <b>24</b>, <b>26</b> queries its realm manager <b>18</b> to map the storage manager with device_ID=α and the object group with object-group_ID=A onto the image domain with image_ID=S. That is, the client <b>24</b>, <b>26</b> queries the RM <b>18</b> to obtain the mapping of the image_ID to object-group_ID for a given storage manager. The pseudo command for such mapping request and the mapping received from the realm manager <b>18</b> may look like: map [(α, A)→S]→(α, A′), where A′ is the object-group_ID for the cloned version of the object group with object-group_ID=A.
P-0104[0104] The client will then continue doing its normal name resolution (discussed above with reference to FIG. 14) to locate the next component in the file path, except that the client will replace the object-group_ID of each live object with the corresponding object-group_ID in the image (S) obtained by querying the RM <b>18</b>. For example, in FIG. 15, when the client reads the image object in the root (with triplet {α, A′, 3}), the client still sees the name “a” at object_ID=3 because the image object is a perfect copy of the original, live object. That is, the image object points to the corresponding live object. Upon reading the triplet values in the live domain for the directory name “a”, the client receives the values {α, B, 37}. Again, as the client is interested in accessing the object “c.doc” in the image domain, the client further queries the RM <b>18</b> to map the storage manager with device_ID=α and the object group with object-group_ID=B onto the image domain with image_ID=S. The pseudo command for such mapping request and the mapping received from the realm manager <b>18</b> may look like: map [(α, B)→S]→(α, B′), where B′ is the object-group_ID for the cloned version of the object group with object-group_ID=B. Upon reading the directory entries at object_ID=37, the client locates the name “b” with identifier {(α, C, 53} in the live domain. The client continues the mapping process (shown in FIG. 15) because it knows that it has to locate the “c.doc” object in the image domain. Finally, as the last row in the table in FIG. 15 illustrates, the client queries the RM <b>18</b> to map the live object-group_ID=B and device_ID=α to obtain the values (α, B′) in the image domain. The client then accesses the image version of the object “c.doc” by replacing the object-group_ID=B with object-group_ID=B′ to obtain the triplet {α, B′, 12} and then to locate the object with object_ID=12 in the object group with object-group_ID=B′ on the storage manager α. In this manner, the client <b>24</b>, <b>26</b> can traverse the path in the image space rather than in the normal, live space. As noted earlier and as can be seen in FIG. 15, two different objects (here, objects “a” and “c.doc”) may still have the same object-group_ID B).
P-0105[0105] The above describes a mapping scheme (block <b>86</b>, FIG. 10) where a client <b>24</b>, <b>26</b> dynamically (i.e., at run time) substitutes image (or cloned) object-group_ID queried from the RM for the object-group_ID found in the (live) directory entries. Because of such dynamic mapping, the re-writing of metadata for each object in an image does not need to be performed when the image is taken. Also, because of the client's ability to establish the correspondence between the live objects and the image objects, the live directory objects are not updated either during or after the image-taking to reflect such correspondence therethrough. In other words, the original (live) directory objects maintain their original object-group_ID's and object_ID's during and after the image. Similarly, no original directory object is updated either during or after the image in the embodiment illustrated in FIG. 4 because, in that embodiment also, no original (or live) object is written (e.g., in the leaf node approach at block <b>54</b> in FIG. 4) to maintain the image. Further, in the image mechanisms according to the present invention, there is no static name translation between various object groups. Instead, new ID's for object groups are created randomly. As the original (live) objects are not updated during or after the image, it becomes easier to delete the image versions of objects contained in a partially complete image when an image master fails.
P-0106[0106] On the other hand, in the prior art systems, one of two approaches are taken during image. In the first approach, the new directory object (created during image) becomes the live object itself and updates the corresponding previous live directory to point to this new object. In the second approach, the new object becomes an image object, in which case it updates the image directory (or directories if there is more than one) to point to this new image object. Thus, in both of these approaches, one or more directories are updated to store the ID of the object newly created during the image.
P-0107[0107] The foregoing describes various schemes to accomplish a system-wide file system image in a distributed object-based data storage system. In one embodiment, a realm manager in the system is elected using the Distributed Consensus Algorithm to function as an image master to coordinate the image taking process. When the image master receives a request for the image, it prepares the system for image-taking. In one embodiment, the image master quiesces the realm by informing the file managers that an image is under way, and the file managers, in turn, invalidate all write capabilities pending in the system. A client requires a write capability to perform a data write operation on an object-based secure disk (OBD) in the system. The system is quiesced when all pending write capabilities are invalidated. In one embodiment, as part of taking the system-wide image, the image master instructs each storage manager in the system to clone corresponding object groups contained therein, without cloning any objects contained in the object groups. To preserve the image immutability, all objects stored in the system are marked copy-on-write during image-taking. Neither the cloning operation nor the copying operation update any directory object in the system. The correspondence between live directory objects and image objects is also established without updating the live directory objects to provide such correspondence. In a still further embodiment, in order to traverse the image domain, the client, at run time, replaces the non-image identity of each object group in the file path for an object with that object group's image identity received from a realm manager when doing the path name resolution in the image domain. This allows the client to access the image version of the object without the need to rewrite metadata for the object when the image is taken.
P-0108[0108] It is noted that various managers (e.g., file managers <b>14</b>, storage managers <b>16</b>, etc.) shown and described with reference to FIG. 1 (and FIG. 2), the program code implemented to operate the systems in FIGS. 1 and 2, and the program code for performing the system-wide file system image according to an embodiment of the present invention may reside on a computer-readable, tangible storage medium (e.g., a compact disc, an optical disc, a magnetic storage medium such as a computer hard drive, etc.) allowing ease of software portability and system management. The program code on the storage medium can be executed by a computer system and upon execution, the program code performs various operations described hereinabove with reference to individual components constituting the program code. Similarly, a client application (e.g., any of the client applications <b>24</b>, <b>26</b> in FIG. 1) can also be stored on a computer-readable data storage medium and executed therefrom.
P-0109[0109] The image methodology according to the present invention allows taking system-wide file system images without time smear and without the need to pre-schedule the images (because of the implementation of capability invalidation). Further, there is no hard limit on the number of images that can be kept concurrently on line. The images are performed without a significant overhead on system I/O operations. Because no directory objects are updated either during or after the image, the failure-handling procedures in the system are also simplified.
P-0110[0110] While the invention has been described in detail and with reference to specific embodiments thereof, it will be apparent to one skilled in the art that various changes and modifications can be made therein without departing from the spirit and scope thereof. Thus, it is intended that the present invention cover the modifications and variations of this invention provided they come within the scope of the appended claims and their equivalents.
Contents5
8 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US8423591B2 | Cited by | United States of America | Search report |
| US8359366B2 | Cited by | United States of America | Search report |
| US9507579B2 | Cited by | United States of America | Applicant |
| US10095559B2 | Cited by | United States of America | Applicant |
| US2011072069A1 | Cited by | United States of America | Pre-grant |
| US8671166B2 | Cited by | United States of America | Applicant |
| US7721157B2 | Cited by | United States of America | Applicant |
| US9971787B2 | Cited by | United States of America | Search report |
| US2007234302A1 | Cited by | United States of America | Pre-grant |
| US9804901B2 | Cited by | United States of America | Applicant |
| US7809778B2 | Cited by | United States of America | Applicant |
| US2007083473A1 | Cited by | United States of America | Pre-grant |
| US10757104B1 | Cited by | United States of America | Applicant |
| US7958200B2 | Cited by | United States of America | Search report |
| US2009049153A1 | Cited by | United States of America | Pre-grant |
| US2008065637A1 | Cited by | United States of America | Pre-grant |
| US8042163B1 | Cited by | United States of America | Search report |
| US2007282868A1 | Cited by | United States of America | Pre-grant |
| US2007214146A1 | Cited by | United States of America | Pre-grant |
| US7036039B2 | Cited by | United States of America | Search report |
| US2010057787A1 | Cited by | United States of America | Pre-grant |
| US2008001791A1 | Cited by | United States of America | Pre-grant |
| US2004078633A1 | Cited by | United States of America | Pre-grant |
| US10055415B2 | Cited by | United States of America | Applicant |
| US2008033902A1 | Cited by | United States of America | Pre-grant |
| US2007214255A1 | Cited by | United States of America | Pre-grant |
| US10152530B1 | Cited by | United States of America | Applicant |
| US8266182B2 | Cited by | United States of America | Applicant |
| US12079162B2 | Cited by | United States of America | Applicant |
| US2007214183A1 | Cited by | United States of America | Pre-grant |
| US8478873B2 | Cited by | United States of America | Applicant |
| US7996421B2 | Cited by | United States of America | Applicant |
| US2009043890A1 | Cited by | United States of America | Pre-grant |
| US8306918B2 | Cited by | United States of America | Search report |
| US10613914B2 | Cited by | United States of America | Search report |
| US2012117126A1 | Cited by | United States of America | Pre-grant |
| US8954789B2 | Cited by | United States of America | Applicant |
| US12399869B2 | Cited by | United States of America | Applicant |
| WO2008079955A3 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| US2012226773A1 | Cited by | United States of America | Pre-grant |
| US2008162582A1 | Cited by | United States of America | Pre-grant |
| US2011106764A1 | Cited by | United States of America | Pre-grant |
| US8473777B1 | Cited by | United States of America | Search report |
| US12373397B2 | Cited by | United States of America | Applicant |
| US2007276916A1 | Cited by | United States of America | Pre-grant |
| US9208181B2 | Cited by | United States of America | Search report |
| US10769116B2 | Cited by | United States of America | Applicant |
| US2007214285A1 | Cited by | United States of America | Pre-grant |
| CN106204794A | Cited by | China | Search report |
| US2006235847A1 | Cited by | United States of America | Pre-grant |
| US11704035B2 | Cited by | United States of America | Applicant |
| US7826161B2 | Cited by | United States of America | Applicant |
| US9547485B2 | Cited by | United States of America | Applicant |
| US10296879B2 | Cited by | United States of America | Applicant |
| US2008065663A1 | Cited by | United States of America | Pre-grant |
| US2007214105A1 | Cited by | United States of America | Pre-grant |
| US11194635B2 | Cited by | United States of America | Applicant |
| US7873799B2 | Cited by | United States of America | Search report |
| US2014025711A1 | Cited by | United States of America | Pre-grant |
| US2007234337A1 | Cited by | United States of America | Pre-grant |
| US11727376B2 | Cited by | United States of America | Applicant |
| US10152485B2 | Cited by | United States of America | Search report |
| US2007106851A1 | Cited by | United States of America | Pre-grant |
| WO2008079955A2 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| US9887978B2 | Cited by | United States of America | Applicant |
| US9971788B2 | Cited by | United States of America | Applicant |
| US2008016124A1 | Cited by | United States of America | Pre-grant |
| US7657579B2 | Cited by | United States of America | Search report |
| US7877357B1 | Cited by | United States of America | Search report |
| US2007226224A1 | Cited by | United States of America | Pre-grant |
| US8949312B2 | Cited by | United States of America | Search report |
| US2014304398A1 | Cited by | United States of America | Pre-grant |
| US12235799B2 | Cited by | United States of America | Applicant |
| US2010312805A1 | Cited by | United States of America | Pre-grant |
| US10515058B2 | Cited by | United States of America | Applicant |
| US2008168568A1 | Cited by | United States of America | Pre-grant |
| US2007266198A1 | Cited by | United States of America | Pre-grant |
| US8135674B2 | Cited by | United States of America | Search report |
| US9645811B2 | Cited by | United States of America | Applicant |
| US2008002272A1 | Cited by | United States of America | Pre-grant |
| US2009307329A1 | Cited by | United States of America | Pre-grant |
| US2015046502A1 | Cited by | United States of America | Pre-grant |
| US7941455B2 | Cited by | United States of America | Applicant |
| US7975312B2 | Cited by | United States of America | Search report |
| US2002107945A1 | Cites | United States of America | Pre-grant |
| US2002161889A1 | Cites | United States of America | Pre-grant |
| US2002174168A1 | Cites | United States of America | Pre-grant |
| US5303379A | Cites | United States of America | Pre-grant |
| US5634124A | Cites | United States of America | Pre-grant |
| US5680619A | Cites | United States of America | Pre-grant |
| US5764877A | Cites | United States of America | Pre-grant |
| US5845082A | Cites | United States of America | Pre-grant |
| US5946686A | Cites | United States of America | Pre-grant |
| US5956713A | Cites | United States of America | Pre-grant |
| US5956734A | Cites | United States of America | Pre-grant |
| US5960446A | Cites | United States of America | Pre-grant |
| US5987477A | Cites | United States of America | Pre-grant |
| US5991763A | Cites | United States of America | Pre-grant |
| US5999179A | Cites | United States of America | Pre-grant |
| US6014651A | Cites | United States of America | Pre-grant |
9 members in 1 office; this record represents the family
Priority claims8
| Document | Office | Kind | Date |
|---|---|---|---|
| 36878502 | United States of America | P | |
| 37202702 | United States of America | P | |
| 27394502 | United States of America | A | |
| 60368785 | – | – | – |
| 60372027 | – | – | – |
| US20020273945 | – | – | – |
| US20020368785P | – | – | – |
| US20020372027P | – | – | – |
Members9
| Document | Office | Kind | |
|---|---|---|---|
| US2003187859A1 | United States of America | A1 | |
| US2003187860A1 | United States of America | A1 | |
| US2003187883A1 | United States of America | A1 | |
| US2004003055A1 | United States of America | A1 | |
| US7007047B2 | United States of America | B2 | |
| US2006123057A1 | United States of America | A1 | |
| US7155464B2 | United States of America | B2 | |
| US7194467B2 | United States of America | B2 | |
| US8095503B2 | United States of America | B2 |
14 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Fee payment procedure11.5 YR SURCHARGE- LATE PMT W/IN 6 MO, SMALL ENTITY (ORIGINAL EVENT CODE: M2556)FEPP | FEPP | |
| Maintenance fee paymentMAFP | MAFP | |
| Fee payment procedureMAINTENANCE FEE REMINDER MAILED (ORIGINAL EVENT CODE: REM.)FEPP | FEPP | |
| AssignmentAS | AS | |
| AssignmentAS | AS | |
| AssignmentAS | AS | |
| AssignmentAS | AS | |
| AssignmentAS | AS | |
| Fee paymentFPAY | FPAY | |
| AssignmentAS | AS | |
| Fee paymentFPAY | FPAY | |
| AssignmentAS | AS | |
| Information on status: patent grantGrantedPATENTED CASESTCF | STCF | |
| AssignmentAS | AS |
Numbers
- Publication, DOCDB
- 2003187883
- Publication, EPODOC
- US2003187883
- Application
- 10273945
- Application, DOCDB
- 27394502
- Application, EPODOC
- US20020273945
Titles
- English
- Internally consistent file system image in distributed object-based data storage
Classification
- CPC, 5
- G06F17/30088
- G06F16/128
- Y10S707/99952
- Y10S707/99953
- Y10S707/99955
- IPC, 1
- G06F17 30
- USPC, 4
- 001001000
- 707999201
- 707E17007
- 707E17010