File data protection apparatus and method
Abstract
Embodiments of this invention provide primary magnetic disk data storage capacity to clients while at the same time making sure that client data is replicated locally and at an offsite location to protect from all forms of data loss.

Term
Term ended
Expired 10 September 2023, 3 years ago.
- Priority
- Filed
- Granted
- Expired
- Today
16 claims: 16 independent, 0 dependent
- 1A method for transferring a protected set of files within a system comprising a fileserver (27) having:a file system (32) operative to store client files;a fileserver API (37) operative to communicate with a repository (28);a fileserver file transfer module (40) in communication with the file system (32) and operative to receive files for the file system from at least onerepository (28);anda recovery service (45) in communication with the fileserver API (37) and with the file system (32) and operative to transfer a set of files, the method comprising: receiving (63), at a destination fileserver (27), metadata and stub files associated with the set of files;maintaining a list of repository nodes that are associated with each file in the set of files by updating (64) a location component (36) in the fileserver (27);andreplacing each stub file with the full content of the file associated with the stub file;wherein replacing each stub file includes: receiving a client request for a specified file in the set of files, andreplacing the stub file for the specified file with the full content of the specified file, replacement of said stub file for said specified file being a higher priority task than replacement of the stub files for non-requested files. Verfahren zum Übertragen eines gesicherten Satzes von Dateien innerhalb eines Systems, umfassend: einen Dateiserver (27) mit: einem zum Speichern von Klienten-Dateien betreibbaren Dateisystem (32);einem zur Kommunikation mit einem Magazin (28) betreibbaren Dateiserver-API (37);einem Dateiserver-Übertragungsmodul (40), welches mit dem Dateisystem (32) in Verbindung steht und betreibbar ist, um Dateien für das Dateisystem von mindestens einem Magazin (28) entgegenzunehmen;und miteinem Wiederheratellungsdienst (45), welcher mit der Dateiserver-API (37) und mit dem Dateisystem (32) in Verbindung steht und zum Übertragen eines Satzes von Dateien betreibbar ist;wobei das Verfahren umfasst: Empfangen (63) von Metadaten und Stummel-Dateien, die mit dem Satz von Datelen verknüpft sind, auf einem Ziel-Dateiserver (27);Unterhalten einer Liste von mit jeder Datei aus dem Satz von Datelen verknüpften Magazin-Knoten durch Aktualisieren (64) einer Örtlichkeitskomponente (36) in dem Dateiserver (27);undErsetzen jeder Stummel-Datei durch den vollen Inhalt der mit der Stummel-Datei verknüpften Datei;wobei das Ersetzen jeder Stummel-Datei beinhaltet: Empfangen einer Klienten-Anforderung für eine spezifizierte Datei aus dem Satz von Dateien, undErsetzen der Stummel-Datei für die spezifizierte Datei durch den vollen Inhalt der spezifizierten Datei, wobei der Ersatz der Stummel-Datei durch die spezifizierte Datei eine Aufgabe mit höherer Priorität ist als das Ersetzen der Stummel-Dateien für nicht angeforderte Dateien. procédé de transfert d'un ensemble de fichiers protégé dans un système comprenant un serveur de fichiers (27) comportant: un système de fichiers (32) apte à agir pour stocker des fichiers clients;une interface API de serveur de fichiers (37) apte à agir pour communiquer avec un entrepôt (28);un module de transfert de fichiers de serveur de fichiers (40) en communication avec le système de fichiers (32) et apte à agir pour recevoir des fichiers pour le système de fichiers à partir d'au moins un entrepôt (28);etun service de récupération (45) en communication avec l'interface API de serveur de fichiers (37) et avec le système de fichiers (32) et apte à agir pour transférer un ensemble de fichiers,lequel procédé comprend: la réception (63), au niveau d'un serveur de fichiers de destination (27), de métadonnées et de fichiers souches associés à l'ensemble de fichiers;la gestion d'une liste de noeuds d'entrepôt qui sont associés à chaque fichier de l'ensemble de fichiers par l'actualisation (64) d'un organe d'emplacements (36) dans le serveur de fichiers (27);etle remplacement de chaque fichier souche par le contenu complet du fichier associé au fichier souche;dans lequel le remplacement de chaque fichier souche comprend: la réception d'une demande de client concernant un fichier spécifié de l'ensemble de fichiers, etle remplacement du fichier souche correspondant au fichier spécifié par le contenu complet du fichier spécifié, le remplacement dudit fichier souche correspondant audit fichier spécifié constituant une tâche de priorité plus grande que le remplacement des fichiers souches correspondant à des fichiers non demandés.
- 2Procédé selon la revendication 1, dans lequel les métadonnées sont reçues au niveau d'un serveur de fichiers de destination (27) à partir d'un noeud d'entrepôt (28). The method of claim 1, wherein the metadata is received at a destination fileserver (27) from a repository node (28). Verfahren nach Anspruch 1, wobei die Metadaten aus einem Magazin-Knoten (28) auf einem Ziel-Deteiserver (27) empfangen werden.
- 3Procédé selon la revendication 1 ou 2, comprenant également, avant la réception de métadonnées, la sélection (61) du serveur de fichiers de destination (27) qui doit recevoir les métadonnées et les fichiers souches. The method of claim 1 or 2, further comprising:prior to receiving metadata, selecting (61) the destination fileserver (27) that is to receive the metadata and the stub files. Verfahren nach Anspruch 1 oder 2, ferner umfassend: Auswählen (61) des Ziel-Dateiservers (27), der die Metadaten und die Stummel-Datelen zu empfangen hat, vor dem Empfangen von Metadaten.
- 4Procédé selon l'une quelconque des revendications précédentes, comprenant également la sélection (61) d'un système de partage de données en vue de leur réception au niveau du serveur de fichiers de destination (27). The method of any preceding claim, further comprising:selecting (61) a share of data for receipt at the destination fileserver (27). Verfahren nach irgendeinem der vorstehenden Ansprüche, ferner umfassend: Auswählen (61) eines Daten-Teilhaberschaftsbereiches zum Empfang auf dem Ziel-Datelserver (27).
- 5The method of any preceding claim, wherein the set of files is the set of files that have been accessed during a specified period and wherein replacing each stub file comprises recursively replacing the stub file associated with the file that was most-recently accessed until all the stub files in the set of files have been replaced. Verfahren nach irgendeinem der verstehenden Ansprüche, wobei der Satz von Dateien derjenige Satz von Dateien ist, auf die während einer spezifizierten Periode zugegriffen worden ist, und wobei das Ersetzen jeder Stummel-Datei das rekursive Ersetzen der Stummel-Dateien umfasst, die mit der Datei verknüpft sind, auf die zuletzt zugegriffen worden war, bis alle Datei-Stummel aus dem Satz von Dateien ersetzt worden sind. procédé selon l'une quelconque des revendications précédentes, dans lequel l'ensemble de fichiers est l'ensemble des fichiers qui ont fait l'objet d'un accès au cours d'une période spécifiée, et dans lequel le remplacement de chaque fichier souche comprend le remplacement récurrent du fichier souche associé au fichier qui a fait l'objet d'un accès le plus récemment jusqu'à ce que tous les fichiers souches de l'ensemble de fichiers aient été remplacés.
- 6Procédé selon la revendication 5, dans lequel la période spécifiée est une période très récente. The method of claim 5, wherein the specified period is a most-recent period. Verfahren nach Anspruch 5, wobei die spezifizierte Periode eine zuletzt aufgetretene Periode Ist.
- 7Procédé selon l'une quelconque des revendications précédentes, dans lequel l'organe d'emplacements est une antémémoire d'emplacements (36). The method of any preceding claim, wherein the location component is a location cache (36). Verfahren nach irgendeinem der vorstehenden Ansprüche, wobei die Örtlichkeitskomponente ein Orts-Cache (36) ist.
- 8A method according to any of Claims 1 to 7, wherein the fileserver (27) includes:a file system (32) operative to store client files;a policy component (34) operative to store a protection policy associated with the set of files;a mirror service (33) in communication with the policy component (34), the mirror service (33) operative to prepare modified and created files in the set of files to be written to a repository (28) as specified in the protection policy associated with the set of files;a fileserver API (37) coupled to the mirror service (33) and operative to communicate with a repository (28);anda fileserver file transfer module (40) in communication with the file system (32) and operative to transfer files from the file system to at least one repository;the method further comprising the steps of: determining a caching level as stored in the policy component;andrecursively, determining a utilization of the fileserver,comparing (90) the caching level against the utilization;andcreating (91) a file migration candidate list when the utilization exceeds the caching level;staging (92) out one candidate file;replacing (92) the candidate file with a stub file;anddetermining (93) whether the utilization of the fileserver still exceeds the caching level. Procédé selon l'une quelconque des revendications 1 à 7, dans lequel le serveur de fichiers (27) comprend: un système de fichiers (32) apte à agir pour stocker des fichiers clients;un organe de politique (34) apte à agir pour stocker une politique de protection associée à l'ensemble de fichiers;un service de mise en miroir (33) en communication avec l'organe de politique (34), le service de mise en miroir (33) étant apte à agir pour préparer des fichiers modifiés et créés dans l'ensemble de fichiers et qui doivent être écrits dans un entrepôt (28) spécifié dans la politique de protection associée à l'ensemble de fichiers;une interface API de serveur de fichiers (37) reliée au service de mise en miroir (33) et apte à agir pour communiquer avec un entrepôt (28);etun module de transfert de fichiers de serveur de fichiers (40) en communication avec le système de fichiers (32) et apte à agir pour transférer des fichiers du système de fichiers dans au moins un entrepôt;le procédé comprenant également les étapes de: détermination d'un niveau de mise en antémémoire stocké dans l'organe de politique;etdétermination récurrente d'une utilisation du serveur de fichiers;comparaison (90) du niveau de mise en antémémoire avec l'utilisation;etcréation (91) d'une liste de fichiers candidats à la migration lorsque l'utilisation est supérieure au niveau de mise en antémémoire;transfert (92) de l'un des fichiers candidats;remplacement (92) du fichier candidat par un fichier souche;etdétermination (93) du fait que l'utilisation du serveur de fichiers est encore supérieure ou non au niveau de mise en antémémoire. Verfahren nach irgendeinem der Ansprüche 1 bis 7, wobei der Dateiserver (27) beinhaltet;ein zum Speichern von Klienten-Dateien betreibbares Dateisystem (32);eine zum Speichern einer Sammlung von Sicherungsverhaltensregeln, die mit dem Satz von Dateien verknüpft ist, betreibbare Komponente (34) für eine Sammlung von Regeln:einen Spiegelungsdienst (33), welcher mit der Komponente (34) für eine Sammlung von Regeln in Verbindung steht, wobei der Spiegelungsdienst (33) betreibbar ist um vorzubereiten, dass modifizierte und erzeugte Dateien aus dem Satz von Dateien wie in der mit dem Satz von Dateien verknüpften Sammlung von Sicherungsverhaltensregeln spezifiziert in ein Magazin (28) geschrieben werden;eine mit dem Spiegelungsdienst (33) gekoppelte und zur Kommunikation mit einem Magazin (28) betreibbare Dateiserver-API (37);undein Dateiserver-Übertragurigsmodul (40), welches mit dem Dateisystem (32) in Verbindung steht und betreibbar ist, um Dateien aus dem Dateisystem zu mindestens einem Magazin zu übertragen;wobei das Verfahren ferner folgende Schritte umfasst: Festlegen eines Caching-Levels wie in der Komponente für eine Sammlung von Regeln gespeichert;undrekursives Bestimmen einer Nutzung des Dateiservers;Vergleichen (90) des Caching-Levels mit der Nutzung;undErzeugen (91) einer Dateimigrations-Kandioiatenliste, wenn die Nutzung den Caching-Level überschreitet;Bereitstellen (92) eines Datei-Kandldaten;Ersetzen (92) des Datei-Kandidaten durch eine Stummel-Datei;undBestimmen (93), ob die Nutzung des Dateiservers immer noch den Caching-Level überschreitet.
- 9Procédé selon la revendication 8, dans lequel la détermination du fait que l'utilisation du serveur de fichiers est encore supérieure ou non au niveau de mise en mémoire comprend également le transfert (92) d'un autre fichier candidat de la liste de candidats et l'étape qui consiste à déterminer (93) à nouveau si l'utilisation du serveur de fichiers est supérieure au niveau de mise en antémémoire. The method of claim 8, wherein determining whether the utilization of the fileserver still exceeds the caching level further comprises:staging out (92) another candidate file on the candidate list and again determining (93) if the utilization of the fileserver exceeds the caching level. Verfahren nach Anspruch 8, wobei das Bestimmen, ob die Nutzung des Dateiservers immer noch den Carhing-Level überschreitet, ferner umfasst: Bereitstellen (92) eines anderen Datei-Kandidaten aus der Kandidatenliste und erneutes Bestimmen (93), ob die Nutzung des Dateiservers den Caching-Level überschreitet.
- 10A file data protection system comprising:the recovery service (45) having: a receiving component (94) operative to receive metadata and stub files associated with the set of files at the fileserver (27);a location updating component (95) in communication with the receiving component (94) and operative to maintain a list of repository nodes (28) that are associated with each file in the set of files;anda stub file replacement component (96) in communication with the receiving component (94) and operative to replace each stub file with the full content of the file associated with the stub file, said stub file replacement component (96) being configured to be responsive to a client request for a specified file in said set of files to undertake replacement of the stub file for said specified file as a higher priority task than replacement of stub files for non-requested files. Datensicherungssystem, umfassend;einen Dateiserver (27) mit: einem zum Speichern von Klientan-Dateien betreibbaren Dateisystem (32);einem zur Kommunikation mit einem Magazin (28) betreibbaren Dateiserver-API (37);einem Dateiserver-Übertragungsmodul (40), welches mit dem Dateisystem (32) in Verbindung steht und betreibbar ist, um Dateien für das Dateisystem von mindestens einem Magazin (28) entgegenzunehmen;undeinem Wiederherstellungsdienst (45), welcher mit dem Dateiserver-API (37) und mit dem Dateisystem (32) in Verbindung steht und zum Übertragen eines Satzes von Dateien betreibbar ist;wobei der Wiederherstellungsdienst (45) aufweist: eine Empfangskomponente (94), die zum Empfang von mit dem Satz von Dateien auf dem Dateiserver (27) verknüpften Metadaten und Stummel-Dateien betreibbar ist;eine Örtlichkeits-Aktualisierungskomponente (95), welche in Verbindung mit der Empfangskomponente (94) steht und zum Unterhalten einer Liste von mit jeder Datei aus dem Satz von Dateien verknüpften Magazin-Knoten (28) betreibbar ist;undeine Stummel-Detei-Ersetzungskomponente (96), die mit der Empfangskomponente (94) in Verbindung steht und die zum Ersetzen jeder Stummel-Datei durch den vollen Inhalt der mit der Stummel-Datei verknüpften Datei betreibbar ist, wobei die Stummel-Datei-Ersetzungskomponente (96) konfiguriert ist, auf eine Klienten-Anforderung für eine spezifizierte Datei aus dem Satz von Dateien anzusprechen, um das Ersetzen der Stummel-Datei für die spezifizierte Datei als eine Aufgabe mit höherer Priorität durchzuführen als das Ersetzen von Stummel-Dateien für nicht angeforderte Dateien. Système de protection de données de fichiers comprenant: un serveur de fichiers (27) comportant: un système de fichiers (32) apte à agir pour stocker des fichiers clients;une interface API de serveur de fichiers (37) apte à agir pour communiquer avec un entrepôt (28);un module de transfert de fichiers de serveur de fichiers (40) en communication avec le système de fichiers (32) et apte à agir pour recevoir des fichiers pour le système de fichiers à partir d'au moins un entrepôt (28);etun service de récupération (45) en communication avec l'interface API de serveur de fichiers (37) et avec le système de fichiers (32) et apte à agir pour transférer un ensemble de fichiers, le service de récupération (45) comportant: un organe de réception (94) apte à agir pour recevoir des métadonnées et des fichiers souches associés à l'ensemble de fichiers au niveau du serveur de fichiers (27);un organe d'actualisation d'emplacements (95) en communication avec l'organe de réception (94) et apte à agir pour gérer une liste de noeuds d'entrepôt (28) associés à chaque fichier de l'ensemble de fichiers;etun organe de remplacement de fichiers souches (96) en communication avec l'organe de réception (94) et apte à agir pour remplacer chaque fichier souche par le contenu complet du fichier associé au fichier souche, ledit organe de remplacement de fichiers souches (96) étant configuré pour être sensible à une demande de client concernant un fichier spécifié dudit ensemble de fichiers afin d'entreprendre le remplacement du fichier souche correspondant audit fichier spécifié comme une tâche de priorité plus grande que le remplacement de fichiers souches correspondant à des fichiers non demandés.
- 11System nach Anspruch 10, ferner aufweisend:einen Filtertreiber (31), der betreibbar ist, um durch Klientendatei-Anforderungen initiierte Eingabe-/Ausgabe-Aktivitäten mitzuverfolgen und um eine Liste der seit dem letzten Augenblicks-Speicherabzug modifizierten oder erstellten Dateien zu unterhalten;einen Cache (34) mit Verhaltensregeln, der betreibbar ist, um eine mit dem Teilhaberschaftsbereich verknüpfte Sammlung von Sieherungsverhaltensregeln zu speichern;einen Spiegelungsdienst (33), der mit dem Filtertreiber (31) und dem Cache (34) mit Verhaltensregeln in Verbindung steht, wobei der Spiegelungsdienst (33) betreibbar ist, um vorzubereiten, dass modifizierte und erzeugte Dateien in einem Teilhaberschaftsbereich gemäß einer durch eine in der mit dem Teilhaberschaftsbereich verknüpften Sammlung von Sicherungsverhaltensregeln festgelegten Spezifikation in ein Magazin (28) geschrieben werden. Système selon la revendication 10, comprenant également: un pilote de filtre (31) apte à agir pour intercepter une activité d'entrée/sortie initiée par des demandes de fichiers clients et pour gérer une liste de fichiers modifiés et créés depuis une sauvegarde antérieure;une antémémoire de politique (34) apte à agir pour stocker une politique de protection associée à un système de partage;un service de mise en miroir (33) en communication avec le pilote de filtre (31) et avec l'antémémoire de politique (34), le service de mise en miroir (33) étant apte à agir pour préparer des fichiers modifiés et créés dans un système de partage et qui doivent être écrits dans un entrepôt (28) spécifié dans la politique de protection associée au système de partage. The system of claim 10, further comprising a filter driver (31) operative to intercept input/output activity initiated by client file requests and to maintain a list of modified and created files since a prior backup;a policy cache (34) operative to store a protection policy associated with a share;a mirror service (33) in communication with the filter driver (31) and with the policy cache (34), the mirror service (33) operative to prepare modified and created files in a share to be written to a repository (28) as specified in the protection policy associated with the share.
- 12System nach Anspruch 11, ferner umfassend:einen Örtlichkeits-Cache (36), der in Verbindung mit dem Spiegelungsdienst (33) steht und der betreibbar ist, um anzugeben, welches Magazin (28) eine aktualisierte Version einer existierenden Datei empfangen sollte;undeinen Örtlichkeits-Manager (35), der mit dem Örtlichkeits-Cache (36) gekoppelt ist und der zum Aktualisieren des Örtlichkeits-Cache (36) betreibbar ist, wenn das System eine neue Datei in einen spezifischen Magazinknoten (28) schreibt. Système selon la revendication 11, comprenant également: une antémémoire d'emplacements (36) en communication avec le service de mise en miroir (33) et apte à agir pour indiquer l'entrepôt (28) qui doit recevoir une version actualisée d'un fichier existant;etun gestionnaire d'emplacements (35) relié à l'antémémoire d'emplacements (36) et apte à agir pour actualiser l'antémémoire d'emplacements (36) lorsque le système écrit un nouveau fichier dans un noeud d'entrepôt spécifique (28). The system of claim 11, further comprising: a location cache (36) in communication with the mirror service (33) and operative to indicate which repository (28) should receive an updated version of an existing file;anda location manager (35) coupled to the location cache (36) and operative to update the location cache (36) when the system writes a new file to a specific repository node (28).
- 13System nach irgendeinem der Ansprüche 10 bis 12, ferner umfassend:ein lokales Magazin (28), aufweisend: ein lokales Magazinknaten-API (38), das zur Kommunikation mit der Dateiserver-API (37) eingerichtet ist;ein Lokaimagazin-Deteiübertragungsmodul (41), das mit dem Dateiserver-Dateiübertragungsmedul (40) in Verbindung steht und das zum Übertragen von Dateien auf das Dateiserver-Dateiübertragungsmadul (40) eingerichtet ist;undeinen Datentransporteur (39), der in Verbindung mit dem lokalen Magazin-API (38) steht und der zur Überwachung der Replikation von Dateien aus dem lokalen Magazin (28) auf den Dateiserver (27) betreibbar ist. Système selon l'une quelconque des revendications 10 à 12, comprenant également: un entrepôt local (28) comportant: une interface API de noeud d'entrepôt local (38) adaptée pour communiquer avec l'interface API de serveur de fichiers (37);un module de transfert de fichiers d'entrepôt local (41) en communication avec le module de transfert de fichiers de serveurs de fichiers (40) et adapté pour transférer des fichiers au module de transfert de fichiers de serveur de fichiers (40);etun dispositif de mouvement de données (39) en communication avec l'interface API d'entrepôt local (38) et apte à agir pour superviser la duplication de fichiers de l'entrepôt local (28) dans le serveur de fichiers (27). The system of any of claims 10 to 12, further comprising a local repository (28) having: a local repository node API (38) adapted for communicating with the fileserver API (37);a local repository file transfer module (41) in communication with the fileserver file transfer module (40) and adapted for transferring files to the fileserver file transfer module (40);anda data mover (39) in communication with the local repository API (38) and operative to supervise the replication of files from the local repository (28) to the fileserver (27).
- 14System nach Anspruch 13, wobei das Dateiserver-ApI (37) zur Kommunikation mit einem Netzwerk (30) betreibbar ist, und wobei das System ferner umfasst:ein entfernt angeordnetes Magazin (28) mit: einem API (38) betreffend entfernt angeordnete Magazinknoten, eingerichtet zum Kommunizieren mit dem Netzwerk (30);ein Dateiübertragungsmodul (41) betreffend entfernt angeordnete Magazine, welches mit dem lokalen Dateiübertragungsmodul (40) in Verbindung steht und das zum Übertragen von Dateien auf das Dateiserver-Dateiübertragungsmodul (40) eingerichtet ist;undeinen Detentransporteur (39), der mit dem API (28) betreffend entfernt angeordnete Magazins in Verbindung Steht und der zur Überwachung der Replikation von Dateien von dem entfernt angeordneten Magazin (28) auf dem Dateiserver (27) betreibbar ist. Système selon la revendication 13, dans lequel l'interface API de serveur de fichiers (37) est apte à agir pour communiquer avec un réseau (30), et dans lequel le système comprend également: un entrepôt distant (28) comportant: une interface API de noeud d'entrepôt distant (38) adaptée pour communiquer avec le réseau (30);un module de transfert de fichiers d'entrepôt distant (41) en communication avec le module de transfert de fichiers de serveur de fichiers (40) et adapté pour transférer des fichiers au module de transfert de fichiers de serveur de fichiers (40);etun dispositif de mouvement de données (39) en communication avec l'interface API d'entrepôt distant (38) et apte à agir pour superviser la duplication de fichiers de l'entrepôt distant (28) dans le serveur de fichiers (27). The system of claim 13, wherein the fileserver API (37) is operative to communicate with a network (30) and wherein the system further comprises: a remote repository (28) having: a remote repository node API (38) adapted for communicating with the network (30);a remote repository file transfer module (41) in communication with the local file transfer module (40) and adapted for transferring files to the fileserver file transfer module (40);anda data mover (39) in communication with the remote repository API (28) and operative to supervise the replication of files from the remote repository (28) to the fileserver (27).
- 15A computer program comprising computer program elements for implementing the method of any one of claims 1-9. Computerprogramm, umfassend Computerprogrammelemente zum Implementieren des Verfahrens nach irgendeinem der Ansprüche 1 bis 9. programme d'ordinateur comprenant des éléments de programme d'ordinateur pour mettre en oeuvre le procédé selon l'une quelconque des revendications 1 à 9.
- 16A computer usable carrier medium carrying a computer program according to claim 15. Computerverwendbares Trägermedium, ein Computerprogramm gemäß Anspruch 15 tragend. Support utilisable sur ordinateur et porteur d'un programme d'ordinateur selon la revendication 15.
Independent claims16
44 paragraphs in 4 sections, as filed
BACKGROUND OF THE INVENTION
The present invention is related to computer primary data storage systems and methods that provide comprehensive data protection. As background to the invention, hierarchical storage management (HSM) is a data management technique that is deployed with multiple types of data storage technologies like magnetic disk drives, magnetic tape drives, and optical disk drives. HSM has been used to transparently move the contents of files that are least recently accessed by client applications from higher speed, higher cost magnetic disk technology to lower cost, lower speed media like magnetic tape and optical disk. These various storage technologies can be arranged in a multi-level hierarchy from fastest to slowest and/or from most costly to least costly. For the following example, assume that magnetic disks are employed as the fast, more costly primary storage device and magnetic tape is employed as the lower cost, slower access storage technology within an HSM hierarchy. With HSM in operation, the files that are most-recently accessed by client applications are retained in their complete form on magnetic disk and least recently accessed files are migrated out to magnetic tape. HSM software enables client applications to transparently access files that have been migrated out to magnetic tape. When a file that is least-recently used has been migrated out to magnetic tape, a much smaller stub file remains on magnetic disk, typically 1KB to 4KB in size. This stub file contains all of the basic operating system file information as well as an indication that this file has been migrated from the magnetic disk to another medium. The stub file also includes information that HSM uses to acquire the entire contents of a file from a magnetic tape and restore the complete content of a file to magnetic disk transparently when it has been requested by a client application.
Thus, HSM can be an effective way to manage large amounts of infrequently accessed data in a cost effective manner. HSM is also effective in eliminating the administrative alerts and subsequent operational actions associated with filesystems running out of available capacity for additional client data. Without HSM, when a filesystem runs out of available capacity, a storage administrator must quickly employ one of the following manual data management techniques to allow client applications to remain operational: <ul id="ul0001" list-style="bullet" compact="compact"><li>Use a third-party archiving program to identify unwanted files and commit these to magnetic tape or optical disk. Archiving deletes all information about files from the fileserver's filesystem and applications cannot access these files without having them recovered from archive tapes or optical disks.</li><li>If the fileserver can continue to accommodate additional magnetic disk drives, these are added to the server and associated volume and filesystem are expanded. This is a time consuming process and may result in client application downtime.</li><li>If a fileserver has been expanded to its limits of disk storage capacity, the administrator must migrate some number of shares of data (a share being a directory or folder of storage capacity) to another server that has available capacity.</li></ul>
Contrasting this with HSM-based fileservers, as filesystems fill up, least recently used data is automatically migrated out to another storage medium like tape. Unlike archiving, HSM provides client applications with transparent continual access to their files, even when they've been migrated to magnetic tape. There is no need to add additional magnetic disk storage to a fileserver that supports HSM since storage is expanded by adding more magnetic tapes to the lower-level in the storage hierarchy. Finally, migrating shares from one server to another does not have to be performed specifically for reasons of balancing capacity across multiple fileservers.
In <patcit id="pcit0001" dnum="US5537585A"><text>U.S. Patent 5,537,585, Blickenstaff, et al.</text></patcit> describe an HSM system that migrates least recently used data from a fileserver to a lower-cost back-end removable storage system. The HSM servers are monitored for filesystem capacity utilization, and when they fill up, least recently used files are identified and migrated out to back end storage.
<patcit id="pcit0002" dnum="US5276860A"><text>U.S. patent 5,276,860</text></patcit> describes a system that employs both HSM and backup technology using removable storage media like optical disks and magnetic tape. HSM is used as a means of reducing the amount of data that had to be committed to backups since data that is staged out to an optical disk is considered backed-up. However a need remains for data protection systems that provide timely and cost-effective disaster recovery and/or data migration.
<patcit id="pcit0003" dnum="WO0235359A"><text>PCT Publication WO 02/35359</text></patcit> relates to the creation of a file system that separates its directory presentation from its data store. This document discloses processing, division, distribution, managing, synchronizing, and reassembling of file system objects without delay to the presentation of content to the user. The system disclosed uses a reduced amount of storage space, and it is possible to manage and control the integrity of the files distributed across the network, and to serve and reconstruct files in real time using a Virtual File Control System.
SUMMARY OF THE INVENTION
Preferred embodiments of the invention provide a method as claimed in Claim 1 and a system as claimed in Claim 10.
Embodiments of the invention relate to computer primary data storage systems and methods that provide comprehensive data protection. Embodiments of data protections systems according to the present invention work in reverse to <patcit id="pcit0004" dnum="US5276860A"><text>U. S Patent 5,276, 860</text></patcit> in that the systems regularly back up data and, when they must migrate data, they select files that have already been backed up as migration candidates. When a filesystem fills up, migration candidates are quickly and efficiently replaced by smaller stub files.
One aspect of the present invention provides data management systems and methods where hierarchical storage management accelerates the time to complete disaster recovery operations that traditionally could take hours or days to complete. This recovery time factor is relevant since, while these operations are taking place, client applications cannot access their files.
Embodiments of the present invention provide data management scenarios where hierarchical storage management accelerates the time to complete operations that traditionally could take hours or days to complete. This recovery time factor is relevant since while these operations are taking place, client applications cannot access their files. <ul id="ul0002" list-style="none" compact="compact"><li>1. Moving one or more shares of data from one fileserver to another. Even though HSM based fileservers do not require that this be performed since their filesystems do not fill up, storage administrators perform this activity to spread the client application load across more fileservers.</li><li>2. Moving the entire contents of a NAS fileserver to another Fileserver. Administrators perform this service in the following situations: <ul id="ul0003" list-style="bullet" compact="compact"><li>Performance load balancing</li><li>Loss of a fileserver</li><li>Fileserver hardware and operating system upgrades</li><li>Site Disaster Recovery - loss of an entire site's worth of fileservers</li></ul></li></ul>
BRIEF DESCRIPTION OF THE DRAWINGS
<ul id="ul0004" list-style="none" compact="compact"><li><figref idref="f0001">FIG 1</figref> is a diagram of a deployment of one embodiment of the invention across three data centers.</li><li><figref idref="f0001">FIG 2</figref> illustrates how one embodiment of a protection policy creates a relationship between a fileserver share and associated repositories such as those shown in <figref idref="f0001">FIG. 1</figref>.</li><li><figref idref="f0002">FIG 3</figref> shows one embodiment of a first step in recovering from the loss of a fileserver or an entire site.</li><li><figref idref="f0002">FIG 4</figref> shows one embodiment of a second step in recovering from the loss of a fileserver or an entire site.</li><li><figref idref="f0003">FIG5</figref> shows one embodiment of a first step in migrating shares of data from one fileserver to another.</li><li><figref idref="f0004">FIG6</figref> shows one embodiment of a second step in migrating shares of data from one fileserver to another.</li><li><figref idref="f0005">FIG7</figref> shows one embodiment of a user interface that a storage administrator uses to initiate a migration of one or more shares from the selected fileserver to another fileserver.</li><li><figref idref="f0006">FIG8</figref> shows a screenshot of one embodiment of a user interface for the protection policy of <figref idref="f0001">FIG. 2</figref>.</li><li><figref idref="f0007">FIG9</figref> shows one embodiment of the apparatus and the software components that are used to protect new client data to a local repository node.</li><li><figref idref="f0008">FIG10</figref> shows one embodiment of the apparatus that replicates data among repositories.</li><li><figref idref="f0009">FIG11</figref> is a flow chart illustrating one embodiment of a two stage share migration or server recovery operation.</li><li><figref idref="f0010">FIG12</figref> is a flow chart illustrating one embodiment of an HSM stage out process.</li><li><figref idref="f0011">FIG13</figref> is a schematic block diagram of one embodiment of the recovery service of <figref idref="f0007">FIG9</figref>.</li></ul>
DETAILED DESCRIPTION OF THE DRAWINGS
<figref idref="f0001">FIG1</figref> is a diagram that illustrates one embodiment of an integrated primary data storage and data protection system according to the invention. Fileservers 4 provide primary data storage capacity to client systems <b>5</b> via standard network attached storage (NAS) protocols like network file system (NFS), common Internet file system (CIFS) and file transfer protocol (FTP). The apparatus is designed to operate among two or more data centers <b>1.</b> Two or more repositories 3 deployed across these data centers provide storage capacity and data management processing capability to deliver complete data protection for their associated fileserver primary storage systems. The apparatus leverages metropolitan or wide area internet protocol (IP) networking 2 to allow repositories to send and receive data for replication. By having data replicated to a local and at least one remote repository from the originating fileserver, these repositories act as a replacement for traditional on-site and off-site tape storage systems and tape vaulting services. According to one embodiment, in the event of a site disaster, all fileservers that were lost are recovered by deploying new fileservers at a surviving site and recreating the content of the failed fileservers from the content in the surviving repositories.
<figref idref="f0001">FIG2</figref> is a diagram that illustrates an association between a fileserver <b>6</b> and two repositories <b>8</b> that are deployed across data centers. All primary data storage activity occurs between one or more clients and one or more fileservers through a NAS share 7. A fileserver is typically configured to have tens of shares. These shares allow the primary storage capacity of the fileserver to be shared and securely partitioned among multiple client systems.
A share is created on a fileserver as a directory or folder of storage capacity. The contents of this shared directory or folder is accessible by multiple clients across a local area network. For example, in the Microsoft Windows environment, CIFS shares appear as storage folders within LAN-connected servers under "My Network Places" of the Windows Explorer user interface. For UNIX environments, shares are accessed through mount points which define the actual server and folder where data is stored as well as a virtual folder that appears to be part of the local client system's filesystem.
Because one embodiment of a system according to the invention is both a primary data storage system and a data protection system, a storage administrator defines how the system protects each share of a fileserver across two or more repositories through the creation of a unique protection policy <b>9</b> for that share. In one embodiment, this protection policy defines which repositories the system will use to protect each share's data. In one embodiment it also defines how often data protection will occur, how many replicas will be maintained within each repository based on the criticality of a share's data, and how updates and modifications to share data should be maintained. On a periodic basis, each fileserver examines the protection policy for its shares and when appropriate, the fileserver captures all recent changes to a share's files and stores/protects these files within two or more repositories.
<figref idref="f0002">FIG3</figref> shows that client data is protected among multiple repositories, as well as the metadata associated with that client data. By maintaining this metadata in replicated repositories, fileservers and/or their shares may be easily migrated to another fileserver. In one embodiment metadata for a file includes the following information: <ul id="ul0005" list-style="bullet"><li>The fileserver name where the file was created</li><li>The size of the file</li><li>A list of all of the repository nodes that maintain a replica of this file</li><li>A computed MD5 content checksum of the file when it was first created or last modified.</li></ul>
Metadata files are maintained within repository nodes and cached on fileservers to assist in managing the protection of these files across multiple repository nodes.
HSM stub files, in contrast, are typically only contained on the fileserver and are used to assist the fileserver's HSM software in transparently providing file access to client applications, regardless of the physical location of a file.
<figref idref="f0002">FIG4</figref> displays server recovery. Fileserver Y <b>14</b> has its client data and metadata regularly protected within two repositories. When fileserver Y becomes unavailable because of hardware failure or the loss of a site, the redundantly replicated fileserver configuration and file metadata from the last backup period is applied to either a surviving fileserver or to a new fileserver, shown as Z 15. This server recovery action is initiated by a storage administrator.
The actions performed by a storage administrator using embodiments of this invention for server recovery are similar to the actions performed when shares of data must be migrated from one operational fileserver to another. Share migration is typically performed in order to improve client application performance by shifting the load created by clients for one or more shares on a heavily accessed fileserver to an underutilized or more powerful existing server, or to a new fileserver. Share migration may also be performed, for example, when client applications are physically relocated to different offices or data centers and expect continued high-performance LAN access to their shares of data instead of accessing the same fileserver across a metropolitan or wide-area network. In this case, all share data from the previous facility can be quickly migrated to one or more servers at the new location.
<figref idref="f0003">FIG5</figref> and <figref idref="f0004">FIG6</figref> represent a time-sequenced pair of diagrams that illustrate the activities that occur when one or more fileserver shares are migrated from one fileserver to another. In this scenario, there has been no failure of fileserver Y, but the storage administrator wishes to migrate Y<sub>a</sub> and Y<sub>b</sub> to a new server called fileserver Z.
In <figref idref="f0003">FIG5</figref>, fileserver Y <b>16</b> has three shares called Y<sub>a</sub>, Y<sub>b</sub> and Y<sub>c</sub>. In one embodiment, all of the metadata associated with these three shares is replicated to two repositories as Y Metadata <b>17,18</b> in the diagram.
In <figref idref="f0004">FIG6</figref>, a storage administrator has decided to migrate shares Y<sub>a</sub> and Y<sub>b</sub> from fileserver Y to fileserver Z 19. The following procedure is performed by the system administrator: <ul id="ul0006" list-style="bullet" compact="compact"><li>Notify clients that there will be a period where their files will not be accessible during the migration.</li><li>Initiate a share migration using the invention's management application interface to initiate the migration of shares Y<sub>a</sub> and Y<sub>b</sub> from fileserver Y to fileserver Z <b>19.</b></li><li>All of the metadata associated with these two shares is replicated to two repositories as Z Metadata <b>20, 21</b> in the diagram.</li><li>When the management interface reports that a migration is completed, client applications are re-associated with their files now located on fileserver Z. This share migration occurs quickly relative to conventional migration because only the much smaller 4KB stub files need to be present on the fileserver to allow clients to regain access to their files.</li><li>As client requests are made for data that is only represented as a stub file on the fileserver, the stub file is replaced with the entire contents of a file transparently through a "stage-in" process. In other words, staging-in a file involves replacing a stub file with the entire contents of the file. Since data to be staged-in is sent from fast networked repository node magnetic disk drives instead of traditional magnetic tapes or optical disks, the time to stage-in is typically reduced from tens of seconds or minutes to under a second for moderately sized files (10-50MB).</li><li>As a background task, the Recovery service <b>45</b> shown in <figref idref="f0007">FIG9</figref> also repopulates fileserver Z with all of the files that were present in their full form on fileserver Y before the share migration was requested.</li><li>Finally, the share or shares that were migrated from fileserver Y to fileserver Z are deleted from fileserver Y since they are no longer referenced by client applications from that server.</li></ul>
<figref idref="f0005">FIG7</figref> displays one embodiment of a user interface that is used to initiate a file migration of a share from one fileserver to another. The illustrated screenshot displays two shares <b>22</b>, TEST1 and test2, existing on a single fileserver, GPNEON 23. In this example, a storage administrator recently chose to migrate the share called TEST1 from the GPNEON fileserver to the GPXEON <b>24</b> fileserver. The status column of the user interface shows that the TEST1 share is currently in a migrate state <b>25.</b> The status column of the user interface also shows that the share test2 on the GPNEON fileserver is online <b>26</b> and available for migrating to another fileserver if desired.
<figref idref="f0006">FIG8</figref> is a screenshot of one embodiment of the present invention's protection policy. There is a unique protection policy defined by a storage administrator for each share of each fileserver. Before arriving at the protection policy screen, a storage administrator creates a share and allows it to be accessible by CIFS and/or NFS and/or FTP. Once a new share is created, the protection policy screen is displayed. Within this screen, the storage administrator can specify the following data protection parameters: <ul id="ul0007" list-style="bullet" compact="compact"><li>Protect this share 78 - this checkbox is normally checked indicating the data in this share should be protected by repositories. There are certain client applications that might choose to use a fileserver for primary storage, yet continue to protect data using third party backup or archiving products. If this checkbox is left unchecked, all other options in the protection policy user interface are disabled. One embodiment of the system can only perform share migration and server recovery with shares that are protected.</li><li>Protection Management - Backup Frequency <b>79</b> - this option determines how often a fileserver share's data will be protected in the local and remote repositories. In one embodiment, the backup frequency intervals can be selected from a list of time intervals which include: 15 minutes, 30 minutes, 1 hour, 2 hours, 3 hours, 4 hours, 6 hours, 8 hours, 12 hours and 24 hours. All backup frequency intervals are anchored to 12:00 midnight of each fileserver's local time-zone. Setting the backup frequency to 24 hours is similar to performing traditional daily incremental backups. Setting this interval to 15 minutes allows versions of files that change often during the day to be protected on 15 minute intervals. Only files that have changed since the last backup occurred are saved in repositories.</li><li>Protection Management - Number of replicas per repository. This feature allows a storage administrator to determine how many replicas <b>80</b> of data to create within each repository <b>81</b> when a share is protected. In one embodiment, there must be at least one replica stored in a repository that is local to the share's fileserver. It's possible to maintain multiple replicas within a single repository. In this case, replicas are maintained on different repository nodes of a repository to ensure continued access to a replica in the event of a single repository node failure or network failure. The location and number of replicas can be changed over time. To increase data availability for data that is increasing in criticality, more replicas per repository and additional repositories may be specified. For data that is decreasing in importance, fewer replicas may be maintained in the repositories, which makes more storage capacity available to other shares that are also using those repositories.</li><li>Version Management - Keep Version History <b>82</b> - this checkbox should be checked for shares whose file content is regularly being updated. When checked, the specified repositories will maintain a version chain of all changes that were identified at each backup frequency interval. For shares of data that have unchanging file content, this checkbox can be unchecked.</li><li>Version Management - Version Compression <b>83</b> - the three compression options are to not compress, to reverse delta compress or to apply simple file compression to a share's files. Delta compression typically provides the highest compression ratio for shares whose files are regularly being modified.</li><li>Version Management - Version Compaction <b>84</b> - compaction provides a means of removing versions of files based on each version's age. For example, the version compaction option for a file share may be configured to maintain only one monthly version of a file after a year, one weekly version of a file that's older than 6 months and one daily version of a file that's older than 1 month. All "extra" versions can be automatically purged from repositories, which in turn, makes more storage capacity available for new versions of files as well as new files.</li><li>Advanced Options - Purge on Delete <b>85 -</b> this option, when checked will cause files that are deleted from a fileserver's share to also be purged from repositories as well. This feature is effective with applications like third party backup, where some of the replicas and versions that are being retained by repositories are no longer needed to satisfy that application's recovery window and may be purged from all repositories.</li><li>Advanced Options - Caching Level <b>86</b> - this allows the storage administrator to control the level at which the Mirror service <b>33</b> (<figref idref="f0007">FIG9</figref>) looks for candidate files to be staged out to repository nodes. It represents the approximate percentage of client data that will be cached on a fileserver. Normally, this option is set to "Optimize for Read" to allow the maximum number of most-recently accessed files to be available to client applications at the highest performance levels. All least recently used data is maintained in two or more repositories. Conversely, the caching level can be set to "Optimize for Write", which reduces the amount of cached data available to clients but provides consistently high levels of available storage capacity for write operations on the fileserver, i.e., for receiving data. One uses this mode/setting -for applications such as third party backup applications. In this mode, by aggressively moving data off of a fileserver into repositories, the application sees the fileserver as a storage device with virtually infinite capacity.</li></ul>
<figref idref="f0007">FIG9</figref> and <figref idref="f0008">FIG10</figref> illustrate modules used to protect data files created by a client using a local repository and a remote repository.
<figref idref="f0007">FIG9</figref> displays one embodiment of the apparatus and software modules of the present invention that are associated with protecting client files to a local repository. The apparatus includes a fileserver <b>27</b> and a single local repository node 28. Clients access a fileserver via the client IP-based (Internet Protocol) network <b>29</b> and communicate with the fileserver using NFS, CIFS or FTP protocols. All fileservers and all repository nodes are interconnected by an internal IP-based (Internet Protocol) network <b>30.</b> Current client files reside on a fileserver's filesystem <b>32.</b>
According to embodiments of the invention, all input/output activity initiated by client file requests is intercepted by the filter driver <b>31</b>. The fileserver software maintains a list of all modified or created files since this last snapshot occurred. In one embodiment, snapshot intervals can range from 15 minutes to 24 hours, based on the backup frequency <b>19</b> of the protection policy. On the schedule of the backup frequency, the mirror service <b>33</b> prepares all modified files in a share to be put into the repositories <b>3</b> (shown in <figref idref="f0001">Fig. 2</figref>) that are specified in that share's protection policy. The protection policies are stored and replicated across multiple repositories, and they are cached and regularly updated within each fileserver in the protection policy cache <b>34.</b> For example, if a share's protection policy has its backup frequency set to one hour, on the transition to the next hour, the mirror service <b>33</b> initiates a backup of all changed files in the last hour to a local repository <b>28.</b> For all new files, any repository node of the local repository can be used to hold a replica of a file. For files that have been modified, the mirror service directs new versions of the existing file to the same repository node as prior versions of that file. The mirror service queries the location cache <b>36</b> to determine which repository node should receive an updated version of an existing file. This location cache is updated regularly by the location manager 35 when the fileserver writes files to specific repository nodes. Once the location manager identifies all destination repository nodes for each file of a share for the latest collection of updated or created files, the fileserver communicates to each local repository via a fileserver API <b>37</b> and a repository node API <b>38.</b> Each repository node's data mover <b>39</b> supervises the replication of files from the fileserver to its repository node. The fileserver file transfer module <b>40</b> transfers files from the fileserver filesystem to each repository node's file transfer <b>41</b> module. Once the files are replicated to specific disk drives within a repository node, its location manager <b>42</b> updates its location cache <b>43</b> with repository node location information. For all files that arrive at a repository node that are modified versions of existing files, the share's protection policy <b>44</b> version management settings are reviewed to determine whether new versions should be compressed and whether older versions should be maintained.
In <figref idref="f0007">FIG9</figref>, the recovery service <b>45</b> on the fileserver is responsible for managing the movement of one or more shares of data either from an existing fileserver (migration) or to a new server when an existing fileserver has failed or was lost in a site disaster.
At this point in the description, client data is only replicated to a local repository. <figref idref="f0008">FIG10</figref> illustrates one embodiment of modules that implement a process that protects data to one or more remote repositories to completely protect client data from site disaster. <figref idref="f0008">FIG10</figref> displays a local repository node 46 that, from the actions described in <figref idref="f0007">FIG9</figref>, holds the first replica of data. <figref idref="f0008">FIG10</figref> also shows a remote repository node <b>47.</b> These are connected to each other across a metropolitan or wide-area network <b>48.</b> In one embodiment, all data that is transferred between local and remote repositories may be secured by virtual private networking (VPN) <b>49</b> encryption. The local repository node's replication service <b>50</b> is responsible for reviewing the protection policy <b>51</b> for all files that were just created as part of the recent fileserver backup. Each repository node acts as a peer of other repository nodes. Based on the protection policy each repository node manages the movement of files among all repository nodes using repository node APIs <b>52, 53,</b> data movers <b>54,</b> and file transfer modules <b>55, 56.</b> Once the data is replicated to remote repositories, the location manager <b>57</b> of each repository node updates the location cache <b>58</b> to track where files are maintained within that repository node. The version service <b>59</b> of the remote repository node manages file version compression, and compaction according to the protection policy.
<figref idref="f0009">FIG11</figref> displays the method by which a share migration or fileserver recovery operation is performed. A storage administrator initiates <b>60</b> the migration/recovery activity through the secure web-based user interface displayed in <figref idref="f0005">FIG7</figref>. From this interface, the storage administrator selects <b>61</b> the share or shares that should be included in the share migration or fileserver recovery, and also names the destination fileserver to receive the share(s). This action initiates a share migration or recovery service job on the destination fileserver.
The destination fileserver's recovery service collects <b>62</b> the share's metadata from one of the two or more repositories that may have a replicated copy of that metadata. In addition to share metadata, stub files also populate <b>63</b> the destination fileserver. Stub files are files that appear to a user as a normal file but instead include a pointer to the actual location of the file in question.
During the load of the relevant stub files, the recovery/migration service updates <b>64</b> the fileserver's location cache to maintain a list of the repository nodes that are associated with each recovered file. Once these stub files are in place, clients can be connected to the new fileserver to allow them to immediately view all of their files. To the clients, the recovery operation appears to be complete, but accesses to each file will be slightly slower since a file requested by a client initiates a transfer from a repository node back into the fileserver.
<figref idref="f0009">FIG11</figref> also shows the second phase of the two phases of migration/recovery. In the second phase, clients have access <b>65</b> to their files. The recovery service responds to client requests for files as a high priority task and replaces the stub file. e.g., a 4KB stub file, with the full content of each file requested. At a lower priority, the recovery service begins/continues loading the files that were most-recently accessed on the previous fileserver back into the destination fileserver. This process expedites client requests for data while the migration/recovery process proceeds.
More specifically, the process determines <b>66</b> if all of the files that were most recently accessed on the prior fileserver have been transferred in full to the destination fileserver. If so, the process determines that the migration/recovery process is complete. If not, the process determines <b>67</b> if a client file request for a yet-to-be-transferred file has been received. If so, the process replaces <b>70</b> the stub file associated with the requested file with the full content of the file. If not, the process transfer <b>68</b> in a file that have yet to be transferred and returns to step <b>66,</b> i.e., determining whether all specified files have been transferred.
<figref idref="f0011">FIG13</figref> shows one embodiment of the recovery service 45 of <figref idref="f0007">FIG9</figref>. With reference to <figref idref="f0011">FIG 13</figref>, a receiving component <b>94</b> receives metadata and stub files associated with the set of files at a destination fileserver. A location updating component <b>95,</b> in communication with the receiving component, maintains a list of repository nodes that are associated with each file in the set of files. In addition, a stub file replacement component <b>96,</b> in communication with the receiving component, replaces the stub files with the full content of the file associated with the stub file as described above.
<figref idref="f0010">FIG12</figref> illustrates the flowchart that describes how fileserver files are migrated out to repositories. The Mirror service <b>33</b> (<figref idref="f0007">FIG9</figref>) constantly monitors the consumption of space in the fileserver's filesystem. It compares the current filesystem consumption against the storage administrator's selection for caching level 86 in the protection policy (<figref idref="f0006">FIG8</figref>). The three caching levels available to storage administrators in the protection policy are "Optimize for Read", "Optimize for Write" and "Balanced". The translation from these three settings to actual percentages of filesystem utilization are maintained in configurable HSM XML tables, but in one embodiment by default, optimize for read attempts to maintain about 70% of the total filesystem for caching client application files, balanced attempts to maintain about 50%, and optimize for write attempts to maintain about 30% of the data in the filesystem. Optimize for write, more importantly, tries to keep 70% of the fileserver's space available for incoming data for client applications that mostly write data and do not access this data regularly.
Based on the comparison <b>90</b> of filesystem utilization and the caching level of the protection policy, the mirror service either continues to monitor filesystem utilization or it initiates a "stage out". The first step in preparing for a stage-out is to identify files 91 that make good candidates for staging. In one embodiment, since the stub file that replaces the real file consumes 4KB, all files less than or equal to 4KB are not considered as candidates for stage-out. Larger files are preferred for stage-out as are files that have not been accessed by client applications recently. Finally, the preferred candidate file is one that has already been backed up into the repository by the periodic backups that take place in the present invention. Once the list of candidate files is made, a candidate file is staged out 92 according to the list. Once, the candidate file is staged out, the system determines 93 if the HSM caching threshold is still exceeded. If not, the stage-out of more files takes place.
Since stub files may be replacing very large files, filesystem space quickly frees up to bring the filesystem utilization back under the threshold. Traditional HSM systems must quickly stage-out entire files out to magnetic tape or optical disk when an HSM filesystem fills up. But embodiments of the present invention select candidates that have already been protected in two repositories as part of its ongoing backup. So when it needs to stage a file out, it only needs to replace the full file on the fileserver with a stub file that points to the full version of the file stored during backups. This approach eliminates the time-consuming process of moving hundreds or thousands of files from the fileserver to the repository. This approach creates a system that can more quickly react to rapid changes in filesystem utilization. Once the filesystem consumption falls below the HSM caching level, the mirror service returns to monitoring the filesystem utilization. In this way, the Mirror service continues to stage data to a value below the threshold, as defined in the configurable HSM XML table.
The present invention makes data available to clients after a migration/recovery operation many times faster than traditional data storage and protection systems. The present invention's use of hierarchical storage management allows it to present data to clients as soon as the file metadata and the 4KB HSM stub files are loaded onto the destination fileserver. Since a stub file that has recently either been migrated or recovered to a different fileserver acts as a proxy for the complete file maintained in one or more repositories, the migration / recovery activity appears to be completed from the client application perspective. Also, in one embodiment, only the files that are most-recently used, as defined by a filesystem full threshold, must be reloaded into the new fileserver as a background task. With traditional share migration or server recovery, all files would have to be reloaded back onto a server which could extend the recovery or migration time significantly.
Thus, one embodiment of the present invention applies hierarchical storage management (HSM) in order to accelerate the migration of files from one fileserver to another and to reduce the time necessary for clients to regain access to their data after the complete loss of a fileserver.
One embodiment of the present invention is an integrated data storage and data protection system that is physically deployed across two or more data centers. One deployment of an apparatus according to the invention within each data center includes one or more fileservers and one or more repositories. The fileservers provide primary disk storage capacity to IP-networked clients via NFS, CIFS or FTP protocols. Each repository is a virtualized pool of disk storage capacity that: <ul id="ul0008" list-style="bullet" compact="compact"><li>Acts as a second level of data storage to a fileserver's primary level of data storage as part of the hierarchical storage management system.</li><li>Acts as a replacement for magnetic tape backup systems by regularly storing and maintaining versions of new and modified files.</li><li>Acts as a replacement for offsite media storage and offsite disaster recovery systems by replicating all data that is stored in a repository that's local to a fileserver to one or more offsite repositories.</li><li>Acts as a stable replicated storage environment for replicating all fileserver configuration and file metadata information.</li></ul>
A protection policy defines how client data within the fileserver's primary disk storage system will be protected among the collection of repositories located in multiple data centers. Each share's protection policy specifies, among other things, the number of replicas to maintain within multiple repositories. In one embodiment, replicas reside in at least two repositories. These repositories should be located a disaster-safe distance from each other.
The backup frequency defines the periodic time interval at which the fileserver protects new and updated client data into the specified repositories. As part of periodic backup, each fileserver's configuration and file metadata information is also protected in multiple repositories. The configuration and metadata information that is stored in repositories is used for recovering from the loss of one or more fileservers and for migrating shares of data among a collection of fileservers.
Contents4
11 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| WO0235359A | Cites | World Intellectual Property Organization (WIPO) |
| US5564037A | Cites | United States of America |
| US5659614A | Cites | United States of America |
| US6023709A | Cites | United States of America |
| US2002095416A1 | Cites | United States of America |
| US6330572B1 | Cites | United States of America |
| US6389427B1 | Cites | United States of America |
| US6453339B1 | Cites | United States of America |
50 members in 8 offices
Priority claims7
| Document | Office | Kind | Date |
|---|---|---|---|
| 409684P | United States of America | – | |
| 40968402 | United States of America | P | |
| 0328250 | United States of America | W | |
| 409684P | – | – | – |
| US20020409684P | – | – | – |
| US2003028250 | – | – | – |
| WO2003US28250 | – | – | – |
Members50
| Document | Office | Kind | |
|---|---|---|---|
| CA2497305A1 | Canada | A1 | |
| CA2497306A1 | Canada | A1 | |
| CA2497625A1 | Canada | A1 | |
| CA2497825A1 | Canada | A1 | |
| WO2004025404A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2004025470A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2004025498A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2004025517A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU2003268572A1 | Australia | A1 | |
| AU2003273312A1 | Australia | A1 | |
| AU2003278779A1 | Australia | A1 | |
| AU2003282795A1 | Australia | A1 | |
| US2004088331A1 | United States of America | A1 | |
| US2004088382A1 | United States of America | A1 | |
| US2004093361A1 | United States of America | A1 | |
| US2004093555A1 | United States of America | A1 | |
| WO2004025404A3 | World Intellectual Property Organization (WIPO) | A3 | |
| WO2004025404A9 | World Intellectual Property Organization (WIPO) | A9 | |
| EP1537496A1 | European Patent Office (EPO) | A1 | |
| EP1540441A2 | European Patent Office (EPO) | A2 | |
| EP1540478A1 | European Patent Office (EPO) | A1 | |
| EP1540510A1 | European Patent Office (EPO) | A1 | |
| JP2005538467A | Japan | A | |
| JP2005538469A | Japan | A | |
| JP2005538470A | Japan | A | |
| JP2005538471A | Japan | A | |
| EP1537496A4 | European Patent Office (EPO) | A4 | |
| EP1540441A4 | European Patent Office (EPO) | A4 | |
| EP1540510A4 | European Patent Office (EPO) | A4 | |
| EP1540478A4 | European Patent Office (EPO) | A4 | |
| US7246140B2 | United States of America | B2 | |
| US7246275B2 | United States of America | B2 | |
| EP1537496B1 | European Patent Office (EPO) | B1 | |
| EP1540441B1This record | European Patent Office (EPO) | B1 | |
| AT400026T | Austria | T | |
| AT400027T | Austria | T | |
| ATE400026T1 | Austria | T1 | |
| ATE400027T1 | Austria | T1 | |
| DE60321927D1 | Germany | D1 | |
| DE60321930D1 | Germany | D1 | |
| EP1540478B1 | European Patent Office (EPO) | B1 | |
| AT429678T | Austria | T | |
| ATE429678T1 | Austria | T1 | |
| DE60327329D1 | Germany | D1 | |
| EP1540510B1 | European Patent Office (EPO) | B1 | |
| AT439636T | Austria | T | |
| ATE439636T1 | Austria | T1 | |
| US7593966B2 | United States of America | B2 | |
| DE60328796D1 | Germany | D1 | |
| US7925623B2 | United States of America | B2 |
67 legal events, as 6 offices reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | Office | |
|---|---|---|---|
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Patent expired after termination of 20 yearsExpiredPE20 | PE20 | GB | |
| Expiry of rightR071 | R071 | DE | |
| Opt-out of the competence of the unified patent court (upc) registeredP01 | P01 | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Amendment of ipc main classPREVIOUS MAIN CLASS: G06F0017300000R079 | R079 | DE | |
| Fee paymentPLFP | PLFP | FR | |
| Fee paymentPLFP | PLFP | FR | |
| Change of representativeR082 | R082 | DE | |
| Fee paymentPLFP | PLFP | FR | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| No opposition filedOpposition26N | 26N | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| No opposition filed within time limitOppositionORIGINAL CODE: 0009261PLBE | PLBE | EP | |
| Information on the status of an ep patent application or granted ep patentGrantedSTATUS: NO OPPOSITION FILED WITHIN TIME LIMITSTAA | STAA | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Patent ceasedCeasedPL | PL | CH | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Nl: lapsed or annulled due to failure to fulfill the requirements of art. 29p and 29m of the patents actLapsedNLV1 | NLV1 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| European patents granted designating irelandGrantedFG4D | FG4D | IE | |
| Corresponds to:REF | REF | EP | |
| European patent takes effect as a national patent in ch/liEP | EP | CH | |
| Designated contracting statesAK | AK | EP | |
| European patent grantedGrantedFG4D | FG4D | GB | |
| (expected) grantORIGINAL CODE: 0009210GRAA | GRAA | EP | |
| Grant fee paidORIGINAL CODE: EPIDOSNIGR3GRAS | GRAS | EP | |
| Title (correction)FILE DATA PROTECTION APPARATUS AND METHODRTI1 | RTI1 | EP | |
| Despatch of communication of intention to grant a patentORIGINAL CODE: EPIDOSNIGR1GRAP | GRAP | EP | |
| First examination report despatched17Q | 17Q | EP | |
| Supplementary search report drawn up and despatchedA4 | A4 | EP | |
| Information provided on ipc code assigned before grantRIC1 | RIC1 | EP | |
| Information provided on ipc code assigned before grantRIC1 | RIC1 | EP | |
| Information on inventor provided before grant (corrected)RIN1 | RIN1 | EP | |
| Information on inventor provided before grant (corrected)RIN1 | RIN1 | EP | |
| Information on inventor provided before grant (corrected)RIN1 | RIN1 | EP | |
| Information on inventor provided before grant (corrected)RIN1 | RIN1 | EP | |
| Information on inventor provided before grant (corrected)RIN1 | RIN1 | EP | |
| Information on inventor provided before grant (corrected)RIN1 | RIN1 | EP | |
| Information on inventor provided before grant (corrected)RIN1 | RIN1 | EP | |
| Request for extension of the european patent (deleted)DAX | DAX | EP | |
| Request for examination filed17P | 17P | EP | |
| Designated contracting statesAK | AK | EP | |
| Request for extension of the european patentAX | AX | EP | |
| Public reference made under article 153(3) epc to a published international application that has entered the european phaseORIGINAL CODE: 0009012PUAI | PUAI | EP |
Numbers
- Publication
- 1540441
- Publication, DOCDB
- 1540441
- Publication, EPODOC
- EP1540441
- Application
- 3749544
- Application, DOCDB
- 03749544
- Application, EPODOC
- EP20030749544
Titles3
- German
- VERFAHREN UND VORRICHTUNG ZUR DATEISICHERUNG
- English
- FILE DATA PROTECTION APPARATUS AND METHOD
- French
- PROCEDE ET DISPOSITIF DE PROTECTION DE FICHIERS
Classification
- CPC, 12
- G06F11/1448
- G06F11/1458
- G06F11/1461
- G06F11/1469
- G06F11/1662
- G06F11/2038
- G06F11/2048
- G06F11/2094
- G06F2201/84
- G06F11/1464
- Y10S707/99953
- Y10S707/99943
- IPC, 12
- G06F17 30
- G06F11 20
- G06F13 10
- G06F
- G06F1 00
- G06F3 06
- G06F7 00
- G06F11 00
- G06F12 00
- G06F15 16
- G06F17 00
- G11C29 00
Designated states27
- Contracting states, 27
- Austria
- Belgium
- Bulgaria
- Switzerland
- Cyprus
- Czechia
- Germany
- Denmark
- Estonia
- Spain
- Finland
- France
- United Kingdom
- Greece
- Hungary
- Ireland
- Italy
- Liechtenstein
- Luxembourg
- Monaco
- Netherlands (Kingdom of the)
- Portugal
- Romania
- Sweden
and 3 moreShow fewer
- Slovenia
- Slovakia
- Türkiye