Disaster recovery system suitable for database system
Summary by NHIP
Database disaster recovery and takeover
The method copies database update logs to a remote site and applies them continuously until a failover instruction arrives. Upon receiving that instruction, the system interrupts the log application, marks an oldest uncommitted transaction as a checkpoint, and starts a second server with a second DBMS to take over the recovered database.
Claim Score by NHIP
Abstract
To reduce operational and management costs during normal operations while recovering a database without loss and maintaining on-line performance on a site. A first system includes a primary storage system (103) that stores a DB (107) and a main computer (100) that executes a primary DBMS (101), which provides a DB. A second system includes a secondary (113) that receives from the primary storage system (103) a copy of a log, which shows update differences of the DB (107), and stores a secondary DBMS (117), and a subset (500) that recovers the secondary DB (117) according to the log that is copied from the primary storage system (103). When a failure occurs in the first system, the first system is switched to the second system. A second computer (110) that executes a second DBMS (111) is added to the second system, and the secondary DB (117) that is recovered or is being recovered in the subset (500) is taken over to the second computer (110).

Term
Term ended
Expired 12 September 2026, 0 years ago.
- Priority
- Filed
- Granted
- Expired
- Today
16 claims: 4 independent, 12 dependent
- 1A data recovery and takeover method in a disaster recovery system including a primary site having a first storage system storing a first database and a first server provided with a first database management system (DBMS) managing the first database, and a secondary site, located remotely from the primary site, having a second storage system storing a second database which duplicates the first database and a recovery module, the method comprising:copying a log indicating an update difference in said first database into a log file formed in the second storage system by a remote copy via an inter storage network which connects the first storage system and the second storage system, when the update difference is generated in said first database;when the log indicating the update difference is copied by the remote copy, reading out the copy of the log from the log file formed in the second storage system by the recovery module and applying the log on said second database by the recovery module, in a log applying process which is kept alive until the second site receives a failover instruction;upon a time when the second site receives the failover instruction, interrupting the log applying operation of the data recovery process and marking a log record indicating a start of an oldest uncommitted transaction at that time as a check point;according to the failover instruction, starting in the secondary site a second server which is provided with a second DBMS which is to take over the first DBMS;informing the second server of execution condition of the interrupted data recovery process including information of the check point from the recovery module, taking over by the second DBMS provided in the second server a data recovery of the second database.
- 3A data recovery and takeover method in a disaster recovery system including a primary site having a first storage system storing a first database and a first server provided with a first database management system (DBMS) managing the first database, and a secondary site, located remotely from the primary site, having a second storage system storing a second database which duplicates the first database and a recovery module, the method comprising:copying a log indicating an update difference in said first database into a log file formed in the second storage system by a remote copy via an inter storage network which connects the first storage system and the second storage system, when the update difference is generated in said first database;when the log indicating the update difference is copied by the remote copy, reading out the copy of the log from the log file formed in the second storage system by the recovery module, and applying the log on said second database by the recovery module, in a log applying process which is kept alive until the second site receives a failover instruction;when the second site receives the failover instruction, continuing the log applying operation until a completion up to an end of the logs copied in said log file, and undoing logs concerned in uncommitted transactions by the recovery module;according to the failover instruction, starting in the secondary site a second server which is provided with a second DBMS which is to take over the first DBMS;and after confirming a completion of the log applying operation and undoing the logs concerned in the uncommitted transactions, informing the second server of a completion of the data recovery process from the recovery module.
- 5Broadest claimClaim Score 35, narrow(NHIP)A data recovery and takeover method in a disaster recovery system including a primary site having a first storage system storing a first database and a first server provided with a first database management system (DBMS) managing the first database, and a secondary site, located remotely from the primary site, having a second storage system storing a second database which duplicates the first database and a recovery module, the method comprising:copying a log indicating an update difference in said first database into a log file formed in the second storage system by a remote copy via an inter storage network which connects the first storage system and the second storage system, when the update difference is generated in said first database;when the log indicating the update different is copied by the remote copy, reading out the copy of the log from the log file and buffering the log;analyzing buffered logs to find committed transactions;applying the logs concerned in the committed transactions on said second database by the data recovery process;when the second site receives the failover instruction, continuing the log applying operation until a completion up to an end of logs copies in said log file by said recovery module;according to the failover instruction, starting in the secondary site a second server which is provided with a second DBMS which is to take over the first DBMS;and after confirming a completion of the log applying operation, informing the second server of a completion of the data recovery process from the recovery module.
- 16A data recovery and takeover method in a disaster recovery system including a primary site having a first storage system storing a first database and a first server provided with a first database management system (DBMS) managing the first database, and a secondary site, located remotely from the primary site, having a second storage system storing a second database which duplicates the first database and normally not having an operational second server, the method comprising:copying a log indicating an update difference in said first database into a log file formed in the second storage system by a remote copy via an inter storage network which connects the first storage system and the second storage system, when the update difference is generated in said first database;using a recovery module at the secondary site, where the recovery module is a surrogate component different from, and providing fewer DBMS services than, the first and second servers, the recovery module being operational full-time to read out the copy of the log from the log file formed in the second storage system by the recovery module, and apply the log on said second database by the recovery module, in a log applying process, until the second site receives a failover instruction;when the second site receives the failover instruction, continuing the log applying operation until a completion up to an end of the logs copied in said log file, and undoing logs concerned in uncommitted transactions by the recovery module;according to the failover instruction, starting in the secondary site a second server which is provided with a second DBMS which is to take over the first DBMS;and after confirming a completion of the log applying operation and undoing the logs concerned in the uncommitted transactions, informing the second server of a completion of the data recovery process from the recovery module.
Independent claims4
156 paragraphs in 6 sections, as filed
CLAIM OF PRIORITY
The present application claims priority from Japanese application P2004-179433 filed on Jun. 17, 2004, the content of which is hereby incorporated by reference into this application.
CROSS-REFERENCE TO RELATED APPLICIONS
This application is related to following co-pending applications: <ul><li id="ul0001-0001" num="0000"><ul><li id="ul0002-0001" num="0003">U.S. application Ser. No. 10/849,006, filed May 20, 2004,</li><li id="ul0002-0002" num="0004">U.S. application Ser. No. 10/930,832, filed Sep. 1, 2004,</li><li id="ul0002-0003" num="0005">U.S. application Ser. No. 10/184,246, filed Jun. 26, 2002,</li><li id="ul0002-0004" num="0006">U.S. application Ser. No. 10/819,191, filed Apr. 7, 2004,</li><li id="ul0002-0005" num="0007">U.S. application Ser. No. 10/910,580, filed Aug. 4, 2004.</li></ul></li></ul>
BACKGROUND
This invention relates to an improvement of a disaster recovery system, and more particularly to a disaster recovery system suitably applicable for a database.
In recent years, information technology (IT) systems are so indispensable for business that continuing business despite failures or disasters becomes more and more important. The opportunity loss accompanied by stopping IT systems will be so huge that it is said to be, for example, millions of dollars in the financial sectors or major enterprises. Against such a background, attentions have been focused on a disaster recovery (hereinafter called “DR”) technology, which provides two (primary/secondary) sites to back up business data into the secondary site in normal times, and to continue the business on the secondary site in the event of a disaster.
It is the most important requirement that a database (hereinafter called “DB”) is recovered without any loss even in the event of a disaster in DR because data of enterprises are usually stored in such a DB. Because backups are executed continuously in normal times, it is also required that the influences on on-line business in the primary site should be minimized. Further, in recent years, preparing for a wide area disaster such as an earthquake, that is, a DR system locating a secondary site at a remote area hundreds to thousands of kilometers away is demanded.
In the DR system, in the event of a disaster, the DB is recovered in the secondary site, and the business is resumed by a database management system (DBMS) in the secondary site. A conventional method of updating a DB of a DBMS and recovering a DB in the event of a failure or a disaster is described.
First, the method of updating a DB in normal operations is described. The DBMS manages logs which record differences of data in addition to a DB wherein data is stored. When a data change is instructed to the DBMS, the differences created by the change are recorded on a log file wherein the log is recorded on the storage system. However, such a change is outputted to a storage system at certain times without being reflected on the DB soon, for improving performance. About the certain times, the process called a checkpoint (hereinafter called “CP”) is well known. The CP is generally issued taking the opportunity where a predetermined period of time has passed, or, a predetermined number of transactions are conducted. Here, the logs representing updated differences have serial numbers (log sequence numbers or LSN), because the logs are added each time an update is executed. At the time of CP, CP information including CP acquired time or the LSN of the log in order to indicate up to which log is applied to the DB, is recorded. A header of a log file or an exclusive file is considered as the destination wherein CP is saved. The following describes a case where an exclusive file is used.
A DB recovery process in the event of a failure is executed with the CP as a starting point. The procedure is described with reference to <figref idrefs="DRAWINGS">FIG. 16</figref>. First, CP information is read in a step <b>201</b>, and then a log reading position is decided in a step <b>202</b>. In other words, reading may be started from the log after the LSN which is recorded in the CP information (ensuring that the log has been reflected to the DB at the time of CP). Next, in a step <b>203</b>, log files are read to the end and the logs are applied to the DB in a sequence of reading. The log is recorded with an image after the update is applied and a destination where the update is applied (in some cases, operations are recorded instead of images). Applying the image to the corresponding destination allows reproducing the added updates. Such a log application process is called a redo process or a roll forward process.
Here, redo processes are executed while managing the commitment of transactions (hereinafter called “Tr”), that is, the transaction is committed or uncommitted are managed. Tr management can be executed, for example, by using the management table illustrated in <figref idrefs="DRAWINGS">FIG. 17</figref>. For example, a transaction table <b>300</b> may have items including a transaction identification (Tr-id) <b>301</b> which is uniquely assigned to each Tr, and an LSN <b>302</b> of the starting log of the Tr. When a starting log that belongs to a new Tr is read, its Tr-id and LSN are registered in the transaction management table. Otherwise, as shown in a transaction management table <b>310</b> in <figref idrefs="DRAWINGS">FIG. 7</figref>, the chain of information <b>313</b> (an LSN <b>320</b>, a log type <b>321</b>, a usage resource <b>323</b>) of the log which forms the Tr, for each Tr (a Tr-id <b>311</b>) may be managed. When the Tr is committed, the corresponding Tr is deleted from the table. Through this operation, when all logs are applied, the list of uncommitted Tr is obtained.
When logs are all read and there is no unapplied log, uncommitted transactions are deleted by using this Tr management table, that is, reading back the logs of relating Tr and canceling the transactions. This canceling process is called as an undo process or a rollback process. When the undo process is completed, the DB has consistency where only the updates included in committed Tr are reflected, and it is allowed to resume the business to add new updates.
It is possible, by resuming the business in parallel with the undo process, to reduce the time to restart. In other words, at the time the redo process is completed, a resource, which is used for an uncommitted Tr, is deduced, and the resource is locked while resuming the business.
In recent years, storage systems become more and more sophisticated, and storage systems including a remote copy function allowing transmission between the sites without a server have been developed. A DR system adopting this remote copy function is shown in <figref idrefs="DRAWINGS">FIG. 15</figref>. A primary site <b>1</b> and a secondary site <b>9</b> have, respectively, a primary server <b>100</b> and a secondary server <b>110</b>, and a primary storage system <b>103</b> and a secondary storage system <b>113</b>. The servers and storage systems are respectively connected via a server-to-server network <b>150</b> and a storage-to-storage network <b>120</b>. Further, a primary DBMS <b>101</b> and a secondary DBMS <b>111</b> are installed in the primary server <b>100</b> and the secondary server <b>110</b>, respectively.
In normal operations, business is conducted in the primary DBMS <b>101</b> of the primary site (hereinafter called “primary DBMS”). The storage systems <b>103</b> and <b>113</b> include a primary DB <b>107</b> and a secondary DB <b>117</b>, and the volumes which store a primary log <b>106</b> and a secondary log <b>116</b>, respectively. The primary log <b>106</b> and the primary DB <b>107</b> are respectively copied to the secondary log <b>116</b> and the secondary DB <b>117</b> by a remote copy <b>140</b>.
In this structure, the secondary server <b>110</b> is unnecessary in normal operations because backups in normal operations are feasible only with a storage system. Only when business is conducted at the secondary site <b>9</b> because of a disaster and the like, the secondary server <b>110</b> is necessary. In that case, the process starts the DBMS <b>111</b> on the secondary site <b>9</b>, executes the recovery process (redo/undo) as described in <figref idrefs="DRAWINGS">FIG. 16</figref>, and resumes receiving business. For copying method of the logs <b>106</b>, <b>116</b>, and the DBs <b>107</b>, <b>117</b>, the method of transmitting both the logs and the DBs by synchronous remote copy is well known. In the synchronous copy, writing to the primary site <b>1</b> is not completed until copying to the secondary site <b>9</b> is completed. Therefore, it is ensured that the update in the primary site <b>1</b> has been transmitted to the secondary site <b>9</b>, and it is possible to recover the DB without any loss even in the event of a disaster. However, there is a problem in that the on-line performance in the primary site <b>1</b> is degraded due to delays added each time the writing is executed when the secondary site is located hundreds of kilometers away from the primary site in order to prepare for a wide area disaster.
In view of the above, there is a known method in which a log including differences to update is transmitted synchronously and a database, which is recoverable from the log, need not be transmitted in order to realize recovering without loss and sustain the on-line performance simultaneously. In other words, a method of recovering a DB in a secondary site is realized by redoing a log that is copied from a primary site is known (for example, U.S. Pat. No. 5,640,561).
Likewise, a method in which log transfer function is incorporated to a DBMS is also well known (“Oracle Data Guard”, “Overview of Oracle Data Guard Functional Components”, [online], “Searched Apr. 27 2004”, <http://otn.oracle.com/deploy/availability/htdocs/DataGuardOverview.html>).
The system construction is described referring to <figref idrefs="DRAWINGS">FIG. 18</figref>. It is the same that a primary/a secondary site comprises, respectively, a server <b>100</b>/<b>110</b> and a storage system <b>103</b>/<b>113</b>, and a DBMS <b>101</b>/<b>110</b> is installed in the corresponding server <b>100</b>/<b>110</b>. However, it is unnecessary for the storage system <b>103</b>/<b>113</b> to have a remote copy function, the connection between the sites is connected only via a network <b>110</b>. Copying a log <b>106</b> is executed by the DBMS <b>101</b>/<b>111</b>, that is, the primary DBMS <b>101</b>, simultaneously, writes in the primary storage system <b>103</b> and forwards the log to the secondary DBMS <b>111</b>. The secondary DBMS <b>111</b> simultaneously writes the received log to the secondary storage system <b>113</b> and updates the DB by applying the read log to the DB <b>117</b>.
SUMMARY
However, in the conventional example descried above referring to <figref idrefs="DRAWINGS">FIG. 18</figref> (“Oracle Data Guard”, “Overview of Oracle Data Guard Functional Components”, [online], “Searched Apr. 27, 2004”, <http://otn.oracle.com/deploy/availability/htdocs/DataGuardOverview.html>), the DBMS must be operated in the normal time and there is a problem in that the business handling is subjected to pressure due to the transmitting load because transmitting is executed between the servers. Further, a computer equivalent to the primary site must be operated in the secondary site in order to continue the business in the event of a disaster. However, the actual possibility of the occurrence of a wide area disaster is only several percents of the whole stopping reason. Therefore, there is a problem in that operating the secondary server even in the normal operations increases the operating and construction cost.
In consideration of the above problems, this invention aims to reduce the operation management cost in the normal operations while recovering a DB without loss and maintaining on-line performance.
According to this invention, a first system includes a first storage system that stores a first database and a first computer that executes a first DBMS, which provides a first database. This invention also includes a second system that has a second storage system, which receives a copy of information related to the first database (for example, a log indicating updated differences) from the first storage system and stores a second database, and a database recovery module that recovers the second database, according to the information copied from the first storage system. If a failure occurs in the first system, the first system is switched to the second system, a second computer in which a second DBMS executes is added to the second system, and the second database that is recovered or being recovered by the database recovery module is taken over to the second computer (the second DBMS).
Therefore, this invention recovers a database without loss by recovering a second database (for example, a log applying) according to information related to a first database. A first system and a second system only copy between storage systems, so that on-line performance in a first system (site) is ensured. During normal operations, only a database recovery module may be executed, and a second computer that executes a second DBMS is not necessary, so that operational and management costs are reduced. In the event of failure, a second DBMS executes by adding a second computer to a second system, and a second database is taken over from a database recovery module, so that a second database service can start.
BRIEF DESCRIPTION OF THE DRAWINGS
<figref idrefs="DRAWINGS">FIG. 1</figref> shows a first embodiment of this invention and a system block diagram providing a disaster recovery at two sites.
<figref idrefs="DRAWINGS">FIG. 2</figref> is a flowchart for a process at a subset and a secondary DBMS at fail-over.
<figref idrefs="DRAWINGS">FIG. 3</figref> is a flowchart for a process of a CP information file at a subset.
<figref idrefs="DRAWINGS">FIG. 4</figref> is an explanatory view showing a case where a pseudo-CP is generated while a transaction executing status is considered and a CP information file is updated.
<figref idrefs="DRAWINGS">FIG. 5</figref> is a flowchart for a process of a CP information file at a subset.
<figref idrefs="DRAWINGS">FIG. 6</figref> is a flowchart for takeover process at a subset.
<figref idrefs="DRAWINGS">FIG. 7</figref> is an explanatory view showing a constitution of a transaction management table.
<figref idrefs="DRAWINGS">FIG. 8</figref> shows a process of takeover at a subset and is a flowchart for takeover to the secondary server via a file.
<figref idrefs="DRAWINGS">FIG. 9</figref> is a flowchart for starting a secondary DBMS at a secondary site.
<figref idrefs="DRAWINGS">FIG. 10</figref> shows a process at a subset and a secondary DBMS at fail-over, and is a flowchart for completing DB recovery at the subset.
<figref idrefs="DRAWINGS">FIG. 11</figref> is an explanatory view showing redo only committed transactions.
<figref idrefs="DRAWINGS">FIG. 12</figref> is a flowchart for recovery and takeover processes at a subset.
<figref idrefs="DRAWINGS">FIG. 13</figref> is a flowchart for another process of recovery and takeover at a subset.
<figref idrefs="DRAWINGS">FIG. 14</figref> shows a second embodiment and is a system block diagram, in which a subset is stored in a storage system, showing disaster recovery at two sites.
<figref idrefs="DRAWINGS">FIG. 15</figref> shows a prior art and a system block diagram of a disaster recovery system with a remote copy.
<figref idrefs="DRAWINGS">FIG. 16</figref> shows a prior art and a flowchart of a DB recovery process at a checkpoint as a starting point.
<figref idrefs="DRAWINGS">FIG. 17</figref> is an explanatory view of a transaction management table.
<figref idrefs="DRAWINGS">FIG. 18</figref> shows a prior art and a system block diagram in which a log transfer function is incorporated in a DBMS.
DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
Hereinafter, a first embodiment of this invention will be explained referring to the accompanying drawings as follows.
<figref idrefs="DRAWINGS">FIG. 1</figref> is a figure of typical system used in this invention, and shows an example to execute a disaster recovery between two sites.
<figref idrefs="DRAWINGS">FIG. 1</figref> shows an example that a subset <b>500</b> of a database management system (DBMS) function which is off-loaded (divided) from a secondary DBMS in a secondary site <b>2</b> operates in an exclusive unit (an intermediate unit) <b>501</b> located between a secondary server <b>110</b> and a secondary storage system <b>113</b>.
A primary site <b>1</b> includes a primary server <b>100</b> and a primary storage system <b>103</b>. The primary server <b>100</b> is installed with a primary DBMS <b>101</b> that receives business. The primary DBMS <b>101</b> manages a DB (database) <b>107</b> by using a DB buffer <b>102</b> on a memory of the primary server <b>100</b> and the primary storage system <b>103</b>. The DB <b>107</b> is located onto the storage <b>103</b>, but updates are executed on the DB buffer <b>102</b> in the normal operations because accessing the primary storage system <b>103</b> for each update occurs degrades the performance. At the predetermined timing such as when a buffer <b>102</b> overflows, updates are reflected to the storage system <b>113</b>.
The primary storage system <b>103</b> includes a cache <b>105</b>, a disk control program <b>104</b>, a control unit, and a plurality of volumes. Each volume has the DB <b>107</b> that stores a data body, a log <b>106</b> that stores update differences to the data (DB <b>107</b>), and a setting file <b>502</b> that stores a checkpoint (CP) information file and the like. A CP is the same as the conventional example as described above.
The disk control program <b>104</b> controls a remote copy between the primary storage system <b>103</b> in the primary site <b>1</b> and the secondary storage system <b>113</b> in the secondary site <b>2</b> via an inter-storage network <b>120</b>. Only the log <b>106</b> is forwarded using a synchronous or asynchronous remote copy.
The secondary site <b>2</b> has: the exclusive unit <b>501</b> including the subset <b>500</b>, which is composed of functions of a part of the secondary server <b>110</b>, the secondary storage system <b>113</b>, and a secondary DMBS <b>111</b> (for example, a log applying function); the secondary server <b>110</b> operating only when a fail-over occurs; and the secondary DBMS <b>111</b> operating in the secondary server <b>110</b>. The servers <b>110</b> and <b>111</b> in the primary site <b>1</b> and the secondary site <b>2</b>, respectively, are connected via a network <b>150</b> between the servers. This network <b>150</b> is also connected to a management unit <b>200</b> that detects failures in the primary site <b>1</b>.
The management unit <b>200</b> is composed of a computer, such as a server wherein a monitoring program and the like are executed. The management unit <b>200</b> detects failures of the primary site <b>1</b> by heart beats from the primary server <b>100</b> or such as failure information of the primary storage system <b>103</b>. If any failure is detected, the management unit <b>200</b> instructs a fail-over to the subset <b>500</b> or the secondary server <b>110</b> in the secondary site <b>2</b>.
The exclusive unit <b>501</b> may be any computer provided with the processing performance to execute the subset <b>500</b>, and is not required to have the processing performance to execute the DBMS <b>111</b>.
The secondary server <b>110</b> is not required to operate as the secondary site <b>2</b> in the normal operations, and simply operates as the secondary site <b>2</b> at the starting point of the fail-over process to start the DBMS <b>111</b>.
The secondary storage system <b>113</b> has a remote copy function. In the secondary storage system <b>113</b>, the log <b>116</b>, the DB <b>117</b>, and the setting file <b>512</b> including the CP information file, are respectively located in the separated volumes. Among them, only the log <b>116</b> is forwarded from the primary storage <b>103</b> to the secondary storage system <b>113</b> by the remote copy function. As described above, the remote copy may be executed synchronously or asynchronously. In the synchronous remote copy, writing to the primary storage system <b>103</b> is not completed until copying to the secondary storage system <b>113</b> is completed. Therefore, it is ensured to copy data without any loss even in the event of a disaster or failure. However, there are considerable influences on the business conducted in the primary DBMS <b>101</b> due to the line delays added each time the writing is executed. On the other hand, when asynchronous copy is used, writing to the primary storage system <b>103</b> is thought to be completed before the forwarding writing to the secondary storage system <b>113</b> is executed and the processing continues. Accordingly, the influence on the primary DBMS <b>101</b> is minimized. However, if a disaster occurs before the completion of the writing to the secondary storage system <b>113</b>, some losses are caused.
The primary DB <b>107</b> and the secondary DB <b>117</b> and the setting files <b>502</b>/<b>512</b> including respectively the primary/the secondary CP information are synchronized only when a disaster recovery (hereinafter, called “DR”) system starts. The first synchronization may use a remote copy, a forward function if the primary DMBS <b>101</b> and the secondary DMBS <b>111</b> have the function, and a copy program such as ftp software. Consistency must be kept between the primary DB <b>107</b> and the secondary DB <b>117</b>, and between the setting files <b>502</b> and <b>512</b>, including respectively the primary/secondary CP information. In the primary/secondary setting files <b>502</b> and <b>512</b>, a serial number of the log (log sequence number or LSN) is recorded, and this LSN should be reflected in the DBs <b>107</b>/<b>117</b> at least.
After the starting point of recovery, the secondary DB <b>117</b> will be recovered by applying a log not by a remote copy. The log application is executed by the exclusive unit <b>501</b>, which is located between the secondary storage system <b>113</b> and the secondary server <b>110</b>.
In the exclusive unit <b>501</b>, a log recovery function, one of the functions of the secondary DBMS <b>111</b>, is off-loaded as the subset <b>500</b>. The exclusive unit <b>501</b> can be realized with a low-performance, low-price computer because the exclusive unit <b>501</b> is only required to execute applying logs. Therefore, the DR system can be built at low cost because the secondary server <b>110</b> does not need to be established, so only the exclusive unit <b>501</b> is necessary in the normal operations.
Besides, cost reduction in the phase of operating and management is possible because the subset <b>500</b> has limited functions, so complicated management is not necessary compared with the DMBS <b>111</b> including full functions.
Furthermore, the exclusive unit <b>501</b> exists only in the secondary site <b>2</b> in <figref idrefs="DRAWINGS">FIG. 1</figref> described above. The exclusive unit <b>501</b> can also be located in the primary site <b>1</b>. For example, switching system process may be executed not only in the event of a disaster, but also according to an organized plan. In that case, the rolls of the primary site <b>1</b> and the secondary site <b>2</b> are replaced while doing business in the secondary site <b>1</b> and executing recovery processes in the primary site <b>2</b>. In that case, it is necessary to replace the copy source by the copy destination of the remote copy of the logs <b>106</b> and <b>107</b> and recover the DB <b>107</b> at the exclusive unit <b>501</b> in the primary site.
In the embodiment described above, the subset <b>500</b> may be the secondary DBMS <b>111</b> itself, when processing performance of the exclusive unit <b>501</b> is high enough.
<Switching from the Subset to the Secondary Server—<b>1</b>>
There are two procedures for switching in a secondary site <b>2</b>: taking over to an upper-level server (a secondary server <b>110</b>) while a recovery at a subset <b>500</b> (fail-over in the site), and taking over to an upper-level server after an entire recovery at the subset <b>500</b> is completed.
First, the procedure for taking over to an upper-level server while the recovery at the subset is described.
In the process, an upper-level server (the secondary server <b>110</b>) that has been taken over the process from the subset <b>500</b> continues a recovery process. Business will be resumed after all recovery process is completed.
<figref idrefs="DRAWINGS">FIG. 2</figref> shows a flowchart of the subset <b>500</b> and the upper-level server (the secondary server <b>110</b>), and the procedure that immediately takes over to the secondary server during the recover of the subset.
The subset <b>500</b> usually executes applying logs as described above and it is possible that the subset <b>500</b> can receive F.O. (fail-over) instructs from the management unit <b>200</b> and the like, by interrupt or other means.
First, in a step <b>700</b>, the subset <b>500</b> receives an F.O. instruction from a user (a DB administrator or a DB <b>101</b>), monitoring software or the like.
After receiving the F.O. instruction, in a step <b>711</b>, the secondary DBMS <b>111</b> starts if it is not started, and becomes stand-by before disk-access. The secondary DBMS <b>111</b> cannot be started when the F.O. is due to a disaster, but if the F.O. is due to a planned incident, the secondary DBMS can be started in advance.
On the other hand, in a step <b>702</b>, the subset <b>500</b> stops applying logs when the F.O. instruction is received. Then, in a step <b>703</b>, the execution conditions of the DB recovery process are outputted. The output may be performed to a file on the secondary storage system <b>113</b> or to a memory in the subset <b>500</b>.
Next, in a step <b>705</b>, the subset <b>500</b> discards a disk control privilege of the secondary storage system <b>113</b> that has been held by the subset <b>500</b>, and in a step <b>706</b>, notices a completion of the recovery processing at the subset <b>500</b> to the secondary server <b>110</b>.
The notice can be sent directly to the secondary DBMS <b>111</b>, or to a DB administrator (DBA) or the management unit <b>200</b>. Besides, when recovery conditions of the DB (the secondary DB <b>117</b>) in the step <b>703</b> are outputted to the memory in the subset <b>500</b>, the DB recovery conditions may be transmitted at the same time when the memory completion is notified to the secondary DBMS <b>111</b>.
The DB administrator (DBA) or the secondary DBMS <b>111</b> that receives the notice obtains the disk control privilege in a step <b>713</b>, and then reads respective files in a step <b>714</b>. The starting point of reading logs is determined by, for example, reading the setting file <b>512</b> including the CP information.
In a step <b>715</b>, DB recovery conditions are reconstructed according to information obtained by reading files or communication between processes. Sequentially, in a step <b>716</b>, the location to be resume recovery is specified according to the reconstructed recovery conditions to resume the DB recovery processes (log applications). After the log <b>116</b> is read and applied to the end, undo processes are executed on the uncommitted transaction (hereinafter called “Tr”) to complete the DB recovery process. The undo process is executed similarly to the conventional example.
After the DB process completion, in a step <b>717</b>, receiving business is resumed. Receiving business with some limitations can be conducted at the same time of resuming undo.
In this procedure, recovery time can be reduced because: the recovery processes after the stop instruction of the subset <b>500</b> can be executed in the upper-level server (the DBMS <b>111</b> in the secondary server <b>110</b>); and business with some limitations can be conducted.
As described above, in the method of taking over to an upper-level server while the recovery is committed at the subset <b>500</b>, the secondary DBMS <b>111</b> should take over the DB recovery conditions executed at the subset <b>500</b> to complete the DB recovery processes.
<Taking Over the DB Recovery>
Regarding to procedures for taking over DB recovery conditions between the subset <b>500</b> and the secondary server <b>110</b>, the following are considered: taking over with files; and taking over with communication between processes. The procedure for taking over with files is described below. With regard to files, there are methods of using CP information file (the setting file <b>512</b>) conventionally used in the DBMS and of defining new files.
Firstly, the procedure for taking over DB recovery conditions using a checkpoint (CP) information file is described. A usual CP information file is updated in the event that a predetermined time has passed or that a predetermined number of Tr has been processed. At this point, while data not yet reflected upon the DB <b>107</b> (on the storage system <b>103</b>) on the DB buffer <b>102</b> is reflected upon the DB <b>107</b>, an LSN indicating up to which part of the log <b>106</b> is reflected upon a DB <b>107</b> and time stamp information are recorded in the CP information file (the setting file <b>502</b>). Through this CP process, the CP information file becomes consistent with the DB <b>107</b>.
In this invention, in the secondary site <b>2</b>, a CP is created pretentiously, independently of the CP issued in the primary site <b>1</b> to update a CP information file.
In other words, the subset <b>500</b> notifies DB recovery conditions to the secondary DBMS <b>111</b> using the CP information file including the pseudo-CP information generated independently of the primary site <b>1</b>. The timing generating pseudo-CP is considered to be the timing when the DB administrator or the secondary DBMS <b>111</b> receives F.O. instructions or stop instructions. Otherwise, a pseudo-CP may be generated when a predetermined time has passed or a predetermined number of logs have been processed. Alternatively, if a CP generated in the primary site <b>1</b> is recorded on the log <b>116</b> as event information, a pseudo-CP may be generated when the event log is read in the subset <b>500</b>.
As procedures for generating a pseudo-CP in the subset <b>500</b> in the secondary site <b>2</b>, there are not considering Tr executing conditions and considering executing conditions.
<Generating Pseudo-Checkpoint—<b>1</b>>
First, a procedure for generating a pseudo-CP not considering Tr executing conditions and updating a CP information file is described.
If Tr executing conditions are not considered, an LSN of the log <b>116</b>, last applied in the subset <b>500</b> or time stamp information may be written in a CP information file (step <b>1001</b> in <figref idrefs="DRAWINGS">FIG. 3</figref>) at the point when an F.O. instruction or a stop instruction is received, or a CP event log is read. At the startup of the secondary server <b>110</b>, it is possible to know up to which point of the log <b>116</b> is applied to the DB <b>117</b> in the secondary site <b>2</b> by reading the file. Therefore, after the log <b>116</b> is read, DB recovery can be completed by applying logs that occur after the LSN and executing undo finally. However, when the procedure is used, resuming business with some limitations is impossible in the secondary DBMS <b>111</b>.
The secondary DBMS <b>111</b> manages resource (resource that is used for a Tr is locked until the Tr is committed) during a recovery process. When resuming process in the secondary DBMS <b>111</b>, this resource information cannot be taken over, so that complete resource information is not obtained at the end of redo. As a result, if a Tr, which is uncommitted during takeover, is still not committed at the subsequent redo by the secondary DBMS <b>111</b>, the resource that should be locked by the Tr cannot be reproduced completely. Besides, in order to undo the Tr, reading back to logs that are issued at earlier reading start point (an LSN recorded in a CP information file) where the secondary DBMS <b>111</b> starts to read is necessary.
<Generating Pseudo-Checkpoint—<b>2</b>>
Second, a procedure for generating a pseudo-CP considering Tr executing conditions and updating a CP information file is described. <figref idrefs="DRAWINGS">FIG. 4</figref> shows a schematic view of a CP considering Tr executing conditions. <figref idrefs="DRAWINGS">FIG. 4</figref> shows the conditions of applying logs in the subset <b>500</b>. The solid and the dotted lines with circles show transactions, and the solid lines show committed Trs, and the dotted lines indicate uncommitted Trs at a point <b>802</b> when a disaster occurs.
As described above, the method of recording the LSN of logs last applied at the point <b>802</b> onto a CP information file <b>804</b> is possibly used to inform the secondary DBMS <b>111</b> of the applying logs conditions. However, when the method is used, the secondary DBMS <b>111</b> should read back to a point prior to the point <b>802</b> for undo. Besides, there is a problem in that resuming the DBMS <b>111</b> with some limitations is impossible because the applying logs conditions in the subset <b>500</b> at the point <b>802</b> cannot be reproduced completely. The solutions to the problem are as follows: preventing from reading back earlier to the reading log point at the time of the secondary DBMS <b>111</b> startup; and in order to correctly generate the resource information of the finally in the secondary DBMS <b>111</b>, the start log of Tr that may be finally uncommitted at the point of receiving an F.O. or a stop instruction is regarded as a pseudo-CP.
That is, registering the LSN of a start log (<b>808</b>) of a Tr (<b>807</b>) that has been started in the earliest time among uncommitted Trs (<b>805</b>, <b>806</b>, <b>807</b>) at the point <b>802</b> when an instruction is received, to the pseudo-CP file, is acceptable. As a result, reading back earlier to the start point is not necessary because the secondary DBMS <b>111</b> reads information of the Tr that may be finally undone. Resuming with some limitations also becomes possible. However, the following processes are required to avoid applying the same log <b>116</b> more than once: the log LSN applied to data pages or the like has been recorded; and when applying logs, the log LSN to be applied and the LSN to be recorded on data pages are compared.
To realize such a pseudo-CP, it is required to manage uncommitted Tr information in the process of applying logs in the subset <b>500</b>. For example, managing a Tr table <b>803</b> is acceptable. According to the Tr table <b>803</b>, uncommitted Tr can be informed at the time of receiving an F.O. or a stop instruction.
<figref idrefs="DRAWINGS">FIG. 5</figref> shows a flowchart of generating a pseudo-CP in the subset <b>500</b> in the case of considering Tr. After an F.O. or a stop instruction is received, the Tr table <b>803</b> is scanned to sort the LSN of the start log in a step <b>1004</b>; the Tr uncommitted and started in the earliest time is found in a step <b>1005</b>; and finally, the LSN of the uncommitted Tr detected are registered onto CP information file <b>804</b> in a step <b>1006</b>. Through such processes, in the secondary DBMS <b>111</b>, it is ensured that the Tr that has a possibility of finally being uncommitted can be read from the start log. As a result, reading back to the log reading point is not required, and resuming business with some limitations at the stage where redo is completed will be possible because the resources used by uncommitted Tr can be managed completely.
<Switching from the Subset to the Secondary Server—<b>2</b>>
<figref idrefs="DRAWINGS">FIG. 6</figref> shows a processing flowchart of the subset that enables to take over the recovery process while being committed in the subset <b>500</b>, by using such a pseudo-CP described above.
In a step <b>901</b>, a CP information file is read. In a step <b>902</b>, a log reading point is decided. The CP information file may be a CP information file generated in the primary site <b>1</b> or a CP information file (including pseudo-CP) generated in the subset <b>500</b>.
The pseudo-CP can be generated when an instruction is issued from a DB administrator or a monitoring program. The pseudo-CP can be also generated when a predetermined period of time has passed or in such an event that a predetermined amount of process is committed. By generating such a pseudo-CP periodically, applying logs can be started from the nearby point where a failure occurred, even when a failure occurs on the subset <b>500</b> and the subset <b>500</b> needs to be restarted.
The following describes the example that the pseudo-CP is generated when a predetermined amount of log is applied. The variable is set as blk_num, while the threshold is set as P. The blk_num is initialized in a step <b>903</b>, and i flag described later is also initialized.
The subset <b>500</b> has an I/F for receiving instructions of such as interruption, return, F.O., or stop from the management unit <b>200</b> etc. The interruption instruction from the management unit <b>200</b> etc. means to interrupt applying logs to achieve a standby state, and the return instruction from the management unit <b>200</b> etc. means to return from the standby state to resume applying logs. As shown in <figref idrefs="DRAWINGS">FIG. 6</figref>, interruption/return is controlled with the variable i flag. In steps <b>905</b> and <b>907</b>, the control variable i flag is changed (<b>906</b>, <b>908</b>) according to the instruction received, then returning to S<b>1</b> is conducted to execute interruption or return.
In a step <b>909</b>, the existence of an F.O. instruction is judged. If an F.O. instruction exists, in a step <b>910</b>, a CP information file is generated with a pseudo-CP as described above. In a step <b>911</b>, the process is terminated. The CP information file generated enables the secondary DBMS <b>111</b> to take over the DB recovery from the subset.
If the F.O. is not instructed, i flag is checked in a step <b>912</b>. If suspension is instructed and return is not instructed, loop is formed in a step <b>913</b>.
If interrupt is not instructed, the existence of the log <b>116</b> to be applied is checked in a step <b>914</b>. If an unapplied log <b>116</b> does not exist, the process returns to a step <b>904</b> to form loop. If an unapplied log is existent, the log <b>116</b> and the transaction table <b>300</b> (refer to <figref idrefs="DRAWINGS">FIG. 17</figref>) are read in a step <b>915</b>, and the log <b>116</b> is applied in a step <b>916</b>. The log is applied while the transaction table <b>300</b> is managed.
In a step <b>917</b>, the existence of a stop instruction is judged. If stop is instructed due to maintenance or the like, a pseudo-CP is processed to terminate the subset <b>500</b> in a step <b>918</b>. If taking over to an upper-level server is performed during the recovery, undo is not executed. Therefore, the process executed by the F.O. instruction and the process executed by the stop instruction are the same.
If stop is not instructed, the blk_num is checked in a step <b>920</b>. If the log <b>116</b> whose blk_num is equal to or larger than the threshold P specified in advance has been processed, a pseudo-CP is processed in a step <b>921</b>, blk_num is initialized in a step <b>922</b>, and loop is formed in a step <b>923</b> by returning to the step <b>904</b>. The pseudo-CP is executed probably because the subset has a failure itself and restarting is required. In this case, when the subset <b>500</b> is restarted, in the step <b>901</b>, the CP information file generated by the subset <b>500</b> before the failure occurred, not the CP information file generated in the primary site <b>1</b>, is read.
If blk_num did not exceed the threshold P, blk_num is incremented in a step <b>924</b>, and the process returns to the step <b>904</b> (S<b>1</b>) in a step <b>925</b>.
Either procedure shown in <figref idrefs="DRAWINGS">FIG. 3</figref> or <figref idrefs="DRAWINGS">FIG. 5</figref> may be used as a method of processing a pseudo-CP. However, if the procedure (shown in <figref idrefs="DRAWINGS">FIG. 3</figref>) in which a Tr is not considered is used, business with some limitations cannot be resumed in the secondary DBMS <b>111</b>, so the resuming business time will be delayed. On the other hand, if a Tr is considered as shown in <figref idrefs="DRAWINGS">FIG. 5</figref>, business can be resumed with some limitations, so it is possible to resume the business immediately.
<Switching from the Subset to the Secondary Server—<b>3</b>>
As a procedure for takeover of processing to a secondary DBMS in the middle of recovery at a subset <b>500</b>, there is a method involving defining a completely new I/F (interface), not using a CP information file. For example, a status of the subset <b>500</b> at the time of receiving an F.O./a stop instruction may be taken over to an upper-level server (a secondary server <b>110</b>). The status should include an uncommitted Tr and information about the resource that the uncommitted Tr uses. For example, if uncommitted Tr information is managed in the Tr management table <b>310</b> as shown in <figref idrefs="DRAWINGS">FIG. 7</figref>, it is acceptable to take over the Tr management table <b>310</b>. A takeover method may be a file or a communication between processes. The transaction management table <b>310</b> manages a chain of log information <b>313</b> (an LSN <b>320</b>, a log type <b>321</b>, resource being used <b>323</b>) which constitutes the Tr, for each Tr-id <b>311</b> which is the identifier of Tr. If the Tr is committed, the corresponding Tr is deleted from the table. Through this operation, a list of the uncommitted Tr is obtained when all logs are applied.
<figref idrefs="DRAWINGS">FIG. 8</figref> shows a flowchart of the process executed in the subset <b>500</b> for taking over the subset <b>500</b> status with a file. The basic flow is the same as that when a pseudo-CP is used (shown in <figref idrefs="DRAWINGS">FIG. 6</figref>). However, if it is judged that an F.O. is instructed in the step <b>909</b>, the contents of the Tr table are outputted to a file in a step <b>1500</b>. In the case of a stop instruction, takeover to an upper-level server is not executed. Because the major purpose is to restart the subset, a pseudo-CP is generated.
If a new I/F is defined, the startup flow of the secondary DBMS <b>111</b> should be modified. The startup flow is shown in <figref idrefs="DRAWINGS">FIG. 9</figref>. In the steps <b>201</b> and <b>202</b>, a log reading position is determined according to a CP information file in the setting file <b>512</b> in the secondary storage system <b>113</b>. In a step <b>210</b>, if a file (including information on the Tr table) outputted by the subset <b>500</b> exists, the file is read. In a step <b>211</b>, information on the uncommitted Tr is reconstructed according to the file. That is, the resource that is used for the uncommitted Tr is clarified and locked. In the step <b>203</b> or later, applying log is started. Logs are applied while the Tr table is managed or exclusive information on an uncommitted Tr is managed. When logs are applied up to the end of the log <b>116</b>, business is resumed with some limitations based on exclusive-information that has been constructed so far. That is, accessing the uncommitted Tr is prohibited to process other transactions. In a step <b>205</b> or later, undo is processed in parallel with resuming business. When all undo processes are completed, the limitations are removed to resume business completely.
<Switching from the Subset to the Secondary Server—<b>4</b> After Complete Recovery>
As a procedure for switching from a subset <b>500</b> to a secondary server <b>110</b> in the secondary site <b>2</b>, takeover to an upper-level server in the middle of the recovery at the subset <b>500</b> is described above. In addition, there is another method involving taking over a process to an upper-level server (the secondary server <b>110</b>), after entire recovery is completed in the subset <b>500</b>. The following describes the case of taking over a process to an upper-level server after entire recovery in the subset <b>500</b> is completed.
<figref idrefs="DRAWINGS">FIG. 10</figref> shows a flow of takeover between the subset <b>500</b> and the secondary server <b>110</b>.
First, in the step <b>700</b>, an F.O. instruction is decided in the management unit <b>200</b> or the like. This decision can be made by a DB administrator or automatically made by a monitoring program. In a step <b>701</b>, the subset <b>500</b> receives the F.O. instruction. Receiving can be made via a communication between processes or a file.
In a step <b>1100</b>, Trs with consistency in a DB <b>117</b> and all recovery process completions are confirmed. In a step <b>1101</b>, a CP information file in the setting file <b>512</b> is initialized. This is because all recovery is completed and the upper-level server does not have to continue the recovery process.
In the step <b>705</b>, a disk control privilege of the subset <b>500</b> is abandoned. In the step <b>706</b>, completion of the recovery is notified to the upper-level server (the secondary server <b>110</b>).
In the secondary server <b>110</b>, an F.O. is decided, and if the DBMS <b>111</b> has not started yet, the DBMS <b>111</b> first starts in the step <b>711</b>. By this startup, startup is made up to the point prior to the disk accessing. In a step <b>712</b>, the DBMS waits for the recovery completion notice from the subset. A method of this notice may be a file connection, a communication between processes, or an instruction by a DB administrator. After the recovery completion is confirmed, in the step <b>713</b>, a disk control privilege is obtained. In the step <b>714</b>, the secondary storage system <b>113</b> is mounted on the secondary server <b>110</b> and each file is read. In the step <b>716</b>, business starts to be received.
In order to complete the DB recovery in the subset <b>500</b> described above, a following method is considered: after redo is completed up to the log end as in the normal DB server, undo is executed. Alternatively, logs are applied while Trs are considered in the process of applying logs in the subset <b>500</b>. That is, committed/uncommitted Trs are managed and only committed Trs are redo, so that undo is not necessary.
<The DB Recovery Process at the Subset <b>1</b>>
<figref idrefs="DRAWINGS">FIG. 11</figref> is the schematic drawing in the case where only committed Trs are redo. <figref idrefs="DRAWINGS">FIG. 11</figref> shows a process of applying logs at the subset <b>500</b> in order of occurrence. The lines with circles indicate Trs, each of which is concurrently processed. The subset <b>500</b> buffers a log <b>116</b>, analyses the buffered log <b>116</b>, and only committed Trs are redo.
<figref idrefs="DRAWINGS">FIG. 11</figref> shows the state where logs <b>116</b> from the time <b>1200</b> to the time <b>1201</b> are buffered. The buffered logs <b>116</b> are analyzed per Tr. Only Trs <b>1202</b>, <b>1203</b>, and <b>1204</b> committed during this period are redo. Trs committed at the time <b>1201</b> or later are redo by a subsequent buffering and a subsequent analysis.
The following describes a process flow in the subset <b>500</b> when DB recovery is completed in the subset <b>500</b>.
Firstly, <figref idrefs="DRAWINGS">FIG. 12</figref> shows a process flow showing up to the end of a log <b>116</b> is redo as usual, and then undo.
In the steps <b>901</b> and <b>902</b>, a log reading point is decided by accessing a CP information file. In the step <b>903</b>, a variable for processing a pseudo-CP etc. is initialized. The following shows an example that a pseudo-CP is issued when a certain amount of log is processed. The pseudo-CP may be issued at another timing, for example, when a certain period of time has passed.
In the steps <b>905</b> and <b>907</b>, the existence of a suspension/return instruction is checked. If the suspension instruction exists, a value is set to i flag. While the i flag is ON (i flag=1), loop is formed as in <figref idrefs="DRAWINGS">FIG. 6</figref> and a log applying process is suspended. If return is instructed, the i flag is set to 0, so the loop ends and a subsequent process is executed.
If the state is not a suspended state, the existence of an unapplied log <b>116</b> is checked in a step <b>1300</b>. If an unapplied log exits, the log is read in a step <b>1301</b> and a DB is updated while the log is applied in a step <b>1302</b>. A Tr table is updated at the same time. In a step <b>1303</b>, the existence of a stop instruction is checked. If stop is instructed, a pseudo-CP is processed in a step <b>1304</b>. In a step <b>1305</b>, the process at the subset is terminated. This stop is made when a subset temporarily needs to be shutdown because of an operational reason such as replacing hardware in which the subset operates.
If stop is not instructed, it is checked whether a certain amount of log is applied or not in a step <b>1306</b>. If a certain amount of log is applied, a pseudo-CP is processed in a step <b>1307</b>. In a step <b>1308</b>, a variable is initialized. If a certain amount of log is not applied, the variable is incremented in a step <b>1310</b>, and the process returns to the step <b>904</b>.
If an unapplied log does not exist and an F.O. instruction is confirmed in a step <b>1312</b>, the completion of applying log up to the end of the log <b>116</b> is judged in a step <b>1313</b>. If the application is completed, uncommitted Trs are undo in a step <b>1314</b>. After the completion of the undo process, a CP information file is initialized. If a log is not applied to the end of the log <b>116</b>, the process returns to the step <b>1301</b> and all logs are applied.
<The DB Recovery Process at the Subset <b>2</b>>
As a procedure for completing DB recovery at a subset, <figref idrefs="DRAWINGS">FIG. 13</figref> shows a process flowchart in which a committed/uncommitted state of a Tr is managed in the process of applying logs and only committed Trs are redo. The basic flow is the same as that shown in <figref idrefs="DRAWINGS">FIG. 12</figref>. The difference from <figref idrefs="DRAWINGS">FIG. 12</figref> is the process where an unapplied log exists: in the step <b>1301</b>, a log is read (and buffered); in a step <b>1400</b>, a Tr is analyzed and a committed Tr is judged; in a step <b>1401</b>, only a committed Tr is redo.
A process at the time when an F.O. instruction is received also differs: in the step <b>1312</b>, an F.O. instruction is confirmed; in the step <b>1313</b>, whether up to the end of a log <b>116</b> is read or not is judged; in a step <b>1402</b>, a CP information file is initialized to terminate the process. Regarding the log <b>116</b> that has been already read, the logs relating to the committed Tr at that time are applied and the logs relating to the uncommitted Tr are not applied. Therefore, it is possible to immediately terminate the process without undo. Because a DB recovery process is not required in the secondary DBMS <b>111</b>, the CP information file may be left initialized.
<Conclusion>
As described above, the two sites <b>1</b> and <b>2</b> (primary and secondary) are provided. In normal operations, through the disaster recovery system which enables to copy data in the primary site <b>1</b> into the secondary site <b>2</b>, only the log <b>106</b> in the primary site <b>1</b> is copied into the secondary storage system <b>113</b> in the secondary site <b>2</b> using the remote copy function in the storage system <b>103</b>. The secondary site <b>2</b> is provided with the exclusive unit <b>501</b> which controls the secondary storage system <b>113</b>. On the exclusive unit <b>501</b>, the subset <b>500</b> which has only a part of the DBMS, for example, a log application function, is operated. The subset <b>500</b> is recovered with the log <b>116</b> stored in the secondary storage system <b>113</b> in the secondary site <b>2</b> that remote-copied. The DB <b>117</b> in the secondary site is recovered with the log <b>116</b> that copied.
Therefore, the secondary server <b>110</b> is not required in normal operations and only the exclusive unit <b>501</b> including the resource which enables to apply logs must be operated. For that reason, the capacity to process transactions such as the secondary server <b>110</b> is not required and it would be possible to reduce the operational costs in normal operations by constituting the exclusive unit <b>501</b> with a small computer.
When fail-over occurs, the service can be resumed by: adding secondary server <b>110</b> into the secondary site <b>2</b>; and taking over the completely recovered DB <b>117</b> or the recovery process of the DB <b>117</b> from the subset <b>500</b> to the started DBMS <b>111</b> in the secondary server <b>110</b>.
When a disaster occurs, in order to execute in the secondary site <b>2</b> the same process as the process on the business conducted in the primary site <b>1</b> in normal operations, the business is conducted in the secondary DBMS <b>111</b>, which is an upper-level unit, on the secondary server <b>110</b>. Therefore, the conventional switch <b>130</b> between sites and the switch <b>505</b> between the systems in the secondary site are required.
To realize the switch <b>130</b> between sites, the subset <b>500</b> is provided with a switch mechanism <b>505</b> in the site.
Switching from the subset <b>500</b> in the secondary site <b>2</b> to the secondary server <b>110</b> is summarized as follows. <ul><li id="ul0003-0001" num="0129">1. Taking over to an upper-level server (a secondary server <b>110</b>) during recovery at a subset <b>500</b> (fail-over in the site)</li></ul>
1-1 Takeover through a pseudo-CP information file not considering a Tr executing status
1-2 Generating a pseudo-CP while considering a Tr executing status to update a CP information file
1-3 Takeover through a Tr management table <b>310</b><ul><li id="ul0004-0001" num="0133">2. Taking over to an upper-level server after completing entire recovery of the DB <b>117</b> at a subset <b>500</b></li></ul>
2-1 Performing undo after up to the end of a log is redo
2-2 Applying logs while considering Trs and only committed Trs are redo
As described in 1 above, the process is taken over during recovery at the subset <b>500</b> to the secondary server <b>110</b> through the pseudo-CP or the transaction management table <b>310</b>. That is, the process is immediately taken over to the secondary server <b>110</b>, which has higher performance than the exclusive unit <b>501</b> in which the subset <b>500</b> is operated. Therefore, the overhead before resuming the business can be reduced because the recovery of the DB is executed quickly.
As described in 2 above, if logs are applied to the end of the log <b>116</b> when fail-over is instructed, the process can be immediately taken over to the secondary server <b>110</b>.
As described above, protecting the data and maintaining the on-line performance can be both managed even in the event of a wide area disaster because logs are only targeted as the synchronous. Besides, by off-loading the log application unit to the intermediate exclusive unit <b>501</b>, server less is realized in normal operations. Therefore, the DR system suitably corresponding to a wide area disaster can be constructed with lower cost.
A switching mechanism between the subset <b>500</b> and the secondary DBMS can provide reasonable constitution during normal operations. In the event of a disaster or during a planned system shutdown, switching control from the subset to the secondary server can make the secondary site <b>2</b> process the same business amount as that during normal operations.
In the secondary site <b>2</b>, it is possible to take over the process to the secondary server <b>110</b> while the recovery at the subset <b>500</b> is being processed by managing CP information independently of the primary site <b>1</b>. As a result, the rest of the recovery can be executed in the secondary server <b>110</b> which has rich resources. Besides, because the recovery with some limitations is possible, the time before resuming the business can be reduced.
Furthermore, because the Tr management is executed in the subset <b>500</b>, and the recovery is executed so that there is consistency in Trs, undo processing is not necessary at takeover. Since an upper-level server does not have to execute any recovery process, the operation in the upper-level server can be done easily.
Second Embodiment
<figref idrefs="DRAWINGS">FIG. 14</figref> shows a second embodiment in which the subset of the first embodiment is located in a secondary storage system <b>113</b>, not in the intermediate exclusive unit.
A subset <b>600</b> in the figure has also only a part of function (log recovery function) of the DBMS <b>111</b> in the secondary site <b>2</b> as in the first embodiment. The secondary storage system <b>113</b> has a CPU and memory (not shown), and the subset <b>600</b> is operated using such resources.
By locating the subset <b>600</b> in the secondary storage system <b>113</b>, the high speed network (a network between storages <b>120</b>) is available, so speeding up the log application process is possible.
If an exclusive unit <b>501</b> is used to operate the subset as in the first embodiment, it is required to connect between the secondary storage system <b>113</b> and the exclusive unit <b>501</b> via an FC switch or the like. Meanwhile, if the subset <b>600</b> is located in the secondary storage system <b>113</b>, such connection is not required and the constitution and setting up of the system can be simplified. Besides, recovery in which the DB <b>117</b> level is considered can be performed in a storage system, a DB recovery process is not necessary from a user standpoint. Therefore, the operation can be realized easily moreover.
In this case also, because the secondary server <b>110</b> is not necessary in normal operations, low-cost operation is possible. However, in the event of a disaster, continuing business at the subset in the storage system is difficult in order to execute the same amount of process as that before the disaster because the business should be continued in the secondary site <b>2</b>. Besides, if the subset <b>600</b> is limited to application of logs, the business cannot be received. Therefore, in the event of a disaster, it is required to startup the secondary server <b>110</b> equal to the primary site <b>1</b> to continue the business in the secondary server <b>110</b>. For that reason, not only the conventional system switch <b>130</b> between the primary site and the secondary site but also a switch in the secondary site <b>603</b> between the subset <b>600</b> and the secondary server <b>110</b> (the subset is off-loaded to the storage system) are required.
The switch in the site <b>603</b> is similar to that in the first embodiment. As in the 1-1 and 1-2 in 1. described above, takeover to an upper-level server (secondary server <b>110</b>) in the middle of recovery at the subset <b>600</b> (fail-over in the site) or after entire recovery in the DB <b>117</b> is completed as in 2. described above is possible.
In this embodiment, the exclusive unit <b>501</b> in the first embodiment is not necessary. In normal operations, only the secondary storage system <b>113</b> must be operated in the secondary site <b>2</b>. Therefore, the operation cost for the secondary site <b>2</b> can be reduced further.
A Network Attached Storage (NAS) or the like can be adopted as the secondary storage system <b>113</b>.
In this invention according to claim <b>1</b> above, the first system and the second system are connected via the storage network between the first storage system and the second storage system described above, and the server network between the first computer and the intermediate unit or the second computer described above.
As described above, this invention can reduce the operational costs in a secondary site and apply the technology to a disaster recovery system adopted in the financial sectors, major enterprises, and the like.
While the present invention has been described in detail and pictorially in the accompanying drawings, the present invention is not limited to such detail but covers various obvious modifications and equivalent arrangements, which fall within the purview of the appended claims.
Contents6
17 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17
Every citation, both waysCites: the store holds 5 of 6
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US2016110378A1 | Cited by | United States of America | Pre-grant |
| US9720752B2 | Cited by | United States of America | Search report |
| US10379919B2 | Cited by | United States of America | Applicant |
| US11449373B2 | Cited by | United States of America | Applicant |
| US10255138B2 | Cited by | United States of America | Applicant |
| US11928005B2 | Cited by | United States of America | Applicant |
| US2001344141A1 | Cites | United States of America | Applicant |
| JP2001344141A | Cites | Japan | Applicant |
| US5640561A | Cites | United States of America | Applicant |
| US6226651B1 | Cites | United States of America | Search report |
| JPH0962555A | Cites | Japan | Applicant |
| Oracle Technology Network, "Oracle Database 10g Oracle Data Guard," Web:http://otn.oracle.com/deploy/availability/htdocs/DataGuardOverview.html; Apr. 6, 2004, pp. 1-7. | Non-patent | – | Applicant |
4 members in 2 offices
Priority claims4
| Document | Office | Kind | Date |
|---|---|---|---|
| 2004179433 | Japan | A | |
| 2004179433 | Japan | A | |
| 2004179433 | – | – | – |
| JP20040179433 | – | – | – |
Members4
| Document | Office | Kind | |
|---|---|---|---|
| US2005283504A1 | United States of America | A1 | |
| JP2006004147A | Japan | A | |
| JP4581500B2 | Japan | B2 | |
| US7925633B2This record | United States of America | B2 |
93 transactions on the USPTO file
Allowed after 3 non-final rejections, 2 final rejections and 2 RCEs.
- Non-final rejections
- 3
- Final rejections
- 2
- RCEs
- 2
- Appeals
- 0
Over time
Point at a mark for the transactionTransactions
| Event | Code | |
|---|---|---|
| Expire PatentEXP. | EXP. | |
| Maintenance Fee Reminder MailedREM. | REM. | |
| Payment of Maintenance Fee, 8th Year, Large EntityM1552 | M1552 | |
| Recordation of Patent Grant MailedPGM/ | PGM/ | |
| Patent Issue Date Used in PTA CalculationAllowedPTAC | PTAC | |
| Issue Notification MailedAllowedWPIR | WPIR | |
| Dispatch to FDCD1935 | D1935 | |
| Application Is Considered Ready for IssuePILS | PILS | |
| Issue Fee Payment VerifiedN084 | N084 | |
| Issue Fee Payment ReceivedIFEE | IFEE | |
| Mail PUB Notice of non-compliant IDSMM327-B | MM327-B | |
| PUB Notice of non-compliant IDSM327-B | M327-B | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Mail Notice of AllowanceAllowedMN/=. | MN/=. | |
| Notice of Allowance Data Verification CompletedAllowedN/=. | N/=. | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Request for Continued Examination (RCE)RCEX | RCEX | |
| Disposal for a RCE / CPA / R129AbandonedABN9 | ABN9 | |
| Workflow - Request for RCE - BeginBRCE | BRCE | |
| Mail Final Rejection (PTOL - 326)Final rejectionMCTFR | MCTFR | |
| Final RejectionFinal rejectionCTFR | CTFR | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Response after Non-Final ActionA... | A... | |
| Request for Extension of Time - GrantedXT/G | XT/G | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Email NotificationEML_NTF | EML_NTF | |
| Mail Non-Final RejectionNon-final rejectionMCTNF | MCTNF | |
| Non-Final RejectionNon-final rejectionCTNF | CTNF | |
| Mail Notice of Withdrawn ActionMW/AC | MW/AC | |
| Mail Examiner Interview Summary (PTOL - 413)MEXIN | MEXIN | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Withdrawing/Vacating Office Action LetterW/AC | W/AC | |
| Mail Notice of Restarted Response PeriodMNRES | MNRES | |
| Letter Restarting Period for Response (i.e. Letter re References)NRES | NRES | |
| Preliminary AmendmentA.PE | A.PE | |
| Examiner Interview Summary Record (PTOL - 413)EXIN | EXIN | |
| Miscellaneous Incoming LetterLET. | LET. | |
| Mail Non-Final RejectionNon-final rejectionMCTNF | MCTNF | |
| Non-Final RejectionNon-final rejectionCTNF | CTNF | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Disposal for a RCE / CPA / R129AbandonedABN9 | ABN9 | |
| Request for Continued Examination (RCE)RCEX | RCEX | |
| Request for Extension of Time - GrantedXT/G | XT/G | |
| Workflow - Request for RCE - BeginBRCE | BRCE | |
| Mail Final Rejection (PTOL - 326)Final rejectionMCTFR | MCTFR | |
| Final RejectionFinal rejectionCTFR | CTFR | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Response after Non-Final ActionA... | A... | |
| Mail Notice of Informal or Non-Responsive AmendmentNINA | NINA | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Informal or Non-Responsive Amendment after Examiner ActionA.I. | A.I. | |
| Response after Non-Final ActionA... | A... | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Email NotificationEML_NTF | EML_NTF | |
| Mail Non-Final RejectionNon-final rejectionMCTNF | MCTNF | |
| Non-Final RejectionNon-final rejectionCTNF | CTNF | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| IFW TSS Processing by Tech Center CompleteTSSCOMP | TSSCOMP | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Transfer Inquiry to GAUTI1050 | TI1050 | |
| Application Return from OIPEWROIPE | WROIPE | |
| Application Return TO OIPEROIPE | ROIPE | |
| Application Is Now CompleteCOMP | COMP | |
| Application Is Now CompleteCOMP | COMP | |
| Application Return from OIPEWROIPE | WROIPE | |
| Application Return TO OIPEROIPE | ROIPE | |
| Application Return from OIPEWROIPE | WROIPE | |
| Application Return TO OIPEROIPE | ROIPE | |
| Application Dispatched from OIPEOIPE | OIPE | |
| Application Is Now CompleteCOMP | COMP | |
| Request for Foreign Priority (Priority Papers May Be Included)RQPR | RQPR | |
| Additional Application Filing FeesADDFLFEE | ADDFLFEE | |
| A statement by one or more inventors satisfying the requirement under 35 USC 115, Oath of the ApplicOATHDECL | OATHDECL | |
| Notice Mailed--Application Incomplete--Filing Date AssignedINCD | INCD | |
| Cleared by OIPE CSRL194 | L194 | |
| IFW Scan & PACR Auto Security ReviewSCAN | SCAN | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Reference capture on IDSRCAP | RCAP | |
| Request for Foreign Priority (Priority Papers May Be Included)RQPR | RQPR | |
| Information Disclosure Statement (IDS) FiledM844 | M844 | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Initial Exam Team nnIEXX | IEXX |
10 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Lapsed due to failure to pay maintenance feeLapsedFP | FP | |
| Lapse for failure to pay maintenance feesLapsedPATENT EXPIRED FOR FAILURE TO PAY MAINTENANCE FEES (ORIGINAL EVENT CODE: EXP.); ENTITY STATUS OF PATENT OWNER: LARGE ENTITYLAPS | LAPS | |
| Information on status: patent discontinuationPATENT EXPIRED DUE TO NONPAYMENT OF MAINTENANCE FEES UNDER 37 CFR 1.362STCH | STCH | |
| Fee payment procedureMAINTENANCE FEE REMINDER MAILED (ORIGINAL EVENT CODE: REM.); ENTITY STATUS OF PATENT OWNER: LARGE ENTITYFEPP | FEPP | |
| Maintenance fee paymentMAFP | MAFP | |
| Fee paymentFPAY | FPAY | |
| Information on status: patent grantGrantedPATENTED CASESTCF | STCF | |
| Fee payment procedurePAYOR NUMBER ASSIGNED (ORIGINAL EVENT CODE: ASPN); ENTITY STATUS OF PATENT OWNER: LARGE ENTITYFEPP | FEPP | |
| AssignmentAS | AS | |
| AssignmentAS | AS |
Numbers
- Publication
- 07925633
- Publication, DOCDB
- 7925633
- Publication, EPODOC
- US7925633
- Application
- 10989398
- Application, DOCDB
- 98939804
- Application, EPODOC
- US20040989398
Titles
- English
- Disaster recovery system suitable for database system
Patent term adjustment
- A delay
- +566 daysthe office missed an examination deadline
- B delay
- +390 dayspendency past three years
- Applicant delay
- −292 days
- Net adjustment
- 664 days
Classification
- CPC, 3
- G06F11/2028
- G06F11/2025
- G06F11/2097
- IPC, 2
- G06F7 00
- G06F12 00
- USPC, 1
- 707674000