Storage system, method for calculating estimated value of data recovery available time and management computer
Summary by NHIP
Asynchronous Remote Copy Monitoring
The storage system monitors data recovery available time during Asynchronous Remote Copy operations between two sites. A management computer calculates this estimate using sequence numbers and temporal information stored at specific intervals, then displays the result based on buffer quantities or earliest data sequence numbers.
Claim Score by NHIP
Abstract
The invention provides a technology applicable to technologies other than a main frame technology and monitors data recovery available time while suppressing a monitoring error within a certain range in a storage system that performs Asynchronous Remote Copy among storage devices. A management computer in the storage system stores latest or quasi-latest management data corresponding to data staying in a buffer of the first storage device with temporal information at certain monitoring intervals, calculates an estimated value of the data recovery available time which is time of data stored in the second storage device corresponding to data stored in the first storage device, based on the temporal information stored, and based on a certain management data among earliest or quasi-earliest management data or a number of the data staying in the buffer at the certain time and displays the estimated value on a display section.

Term
Projected expiry 29 January 2030.
- Priority
- Filed
- Granted
- Today
- Projected expiry
12 claims: 3 independent, 9 dependent
- 1A storage system, comprising:a first site having a first storage device and a first host computer for reading/writing data from/into the first storage device;a second site having a second storage device and a second host computer for reading/writing data from/into the second storage device;and a management computer for managing the first host computer of the first site, wherein the first storage device performs Asynchronous Remote Copy from the first storage device of the first site to the second storage device of the second site, wherein the management computer stores a sequence number as identification information of a current latest or quasi-latest data stored in a buffer of the first storage device of the first site with temporal information of the sequence number at each of certain monitoring intervals, wherein the management computer calculates an estimated value of data recovery available time, which is a time at which data most recently stored in the second storage device of the second site that corresponds to data stored in the first storage device was stored in the first storage device, based on the temporal information stored by the management computer and based on a certain sequence number of a current earliest or quasi-earliest data or a quantity of the data stored in the buffer at a certain time, and wherein the management computer displays the estimated value on a display section.
- 5A method for calculating an estimated value of data recovery available time of a storage system that performs Asynchronous Remote Copy, the storage system comprising a first site having a first storage device and a first host computer for reading/writing data from/into the first storage device, a second site having a second storage device and a second host computer for reading/writing data from/into the second storage device, and a management computer for managing the host computer of the first site, the method comprising:performing the Asynchronous Remote Copy from the first storage device of the first site to the second storage device of the second site;storing a sequence number as identification information of a current latest or quasi-latest stored in a buffer of the first storage device of the first site with temporal information of the sequence number at each of certain monitoring intervals, calculating an estimated value of data recovery available time, which is a time at which data most recently stored in the second storage device of the second site that corresponds to data stored in the first storage device was stored in the first storage device, based on the temporal information stored by the management computer and based on a certain sequence number of a current earliest or quasi-earliest data or a quantity of the data stored in the buffer at a certain time;and displaying the estimated value on a display section.
- 9Broadest claimClaim Score 46, average(NHIP)A management computer managing a first storage device and a second storage device being a copy destination of Asynchronous Remote Copy from the first storage device, the management computer comprising:a memory storing a sequence number as identification information of a current latest or quasi-latest data stored in a buffer of the first storage device of the first site with temporal information of the sequence number at each of certain monitoring intervals, a processor calculating an estimated value of data recovery available time, which is a time at which data most recently stored in the second storage device of the second site that corresponds to data stored in the first storage device was stored in the first storage device, based on the temporal information stored by the management computer and based on a certain sequence number of a current earliest or quasi-earliest data or a quantity of the data stored in the buffer at a certain time;and a display section providing the estimated value.
Independent claims3
191 paragraphs in 6 sections, as filed
CROSS REFERENCE TO RELATED APPLICATIONS
This application claims the foreign priority benefit under Title 35, United States Code, §119 (a)-(d) of Japanese Patent Application No. 2008-321323, filed on Dec. 17, 2008 in the Japan Patent Office, the disclosure of which is herein incorporated by reference in its entirety.
BACKGROUND OF THE INVENTION
1. Field of the Invention
The present invention relates to a technology for monitoring a data recovery available time in a storage system that performs Asynchronous Remote Copy among a plurality of storage devices.
2. Related Art
An importance of nondisruptive business operation and data protection in a business information system is increasing more and more lately due to globalization of a market and due to provision of services 24 hours a day and 365 days a year through Web. However, there exists a lot of risks such as terrorism and natural disasters that possibly lead to disruption and data loss of the business information system.
One of technologies that relieves these risks is Asynchronous Remote Copy in a storage system. The Remote Copy here is a technology of duplicating data by copying data in a certain volume in a storage device to a volume of another storage device.
There are two main types of Remote Copy. One is Synchronous Remote Copy (RCS: Remote Copy Synchronous) of copying the data in real-time as extension of writing in a host and of the other is Asynchronous Remote Copy of copying the data through processing different from the writing in the host. In the case of the Asynchronous Remote Copy, there are also two methods of storing changed data history in a memory (RCA: Remote Copy Asynchronous) and of storing in the memory in combination with a volume (RCD: Remote Copy asynchronous with Disk). The present invention intends the Asynchronous Remote Copy (the both RCA and RCD).
Note that it is said to be preferable to adopt the Asynchronous Remote Copy if a distance between a primary site (a site normally performing business operations) and a remote site (a different site for supporting the business continuity) is distant, e.g., 100 km or more, even though the Synchronous Remote Copy may be adopted naturally if the distance is close, e.g., less than 100 km, in general for convenience of communication performance and the like. That is, assuming a natural disaster such as a huge earthquake beside a terrorism, it is necessary to largely keep the distance between the primary and remote sites and in that case, it is desirable to adopt the Asynchronous Remote Copy.
That is, the use of the Asynchronous Remote Copy allows the remote site to support the business continuity when the primary site is hit by the terrorism, natural disaster or the like even if the distance between the primary and remote sites is distant. To that end, storage devices are provided on the both primary and remote sites in the Asynchronous Remote Copy to deal with the terrorism, natural disaster or the like that is unpredictable when it occurs. Upon that, the Asynchronous Remote Copy realizes the business continuity while minimizing data loss when the primary site is damaged by conforming contents of data stored in the primary and the remote sites as much as possible.
For instance, there has been a practical technology of duplicating contents of a physical volume or logical volume by using the Asynchronous Remote Copy between the storage devices provided in the primary and remote sites and of continuing business operations by using the storage device of the remote site when the primary site is damaged.
When the Asynchronous Remote Copy is carried out, there is a time lag in writing data respectively into the storage devices of the primary and remote sites. Therefore, it is necessary to monitor that data stored in the storage device of the remote site corresponds to data of which point of time in the primary site to be ready for a case when a failure occurs in the primary site. Here, time when the latest data stored in the storage device of the remote site had been written into the storage device of the primary site will be called as data recovery available time with reference to certain time. The lag of time of the data stored in the storage device of the primary site and the storage device of the remote site will be called as a data loss time period.
For instance, assuming that write data issued at 21:00:00 is written into the storage device of the primary site, and assuming that the write data issued at 20:59:20 is written into the storage device of the remote site and no write data is written after that, the data recovery available time is 20:59:20 and the data loss time period is 40 seconds.
US2002/0103980A discloses a technology of monitoring the data recovery available time by utilizing a time stamp (temporal information) embedded by a host into a write I/O (Input/Output) (write request) in a main frame environment.
US2006/0277384A discloses a technology of monitoring the data recovery available time in a SAN (Storage Area Network) environment.
However, US2002/0103980A presupposes a main frame technology in which the temporal information may be given to the write I/O. However, no temporal information is given to the write I/O in Fibre Channel Protocol used in the SAN environment, so that it is difficult to apply the technology of US2002/0103980A to environments other than the main frame.
US2006/0277384A discloses a technology of indicating the data loss time period that is a time lag in writing data respectively into the storage devices of the primary and remote sites. However, US2006/0277384A discloses no technology of indicating the data recovery available time nor discloses accuracy in monitoring the data loss time period.
In view of the problems described above, the invention seeks to provide a technology that is applicable to technologies other than the main frame technology and that monitors the data recovery available time while suppressing a monitoring error within a certain range in a storage system that performs the Asynchronous Remote Copy among a plurality of storage devices.
SUMMARY OF THE INVENTION
A storage system of the invention includes a first site having a first storage device and a first host computer for reading/writing data from/into the first storage device, a second site having a second storage device and a second host computer for reading/writing data from/into the second storage device and a management computer for managing the first host computer of the first site.
The first storage device performs Asynchronous Remote Copy from the first storage device of the first site to the second storage device of the second site, and the management computer stores latest or quasi-latest management data corresponding to data staying in a buffer of the first storage device of the first site with temporal information at certain monitoring intervals, calculates an estimated value of data recovery available time which is time of data stored in the second storage device of the second site, corresponding to data stored in the first storage device at the time, based on the temporal information stored, and based on a certain management data among earliest or quasi-earliest management data or a number of the data staying in the buffer at a certain time, and displays the estimated value on a display section.
The other means will be described later.
The invention provides a technology that is applicable to technologies other than a main frame technology and that monitors the data recovery available time while suppressing a monitoring error within a certain range in a storage system that performs the Asynchronous Remote Copy among a plurality of storage devices.
BRIEF DESCRIPTION OF DRAWINGS
<figref idrefs="DRAWINGS">FIG. 1</figref> is a block diagram showing an outline of the invention;
<figref idrefs="DRAWINGS">FIG. 2</figref> is a block diagram showing a configuration of a storage system of an embodiment;
<figref idrefs="DRAWINGS">FIG. 3</figref> is a block diagram showing a configuration of a management computer;
<figref idrefs="DRAWINGS">FIG. 4</figref> is a block diagram showing a configuration of a host computer;
<figref idrefs="DRAWINGS">FIG. 5</figref> is a block diagram showing a configuration of a storage device;
<figref idrefs="DRAWINGS">FIG. 6</figref> is one example of a remote copy definition table created in a memory on a disk controller within the storage device;
<figref idrefs="DRAWINGS">FIG. 7</figref> is a state transition diagram of a copy pair of Asynchronous Remote Copy;
<figref idrefs="DRAWINGS">FIG. 8</figref> shows one example of a flow-in section I/O table created by a data recovery available time monitoring program on a memory of the management computer;
<figref idrefs="DRAWINGS">FIG. 9</figref> is one example of a flow-out section I/O table created by the data recovery available time monitoring program on the memory of the management computer;
<figref idrefs="DRAWINGS">FIG. 10</figref> shows a monitoring condition setting screen displayed on a display unit by the management computer to urge a user to set a monitoring interval of the data recovery available time;
<figref idrefs="DRAWINGS">FIG. 11</figref> shows a monitoring screen of the data recovery available time displayed by the management computer on the display unit;
<figref idrefs="DRAWINGS">FIG. 12</figref> is a flowchart showing a processing procedure of the data recovery available time monitoring program;
<figref idrefs="DRAWINGS">FIG. 13</figref> is a graph for explaining calculations of the data recovery available time; and
<figref idrefs="DRAWINGS">FIG. 14</figref> shows a monitoring condition setting screen displayed by the management computer on the display unit to urge the user to set a permissible error of the data recovery available time.
BEST MODE FOR CARRYING OUT THE INVENTION
A best mode for carrying out the invention (referred to as an “embodiment” hereinafter) will be explained with reference to the drawings (see appropriately also the drawings not mentioned). An outline of the invention will be explained to facilitate understanding of the invention.
1. Outline of the Invention
1.1 Asynchronous Remote Copy:
The Asynchronous Remote Copy intended by the invention will be explained with reference to <figref idrefs="DRAWINGS">FIG. 1</figref>. <figref idrefs="DRAWINGS">FIG. 1</figref> is a block diagram showing an outline of the invention. A storage system S of the invention includes a management computer <b>1100</b>, a host computer <b>1200</b>A coupled with the management computer <b>1100</b>, a storage device <b>1300</b>A coupled with the host computer <b>1200</b>A, a storage device <b>1300</b>B coupled with the storage device <b>1300</b>A and a host computer <b>1200</b>B coupled with the storage device <b>1300</b>B.
A primary site <b>1700</b> (first site) is a site where business operations are carried out. A remote site <b>1800</b> (second site) is a site for continuing the business operations when the primary site <b>1700</b> is damaged. Each site is provided with the host computer <b>1200</b> (generic term of the host computers <b>1200</b>A and <b>1200</b>B. The same applies also to other reference numerals) for performing the business operations and with the storage device <b>1300</b> for storing business operation data.
The business operations executed on the host computer <b>1200</b>A provided in the primary site <b>1700</b> use a logical disk <b>1315</b>A on the storage device <b>1300</b>A. The host computer <b>1200</b>A issues a write I/O (Write request) in performing the business operations. The write I/O is composed of a write command and write data. The write data issued by an OS (Operating System) or an application and is divided into a transmission unit to be transferred on a network is called as a frame <b>100</b> (or referred to simply as a “frame” hereinafter). The frame <b>100</b> is a frame of the FCP (Fibre Channel Protocol) in the SAN environment.
Here, the storage device <b>1300</b>A set up in the primary site <b>1700</b> has a buffer <b>1318</b>. The frame <b>100</b> issued from the host computer <b>1200</b>A is written into the both of the logical disk <b>1315</b>A and the buffer <b>1318</b> on the storage device <b>1300</b>A of the primary site. The storage device <b>1300</b>A returns a response that writing is completed to the host computer <b>1200</b>A at the point of time when write data corresponding to the frame <b>100</b> is written into the logical disk <b>1315</b>A. The write data corresponding to the frame <b>100</b> written into the buffer <b>1318</b> is transferred to the remote site <b>1800</b> (by another process) asynchronously from the writing into the logical disk <b>1315</b>A. Then, the storage device <b>1300</b>B of the remote site <b>1800</b> writes the received write data into a logical disk <b>1315</b>B of the device.
Or, the storage device <b>1300</b>A may store the write data received from the host computer <b>1200</b>A in a cache memory of the storage device <b>1300</b>A and may manage the data stored in the cache memory together with the data stored in the buffer <b>1318</b> after completing the storage. Then, the storage device <b>1300</b>A may transfer the write data on the cache memory to the storage device <b>1300</b>B. At this time, rewriting of the data from the cache memory to a disk device corresponding to the logical disk <b>1315</b>A may be carried out at time different from the transfer between the aforementioned storage devices.
Note that the logical disk <b>1315</b>B must keep consistency at a point of time when the host computer <b>1200</b>B reads in order for the host computer <b>1200</b>B to resume the business operations by using the logical disk <b>1315</b>B. Here, the consistency is a concept regarding sequence of data written into the logical disk <b>1315</b>.
When the host computer writes first data A and next data B into the logical disk <b>1315</b> for example, it may be considered that the consistency is kept by arranging so that the host computer writes the data B after receiving a notice that writing of the data A has been completed from the storage device.
A simplest process for keeping this consistency is to write all frames <b>100</b> into the logical disk <b>1315</b>A while keeping the order issued from the host computer <b>1200</b>A. However, it is possible to write the write data into the logical disk <b>1315</b>A by any method, as long as the consistency of data in the disk at the moment of accessing the logical disk <b>1315</b>B is guaranteed, by storing the write data while referring to sequential information given to the frames.
It becomes possible to conform the contents of the data of the logical disks <b>1315</b> of the primary and remote sites <b>1700</b> and <b>1800</b> at most without affecting response of the host computer <b>1200</b>A by asynchronously conducting the writing into the primary site <b>1700</b> and the writing into the remote site <b>1800</b>. It is also possible to reduce the load on the host computer <b>1200</b> because the remote copy process is carried out among the storage devices <b>1300</b>.
It is noted that the host computer <b>1200</b>A is not necessary to be one and there may be a plurality of host computers. Because the storage devices <b>1300</b> conduct the remote copy, it becomes possible to provide data whose consistency is kept throughout the data used by the plurality of host computers <b>1200</b>A when the plurality of host computers <b>1200</b>A realizes a certain process in concert with each other.
There may be a plurality of host computers <b>1200</b>B in terms of its configuration similarly to the host computers <b>1200</b>A.
1.2 Data Recovery Available Time:
Next, the data recovery available time monitored by the invention will be explained.
There exists a time lag in writing data into the primary site <b>1700</b> and into the remote site <b>1800</b> in the Asynchronous Remote Copy as described above. When the primary site <b>1700</b> is damaged, the remote site <b>1800</b> recovers the data while missing the data of that time lag. Accordingly, it is necessary to monitor a length of the time lag, i.e., to monitor that data of which time can be used for the recovery of the data when the primary site <b>1700</b> is damaged. Supposing a case when a failure occurs in the primary site <b>1700</b> at certain time, e.g., 21:00:00, the data recovery available time is, as described above, the time when the latest data stored in the storage device <b>1300</b>B of the remote site <b>1800</b> has been written into the storage device <b>1300</b>A of the primary site <b>1700</b>, e.g., 20:59:20. That is, the data recovery available time is the latest time by which the data can be recovered.
1.3 Outline of Processing:
A time necessary for transferring the frame <b>100</b> from the primary site <b>1700</b> to the remote site <b>1800</b> is normally in the order of millisecond when the Fibre Channel that is normally used in a large-scale storage system is used for the network among the sites. A time necessary for writing the frame <b>100</b> into the logical disk <b>1315</b>B in the remote site <b>1800</b> is also normally in the order of millisecond. Meanwhile, a time during which the frame <b>100</b> stays in the buffer <b>1318</b> is about five minutes at most. Therefore, the second time scale is sufficient as accuracy in monitoring the data recovery available time. That is, it may be considered that the data recovery available time substantially depends on the time during which the frame <b>100</b> stays in the buffer <b>1318</b>.
Then, the invention monitors the data recovery available time by the following method. In this case, the invention presupposes two points as follows.
The first point is that the host computer <b>1200</b>A embeds a sequence number into the frame <b>100</b> to be issued. This presupposition is realized in general in using the Fibre Channel.
The second point is that the storage device <b>1300</b>A has a function of acquiring the sequence numbers of the frames <b>100</b> existing respectively at inlet and outlet of the buffer <b>1318</b>. Suppose a case when a frame <b>100</b> whose sequence number is “3745” (latest or quasi-latest management) is acquired at the inlet of the buffer <b>1318</b> in <figref idrefs="DRAWINGS">FIG. 1</figref> and a frame <b>100</b> whose sequence number is “1150” (earliest or quasi-earliest data) is acquired at the outlet of the buffer <b>1318</b> for example. It is supposed here that FIFO (First In First Out) is adopted for the buffer <b>1318</b>.
When the two presuppositions described above is true, the management computer <b>1100</b> acquires the sequence number of the frame <b>100</b> existing at the inlet of the buffer <b>1318</b> within the storage device <b>1300</b>A repeatedly (preferably periodically) through the host computer <b>1200</b>A. In combination with that, the management computer <b>1100</b> acquires TOD (Time of Day: simply referred to as “time” hereinafter) of the host computer <b>1200</b>A at that moment when it acquires the sequence number and retains their combination as an array.
When it is necessary to acquire the data recovery available time at certain point of time, the management computer <b>1100</b> acquires the sequence number of the frame <b>100</b> existing at the outlet of the buffer <b>1318</b> within the storage device <b>1300</b>A through the host computer <b>1200</b>A. The time when the frame <b>100</b> having this sequence number has existed at the inlet of the buffer <b>1318</b> is the data recovery available time. This time may be acquired approximately by using the array described above (this will be detailed later).
Thus, it becomes possible to monitor the data recovery available time by the time managed by the host computer <b>1200</b>A. It is also possible to display its result in a GUI (Graphical User Interface) like a monitoring screen <b>10600</b> shown in <figref idrefs="DRAWINGS">FIG. 11</figref>.
2. Configuration of Storage System
2.1 Outline of Configuration:
A configuration of the storage system S will be explained with reference to <figref idrefs="DRAWINGS">FIGS. 2 through 5</figref>. <figref idrefs="DRAWINGS">FIG. 2</figref> is a block diagram showing the configuration of a storage system of an embodiment. The storage system S includes the management computer <b>1100</b>, the host computer <b>1200</b>A and the storage device <b>1300</b>A in the primary site <b>1700</b> and the host computer <b>1200</b>B and the storage device <b>1300</b>B in the remote site <b>1800</b>.
In the storage system S, the storage device <b>1300</b> (<b>1300</b>A and <b>1300</b>B) and the host computer <b>1200</b> (<b>1200</b>A and <b>1200</b>B) are coupled from each other through the data network <b>1500</b> (<b>1500</b>A and <b>1500</b>B). It is noted that although the data network <b>1500</b> is the SAN in the present embodiment, it may be IP (Internet Protocol) network or a data communication network other than those networks.
The host computer <b>1200</b> is coupled with the management computer <b>1100</b> through a management network <b>1400</b>. While the management network <b>1400</b> is the IP network in the present embodiment, it may be the SAN or the data communication network other than those networks. Still more, although the management computer <b>1100</b> is not coupled directly with the storage device <b>1300</b> and acquires information through the host computer <b>1200</b>, the invention may be carried out even by arranging such that the management computer <b>1100</b> is directly coupled with the storage device <b>1300</b>. Still more, although the data network <b>1500</b> is different from the management network <b>1400</b> in the present embodiment, those networks may be one and same network. Further, the management computer <b>1100</b> may be also one and same computer with the host computer <b>1200</b>.
It is noted that although the storage system S is composed of the two storage devices <b>1300</b>, the two host computers <b>1200</b> and one management computer <b>1100</b> as shown in <figref idrefs="DRAWINGS">FIG. 2</figref> for convenience of explanation, the invention is limited to those numbers. Still more, although the management computer <b>1100</b> is coupled with the host computers <b>1200</b>A and <b>1200</b>B, the invention may be carried out by coupling the management computer <b>1100</b> only with the host computer <b>1200</b>A of the primary site <b>1700</b>. The reason why the management computer <b>1100</b> is coupled also with the host computer <b>1200</b>B in the present embodiment is that the roles of the primary and remote sites <b>1700</b> and <b>1800</b> are possibly swapped. Specifically, when the primary site <b>1700</b> is damaged, a swapping command is issued to change a data transferring direction from the remote site <b>1800</b> to the primary site <b>1700</b> as described later.
A set of the host computer <b>1200</b>, the storage device <b>1300</b> and the data network <b>1500</b> coupling them will be called as a site in the present embodiment. A plurality of such sites is placed at positions geographically distant from each other in general, so that the other site can continue business operations even if one site is damaged. <figref idrefs="DRAWINGS">FIG. 2</figref> shows an architecture including the primary site <b>1700</b> that performs business operations and the remote site <b>1800</b> that backs up the business operations. Such architecture will be called as a two data center architecture (referred to as “2DC” hereinafter).
In the 2DC architecture, Remote Copy is carried out between the primary site <b>1700</b> and the remote site <b>1800</b> through a remote network <b>1600</b>. The Remote Copy technology allows system operations to be continued by using data stored in a volume of another site even if one site causes a failure in its volume and becomes inoperable. A pair of the two volumes of a copy source and a copy destination involved in the Remote Copy will be called as a copy pair.
Configurations of the management computer <b>1100</b>, the host computer <b>1200</b> and the storage device <b>1300</b> will be explained below.
2.2 Configuration of Management Computer:
<figref idrefs="DRAWINGS">FIG. 3</figref> is a block diagram showing the configuration of the management computer <b>1100</b>. The management computer <b>1100</b> has an input device <b>1130</b> such as a keyboard and a mouse, a CPU (Central Processing Unit) <b>1140</b>, a display unit (display section) <b>1150</b> such as a CRT (Cathode Ray Tube), a memory <b>1120</b>, a local disk <b>1110</b> and a supervisory I/F (Interface) <b>1160</b> for transmitting/receiving data and control commands to/from the host computer <b>1200</b> to manage the system.
The local disk <b>1110</b> is a disk device such as a hard disk coupled with (or mounted within) the management computer <b>1100</b> and stores a data recovery available time monitoring program <b>1112</b>.
The data recovery available time monitoring program <b>1112</b> is loaded into the memory <b>1120</b> of the management computer <b>1100</b> and is executed by the CPU <b>1140</b>. The data recovery available time monitoring program <b>1112</b> is a program for providing the function of monitoring the data recovery available time of the Asynchronous Remote Copy through the input device <b>1130</b> such as the keyboard and the mouse and the display unit <b>1150</b> such as the GUI (Graphical User Interface).
A flow-in section I/O table <b>1124</b> and a flow-out section I/O table <b>1126</b> on the memory <b>1120</b> will be described later.
The management I/F <b>1160</b> is an interface for the management network <b>1400</b> and transmits/receives data and control commands to/from the host computer <b>1200</b>.
2.3 Configuration of Host Computer:
<figref idrefs="DRAWINGS">FIG. 4</figref> is a block diagram showing the configuration of the host computer <b>1200</b>. The host computer <b>1200</b> has an input device <b>1240</b> such as a keyboard and a mouse, a CPU <b>1220</b>, a display unit <b>1250</b> such as a CRT, a memory <b>1230</b>, a storage I/F <b>1260</b>, a local disk <b>1210</b> and a time generator <b>1270</b>.
The storage I/F <b>1260</b> is an interface for the data network <b>1500</b> and transmits/receives data and control commands to/from the storage device <b>1300</b>. The local disk <b>1210</b> is a disk device such as a hard disk coupled with (or mounted within) the host computer <b>1200</b> and stores an application <b>1212</b>, a remote copy operation requesting program <b>1214</b>, a remote copy state acquiring program <b>1215</b>, a sequence number acquisition requesting program <b>1216</b> and a time acquiring program <b>1217</b>.
The application <b>1212</b> is loaded into the memory <b>1230</b> of the host computer <b>1200</b> and is executed by the CPU <b>1220</b>. The application <b>1212</b> is a program that executes processes by reading/writing data from/into volumes on the storage device <b>1300</b> and is a DBMS (Data Base Management System), a file system and the like. It is noted that although <figref idrefs="DRAWINGS">FIG. 4</figref> shows only one application <b>1212</b> for convenience of explanation, the invention is not limited to this number.
The time acquiring program <b>1217</b> is a program for acquiring time managed by the host to return to a request source based on an instruction from the user or another program.
The remote copy operation requesting program <b>1214</b>, the remote copy state acquiring program <b>1215</b>, the sequence number acquisition requesting program <b>1216</b> and the time acquiring program <b>1217</b> are loaded into the memory <b>1230</b> of the host computer <b>1200</b> and are executed by the CPU <b>1220</b>.
The remote copy operation requesting program <b>1214</b> is a program for requesting an operation of Remote Copy on the storage device <b>1300</b> specified based on an instruction from the user or the other program. Description will be made later as to what kinds of operation can be requested.
The remote copy state acquiring program <b>1215</b> is a program for acquiring a state of the Remote Copy from the storage device <b>1300</b> to return to a request source based on an instruction from the user or the other program. Description will be made later as to what kinds of state can be obtained.
The sequence number acquisition requesting program <b>1216</b> invokes the sequence number acquiring program <b>1340</b> (see <figref idrefs="DRAWINGS">FIG. 5</figref>) on the storage device <b>1300</b> based on an instruction from the user or the other program. The sequence number acquisition requesting program <b>1216</b> also invokes the time acquiring program <b>1217</b> to acquire the current time. The sequence number acquisition requesting program <b>1216</b> returns a value returned from the sequence number acquiring program <b>1340</b> on the storage device <b>1300</b> and the current time acquired from the time acquiring program <b>1217</b> to the invoker.
2.4 Configuration of Storage Device:
<figref idrefs="DRAWINGS">FIG. 5</figref> is a block diagram showing the configuration of the storage device <b>1300</b>. The storage device <b>1300</b> has the disk device <b>1310</b> for storing data and a disk controller <b>1320</b> for controlling the storage device <b>1300</b>.
The disk controller <b>1320</b> is provided with a host I/F <b>1328</b>, a remote I/F <b>1326</b>, a disk I/F <b>1325</b>, a memory <b>1321</b>, a CPU <b>1323</b> and a local disk <b>1327</b>.
The host I/F <b>1328</b> is an interface for the data network <b>1500</b> and transmits/receives data and control commands to/from the host computer <b>1200</b>.
The remote I/F <b>1326</b> is an interface for the remote network <b>1600</b> and is used in transferring Remote Copy data across the sites.
The disk I/F <b>1325</b> is an interface for the disk device <b>1310</b> and transmits/receives the data and control commands.
A remote copy definition table <b>1322</b> on the memory <b>1321</b> will be explained later.
The local disk <b>1327</b> is a disk device such as a hard disk coupled with the disk controller <b>1320</b> and stores an I/O processing program <b>1334</b>, a remote copy control program <b>1336</b>, a disk array control program <b>1338</b> and a sequence number acquiring program <b>1340</b>.
The I/O processing program <b>1334</b> is loaded into the memory <b>1321</b> of the disk controller <b>1320</b> and is executed by the CPU <b>1323</b>. The I/O processing program <b>1334</b> accepts write and read requests from the host computer <b>1200</b> and another storage device <b>1300</b>. If the request is a write request, the I/O processing program <b>1334</b> writes data to the disk device <b>1310</b> and if the request is a read request, the I/O processing program <b>1334</b> reads requested data out of the disk device <b>1310</b>.
The remote copy control program <b>1336</b> is loaded into the memory <b>1321</b> of the disk controller <b>1320</b> and is executed by the CPU <b>1323</b>.
The remote copy control program <b>1336</b> controls Remote Copy and acquires a state of Remote Copy by making reference to the remote copy definition table <b>1322</b> based on an instruction from the user, the host computer <b>1200</b> or the other program. As to controls of the copy pair, there are such operations as creation of copy pair for newly creating a copy pair, suspension of copy pair for interrupting a synchronous relationship, resynchronization of copy pair for conforming contents of a remote side volume with that of a primary side volume and others. The acquisition of the state of the copy pair means to grasp that each copy pair is put in which state by which operation. The transition of states of the copy pair will be described later.
The disk array control program <b>1338</b> is loaded into the memory <b>1321</b> of the disk controller <b>1320</b> and is executed by the CPU <b>1323</b>. The disk array control program <b>1338</b> has a function of reconstructing a plurality of hard disk drives <b>1312</b>A, <b>1312</b>B and <b>1312</b>C coupled with the disk controller <b>1320</b> as the logical disk <b>1315</b> for example by controlling them. A disk array control method may be one such as RAID (Redundant Array of Independent Disks), the invention is not limited to such method.
The sequence number acquiring program <b>1340</b> is loaded into the memory <b>1321</b> of the disk controller <b>1320</b> and is executed by the CPU <b>1323</b>. The sequence number acquiring program <b>1340</b> is a program that acquires the sequence numbers of the frames <b>100</b> staying at the inlet and the outlet among the frames <b>100</b> staying in the buffer <b>1318</b> of the storage device <b>1300</b> and returns the sequence numbers to a request source based on an instruction from the user, the host computer <b>1200</b> or the other program. The sequence number is an ID given to the frame <b>100</b>. For instance, when the frame <b>100</b> is issued from the host computer, the sequence number is given to the individual frame <b>100</b> in the storage system using the Fibre Channel Protocol.
It is noted that although the programs <b>1334</b>, <b>1336</b>, <b>1338</b> and <b>1340</b> are stored in the local disk <b>1327</b> on the disk controller <b>1320</b> in the present embodiment, the invention is not limited to such configuration. For instance, it is also possible to provide a flash memory or the like on the disk controller <b>1320</b> and to store those programs <b>1334</b>, <b>1336</b>, <b>1338</b> and <b>1340</b> in the flash memory or to store in an arbitrary disk within the disk device <b>1310</b>.
The disk device <b>1310</b> is composed of the plurality of hard disk drives (HDD) <b>1312</b>A, <b>1312</b>B and <b>1312</b>C. It is noted that the disk device may be composed of other drives such as a solid state drive (SSD) other than the hard disk drive. It is also noted that <figref idrefs="DRAWINGS">FIG. 5</figref> shows only three volumes for convenience of explanation, the invention is not limited such number. Note that the logical disk <b>1315</b> is constructed within the disk device <b>1310</b> as described above.
The buffer <b>1318</b> (<b>1318</b>A and <b>1318</b>B) is constructed in the logical disk <b>1315</b> within the disk device <b>1310</b>. During Asynchronous Remote Copy, writing into the disk device <b>1310</b> within the storage device <b>1300</b>B of the remote site <b>1800</b> is carried out asynchronously from writing into the disk device <b>1310</b> within the storage device <b>1300</b>A of the primary site <b>1700</b>. The buffer <b>1318</b> is used in order to temporarily store data to asynchronously write into the disk device <b>1310</b> within the storage device <b>1300</b>B of the remote site <b>1800</b>.
It is noted that although not shown in <figref idrefs="DRAWINGS">FIG. 5</figref>, the storage devices <b>1300</b>A and <b>1300</b>B have cache memories for temporarily storing a part or whole of write data from each host computer <b>1200</b> and a part or whole of read data transferred to each host computer <b>1200</b> due to a read request in the past. The storage device <b>1300</b> transfers the data stored in the cache memory corresponding to the read request from the host computer <b>1200</b> (i.e., the part or whole of past write data and the part or whole of past read data) to the host computer <b>1200</b>. It is noted that the memory <b>1321</b> may assume the role of the cache memory.
It is noted that the buffer <b>1318</b> may straddle the plurality of logical disks <b>1315</b>. Still more, the buffer <b>1318</b> needs not be always provided within the disk device <b>1310</b>. The buffer <b>1318</b> may be provided on the local disk <b>1327</b> and the memory <b>1321</b> of the disk controller <b>1320</b> or in the cache memory. It is also possible to use those memory areas together. In the case of the cache memory in particular, an area used as a buffer of the Asynchronous Remote Copy needs not be physically partitioned or needs not be continuous addresses. It is just required to store data to be copied as Remote Copy.
Still more, although <figref idrefs="DRAWINGS">FIG. 5</figref> shows the two buffers <b>1318</b> for convenience of explanation, the invention is not limited to that number.
3. Asynchronous Remote Copy
A concrete example of the Asynchronous Remote Copy that is an object of the present invention will be explained with reference to <figref idrefs="DRAWINGS">FIGS. 6 and 7</figref>.
3.1 Remote Copy Definition Table:
<figref idrefs="DRAWINGS">FIG. 6</figref> is one exemplary table of a remote copy definition table <b>1322</b> created in the memory <b>1321</b> on the disk controller <b>1320</b> within the storage device <b>1300</b>.
The remote copy definition table <b>1322</b> is roughly divided into a Primary Site field <b>6100</b> and a Remote Site field <b>6200</b>. Each field stores information of the logical disk <b>1315</b> of the primary site <b>1700</b> and of the logical disk <b>1315</b> of the remote site <b>1800</b> composing the copy pair.
The Primary Site field <b>6100</b> and the Remote Site field <b>6200</b> include Storage Device fields <b>6110</b> (<b>6110</b>A and <b>6110</b>B) and Logical Disk Number fields <b>6120</b> (<b>6120</b>A and <b>6120</b>B), respectively. The Storage Device filed <b>6110</b> stores an ID of the storage device <b>1300</b> in which the logical disk <b>1315</b> is stored. The storage device <b>1300</b> may be identified uniquely by the ID. The ID includes an IP address for example. The Logical Disk Number field <b>6120</b> stores a number of the logical disk <b>1315</b>. This number is unique within the storage device <b>1300</b>.
Rows <b>6010</b> through <b>6050</b> indicate information of copy pairs. For example, the row <b>6010</b> indicates that the logical disk <b>1315</b> “LDEV<b>01</b>” stored in the storage device <b>1300</b>A of the primary site <b>1700</b> composes a copy pair with the logical disk <b>1315</b> “LDEV<b>01</b>” stored in the storage device <b>1300</b>B of the remote site <b>1800</b>. Therefore, the remote copy definition table <b>1322</b> shows that the five logical disks <b>1315</b> stored in the storage device <b>1300</b>A of the primary site <b>1700</b> compose five copy pairs with the five logical disks <b>1315</b> stored in the storage device <b>1300</b>B of the remote site <b>1800</b>. These five copy pairs are handled as a group that is called as a copy group.
It is noted that although the present embodiment exemplifies the structure in which there is only one each storage device <b>1300</b> in the primary site <b>1700</b> and the remote site <b>1800</b>, there may be two or more storage devices in each site. In this case, the remote copy definition table <b>1322</b> may be defined across the plurality of storage devices <b>1300</b>B of the primary site <b>1700</b> or the remote copy definition table <b>1322</b> may be defined across the plurality of storage devices <b>1300</b>B of the remote site <b>1800</b>. Further, although the storage device <b>1300</b> has one remote copy definition table <b>1322</b> in the present embodiment, the storage device <b>1300</b> may have a plurality of remote copy definition tables <b>1322</b>. That is, it is possible to have a plurality of copy groups.
3.2 State Transition of Copy Pair:
<figref idrefs="DRAWINGS">FIG. 7</figref> is a state transition diagram of a copy pair of the Asynchronous Remote Copy.
As described in a legend shown in a lower part of <figref idrefs="DRAWINGS">FIG. 7</figref>, rectangles indicate states of the copy pair and hexagons indicate operations instructed from the user or programs and behaviors of the copy pair caused by the operations.
The copy pair of the Asynchronous Remote Copy has four states.
Pair <b>7110</b> is a state in which the Asynchronous Remote Copy is conducted between the logical disk <b>1315</b> of the primary site <b>1700</b> and the logical disk <b>1315</b> of the remote site <b>1800</b>. That is, the frame <b>100</b> written into the logical disk <b>1315</b> of the primary site <b>1700</b> is written asynchronously (by another process) into the logical disk <b>1315</b> of the remote site <b>1800</b>.
Simplex <b>7120</b> indicates a state in which there is no relationship of copy pair between the logical disk <b>1315</b> of the primary site <b>1700</b> and the logical disk <b>1315</b> of the remote site <b>1800</b>. That is, the frame <b>100</b> written into the logical disk <b>1315</b> of the primary site <b>1700</b> is not written into the remote site <b>1800</b>.
Suspend <b>7130</b> is a state in which the Asynchronous Remote Copy between the logical disk <b>1315</b> of the primary site <b>1700</b> and the logical disk <b>1315</b> of the remote site <b>1800</b> is temporarily interrupted. That is, the frame <b>100</b> written into the logical disk <b>1315</b> of the primary site <b>1700</b> is not written into the remote site <b>1800</b> and is kept as differential information.
Swapped Pair <b>7140</b> is a state in which the Asynchronous Remote Copy is conducted in a state in which the role of the logical disk <b>1315</b> of the primary site <b>1700</b> is switched with that of the logical disk <b>1315</b> of the remote site <b>1800</b>. That is, the frame <b>100</b> written into the logical disk <b>1315</b> of the remote site <b>1800</b> is written asynchronously (by another process) into the logical disk <b>1315</b> of the primary site <b>1700</b>.
It is noted that the data recovery available time may be acquired only in the states of Pair <b>7110</b> and Swapped Pair <b>7140</b>. That is, no Asynchronous Remote Copy is conducted from the primary site <b>1700</b> to the remote site <b>1800</b> in the states of Simplex <b>7120</b> and Suspend <b>7130</b>, so that there is no meaning of monitoring the data recovery available time.
Next, the operations instructed from the user or the programs and the transition of the states of the copy pair caused by the operations will be explained.
Initial Copy starts when a Create command of forming a copy pair is issued in the state of Simplex <b>7120</b> (<b>7210</b>). The Initial Copy is to copy all data of the logical disk <b>1315</b> of the primary site <b>1700</b> to the logical disk <b>1315</b> of the remote site <b>1800</b> before starting the Asynchronous Remote Copy. When the Initial Copy ends, the contents of the logical disk <b>1315</b> of the primary site <b>1700</b> coincide with that of the logical disk <b>1315</b> of the remote site <b>1800</b>, thus forming the Pair <b>7110</b>. Still more, the buffer <b>1318</b> is allotted at this time. It is noted that the buffer <b>1318</b> is allotted in unit of the copy group.
Suspending starts when a Suspend command is issued in the state of the Pair <b>7110</b> (<b>7240</b>). Suspending is a performance for temporarily interrupting the Asynchronous Remote Copy. Specifically, it is to prepare a differential management table for managing that writing has been made to which address within the logical disk <b>1315</b> during the Suspend <b>7130</b> in the memory <b>1321</b> within the disk controller <b>1320</b>. This differential management table manages addresses within the disk device <b>1310</b> that are changed by the frame <b>100</b>. It is noted that the differential management table may be created in the disk device <b>1310</b>.
Differential Copy starts when a Resynchronize command is issued in the state of the Suspend <b>7130</b> (<b>7230</b>). The Differential Copy is to transfer the difference generated during the Suspend <b>7130</b> from the logical disk <b>1315</b> of the primary site <b>1700</b> to the logical disk <b>1315</b> of the remote site <b>1800</b> based on the differential management table. When the transfer ends, the contents of the logical disk <b>1315</b> of the primary site <b>1700</b> is synchronized with that of the logical disk <b>1315</b> of the remote site <b>1800</b>, thus forming the Pair <b>7110</b>.
Swapping starts when a Takeover command is issued in the state of the Pair <b>7110</b> or the Swapped Pair <b>7140</b>. Swapping is a behavior of reversing the data transferring direction between the sites. When the Takeover command is issued in the state of the Pair <b>7110</b>, the data transferring direction is changed in a direction from the remote site <b>1800</b> to the primary site <b>1700</b>. This operation is used when a failure occurs in the primary site <b>1700</b> for example. This operation allows the site conducting the business operations to be switched from the primary site <b>1700</b> to the remote site <b>1800</b>.
Deleting starts when a Delete command is issued in the state of the Pair <b>7110</b> or the Suspend <b>7130</b> (<b>7220</b>). Deleting is to cancel the relationship of the Asynchronous Remote Copy. When Deleting ends, the state changes to that of the Simplex <b>7120</b> and the frame <b>100</b> written into the logical disk <b>1315</b> of the primary site <b>1700</b> is not transferred to the remote site <b>1800</b>. It is noted that the buffer <b>1318</b> allotted during the Initial Copy is released at this time.
3.3 Processing of Remote Copy in Pair:
The Asynchronous Remote Copy of the present embodiment realizes duplication of data of the logical disk <b>1315</b> by the following process. It is noted that the Asynchronous Remote Copy may be realized by a method other than this method.
When the host computer <b>1200</b>A transmits the frame <b>100</b> to the storage device <b>1300</b>A of the primary site <b>1700</b>, the host computer <b>1200</b>A gives the sequence number to the frame <b>100</b>.
The storage device <b>1300</b>A of the primary site <b>1700</b> received the frame <b>100</b> from the host computer <b>1200</b>A stores the received frame <b>100</b> in the memory. The storage device <b>1300</b>A also stores the received frame <b>100</b> in the memory <b>1321</b> and the logical disk <b>1315</b>A within the device <b>1300</b>A and transmits a response indicating that receiving of the frame <b>100</b> has been completed to the host computer <b>1200</b>A.
Meanwhile, the storage device <b>1300</b>A of the primary site <b>1700</b> transmits the frame <b>100</b> stored in the memory <b>1321</b> or the logical disk <b>1315</b> to the storage device <b>1300</b>B of the remote site <b>1800</b> together with its given sequence number.
The storage device <b>1300</b>B of the remote site <b>1800</b> that received the frame <b>100</b> writes the frame <b>100</b> into the logical disk <b>1315</b> by making reference to the sequence number given to that frame <b>100</b>. Thereby, even if the data in the logical disk <b>1315</b> of the storage device <b>1300</b>A of the primary site <b>1700</b> is lost, it becomes possible to assure the consistency of the data stored in the logical disk <b>1315</b> of the storage device <b>1300</b>B of the remote site <b>1800</b>.
4. Management Computer:
Tables and processes managed by the management computer <b>1100</b> of the present embodiment will be explained with reference to <figref idrefs="DRAWINGS">FIGS. 8 through 14</figref>.
4.1 Table:
<figref idrefs="DRAWINGS">FIG. 8</figref> is one example of the flow-in section I/O table <b>1124</b> created by the data recovery available time monitoring program <b>1112</b> on the memory <b>1120</b> of the management computer <b>1100</b>. The flow-in section I/O table <b>1124</b> includes a Time field <b>8100</b> and a Sequence Number field <b>8200</b>. The flow-in section I/O table <b>1124</b> may have a plurality of rows (<b>8010</b> through <b>8050</b> for example).
The Time field <b>8100</b> stores current time that is one of return values of the sequence number acquisition requesting program <b>1216</b>. The Sequence Number field <b>8200</b> stores the sequence number of the frame <b>100</b> staying at the inlet among the frames <b>100</b> staying in the buffer <b>1318</b> of the storage device <b>1300</b>. The sequence number is also one of the return values of the sequence number acquisition requesting program <b>1216</b>.
<figref idrefs="DRAWINGS">FIG. 9</figref> is one example of the flow-out section I/O table <b>1126</b> created by the data recovery available time monitoring program <b>1112</b> on the memory <b>1120</b> of the management computer <b>1100</b>.
The flow-out section I/O table <b>1126</b> includes a Time field <b>9100</b> and a Sequence Number field <b>9200</b>. The flow-out section I/O table <b>1126</b> is composed of a single raw (<b>9010</b>).
The Time field <b>9100</b> stores current time that is one of return values of the sequence number acquisition requesting program <b>1216</b>. The Sequence Number field <b>9200</b> stores the sequence number of the frame <b>100</b> staying at the outlet among the frames <b>100</b> staying in the buffer <b>1318</b> of the storage device <b>1300</b>. The sequence number is also one of the return values of the sequence number acquisition requesting program <b>1216</b>.
4.2 GUI:
The GUI arranged so that the management computer <b>1100</b> urges the user to set a data recovery available time monitoring interval and displays its result will be explained with reference to <figref idrefs="DRAWINGS">FIGS. 10 and 11</figref>.
<figref idrefs="DRAWINGS">FIG. 10</figref> shows a monitoring condition setting screen <b>10500</b> displayed on the display unit <b>1150</b> by the management computer <b>1100</b> to urge the user to set the monitoring interval of the data recovery available time. The monitoring condition setting screen <b>10500</b> includes a message portion <b>10510</b>, a monitoring interval input portion <b>10520</b> and a button <b>10530</b>.
The message portion <b>10510</b> indicates a message urging the user to input the monitoring interval.
The monitoring interval input portion <b>10520</b> indicates a text box for accepting an input of the monitoring interval from the user.
The button <b>10530</b> is a control section clicked by the user after inputting the monitoring interval. Thereby, the data recovery available time monitoring program <b>1112</b> keeps the monitoring interval inputted by the user to the monitoring interval input portion <b>10520</b>.
<figref idrefs="DRAWINGS">FIG. 11</figref> shows a monitoring screen <b>10600</b> of the data recovery available time displayed by the management computer <b>1100</b> on the display unit <b>1150</b>. The monitoring screen <b>10600</b> includes Current Time <b>10610</b>, the Data Recovery Available Time <b>10620</b>, Data loss time period <b>10630</b> and a button <b>10640</b>.
The Current Time <b>10610</b> indicates time returned by the data recovery available time monitoring program <b>1112</b>.
The Data Recovery Available Time <b>10620</b> indicates the data recovery available time calculated by the data recovery available time monitoring program <b>1112</b>.
The Data Loss Time Period <b>10630</b> indicates the data loss time period calculated by the data recovery available time monitoring program <b>1112</b>. The button <b>10640</b> closes the monitoring screen <b>10600</b> when clicked.
It is noted that although the monitoring condition setting screen <b>10500</b> and the monitoring screen <b>10600</b> are displayed on the display unit <b>1150</b> of the management computer <b>1100</b> in the present embodiment, these GUI may be displayed on a display of another computer by using a generally known technology such as a Web server.
4.3 Flowchart:
A processing procedure of the data recovery available time monitoring program <b>1112</b> executed by the management computer <b>1100</b> to acquire and display the data recovery available time will be explained by using a flowchart shown in <figref idrefs="DRAWINGS">FIG. 12</figref>. <figref idrefs="DRAWINGS">FIG. 12</figref> is a flowchart showing the processing procedure of the data recovery available time monitoring program <b>1112</b>.
The CPU <b>1140</b> of the management computer <b>1100</b> (the subject of the operation will be denoted as the “management computer <b>1100</b>” hereinafter) starts the processes of the data recovery available time monitoring program <b>1112</b> by receiving an instruction from the user or the other program.
The management computer <b>1100</b> assures an area for the flow-in section I/O table <b>1124</b> and the flow-out section I/O table <b>1126</b> on the memory <b>1120</b> as an initialization process in Step S<b>1</b>.
Next, the management computer <b>1100</b> invokes the sequence number acquisition requesting program <b>1216</b> on the host computer <b>1200</b>A in Step S<b>2</b>. As a result, the management computer <b>1100</b> receives three information of the sequence number of the frame <b>100</b> staying at the inlet of the buffer <b>1318</b> among the frames <b>100</b> staying in the buffer <b>1318</b> of the storage device <b>1300</b>A, the sequence number of the frame <b>100</b> staying at the outlet of the buffer <b>1318</b> and the current time as return values.
Next, the management computer <b>1100</b> updates the flow-in section I/O table <b>1124</b> and the flow-out section I/O table <b>1126</b> in Step S<b>3</b>. Specifically, the management computer <b>1100</b> adds the current time and the sequence number of the frame <b>100</b> staying at the inlet of the buffer <b>1318</b> among the frames <b>100</b> staying in the buffer <b>1318</b> of the storage device <b>1300</b>A to the last row of the flow-in section I/O table <b>1124</b>. The management computer <b>1100</b> also updates the flow-out section I/O table <b>1126</b> by the current time and the sequence number of the frame <b>100</b> staying at the outlet of the buffer <b>1318</b> among the frames <b>100</b> staying in the buffer <b>1318</b> of the storage device <b>1300</b>A.
Next, the management computer <b>1100</b> calculates the data recovery available time by utilizing the flow-in section I/O table <b>1124</b> and the flow-out section I/O table <b>1126</b> in Step S<b>4</b>.
The data recovery available time may be calculated by measuring a time during which the frame <b>100</b> stays in the buffer <b>1318</b> as described above.
This calculation will be concretely explained below by making reference to <figref idrefs="DRAWINGS">FIG. 13</figref>. <figref idrefs="DRAWINGS">FIG. 13</figref> is a graph explaining the calculation of the data recovery available time. In the graph in <figref idrefs="DRAWINGS">FIG. 13</figref>, an axis of abscissas represents time and an axis of ordinate represents sequence numbers of the frames. L<b>1</b> is a line representing the sequence numbers of the frames at the inlet of the buffer <b>1318</b> and L<b>2</b> is a line representing the sequence numbers of the frames at the outlet of the buffer <b>1318</b>. It is noted that although the lines L<b>1</b> and L<b>2</b> are continuous in the figure, they actually assume discrete values.
The explanation will be continued by making reference to the drawings other than <figref idrefs="DRAWINGS">FIG. 13</figref>. The flow-out section I/O table <b>1126</b> stores the sequence number and acquisition time of the latest frame acquired at the outlet of the buffer <b>1318</b>. Suppose here that the current time (this may be certain time other than the current time) is TC and the sequence number acquired from the outlet of the buffer <b>1318</b> at this time is SC. Suppose also an estimated value of time when the frame having the sequence number SC has arrived at the inlet of the buffer <b>1318</b> as TP. That is, the data recovery available time at the time TC is TP. At this time, TC-TP is a data loss time period.
Here, the flow-in section I/O table <b>1124</b> has a plurality of rows. When a number of these rows is denoted by M, time of the respective rows are denoted by T[<b>1</b>], T[<b>2</b>], . . . and T[M] and the sequence numbers of respective rows are denoted by S[<b>1</b>], S[<b>2</b>], . . . and S[M].
At this time, there exist S[X] and S[X+1] that meet S[X]□SC<S[X+1] in the flow-in section I/O table <b>1124</b>.
Therefore, a range of TP may be narrowed down to T[X]□TP<T[X+1]
As a result, TP may be obtained approximately by linear interpolation by using the following equation (1) for example: <br /><i>TP=T[X</i>]+(<i>T[X+</i>1<i>]−T[X</i>])×(<i>SC−S[X</i>]÷(<i>S[X+</i>1<i>]−S[X</i>]) Eq. 1
It is noted that although the linear interpolation is used in the present embodiment, generally known other approximation methods such as polynomial interpolation and spline interpolation may be also used.
The explanation will be continued by returning to <figref idrefs="DRAWINGS">FIG. 12</figref>. Next, the management computer <b>1100</b> displays the current time, data recovery available time and data loss time period found in Step S<b>4</b> on the monitoring screen <b>10600</b> in Step S<b>5</b>.
Due to the process in Step S<b>4</b>, the rows less than X are not referred in the next time and after that (become unnecessary) in the flow-in section I/O table <b>1124</b>. Accordingly, the management computer <b>1100</b> carries out a process for releasing (deleting the unnecessary part) the rows less than X in the flow-in section I/O table <b>1124</b> in Step S<b>6</b>.
Next, the management computer <b>1100</b> checks whether or not the user or the other program instructs to End in Step S<b>7</b>.
When End is instructed, i.e., Yes in Step S<b>7</b>, the management computer <b>1100</b> releases the areas of the flow-in section I/O table <b>1124</b> and the flow-out section I/O table <b>1126</b> from the memory <b>1120</b> in Step S<b>9</b> and ends the process.
Where no End instruction is given, i.e., No in Step S<b>7</b>, the management computer <b>1100</b> stands by for a certain period of time in Step S<b>8</b>. Note that this stand-by interval is time specified by the user on the monitoring condition setting screen <b>10500</b>. After that, the management computer <b>1100</b> returns to Step S<b>2</b>.
It is noted that although the user sets the monitoring interval by using the monitoring condition setting screen <b>10500</b> shown in <figref idrefs="DRAWINGS">FIG. 10</figref> in the present embodiment, the user may also use a monitoring condition setting screen <b>10700</b> shown in <figref idrefs="DRAWINGS">FIG. 14</figref>. <figref idrefs="DRAWINGS">FIG. 14</figref> shows the monitoring condition setting screen <b>10700</b> displayed by the management computer <b>1100</b> on the display unit <b>1150</b> to urge the user to set a permissible error of the data recovery available time. Differences of the monitoring condition setting screen <b>10700</b> shown in <figref idrefs="DRAWINGS">FIG. 14</figref> from the monitoring condition setting screen <b>10500</b> shown in <figref idrefs="DRAWINGS">FIG. 10</figref> will be explained.
A message portion <b>107100</b> indicates a message urging the user to input the permissible error in monitoring the data recovery available time. It is possible to keep the error of time of the data recovery available time under the monitoring interval by estimating the data recovery available time by the method described above. Accordingly, the user may input the permissible error instead of the monitoring interval. In this case, the data recovery available time monitoring program <b>1112</b> handles the permissible error inputted by the user as a monitoring interval. A permissible error input portion <b>10720</b> displays a text box for accepting the input of the permissible error from the user. A button <b>10730</b> is an operation part that is clicked by the user after inputting the permissible error. Its process in the same with the case of <figref idrefs="DRAWINGS">FIG. 10</figref>, so that its explanation will be emitted here.
Thus, according to the storage system S of the embodiment, the management computer <b>1100</b> keeps the data at the inlet and outlet of the frames staying in the buffer <b>1318</b> of the storage device <b>1300</b>A of the primary site <b>1700</b> at the certain monitoring interval, calculates the estimated value of the data recovery available time and displays it on the display unit <b>1150</b>. Thereby, it becomes possible to provide the technology, also applicable to technologies other than the main frame technology, of monitoring the data recovery available time while suppressing the monitoring error within a certain range in the storage system S that performs the Asynchronous Remote Copy among the plurality of storage devices <b>1300</b>. Still more, the management computer <b>1100</b> calculates not only the estimated value of the data recovery available time but also the estimated value of the data loss time period and displays it on the display unit <b>1150</b>.
Further, according to the storage system S of the embodiment, the management computer <b>1100</b> can monitor the data recovery available time by using only the temporal information managed by the host computer <b>1200</b>A of the primary site <b>1700</b>, without using temporal information managed by the host computer <b>1200</b>B and others of the remote site <b>1800</b>, so that it becomes easy to operate and control the storage system S. Still more, because the monitoring error may be kept under the monitoring interval, the user can take an appropriate measure by changing a length of the monitoring interval or by increasing a capacity of the buffer <b>1318</b> when a free space of the buffer <b>1318</b> gets fewer than a certain amount. It is then possible to arrange such that the management computer <b>1100</b> indicates an alarm for example when the free space of the buffer <b>1318</b> gets fewer than the certain amount.
While the embodiment of the invention has been described above, the modes of the invention are not limited to those explained above. The present embodiment may be appropriately modified within a scope of the invention not departing from the gist of the invention with regard to its concrete structure such as hardware, programs and others. Variations of the present embodiment will be explained below.
5. Variations
5.1 Acquisition Information:
Although the sequence numbers of the frames existing at the inlet and outlet of the buffer <b>1318</b> are obtained in the present embodiment, the invention is applicable also to a case when the sequence number of the frame existing at the inlet of the buffer <b>1318</b> and a number of the frame staying within the buffer are obtainable. It is because the sequence number of the frame existing at the outlet of the buffer <b>1318</b> is obtainable through calculation using the two values described above. In the same manner, the invention is applicable also to a case when the sequence number of the frame existing at the outlet of the buffer <b>1318</b> and a number of the frame staying within the buffer are obtainable.
5.2 Dynamic Change of Monitoring Interval:
Although the user sets the monitoring interval in the presents embodiment, the management computer may also decide the monitoring interval.
When the data loss time period is short, an importance of the data recovery available time becomes low in general. Therefore, the monitoring interval may be prolonged when the data loss time period is small and the monitoring interval may be shortened when the data recovery available time is large. Thereby, it is possible to reduce the burden of the storage system S.
5.3 Location of Management Computer:
The management computer <b>1100</b> is coupled with the both host computer <b>1200</b>A of the primary site <b>1700</b> and the host computer <b>1200</b>B of the remote site <b>1800</b> in the present embodiment.
However, the invention may be carried by coupling the management computer <b>1100</b> only with the host computer <b>1200</b>A of the primary site <b>1700</b> as described above. It is because information is acquired only from the host computer <b>1200</b>A of the primary site <b>1700</b> in monitoring the data recovery available time in the state of the Pair <b>7110</b>. This structure is suitable particularly when the primary site <b>1700</b> and the remote site <b>1800</b> are managed by different users (managers). For example, an enterprise system sometimes adopts a structure in which own company holds the primary site <b>1700</b> and a SSP (Storage Service Provider) provides the remote site <b>1800</b>. In this case, the management computer <b>1100</b> can monitor the data recovery available time without collecting information from the remote site <b>1800</b>.
However, it is unable to monitor the data recovery available time in the state of the Swapped Pair <b>7140</b> in the structure described above. Then, the management computers <b>1100</b>A and <b>1000</b>B are provided at the primary and remote sites <b>1700</b> and <b>1800</b>, respectively, for example. Thereby, it becomes possible to monitor the data recovery available time by using the management computer <b>1100</b>A of the primary site <b>1700</b> in the state of the Pair <b>7110</b> and to monitor the data recovery available time by using the management computer <b>1000</b>B of the remote site <b>1800</b> in the state of the Swapped Pair <b>7140</b>.
5.4 Cases when Host Computer is not Necessary:
Although the management computer <b>1100</b> acquires information necessary for calculating the data recovery available time by invoking the sequence number acquisition requesting program <b>1216</b> on the host computer <b>1200</b>A of the primary site <b>1700</b> in the present embodiment, the invention may be also carried out by directly invoking the sequence number acquiring program <b>1340</b> on the storage device <b>1300</b>A without going through the host computer <b>1200</b>A. In this case, it is possible to acquire the temporal information by providing a time generator and a time acquisition program in the storage device <b>1300</b>A and by utilizing them or by providing a time generator and a time acquisition program in the management computer <b>1100</b> and by utilizing them.
5.5 When there Exists a Plurality of Host Computers:
Note that when a plurality of host computers <b>1200</b>A exists in the storage system S, it is conceivable that the data recovery available time may differ depending on each individual host computer <b>1200</b>A. Therefore, when the management computer <b>1100</b> is coupled with the plurality of host computers, the management computer <b>1100</b> may acquire the data recovery available times from the plurality of host computers <b>1200</b>A and may display the data recovery available time of each of the plurality of host computers <b>1200</b>A. In this case, each of the plurality of host computers <b>1200</b>A may be located not only in Japan but also abroad.
5.6 Others:
The certain time that has been the reference for calculating the data recovery available time may not be the current time and may be time at certain point of time in the past.
Contents6
13 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13
Every citation, both waysCites: the store holds 8 of 9
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US2013227353A1 | Cited by | United States of America | Pre-grant |
| US2011216444A1 | Cited by | United States of America | Pre-grant |
| US9367418B2 | Cited by | United States of America | Search report |
| US8351146B2 | Cited by | United States of America | Search report |
| US2002103980A1 | Cites | United States of America | Applicant |
| US2005270930A1 | Cites | United States of America | Search report |
| US2006277384A1 | Cites | United States of America | Applicant |
| US2008005460A1 | Cites | United States of America | Search report |
| US2008209146A1 | Cites | United States of America | Applicant |
| US7240197B1 | Cites | United States of America | Search report |
| US7353240B1 | Cites | United States of America | Search report |
| US7698374B2 | Cites | United States of America | Search report |
| European Search Report mailed Apr. 12, 2010. | Non-patent | – | Applicant |
8 members in 4 offices
Priority claims4
| Document | Office | Kind | Date |
|---|---|---|---|
| 2008321323 | Japan | A | |
| 2008321323 | Japan | A | |
| 2008321323 | – | – | – |
| JP20080321323 | – | – | – |
Members8
| Document | Office | Kind | |
|---|---|---|---|
| US2010153666A1 | United States of America | A1 | |
| EP2199913A1 | European Patent Office (EPO) | A1 | |
| JP2010146198A | Japan | A | |
| JP4717923B2 | Japan | B2 | |
| US8108639B2This record | United States of America | B2 | |
| EP2199913B1 | European Patent Office (EPO) | B1 | |
| AT547757T | Austria | T | |
| ATE547757T1 | Austria | T1 |
47 transactions on the USPTO file
Allowed after 1 non-final rejection.
- Non-final rejections
- 1
- Final rejections
- 0
- RCEs
- 0
- Appeals
- 0
Over time
Point at a mark for the transactionTransactions
| Event | Code | |
|---|---|---|
| Expire PatentEXP. | EXP. | |
| Maintenance Fee Reminder MailedREM. | REM. | |
| Payment of Maintenance Fee, 8th Year, Large EntityM1552 | M1552 | |
| Recordation of Patent Grant MailedPGM/ | PGM/ | |
| Patent Issue Date Used in PTA CalculationAllowedPTAC | PTAC | |
| Email NotificationEML_NTR | EML_NTR | |
| Issue Notification MailedAllowedWPIR | WPIR | |
| Dispatch to FDCD1935 | D1935 | |
| Application Is Considered Ready for IssuePILS | PILS | |
| Issue Fee Payment VerifiedN084 | N084 | |
| Issue Fee Payment ReceivedIFEE | IFEE | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Email NotificationEML_NTF | EML_NTF | |
| Mail Notice of AllowanceAllowedMN/=. | MN/=. | |
| Notice of Allowance Data Verification CompletedAllowedN/=. | N/=. | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Response after Non-Final ActionA... | A... | |
| Request for Extension of Time - GrantedXT/G | XT/G | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Email NotificationEML_NTF | EML_NTF | |
| Mail Non-Final RejectionNon-final rejectionMCTNF | MCTNF | |
| Non-Final RejectionNon-final rejectionCTNF | CTNF | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| PG-Pub Issue NotificationPG-ISSUE | PG-ISSUE | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Reference capture on IDSRCAP | RCAP | |
| Information Disclosure Statement (IDS) FiledM844 | M844 | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Correspondence Address ChangeC.ADB | C.ADB | |
| Correspondence Address ChangeC.ADB | C.ADB | |
| Correspondence Address ChangeC.ADB | C.ADB | |
| Correspondence Address ChangeC.ADB | C.ADB | |
| IFW TSS Processing by Tech Center CompleteTSSCOMP | TSSCOMP | |
| Application Dispatched from OIPEOIPE | OIPE | |
| Sent to Classification ContractorPGPC | PGPC | |
| Filing ReceiptFLRCPT.O | FLRCPT.O | |
| Cleared by OIPE CSRL194 | L194 | |
| IFW Scan & PACR Auto Security ReviewSCAN | SCAN | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Reference capture on IDSRCAP | RCAP | |
| Information Disclosure Statement (IDS) FiledM844 | M844 | |
| Request from applicant for the USPTO to retrieve the Priority DocumentPDREQUST | PDREQUST | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Initial Exam Team nnIEXX | IEXX |
11 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Lapsed due to failure to pay maintenance feeLapsedFP | FP | |
| Lapse for failure to pay maintenance feesLapsedPATENT EXPIRED FOR FAILURE TO PAY MAINTENANCE FEES (ORIGINAL EVENT CODE: EXP.); ENTITY STATUS OF PATENT OWNER: LARGE ENTITYLAPS | LAPS | |
| Information on status: patent discontinuationPATENT EXPIRED DUE TO NONPAYMENT OF MAINTENANCE FEES UNDER 37 CFR 1.362STCH | STCH | |
| Fee payment procedureMAINTENANCE FEE REMINDER MAILED (ORIGINAL EVENT CODE: REM.); ENTITY STATUS OF PATENT OWNER: LARGE ENTITYFEPP | FEPP | |
| Maintenance fee paymentMAFP | MAFP | |
| Fee paymentFPAY | FPAY | |
| Information on status: patent grantGrantedPATENTED CASESTCF | STCF | |
| Notice of allowance mailedORIGINAL CODE: MN/=.ZAAB | ZAAB | |
| Notice of allowance and fees dueORIGINAL CODE: NOAZAAA | ZAAA | |
| AssignmentAS | AS | |
| AssignmentAS | AS |
Numbers
- Publication
- 08108639
- Publication, DOCDB
- 8108639
- Publication, EPODOC
- US8108639
- Application
- 12379032
- Application, DOCDB
- 37903209
- Application, EPODOC
- US20090379032
Titles
- English
- Storage system, method for calculating estimated value of data recovery available time and management computer
Patent term adjustment
- A delay
- +383 daysthe office missed an examination deadline
- Applicant delay
- −31 days
- Net adjustment
- 352 days
Classification
- CPC, 3
- G06F11/2074
- G06F11/2069
- G06F11/2082
- IPC, 1
- G06F13 00
- USPC, 2
- 711162000
- 711151000