Information processing system
Abstract
[Task] In the process of relocating the data stored in the storage to another area, the data transfer between the areas is performed by the host, which puts a load on the host and the channel.
Solution.The host requests data transfer to the storage using the data transfer source / destination area information as a parameter. The storage internally transfers data from the transfer source disk device to the transfer destination disk device. When all the transfer of the target data is completed, the host changes the storage location of the data to the transfer destination area and registers it.
Term
Term ended
Projected expiry passed 28 February 2021, 5.6 years ago.
- Priority and filed
- Published
- Projected expiry
- Today
11 claims: 5 independent, 6 dependent
- 1【特許請求の範囲】 【請求項1】ホストコンピュータと、前記ホストコンピュータに接続され、複数のディスク装置を有する記憶装置とを有する情報処理システムであって、前記ホストコンピュータは、前記複数のディスク装置の物理的な記憶領域と論理的な記憶領域との対応関係の情報が登録される情報記憶手段と、前記複数のディスク装置のうち第一のディスク装置に記録されたデータを第二のディスク装置に移動する場合に、移動先の論理的な記憶領域及び前記情報記憶手段から、前記第二のディスク装置の物理的な記憶領域を示す情報を取得する取得手段と、前記取得手段によって取得された前記第二のディスク装置の物理的な記憶領域を示す情報及び移動対象となるデータが格納されている前記第一のディスク装置の物理的な記憶領域を示す情報を、前記記憶装置に転送する転送手段とを有し、前記記憶装置は、前記転送手段によって転送された前記情報のうち、前記第一のディスク装置の物理的な記憶領域を示す情報を使用して、移動元のデータを読み出す手段と、前記記憶手段に記憶された前記情報のうち、前記第二のディスク装置の物理的な記憶領域を示す情報を使用して、移動先のディスク装置に複写する複写手段と、を有することを特徴とする情報処理システム。
- 2【請求項2】前記記憶装置は、前記複写が終了したことを前記ホストコンピュータに通知する通知手段を有し、前記ホストコンピュータは、前記通知を受け取った後、前記情報記憶手段に登録されたディスク装置の物理的な記憶領域と論理的な記憶領域との対応関係を更新する手段を有することを特徴とする請求項1記載の情報処理システム。
- 3【請求項3】前記転送手段は、前記複数のディスク装置のうちの所定のディスク装置に前記情報を書き込む命令を発行する発行手段を含み、前記記憶装置は、前記発行手段によって発行された命令によって書き込まれた前記情報を当該記憶装置が読み出した場合に、前記読み出し手段及び前記複写手段を実行することを特徴とする請求項1又は2記載の情報処理システム。
- 4【請求項4】前記記憶装置は、前記複写手段によって複写されたデータに対してアクセスがある場合には、当該アクセスがあったことを記録するアクセス記録手段と、前記アクセス記録手段に記録された内容に基づいて、移動元のデータと移動先のデータとのデータの内容を一致させる手段とを有することを特徴とする請求項1、2又は3記載の情報処理システム。
- 5【請求項5】複数のディスク装置を有する記憶装置と接続される接続部と、前記複数のディスク装置の物理的な記憶領域と論理的な記憶領域との対応関係の情報が登録されるメモリと、中央演算部とを有し、前記中央演算部は、前記メモリから、データの移動先となる前記論理的な記憶領域に対応する第一のディスク装置の物理的な記憶領域を示す情報を取得する取得手段と、前記取得手段によって取得された前記第一のディスク装置の物理的な記憶領域を示す情報及び前記データが格納されている第二のディスク装置の物理的な記憶領域を示す情報を、前記接続部を介して前記記憶装置に転送する転送手段と、前記データの移動が終了した後、前記メモリに登録された前記複数のディスク装置の物理的な記憶領域と論理的な記憶領域との対応関係を更新する手段とを有することを特徴とする情報処理装置。
- 6【請求項6】前記取得手段は、前記メモリから論理的な記憶領域が割り当てられていない前記複数のディスク装置の物理的な記憶領域を検索する手段と、前記検索する手段によって検索された前記物理的な記憶領域を前記第一のディスク装置の物理的な記憶領域として取得する手段とを含むことを特徴とする請求項5記載の情報処理装置。
- 7【請求項7】前記転送手段は、前記記憶装置が有する複数のディスク装置のうち、所定のディスク装置に、前記情報をデータとして書き込む命令を発行するものであることを特徴とする請求項6記載の情報処理装置。
- 8【請求項8】ホストコンピュータに接続される接続部と、複数の記憶領域と、前記複数の記憶領域のうち、前記ホストコンピュータに使用されていない記憶領域の情報が登録されるメモリと、前記複数の記憶領域へのデータの入出力を制御する制御部を有し、前記制御部は、前記接続部から入力される情報に基づいて、前記メモリに登録された記憶領域のうちのいずれかを選択する手段と、前記選択された記憶領域に、他の前記記憶領域からデータを複写する複写手段とを有することを特徴とする記憶装置。
- 9【請求項9】前記接続部から入力される情報とは、所定の条件を満たすことを要求する情報であり、前記メモリには、複数の記憶領域の性質が登録され、前記制御部は、前記要求を満たす記憶領域を、前記メモリに登録された前記記憶装置の性質に基づいて検索する検索手段と、前記検索された記憶領域についての情報を前記接続部から出力する出力手段を有することを特徴とする請求項8記載の記憶装置。
- 10【請求項10】ホストコンピュータ及び前記ホストコンピュータに接続され、複数のディスク装置を有する記憶装置を有する情報処理システムにおいて、前記複数のディスク装置内でデータを再配置する方法であって、前記ホストコンピュータにおいて、再配置の対象となるデータが格納されている第一のディスク装置を特定し、再配置先となる第二のディスク装置の情報を取得し、特定した前記第一のディスク装置の情報及び取得した前記第二のディスク装置の情報を前記記憶装置に送信し、前記記憶装置において、送信された前記第一のディスク装置の情報に基づいて前記第一のディスク装置からデータを読み出し、送信された前記第二のディスク装置の情報に基づいて前記第二のディスク装置に前記第一のディスク装置から読み出したデータを格納し、前記読み出されたデータの格納が終了したら、前記ホストコンピュータに格納の完了を通知し、前記通知の後、前記ホストコンピュータにおいて、前記複数のディスク装置と論理的な記憶領域との対応関係を記録したテーブルを変更することを特徴とするデータの再配置の方法。
- 11【請求項11】ホストコンピュータ及び前記ホストコンピュータに接続され、複数のディスク装置を有する記憶装置を有する情報処理システムにおいて、前記複数のディスク装置内でデータを再配置するコンピュータプログラムであって、前記ホストコンピュータにおいて、再配置の対象となるデータが格納されている第一のディスク装置を特定し、再配置先となる第二のディスク装置の情報を取得し、特定した前記第一のディスク装置の情報及び取得した前記第二のディスク装置の情報を前記記憶装置に送信するプログラムと、前記記憶装置において、送信された前記第一のディスク装置の情報に基づいて前記第一のディスク装置からデータを読み出し、送信された前記第二のディスク装置の情報に基づいて前記第二のディスク装置に前記第一のディスク装置から読み出したデータを格納し、前記読み出されたデータの格納が終了したら、前記ホストコンピュータに格納の完了を通知するプログラムとから構成され、前記ホストコンピュータ側のプログラムは、前記通知の後、前記複数のディスク装置と論理的な記憶領域との対応関係を記録したテーブルを変更することを特徴とするデータの再配置を行うコンピュータプログラム。
Independent claims11
283 paragraphs in 1 section, as filed
Description: TECHNICAL FIELD [Detailed description of the invention]
【0001】
[Technical field to which the invention belongs]
The present invention relates to a host computer (host) and a computer system having a storage device connected to the host, and more particularly to a function of supporting the movement of data stored in the storage device of the computer system.
【0002】
[Conventional technology]
Generally, when constructing a computer system, the system is designed so that resources such as a network and a disk device do not become a bottleneck. In particular, an external storage device, which is slower than a processor or the like, tends to become a performance bottleneck, and various measures are taken in system design. One of them is optimization of the data storage format on the storage device. For example, data access performance can be improved by storing frequently accessed data in a higher-speed disk device or by distributing the data to a plurality of disk devices. In addition, when using a so-called RAID (Redundant Array of Independent Disk) device, if the RAID level (redundant configuration) is determined in consideration of the sequential nature of data access of the RAID device, the data suitable for the access characteristics can be obtained. Can be stored.
【0003】
From the viewpoint of system design, it is necessary to determine the capacity of the disk device allocated to each data in consideration of the data storage format used in the system. Specifically, the determination of the area size of the DB table in the database (DB) and the FS size in the file system (FS) corresponds to this. Normally, the amount of data handled increases with the operation of computer systems. For this reason, when designing a system, the rate of increase in the amount of data is predicted based on past performance in related operations, and a free area is secured to withstand the expected increase in the amount of data by the time when maintenance is possible. The capacity of the disk device is determined so that each data area is defined.
【0004】
As described above, when designing a system, it is important to combine a storage device and a data storage format in consideration of an improvement in data access performance and an increase in the amount of data. Logical Volume Manager (LVM) is a means to assist in determining this combination.
【0005】
The LVM is a collection of arbitrary subareas of an actual disk device and logically provided to the host as one volume (this is referred to as a logical volume and hereinafter referred to as "LV"). LVM manages LVs and also creates, deletes, and resizes (enlarges / reduces) LVs. LVM also has a mirroring function that multiplexes LVs and a striping function that divides LVs into small areas (stripes) and distributes them to multiple physical volumes (PV).
【0006】
When LVM is used, the user arranges the area for storing data such as DB table and FS on LV instead of PV. By doing so, it becomes easy to select or manage the data storage format. For example, by arranging FS on the LV, it is possible to distribute and arrange one disk device or FS, which is normally assigned to only one partition, to a plurality of disk devices. In addition, by expanding the LV as the file size increases, it is possible to expand (reconstruct) the FS with a minimum of effort.
【0007】
[Problems to be Solved by the Invention]
As business operations continue on a computer system, it may be necessary to review the data storage format. The triggers for the review were the review of the business model envisioned at the time of system design, such as changes in data access trends and access characteristics, and fluctuations in the amount of data that differed from the initial estimate, and the addition of disk devices and high-speed resources. It is conceivable that it is due to a change in the premise physical resource configuration such as replacement with, or that it inevitably occurs in the data management method such as fragmentation of LV or DB table continuation due to repeated expansion. .. In such a case, the performance of the system can be improved by reviewing the data storage format and rearranging the data.
【0008】
However, in order to rearrange the data stored in the storage device, in the conventional technique, data transfer with the intervention of the host is required.
【0009】
The procedure of data relocation processing when LVs distributed in a plurality of PVs are combined into one PV is shown below. (1) Secure an area on the PV for the capacity of the LV to be processed. (2) The host reads the data from the LV and writes it to the area of the new LV. (3) After repeating (2) and copying all the data, change the LV and PV mapping information.
【0010】
In this way, in order to reconfigure the LV, a large amount of data transfer of the entire LV occurs, so that a large amount of input / output (I / O) occurs in the PV of the data transfer source / destination. At the same time, the host and channel are also heavily loaded, which adversely affects the performance of ongoing processing of data in other LVs.
【0011】
Further, during such data relocation processing, it is necessary to suppress at least data update for access to the data to be relocated. For example, in the case of LV reconfiguration, normally, the LV is taken offline (unmounted on Unix) and online (mounted on Unix) after the reconfiguration is completed to suppress access to data during data relocation. Since the LV that is the target of data access is offline, the business that uses that LV will be suspended during the data relocation process.
【0012】
Therefore, when data relocation such as LV reconstruction is performed, the business of accessing the data must be performed during a time period during which the business can be interrupted for a predetermined period. Therefore, the maintenance work of the computer system is time-constrained.
【0013】
An object of the present invention is to reduce the load on the host and the channel due to data transfer from the movement source data area to the movement destination data area when the data stored in the storage device is rearranged.
【0014】
Another object of the present invention is to reduce the period during which access to data is disabled as much as possible by relocating the data, and to reduce the time during which the business using the data is interrupted.
【0015】
[Means for solving problems]
In order to achieve the above object, in the present invention, the information processing system including a host computer and a storage device connected to the host computer and having a plurality of disk devices has the following configuration. The host computer has a table in which information on the correspondence between the physical storage area and the logical storage area of a plurality of disk devices is registered. When the host computer moves the data recorded in the first disk device to the second disk device, the host computer acquires information indicating the physical storage area of the second disk device from the table. The host computer transfers the acquired information to the storage device. After the data movement is completed, the host computer updates the information registered in the table. Further, the storage device copies the data of the movement source to the disk device of the movement destination by using the information transferred from the host computer.
【0016】
With the above configuration, the host and channel load due to data transfer can be reduced by transferring the data on the storage device side when the data is rearranged.
【0017】
In addition, the data transfer process in the storage device is not simply realized by copying between areas, but a pair is formed in which the contents are temporarily synchronized between the data areas of the transfer source and the transfer destination, and the data is being transferred and the data is transferred. It is also possible to have a configuration in which all data updates made to the transfer source data area after the transfer is completed are reflected in the transfer destination.
【0018】
It is also conceivable to use a storage device instead of the host computer to manage the virtual storage area of the disk device.
【0019】
BEST MODE FOR CARRYING OUT THE INVENTION
FIG. 1 is a configuration diagram showing a configuration in a first embodiment of a computer system to which the present invention is applied.
【0020】
The computer system in this embodiment has a host 100 and a storage device 110. The host 100 and the storage device 110 are connected by a communication line such as a SCSI bus. Communication between the two takes place via a communication line.
【0021】
Host 100 has CPU 101, main memory 102, and channel 103, each of which is connected by an internal bus.
【0022】
CPU101 executes an application program (AP) such as a database. Memory allocation when the AP operates and input / output control to the storage device 110 are performed by the operating system (OS) operating on the CPU 101 and software related to the OS. LVM142 is also OS related software. The LVM142 collectively provides the storage area of the PV provided by the storage device 110 to the AP as an LV. In this embodiment, CPU 101 executes LVM 142. The LVM142 controls the LV by using control information such as the LV-PV correspondence information 141 described later.
【0023】
The main memory 102 stores execution object codes of OS-related software such as application programs, OS, and LVM, data used by each software, control information, and the like.
【0024】
3 and 4 are table diagrams of LV-PV correspondence information 141. LV-PV correspondence information 141 is information indicating PV (or LV corresponding to PV) corresponding to LV. In the table of LV-PV correspondence information 141, LV management information 300 and PV management information 400 are provided for each LV (PV).
【0025】
The LV management information 300 has entries for PV list 301, LE number 302, LE size 303, and LE-PE correspondence information 310. Information indicating the PV assigned to the LV is registered in the PV list 301. LV and PV are divided into small areas of the same size called LE (Logical Extent) and PE (Physical Extent), respectively. By assigning PE to LE, the degree of freedom of physical placement of LV is improved. The number of LEs included in the LV is registered in the number of LEs 302. Information indicating the size of LE is registered in LE size 303. The entry for LE-PE correspondence information 310 has an entry for PV name 312 and PE number 313 corresponding to the LE assigned to LE number 311. Information indicating the PE to which LE is assigned is registered in the LE-PE correspondence information 310.
【0026】
The PV management information 400 indicates the LV information assigned to each PV, which is the opposite of the LV management information.
【0027】
The LV management information 400 has entries for the LV list 401, the number of PEs 402, the PE size 403, and the PE-LE correspondence information 410. Information indicating the LV assigned to the PV is registered in the LV list 401. The number of PEs included in the PV is registered in the number of PEs 402. Information indicating the size of PE is registered in PE size 403. The PE-LE correspondence information 410 has an entry of the LV name 412 and the LE number 413 corresponding to the PE assigned to the PE number 411. Information indicating the LE to which the PE is assigned is registered in the PE-LE correspondence information 410.
【0028】
In addition to the above-mentioned information, the main memory 102 holds information necessary for accessing the PV. For example, as the route information for accessing the PV, the number of the connection channel 103, the number of the port 114 of the storage device 110, and the device number in the storage device 110 (hereinafter referred to as PV number) are stored.
【0029】
The channel 103 is a controller that controls input / output to / from the storage device 110 via a communication line. Channel 103 controls communication protocols such as transmission of request commands on the communication line, notification of completion report, transmission / reception of target data, and communication phase management. If the communication line is a SCSI bus, the SCSI adapter card corresponds to channel 103.
【0030】
The storage device 110 includes a port 114 that controls the connection with the host, a disk device 150, a storage control processor 111, a control memory 112, and a disk cache 113. In order to improve the availability of the storage device 110, each part of the storage device 110 is multiplexed. Therefore, even if a failure occurs in a certain part, it is possible to operate in a degenerate state by the remaining normal constituent parts.
【0031】
When the storage device 110 is a RAID system in which a plurality of disk devices 150 are connected, the plurality of disk devices 150 can be logically set to one or one by using the emulation by the logical physical correspondence management by the storage control processor 111. It can be recognized by the host 100 as a plurality of disk devices. However, in the present embodiment, for the sake of brevity, it is assumed that the PV accessed by the host 100, that is, the logical disk device in the storage device 110 has a one-to-one correspondence with the disk device 150.
【0032】
The storage control processor 111 receives access to the PV from the host 100, transfers data between the disk device 150 and the disk cache 113, transfers data between the disk cache 113 and the host 100, and logic of the disk device 150 in the storage device 110. It manages physical support and manages the area of disk cache 113.
【0033】
The disk cache 113 temporarily holds the write data from the host 100 and the read data from the disk device 150 until it is transmitted to the transfer destination. The data stored in the disk cache 113 is managed by a method such as LRU (Least Recently Used). The disk cache 113 is also used to write write data to the disk device 150 asynchronously with the host I / O request. Since these cache control methods are conventionally known techniques, description thereof will be omitted.
【0034】
The control memory 112 stores a table of various control information used by the storage control processor 111 to control input / output to / from the disk device 150. In the control information table, the cache management information 144 used for area allocation management of the disk cache 113, the storage device management information 143 used for managing the correspondence between the logical disk device and the disk device 150, and the area instructed by the host 100 are displayed. There is data transfer area information 145 used for managing the target range of data transfer processing between data transfer processing and the processing progress status.
【0035】
FIG. 5 is a table diagram of the data transfer area information 145. In the transfer source range information 501 and the transfer destination range information 502, information indicating the range of the data area to be transferred, which is instructed by the host 100, is registered. In the present embodiment, assuming a case where the data areas of the transfer source / destination are discontinuous, the transfer source range information 501 and the transfer destination range information 502 include PV numbers including the partial areas for each continuous partial area. , A list of start positions and sizes indicated by relative addresses in PV is registered. The total capacity of the transfer source / destination data area must be the same.
【0036】
Information indicating the amount of data transferred in the data transfer process is registered in the progress pointer 503. The progress of the data transfer process is managed using the information registered in the progress pointer 503. Since the synchronization state 504 and the difference bitmap 505 are not used in the present embodiment, the description thereof will be omitted.
【0037】
The operation of the CPU 101 and the storage control processor 111 in this embodiment will be described.
【0038】
When the user or the maintenance staff determines that the specific LV needs to be reconfigured from the information such as the allocation status of each LV to the PV, the user or the maintenance staff gives an instruction to reconfigure the LV. In this embodiment, as shown in FIG. 2, when lv0 is stored as lv0_0 and lv0_1 on two PVs, pv0 and pv1, lv0 is newly secured in the area inside pv2 lv0_0'and lv0_1. The case where the instruction to move to'is given will be described.
【0039】
The reconfiguration of the LV is realized by the cooperation of the data relocation process 131 operated by the CPU 101 and the command process 132 operated by the storage control processor 111 of the storage device 110.
【0040】
FIG. 6 is a flow chart of the data relocation process 131 performed by the CPU 101. The data relocation process 131 is executed when the user or the like instructs the relocation of the LV. In the data relocation process 131, the LV name to be relocated and the PV name of the relocation destination are acquired before the process is executed.
【0041】
CPU101 takes the target LV offline in order to suppress access to the target LV. LV offline can be achieved, for example, by unmounting the device (LV) if the OS is UNIX® (step 601).
【0042】
The CPU 101 refers to the LV management information 300 of the LV-PV correspondence information 141 to obtain the PV and PE in which the data of the target LV is stored. CPU101 calculates the target LV size from the number of LEs 302 and the LE size 303. If all or part of the target LV is already stored in the transfer destination PV, the CPU 101 does not transfer the part stored in the transfer destination PV. However, in the following description, it is assumed that there is no part excluded from the transfer target (step 602).
【0043】
CPU101 secures PE for the transfer target area size of the target LV on the transfer destination PV. Specifically, the CPU 101 refers to the PE-LE correspondence information 410 of the PV management information 400, obtains the PE unallocated to the LE, and secures the PE for the capacity required as the transfer destination. The PE may be secured only by storing the transfer destination PV and PE. However, if there is a possibility that PE will be secured for another purpose by other processing, secure PE to be used for transfer in advance. Specifically, a configuration change processing exclusive flag indicating that the change of the PE allocation of PV is prohibited for a certain period is set for each PV, a configuration change processing exclusive flag is set for each PE instead of each PV, or data transfer is completed. Before doing so, a method such as changing the target PE item of the PE-LE correspondence information 410 so that it has already been assigned to the transfer source LV can be considered (step 603).
【0044】
After securing the PV of the transfer destination, the CPU 101 divides the PV area of the transfer source into several subregions, and issues a data transfer processing request for each subregion to the storage device 110. The request for data transfer processing to the storage device 110 is made by using a newly added dedicated command for the request for data transfer processing, instead of the standard input / output command prepared by the existing protocol. The PV area is divided by dividing the transfer source area in order from the beginning with an appropriate size. An appropriate size for division is determined from the time required for the transfer processing of the requested data to the storage device 110 and the time that can be tolerated as the response time to the request at the requesting host 100. Since the transfer processing request command is issued for each PV, if the transfer source LV is distributed and arranged on multiple PVs, it is necessary to divide the partial area so that it does not span the two PVs. is there. The command of the transfer processing request includes the start address and size of the transfer source PV, the PV number of the transfer destination, the start address in the PV, and the data size. After sending the data transfer request command to the storage device 110 via the channel 103, the CPU 101 waits for the completion report from the storage device 110 (step 604).
【0045】
Upon receiving the completion report of the data transfer processing request command transmitted in step 604, the CPU 101 checks whether the data transfer of the entire target area has been completed for the transfer source LV. If there is a partial area for which data transfer has not been completed, CPU101 returns to the process of step 604 (step 605).
【0046】
When all the data in the target area is transferred, the CPU 101 changes the LV-PV correspondence information 141 so that the LV to be transferred is associated with the transfer destination PV. Specifically, the CPU 101 changes the information registered in the PV list 301 of the LV management information 300 to the information indicating the transfer destination PV, and is registered in the PV name 312 and the PE number 313 of the LE-PE correspondence information 310. Change the information so that it indicates the PE of the transfer destination PV. The CPU 101 adds the LV to be transferred to the LV list 401 of the PV management information 400 of the transfer destination PV, and corresponds the LV name 412 and LE413 number of the PE-LE correspondence information 410 to the LE of the transfer destination LV. To change. The CPU 101 deletes the transfer target LV from the LV list 401 of the PV management information 400 of the transfer source PV, and changes the transfer source PE item of the PE-LE correspondence information 410 to LV unassigned (step 606).
【0047】
After that, the CPU 101 releases the offline of the LV to be transferred and ends the process (step 607).
【0048】
FIG. 7 is a flow chart of command processing 132 executed by the storage device 110. The command processing 132 is executed when the storage device 110 receives the command of the host 100.
【0049】
The storage device 110 checks the type of the processing request command issued from the host 100 to the disk device 150 (step 701).
【0050】
If the command is a data transfer processing request, the storage device 110 activates the copy processing 133 and waits for its completion (step 702).
【0051】
When the command is a read request, the storage device 110 checks whether or not the target data exists in the disk cache 113 in response to the read request. If necessary, the storage device 110 secures a cache area, reads the target data from the disk device 150 into the secured cache area, and transfers the data to the host 100 (steps 703 to 707).
【0052】
When the command is a write request, the storage device 110 secures the cache area of the disk cache 113 and temporarily writes the data received from the host 100 to the cache area. Then, write to the disk device 150 (steps 709 to 712).
【0053】
The storage device 110 reports the completion of request processing to the host 100 and ends the processing (step 708).
【0054】
FIG. 8 is a flow chart of the copy process 133 performed by the storage device 110. The copy process 133 is executed when the storage device 110 receives the command of the data transfer request.
【0055】
The storage device 110 that has received the data transfer request checks the validity of the specified information of the transfer source / destination data area for the data transfer request. Specifically, it checks whether the size of the transfer source / destination data area is the same, and whether it has already been set as the transfer source / destination data area of another data transfer request. If the specified information is not valid, storage 110 reports an error to host 100 (step 801).
【0056】
If no error is found, the storage device 110 allocates and initializes an area for storing the data transfer area information 145 corresponding to the data transfer processing request on the control memory 112. Specifically, the transfer source range information 501 and the transfer destination range information 502 are set based on the information included in the received data transfer processing request, and the progress pointer 503 is set to 0, which is the initial value (step 802). ..
【0057】
After the setting is completed, the storage device 110 sequentially reads the data to the disk cache 113 from the beginning of the data area of the transfer source disk device 150 to be transferred, and writes the read data to the transfer destination disk device 150. The amount of data in one data transfer at this time is preferably a data area having a size large to some extent in consideration of the head positioning overhead in the disk device 150. However, if the amount of data at one time is too large, it may adversely affect other processes for accessing data stored in another disk device 150 connected to the same bus as the disk device 150. Therefore, the amount of data in one data transfer needs to be determined in consideration of the processing speed expected for the copy process 133 and the influence on other processes (steps 803 to 804).
【0058】
After finishing writing to the transfer destination in step 804, the storage device 110 updates the contents of the progress pointer 503 according to the transferred data capacity.
【0059】
The storage device 110 refers to the progress pointer 503 and checks whether the copy process is completed for all the target data. If there is a part where the copy process is not completed, the process returns to the process of step 804 (step 805).
【0060】
When the copy process is completed, the storage device 110 reports the completion of the copy process to the command process 132, and ends the process (step 806).
【0061】
According to this embodiment, the host 100 only needs to give an instruction for copy processing, and the storage device performs the actual data transfer, so that the load on the host, the network, and the like is reduced.
【0062】
FIG. 9 is a configuration diagram of a computer system to which the second embodiment of the present invention is applied. The present embodiment is different from the first embodiment in that the synchronization state 504 and the difference bitmap included in the data transfer area information 145 are used and the command volume 900 is provided. Hereinafter, only the part peculiar to the second embodiment will be described.
【0063】
In the synchronization state 504, information indicating the synchronization pair state of the transfer destination / original data area of the data transfer process is registered. Possible values of the synchronized pair state include "unpaired", "pairing", and "paired". The "pair forming" state indicates a state in which data transfer processing is being executed from the transfer source to the transfer destination between the data areas instructed to transfer. The "paired" state means that the copy process between the data areas is completed and a synchronous pair is formed. However, if the data in the transfer source data area is updated in the "pair forming" state, there may be a difference between the data areas of the synchronous pair even in the "pair forming" state. The "unpaired" state means that the data transfer between the data areas is not instructed, or the synchronization pair is released by the instruction of the host 100 after the data transfer is completed. However, in this state, since the data transfer process does not exist or has been completed in the first place, the data transfer area information 145 is not secured on the control memory 112. Therefore, the two actually set to the synchronization state 504 are "pairing" and "pairing completed".
【0064】
The difference bitmap 505 indicates whether or not data is updated in the copy source data area during pair formation and in the paired state. In order to reduce the amount of information, the entire data area of the disk device 150 is divided into, for example, a small area of a specific size of 64 KB, and the small area unit and one bit of the difference bitmap 505 have a one-to-one correspondence. The difference bitmap 505 records whether or not data has been updated for a small area.
【0065】
Similarly, in allocating the disk cache, the disk device 150 is often divided into small areas and the cache is allocated for each small area to simplify the cache management. In this case, the bitmap can be easily set by associating one bit of the difference bitmap 505 with one or more small areas which are cache allocation units.
【0066】
Special processing requests (data transfer processing requests, etc.) that are not included in the standard protocol are written to the command volume 900 as data. In the first embodiment, a dedicated command is added as a command for requesting data transfer processing to the storage device 110. In the present embodiment, the data transfer request is issued to the storage device 100 by writing the data transfer request as data to the command volume 900 by using a normal write request command.
【0067】
When the storage control processor 111 receives the write request for the command volume 900, it analyzes the write data as a processing request and starts the corresponding processing. If the response time is short enough that the response time does not matter even if the requested process is executed as an extension of the write request, the storage control processor 111 executes the request process and reports the completion together with the write request. When the execution time of the required processing is long to some extent, the storage control processor 111 once reports the completion of the write processing. Host 100 periodically checks whether the request processing is completed.
【0068】
Next, the operations of the CPU 101 and the storage control processor 111 in this embodiment will be described.
【0069】
Similar to the first embodiment, the data relocation process 131 operated by the CPU 101 and the command process 132 operated by the storage control processor 111 of the storage device 110 cooperate to reconfigure the LV.
【0070】
FIG. 10 is a flow chart of the data rearrangement process 131 of this embodiment.
【0071】
Since steps 1001 and 1002 are the same processes as steps 602 and 603 of FIG. 6, description thereof will be omitted.
【0072】
The CPU 101 collectively issues a request for data transfer processing of the entire PV area corresponding to the transfer target area of the LV to the storage device 110. For the transfer request included in the write data to the command volume 900, the range information of the entire PV area corresponding to the LV target area (list of position information (each PV number, start address, size) of each partial area) and the transfer destination. Contains parameters such as PV area range information. After issuing the write request command to the storage device 110, the CPU 101 waits for the completion report from the storage device 110 (step 1003).
【0073】
Upon receiving the completion report, CPU101 waits for the specified time to elapse (step 1004).
【0074】
The CPU 101 issues a request for referring to the synchronous pair status of the data transfer area to the storage device 110, and waits for its completion. To refer to the synchronous pair state, specifically, the CPU 101 issues a write request to write data including a preparation request for the synchronous pair state to the command volume 900. After receiving the completion report from the storage device 110, the CPU 101 issues a read request to the command volume 900 (step 1005).
【0075】
CPU101 determines whether or not the obtained synchronous pair state is "paired". When the synchronous pair state is "paired", the CPU 101 performs the process of step 1007. If the synchronous pair state is not "paired", CPU101 returns to the process of step 1004 and waits for the synchronous pair state to change (step 1006).
【0076】
The CPU 101 then takes the LV offline (step 1007), as in step 601.
【0077】
The CPU 101 uses the command volume 900 to issue a release request for the synchronization pair formed in the transfer source area and the transfer destination area of the data transfer (step 1008). When the storage device 110 reports the completion of the write request command that transfers the synchronous release request to the command volume 900, the CPU 101 refers to the synchronous pair status of the data area in the same manner as in step 1004 (step 1009).
【0078】
If the acquired synchronous pair state is not "pair not formed" (step 1010), the CPU 101 waits for the specified time to elapse (step 1011), returns to the process of step 1009, and re-enters the contents of the synchronous state 504. get. If the synchronization state 504 is "pair not formed", the CPU 101 performs the processing after step 1012. The CPU 101 updates the LV-PV correspondence information 141 of the LV, brings the LV online, and completes the LV relocation process.
【0079】
FIG. 11 is a flow chart of command processing 132.
【0080】
The storage device 110 determines whether or not the target of the command received from the host 100 is the command volume 900 (step 1101).
【0081】
If the target of the command is the command volume 900, the storage device 110 starts the copy process and waits for its completion (step 1102).
【0082】
If the target of the command is not the command volume 900, the storage device 110 determines the type of the command. Transition to steps 1104 and 1110 depending on whether the command is read or write (step 1103). When the command type is read, the storage device 110 performs the same read processing as in steps 703 to 707 of FIG. 7 (steps 1104-1108).
【0083】
When the command type is write, the storage device 110 performs the same write process as in steps 709 to 712 of FIG. 7 (steps 1110 to 1113).
【0084】
Steps 1114 and 1115 are parts specific to the present embodiment, and are processing parts when another processing of the host 100 accesses the data of the LV being transferred during the data transfer processing.
【0085】
The storage device 110 determines whether or not the data targeted for writing includes a data area registered as a data transfer area (step 1114). When the registered data area is included, the storage device 110 identifies the updated part of the data transfer area and sets the difference bitmap corresponding to the updated part (step 1115).
【0086】
Storage 110 reports the completion of request processing to host 100 (step 1109).
【0087】
FIG. 12 is a flow chart of the copy process 133.
【0088】
The storage device 110 separates the types of commands sent to the command volume 900 (step 1201).
【0089】
When the command type is write, the storage device 110 analyzes the contents of the write data to the command volume 900 and determines the validity of the processing request contents and the transfer source / destination data area range specification (step 1202). If there is a problem with validity, the storage device 110 reports the error to the higher-level processing and interrupts the processing (step 1203). If there is no problem, the normal termination is reported, and the storage device 110 determines the type of the processing request sent as write data (step 1204).
【0090】
When the processing request is data transfer, the storage device 110 performs the same data transfer processing as in steps 802 to 805 of FIG. 8 (steps 1205 to 1208). However, in the initialization process of the data transfer area information 145 in step 1205, the synchronization state 504 is set to "pair forming" and the difference bitmap 505 is all cleared to 0. When the data transfer process is completed, the storage device 110 changes the synchronization state 504 to "paired" and ends the process (step 1209).
【0091】
When the processing request is unpaired, the storage device 110 refers to the data transfer area information 145 of the data transfer area pair to be released, and checks whether the difference bitmap 505 is set to On (step). 1210). If there is a bit that is O in the difference bitmap 505, that is, there is a part that is not synchronized at the transfer source / destination in the data transfer area, the unreflected data is transferred to the transfer destination data area (step 1212). The storage device 110 returns to step 1210 and redoes the check of the difference bitmap 505.
【0092】
When synchronization is achieved in the data transfer area of the transfer source / destination, the storage device 110 clears the data transfer area information 145 of the data area, releases the allocation of the area of the control memory 112 that stores the data transfer area information 145, and deallocates the area. End the process (step 1213).
【0093】
If the processing request is pair state ready, storage 110 prepares the requested synchronous pair state of the data transfer area (step 1214). If the data transfer area still exists and the memory area is allocated on the control memory 112, the value of the synchronization state 504 is used as the synchronization pair state. If the data transfer area does not exist and the memory area is not allocated, "pair not formed" is used as the synchronous pair state. The synchronous pair state prepared in step 1214 is transferred as the read data of the read request issued to the command volume 900 (steps 1215, 1216).
【0094】
According to this embodiment, it is possible to relocate the LV without imposing a burden on the host or the like after shortening the time for taking the LV offline as compared with the first embodiment.
【0095】
The third embodiment will be described.
【0096】
The system configuration of the third embodiment is basically the same as that of the first and second embodiments. However, in the present embodiment, the storage device 110 manages the storage area of the disk device 150 that is not used by the host 100. Then, the storage device 110 gives instructions such as the conditions of the data area used as the PV of the transfer destination, for example, the area length, the logical unit number in which the area is stored, the number of the connection port 114 used for connection, and the disk type. Or receive from maintenance staff. The storage device 110 selects a data area that matches the user's instruction from the data areas that are not used by the host 100, and presents the data area to the user or the maintenance staff. In this respect, it differs from the first and second embodiments. Hereinafter, a method of presenting information in the storage device 110 will be described.
【0097】
The storage device 110 holds and manages a list of numbers of disk devices 150 that are not used by the host 100 as unused area management information.
【0098】
Upon receiving the instruction for securing the area from the user or the maintenance staff, the storage device 110 performs the following processing.
【0099】
According to the instruction for securing the area, the storage device 110 searches the unused area management information held and selects the unused disk device 150 that satisfies the conditions (step 1-1).
【0100】
The storage device 110 reports the selected disk device 150 number to the user or maintenance personnel (step 1-2).
【0101】
When the user or the maintenance person obtains the number of the unused disk device 150 from the storage device 110, the user or the maintenance person instructs the data transfer according to the following procedure. The user or maintenance personnel sets the OS management information of the unused disk device 150. For example, on UNIX OS, define a device file name for an unused disk device 150 (step 2-1).
【0102】
The user or maintenance personnel defines the disk device 150 in which the OS management information is set as PV so that it can be used by LVM (step 2-2).
【0103】
The user or maintenance personnel designates the newly defined PV as the transfer destination and instructs the storage device 110 to transfer data according to the present invention (step 2-3).
【0104】
When the data transfer is completed, the user or maintenance personnel instructs to update the LV-PV correspondence information 141 (step 2-4).
【0105】
A storage device 110 such that the system to which the present embodiment is applied presents a logical disk device composed of a part or a whole set of storage areas of a plurality of disk devices 150 to the host 100, such as a RAID device. Consider the case where. In this case, the unused area management information is composed of a disk device 150 number to which each unused area belongs, a start area, and a list of area lengths.
【0106】
The storage device 110 secures an unused area by the following procedure.
【0107】
The storage device 110 searches for an unused area of the disk device 150 that meets the conditions of the area reservation instruction (step 3-1).
【0108】
If the capacity for the instruction cannot be secured, the storage device 110 notifies the user that the capacity cannot be secured, and ends the process (step 3-2).
【0109】
When one disk device 150 satisfying the instruction condition is found, the storage device 110 confirms the capacity of the unused area belonging to the disk device 150, and if there is an instruction area, secures the instruction area. Specifically, the storage device 110 cancels the registration of the reserved area from the unused area management information. If the unused area capacity of the disk device 150 is less than the specified amount, reserve the entire unused area, search for another disk device 150, and secure the unused area that meets the conditions (step 3-3). ).
【0110】
The storage device 110 repeats the process of step 3-3 until a data area having a capacity of the indicated amount is secured (step 3-4).
【0111】
The storage device 110 defines a logical disk device composed of the reserved area by registering the reserved area in the logical-physical conversion table of the storage device (step 3-5).
【0112】
The storage device 110 reports the defined logical disk device information to the user or the like (step 3-6).
【0113】
As the instruction for securing the unused area to the storage device 110, a dedicated command may be used as in the first embodiment, or a command may be written to the command volume as in the second embodiment. A service processor connected to the storage device 110 may be provided for maintenance, and a command may be issued from the service processor.
【0114】
It is also possible to script a series of processes from step 1-1 to step 3-4. In that case, the user or the maintenance staff can select the transfer destination data area and automatically execute the data transfer by specifying the transfer destination data area selection condition in more detail. As the condition of the transfer destination data area, the continuity in the storage area of the storage device 110, the physical capacity of the disk device 150 to be stored, the access performance such as the head positioning time and the data transfer speed can be considered. In the case of a RAID device, physical configuration conditions such as RAID level can also be considered as conditions for the transfer destination data area. It is also possible that the specific LV does not share the physical configuration, namely the disk device 150 or the internal path to which the disk device 150 is connected and the storage control processor 111.
【0115】
As a modification of the third embodiment, a method of instructing data transfer is also conceivable in which the user specifies only the selection conditions of the transfer source LV and the transfer destination area without specifying the transfer destination area. In this case, the storage device 110 selects the transfer destination area according to the selection conditions, and transfers the data to the newly generated logical disk device. When the data transfer is completed, the storage device 110 reports the completion of the transfer and the area information selected as the transfer destination to the host 100. Upon receiving the report, the host 100 completes the movement of the LV to the reported transfer destination logical disk device by the processes of steps 2-1, 2-2 and 2-4. At this time, in step 2-2, PV needs to be defined while the data in the logical disk device remains valid.
【0116】
The present invention is not limited to the above embodiment, and many modifications can be made within the scope of the gist thereof.
【0117】
In each embodiment, the logical disk device provided to the PV and the host 100 has a one-to-one correspondence with the actual disk device 150, but the PV is configured by RAID such as RAID 5 level in the storage device 110. You may. In that case, the host 100 issues an I / O to the logical disk device provided by the storage device 110. The storage control processor 111 converts the I / O to the logical disk device into the I / O to the disk device 150 by the logical physical conversion.
【0118】
In each embodiment, the relocation of the LV managed by the LVM is adopted as an example of the relocation of the data, but the process of continuing (garbage collection) the PE to which the LV is not assigned and the DB managed by the DBMS are adopted. The present invention can also be applied to other data rearrangements, such as table rearrangement processing.
【0119】
Further, regarding the data transfer processing request method between the host 100 and the storage device 110, the method of using the dedicated command of the first embodiment and the method of using the command volume 900 of the second embodiment are both embodiments. It can be realized even if it is replaced with.
【0120】
In each embodiment, it is assumed that there is one transfer destination PV, but a plurality of PVs may be used. In that case, it is necessary to specify how to distribute the transfer source data for each of the plurality of PVs. As a method of distributing to a plurality of PVs, when distributing evenly to each PV, it is conceivable to pack as much free space as possible in the specified order of PVs. In addition, in the case of even distribution, further, when the data is continuous in each PV, the data is divided in a specific size, and the divided data is stored in each PV in order like striping in RAID. Be done.
【0121】
The user or maintenance personnel may allocate a continuous area to the PE as a transfer destination in advance and then pass it to the data relocation process 131 as a parameter, or perform garbage collection of the free PE described above in the data relocation process.
【0122】
In each embodiment, it is assumed that access to the transfer destination data area does not occur. That is, the storage device 110 does not consider the access suppression to the transfer destination data area, and when the transfer destination data area is accessed, the data reference / update is performed to the access target area as it is. However, in case there is no guarantee of access suppression on the host 100, a configuration is also conceivable in which the storage device 110 rejects I / O for the data area registered as the transfer destination data area of the data transfer area. On the contrary, without waiting for the completion of the data transfer process, the host 100 updates the LV-PV correspondence information 145 to the state after the rearrangement, and accepts the access to the transfer target LV by the transfer destination PV. Good. In this case, forming a pair of data transfer areas in the storage device 110 and copying data for synchronization is the same as in the second embodiment. However, it is necessary to refer to the data in the transfer source area for a read request to the transfer destination area and reflect the data in the transfer source area for a write request to the transfer destination area.
【0123】
According to this embodiment, the LV relocation process can be performed only by taking the LV offline for a shorter time than usual, and the availability of the system can be improved.
【0124】
[Effect of the invention]
According to the computer system of the present invention, when the data stored in the storage device is moved to another area, the data transfer process is performed in the storage device, so that the load on the host and the channel can be reduced. ..
【0125】
Further, according to the computer system of the present invention, it is possible to receive access to the data during data transfer in the data rearrangement. As a result, it is possible to shorten the downtime of the business of accessing the target data in the data relocation.
[Simple explanation of drawings]
[Figure 1]
It is a block diagram of the computer system which is the object of the 1st Embodiment of this invention.
[Figure 2]
It is a schematic diagram of the data rearrangement processing assumed by this invention.
[Fig. 3]
It is a block diagram of LV management information in this invention.
[Fig. 4]
It is a block diagram of PV management information in this invention.
[Fig. 5]
It is a block diagram of the data transfer area information in this invention.
[Fig. 6]
It is a flow chart of the data rearrangement processing in the 1st Embodiment of this invention.
[Fig. 7]
It is a flow chart of the command processing in the 1st Embodiment of this invention.
[Fig. 8]
It is a flow chart of copy processing in 1st Embodiment of this invention.
[Fig. 9]
It is a block diagram of the computer system in 2nd Embodiment of this invention.
[Fig. 10]
It is a flow chart of the data rearrangement processing in the 2nd Embodiment of this invention.
[Fig. 11]
It is a flow chart of the command processing in the 2nd Embodiment of this invention.
[Fig. 12]
It is a flow chart of copy processing in 2nd Embodiment of this invention.
[Explanation of symbols]
100 ... host, 101 ... CPU, 102 ... main memory, 103 ... channel, 110 ... storage device, 111 ... storage control processor, 112 ... control memory, 113 .. .Disk cache, 141 ... LV-PV correspondence information, 143 ... Storage control management information, 144 ... Cache management information, 145 ... Data transfer area information, 150 ... Disk device, 900 .. .Command volume
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| JP2005202495A | Cited by | Japan | Examiner |
| JP2007537522A | Cited by | Japan | Examiner |
| US8738869B2 | Cited by | United States of America | Applicant |
| JP4750040B2 | Cited by | Japan | Examiner |
| JP2008242736A | Cited by | Japan | Examiner |
| JP4704463B2 | Cited by | Japan | Examiner |
| JP4843604B2 | Cited by | Japan | Examiner |
| US8516204B2 | Cited by | United States of America | Applicant |
| WO2004104845A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| JP2007018455A | Cited by | Japan | Search report |
| JP2007516523A | Cited by | Japan | Examiner |
| WO2007135731A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| JP2008129723A | Cited by | Japan | Examiner |
| JP2007122716A | Cited by | Japan | Search report |
| JPWO2007135731A1 | Cited by | Japan | Examiner |
| JP2007115264A | Cited by | Japan | Examiner |
4 members in 2 offices
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 2001053472 | Japan | A | |
| JP20010053472 | – | – | – |
Members4
| Document | Office | Kind | |
|---|---|---|---|
| JP2002259172AThis record | Japan | A | |
| US2002144076A1 | United States of America | A1 | |
| US6915403B2 | United States of America | B2 | |
| JP4105398B2 | Japan | B2 |
16 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Notification of resignation of power of attorneyJAPANESE INTERMEDIATE CODE: A7424RD04 | RD04 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of appointment of power of attorneyJAPANESE INTERMEDIATE CODE: A7423RD03 | RD03 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Notification of change of attorneyJAPANESE INTERMEDIATE CODE: A7421RD01 | RD01 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 2002-259172
- Publication, DOCDB
- 2002259172
- Publication, EPODOC
- JP2002259172
- Application
- 53472
- Application, DOCDB
- 2001053472
- Application, EPODOC
- JP20010053472
Titles2
- Japanese
- 【発明の名称】情報処理システム
- English
- [Title of Invention] Information Processing System
Classification
- CPC, 3
- G06F3/0617
- G06F3/065
- G06F3/0689
- IPC, 3
- G06F12 00
- G06F12 10
- G06F3 06