Globally-spared distributed storage system
15 claims: 11 independent, 4 dependent
- 1データ記憶システムであって、 アクセス・コマンドを遠隔装置と記憶空間との間で渡すための、ネットワークにより前記遠隔装置に接続可能な仮想エンジンと、 前記記憶空間内の複数のインテリジェント記憶要素(ISE)であって、前記アクセス・コマンドを前記記憶空間の選択された論理的な記憶位置へ通信するために、前記仮想エンジンが、複数の通信経路を介して、各々の前記ISEに対してユニークにアドレスを指定することができ、前記ISEは、 前記記憶空間の記憶装置内で検出された故障に応答して、前記ネットワークを介して前記ISEと通信する前記遠隔装置または他の装置ではなく、前記ISEが前記記憶装置内で前記検出された故障について予測する予防的回復ステップを自発的に実行し、全て前記ISEに基づいて、予備の記憶容量が存在するか否かを決定し、前記ISEは更に、 前記アクセス・コマンドが同時に前記仮想エンジンと第1の論理的な記憶位置との間で前記複数の通信経路のうちの第1の通信経路を介して通信されている間に、 前記決定の結果に基づいて、 記憶されたデータを、前記第1の通信経路を介して前記仮想エンジンがアドレスを指定した前記第1の論理的な記憶位置から、前記複数の通信経路のうちの異なる第2の通信経路を介して前記仮想エンジンがアドレスを指定した第2の論理的な記憶位置に移す、前記ISEと、を備える前記データ記憶システム。
- 2各ISEは複数の回転可能なスピンドルを備え、各スピンドルはそれぞれ独立に動くアクチュエータに近接して記憶媒体を支持し、前記アクチュエータは記憶媒体との間でデータを記憶しまた検索する、請求項1記載のデータ記憶システム。
- 3各ISEは共通の密閉されたハウジング内に収められた複数のスピンドルおよび媒体を備える、請求項2記載のデータ記憶システム。
- 4各ISEは、前記遠隔装置に対して仮想化された前記記憶空間内の仮想記憶ボリュームを前記複数の媒体にマップして管理するプロセッサを備える、請求項2記載のデータ記憶システム。
- 5各ISEプロセッサは、データを冗長方式で記憶して故障に耐える方法でデータを記憶するために前記仮想記憶ボリュームにメモリを割り当てる、請求項4記載のデータ記憶システム。
- 6各ISEプロセッサは複数の異なる独立ドライブの冗長アレイ(RAID)方式の選択された1つにデータを記憶する、請求項5記載のデータ記憶システム。
- 7各ISEプロセッサは異なるISE内に第2の仮想記憶ボリュームを前記記憶空間に割り当てる、請求項 4 記載のデータ記憶システム。
- 8仮想エンジンと複数のインテリジェント記憶要素(ISE)との間のアクセス・コマンドを処理するステップであって、各々の前記ISEは、記憶空間内の離散的 および物理的 な記憶位置を有し、前記仮想エンジンは、遠隔装置と前記記憶空間との間で前記アクセス・コマンドを渡すためにネットワークにより前記遠隔装置に接続され、前記仮想エンジンは、前記仮想エンジンと第1のISEとの間の第1の通信経路を介して前記アクセス・コマンドを前記第1のISE 内の第1の論理的な記憶位置 に渡すために複数の通信経路を介して、前記記憶空間内の各々の前記ISEに対してユニークにアドレスを指定することができる、前記処理するステップと、 前記処理するステップと同時に、 各々の前記ISEが個々に、前記記憶位置のうちの1つで検出された故障に応答して、前記ネットワークを介して前記ISEと通信する前記遠隔装置または他の装置ではなく、前記ISEが前記記憶位置内で前記検出された故障について予測する予防的回復ステップを自発的に実行し、全て前記ISEに基づいて、予備の記憶容量が存在するか否かを決定し、前記ISEは更に、前記決定の結果に基づいて、 前記第1の 論理的な記憶位置 から 前記仮想エンジンがアドレスを指定した異なる 第2の 論理的な記憶位置 に、前記仮想エンジンおよび前記第1のISEの間の第2の通信経路を介してデータを移すステップとを含む方法。
- 9前記処理するステップは前記 第1の ISEが、前記遠隔装置に対して仮想化された前記記憶空間内に対する仮想記憶ボリュームを内蔵の物理的記憶にマップして管理することを特徴とする、請求項 8 記載の方法。
- 10前記移すステップは、前記記憶 位置内 の 前記検出された 故障を検出すると前記ISEが第2の仮想記憶ボリュームを前記記憶空間に割り当てることを特徴とする、請求項 9 記載の方法。
- 11前記移すステップは、前記第2の仮想記憶ボリュームを前記 第1の ISEの内部の前記記憶空間に割り当てることを特徴とする、請求項 10 記載の方法。
- 12前記移すステップは、前記第2の仮想記憶ボリュームを前記 第1の ISEの外側の前記記憶空間に割り当てることを特徴とする、請求項 10 記載の方法。
- 13前記移すステップは、 異なる 第2のISEの前記記憶空間に前記第2の仮想記憶ボリュームを割り当てることを特徴とする、請求項 12 記載の方法。
- 14前記処理するステップは、前記記憶空間に対する前記仮想記憶ボリュームにメモリ 容量 を割り当てて、データを冗長方式で記憶して記憶装置の故障に耐える方法でデータを記憶することを特徴とする、請求項 9 記載の方法。
- 15アクセス・コマンドを処理するステップの処理速度に合わせてデータを移す速度を設定するステップを更に含む、請求項 8 記載の方法。
Independent claims15
56 paragraphs, as filed
The present invention generally relates to the field of distributed data storage systems, and more particularly to devices and methods for arranging storage capacities over a wide area in a distributed storage system for the purpose of moving data, but is limited thereto. is not it.
When the data transfer rate of the industrial standard structure could not keep up with the data access rate of the Intel (registered trademark) 80386 processor, computer networking became widespread rapidly. By strengthening the data storage capacity in the network, the local area network (LAN) has evolved into a storage area network (SAN). By enhancing the equipment in the SAN and the associated data processed by this equipment, users can benefit from enormous benefits, such as being able to process one size larger memory at a reasonable cost than the directly attached storage device. Has been realized.
Recent trends have shifted to a network-centric system of controlling data storage subsystems. That is, the system that controls the function of memory has been moved from the server to the network itself in the same way that memory has been strengthened. For example, host-based software delegates maintenance and management tasks to intelligent switches or to dedicated network storage service platforms. Since the device-based method is used, software running in the host is not required, and this method operates in a computer provided as a node in the company. In any case, the intelligent network scheme can centralize tasks such as storage allocation routines, backup routines, and fault tolerance design independently of the host.
<p> Transferring intelligence from the host to the network solves some of the problems, but it does not solve the inherent problem that it is generally difficult to change the presentation of virtual memory to the host. For example, it is necessary to transfer stored data in order to improve reliability, and it is necessary to add storage capacity in order to adapt to a growing network. In such cases, the host or network must be modified to identify the existence of new or altered storage space. What is needed is an intelligent data storage subsystem that voluntarily determines, allocates, manages, and protects each data storage capacity and presents that capacity to the network as a virtual storage space to adapt to wide-area storage demands. is there. This virtual storage space can be arranged as a multiple storage volume. In a distributed computing environment, such an intelligent storage device is used as a wide area arrangement and a wide area backup in the event of a failure. Embodiments of the present invention relate to this scheme.</p>
<p> Embodiments of the present invention generally relate to a distributed storage system having wide area placement capability. In certain embodiments, a data storage system is provided that includes a virtual engine that can be connected to the remote device via a network for passing access commands between the remote device and the storage space. The data storage system also has multiple intelligent storage elements that the virtual engine can uniquely address to pass access commands. The intelligent storage element moves data from the first intelligent storage element to the second intelligent storage element, independent of the access commands being passed between the virtual engine and the first intelligent storage element at the same time. ..</p><p> One embodiment provides a method of processing access commands between a virtual engine and an intelligent storage element while simultaneously transferring data from the intelligent storage element to another storage space. In certain embodiments, a data storage device is provided that includes a plurality of intelligent storage elements that can be individually addressed by the virtual engine and means for transferring data between the intelligent storage elements. These features and advantages that characterize the invention will become apparent by reading the detailed description below and with reference to the relevant drawings.</p>
FIG. 1 is an exemplary computer system 100 in which embodiments of the present invention are useful. One or more hosts 102 are networked to one or more network-attached servers (NAS) 104 via a local area network (LAN) and / or wide area network (WAN) 106. Preferably, the LAN / WAN 106 uses an Internet Protocol (IP) networking infrastructure to communicate over the World Wide Web. The host 102 resides in the server 104 and accesses applications that routinely require data stored in one or more of a number of intelligent storage elements (ISE) 108. Therefore, SAN110 connects server 104 to ISE108 so that it can access the stored data. The ISE108 includes a block of data storage capacity 109 that stores data by various selected communication protocols such as series ATA and Fiber Channel, including enterprise-class or desktop-class storage media.
FIG. 2 is a simple diagram of the computer system 100 of FIG. Host 102 exchanges information with each other and with a pair of ISE108 (indicated by A and B, respectively) over a network or structure 110. Each ISE108 contains a dual redundant controller 112 (indicated by A1, A2 and B1, B2). Preferably, the controller 112 acts on a data storage capacity 109, which is a set of data storage devices characterized as a redundant array (RAID) of independent drives. Since the controller 112 and the data storage capacity 109 preferably use a fault-tolerant arrangement, the various controllers 112 use parallel redundant links and at least part of the user data stored in the system 100 is of the data storage capacity 109. It is stored in a redundant format in at least one set.
In addition, the A-host computer 102 and A-ISE108 are physically at the first site, the B-host computer 102 and B-ISE108 are physically at the second site, and the C-host computer 102 is further at the second site. It may be on 3 sites. However, this is just an example and is not limited. All entities on a distributed computer system are connected by a type of computer network.
FIG. 3 shows an ISE 108 constructed according to an embodiment of the present invention. The shelf 114 defines a cavity for receiving and engaging the controller 112 and electrically connecting to the central plate 116. Shelf 114 is supported within a cabinet (not shown). The shelves 114 receive and engage a pair of multiple disk assemblies (MDA) 118 on the same side of the central plate 116. On the opposite side of the center plate 116, an emergency power supply, a dual battery 122, a dual AC power supply 124, and a dual interface module 126 are connected. Preferably, in the dual component, one or both of the MDA 118s operate at the same time, so that backup protection can be provided in the event of failure of one component.
FIG. 4 is an enlarged partial assembly disassembled isometric view of the MDA 118 constructed according to an embodiment of the present invention. The MDA 118 has an upper 130 and a lower 132, each supporting five data storage devices 128. The compartments 130, 132 align the data storage device 128 so that it can be connected to a common circuit board 134 having a connector 136 that engages the central plate 116 (FIG. 3). Cover 138 shields electromagnetic interference. This exemplary embodiment of the MDA118 is the subject of patent application 10 / 884,605, "Carrier Device and Method for a Multiple Disc Array". This has been transferred to the assignee of the present invention and is incorporated herein by reference. Another exemplary embodiment of the MDA is the subject of patent application 10 / 817,378 of the same title. This is also transferred to the assignee of the present invention, which is incorporated herein by reference. As will be described later, in another equivalent embodiment, the MDA 118 may be housed in a closed container.
FIG. 5 is an isometric view of an exemplary data storage device 128 in the form of a rotating media disk drive suitable for use in embodiments of the present invention. For the following description, a spindle that rotates with a moving data storage medium is used, but in another equivalent embodiment, a non-rotating medium device such as a solid-state memory device is used. The data storage disk 140 is rotated by a motor 142 to present the data storage position of the disk 140 to a read / write head (simply referred to as a "head") 143. The head 143 is supported by the tip of a rotary actuator 144 that moves the head 143 radially between the tracks inside and outside the disc 140. The head 143 is electrically connected to the circuit board 145 by a flexible circuit 146. The circuit board 145 receives and sends control signals that control the functions of the data storage device 128. The connector 148 is electrically connected to the circuit board 145, and connects the data storage device 128 and the circuit board 134 (FIG. 4) of the MDA 118.
FIG. 6 is a diagram of ISE108 constructed according to an embodiment of the present invention. Controller 112 works with Intelligent Storage Processor (ISP) 150 to manage the reliability of data integrity. The ISP150 may reside anywhere else in the controller 112, in the MDA118, or in the ISE108.
A mode of managed reliability involves creating a reliable data storage format such as a RAID scheme. For example, creating a relatively strong system for data storage by forming a system that selectively uses one of several different RAID formats, and the complexity of the software used to manage the MDA118. Firmware algorithms can be optimized to mitigate and recover from memory failure relatively quickly. These aspects of this multiple RAID format system are described in Patent Application 10 / 817,264, Storage Media Data Structure and Method. This has been transferred to the assignee of the present invention and is incorporated herein by reference.
Managed reliability may also include scheduling diagnostic and correction routines based on monitoring and using the system. The method of data recovery is to copy and reconstruct the data. The ISP150 incorporates with the MDA 118 to facilitate "self-healing" of the entire data storage capacity without losing data. These aspects of controlled reliability considered herein are disclosed in Patent Application 10 / 817,617, "Managed Reliability Storage System and Method". This has been transferred to the assignee of the present invention and is incorporated herein by reference. Other aspects of controlled reliability include the speed of response to predictive failure indications for predetermined rules. This is, for example, patent application 11 / 040,410, "Deterministic Preventive Recovery From a Predicted Failure in a Distributed Storage. It is disclosed in "System)". This has been transferred to the assignee of the present invention and is incorporated herein by reference.
FIG. 7 is a diagram showing an ISP circuit board 152 in which a pair of redundant ISP 150s reside. The ISP150 interfaces the data storage capacity 109 with the SAN structure 110. Each ISP150 may manage various storage services such as routing, volume management, and data movement and replication. The ISP 150 divides the ISP circuit board 152 into two ISP subsystems 154,156 coupled by the bus 158. The ISP subsystem 154 includes the ISP 150 indicated by "B". It is connected to the SAN structure 110 by link 160 and to the data storage capacity 109 by link 162. The ISP subsystem 154 also includes a policy processor 164 that runs a real-time operating system. The ISP 150 and the policy processor 164 communicate via bus 166, and both communicate with memory 168.
FIG. 8 is a diagram of an exemplary ISP subsystem 154 constructed according to an embodiment of the present invention. The ISP150 includes a number of functional controllers (170-180) that communicate with list managers 182,184 via a crosspoint switch (CPS) 186 message crossbar. In this way, the function controller (170-180) generates a CPS message according to a predetermined condition and sends this message to the list managers 182 and 184 through CPS186 to access the memory module and activate the ISP150. You can do it. Similarly, the response from list managers 182,184 may be sent to any of the functional controllers (170-180) via CPS186. The arrangement and related description of FIG. 8 is an example and does not limit the possible embodiments of the present invention.
Policy processor 164 can be programmed to perform the desired operation through ISP150. For example, policy processor 164 may communicate (ie, send and receive messages) with list managers 182,184 via CPS186. The response to policy processor 164 may act as an interrupt to signal a read of memory 168 registers.
FIG. 9 is a diagram demonstrating the great flexibility of the ISE 108 to communicate with the host 102 by one of a plurality of preselected communication protocols (such as FC, iSCSI, or SAS) by the intelligent controller 112. ISE108 may be programmed to check the fetch level of the host command and map the virtual storage volume to the physical storage 109 associated with the command accordingly.
For the purposes of the present invention, the term "virtual memory volume" means a logical entity that generally corresponds to the logical retrieval of physical storage. A "virtual storage volume" is, for example, a record in a continuously addressed address block or count-key-data structure within a fixed block structure (logical). May include the entity being treated. Virtual storage volumes may physically reside on more than one storage element.
FIG. 10 is a diagram showing the types of data management services that ISE108 may perform independently of any host 102. For example, RAID management is locally controlled for fault-tolerant data integrity to allow data striping in the desired number of data storage devices 128.<sub>1</sub>,128<sub>2</sub>,128<sub>3</sub>,...,128<sub>n</sub>You may go inside. Virtualization services may be controlled locally to allocate or deallocate memory capacity to logical entities. Application routines such as the managed reliability schemes described above and the movement of data between logical volumes within the same ISE108 may be controlled locally as well. For the purposes of this description and claim, the term "move" refers to the loss of primitive data as part of the completion of the move by moving the data from source to destination. This is the opposite of "copying" the data, and in the case of copying, the data is copied from the source to the destination. However, the destination has a different name.
FIG. 11 shows the implementation of the data storage system 100 in which the virtual engine 200 communicates with the remote host device 102 via SAN 106 in order to pass access commands (I / O commands) between the host device 102 and multiple ISE 108s. Shows the morphology. Each ISE108 has two ports 202,204 and 206,208 that the virtual engine 200 can uniquely address to pass access commands. In order to facilitate data movement without creating a data transfer bottleneck, the following describes how this embodiment transfers data between ISE108 at the same time, independent of the processing of host access commands. Further, by changing the data transfer rate when transferring data, the influence on the application performance of the system 100 can be optimized.
In ISE108-1, ISP150 creates logical volume 210 for physical data pack 212 of data storage 128. For illustration purposes, assume that 40% of the storage capacity of data pack 212 is allocated to logical disk 214 in logical volume 210. Also for illustration purposes, it is assumed that the data pack 212 and all other data packs below include eight data storage devices 128 and two spare data storage devices 128 for data storage. Furthermore, as can be seen from Figure 11, in ISE108-1, the other data pack 216 allocates 93% to logical disk 218, and in ISE108-2 data packs 220 and 222 allocate 30% and 40% to logical disks 224 and 226, respectively. Suppose you have assigned.
The virtual engine 200 created a logical volume 224 from the logical disk 214 and also created a logical disk 226 in response to a storage request from the host and mapped it to the host 102.
As mentioned above, the ISP150 within each ISE108 voluntarily initiates a deterministic preventive recovery step upon detecting a memory failure. For example, when ISE108-1 detects a failure of storage 128 in data pack 212, it immediately removes the failed storage 128 from the line. Data from the failed storage 128 is copied or rebuilt onto the first 10% of the reserve capacity of data pack 212 to restore redundancy. ISE108-1 then determines if any part of the failed storage 128 can be recovered by voluntary recalibration or remanufacturing.
Assuming that the first failed storage 128 is completely unrecoverable, and if the second storage 128 of ISE108-1 fails, it is also removed from the line and its data packed. Copy or rebuild on the last 10% of 212 reserve capacity.
Assuming that the second failed storage 128 is as unrecoverable as the first, and if the third storage 128 of the ISE108-1 fails, the ISE108-1 needs it. Reserve capacity allocation exceeds 20%. Continuing to operate the ISE108-1 under these conditions may result in partial loss of redundancy. Preferably, the ISE108-1 is planned to be shut down and replaced at an appropriate time to fully restore redundancy.
On the other hand, this embodiment considers that ISE108-1 allocates not only internally but also across different virtual storage volumes. In this case, preferably ISE108-1 checks the other data pack 216 inside for available space. However, in this case, data pack 216 is already 93% allocated and does not have the capacity needed to reserve data pack 212. However, both data packs 220 and 222 in ISE108-2 have the available capacity needed to reserve data pack 212.
FIG. 12 shows that the ISP 150 in ISE108-1 created an external logical disk 230 and moved data from the logical disk 214 associated with the lost storage device 128 to it. As you can see, moving data does not necessarily interrupt access command I / O between host 102 and ISE108-1. When the data movement is complete, the communication with the host 102 is temporarily frozen while changing the data path of the logical disk 230 to the virtual engine 200, and then the virtual engine 200 switches the I / O path as shown in FIG. Then, the I / O route is guided to the newly moved data in ISE108-2. This allows the data pack 212 to be replaced without interrupting the I / O service.
FIG. 14 is a flow chart of the steps of the wide area reserve method 250 according to the embodiment of the present invention. Method 250 starts at block 252 and ISE108 is processing in normal I / O mode. At block 254, determine if the last I / O command has been processed. When finished, this method ends. If not, control proceeds to block 256 to determine if ISE108 has detected a data pack failure. If the judgment of block 256 is no, normal I / O processing is continued in block 252. The same applies hereinafter.
If block 256 is determined yes, control proceeds to block 258 to determine if there is sufficient reserve capacity in the failed data pack. In the case of the above example where the storage device of data pack 212 has failed, block 258 examines data pack 212 itself. In other words, check "locally" for reserve capacity. If the verdict of block 258 is yes, at block 260 ISP150 assigns a local LUN, at block 262 transfers data from the failed data pack to the local LUN, and control returns to block 252.
If the verdict of block 258 is no, control proceeds to block 264 to see if there is spare capacity in another data pack in the same ISE108, in other words, if there is reserve capacity "inside". judge. If the verdict of block 264 is yes, at block 266 the ISP allocates an internal LUN, at block 268 moves data from the failed data pack to the internal LUN, and control returns to block 252.
If the block 264 verdict is no, control proceeds to block 270 to determine if there is spare capacity in the data pack in another ISE108, in other words if the reserve capacity exists "outside". .. If the verdict of block 270 is yes, at block 272 the ISP assigns an external LUN, at block 274 transfers data from the failed data pack to the external LUN, and control returns to block 252.
However, if the judgment of block 270 is no, there is no spare capacity, control proceeds to block 276, the operation of the data pack is shut down, and a maintenance plan is made. Control then returns to block 252.
Finally, FIG. 15 is similar to FIG. 4, but the plurality of data storage devices 128 and the circuit board 134 are housed in a closed container formed by the base 190 and the sealing cover 192 attached thereto. When the data storage device 128 forming the MDA 118A is hermetically engaged, there are various advantages such that the arrangement of the data storage device 128 does not change from the optimum arrangement selected in advance. Also, if the number, size, and type of data storage devices 128 can be clearly defined, such an arrangement allows the MDA118A maker to adjust the system for optimum performance.
Sealing the MDA118A also allows the producer to maximize the reliability and fault tolerance of the group of internal storage media, while at the same time providing little service for the life of the MDA118A. This is done by optimizing the multi-spindle drive. Design optimization reduces costs, improves performance, improves reliability, and generally extends the life of the data in the MDA118A. Furthermore, the design of the MDA118A itself almost eliminates rotational vibration, providing an environment with high cooling efficiency. This is the subject of the pending US patent application 11 / 145,404, "Storage Array with Enhanced RVI." It has been transferred to the assignee of this application. As a result, MDA118<u style="single">A</u>Internal storage media can be manufactured at low cost without compromising reliability, performance, or capacity. Sealing the MDA118A in this way eliminates single point of failure, eliminates rotational vibration and makes cooling efficiency almost complete. This allows the MDA118A to be designed for optimal disk media characteristics, reducing costs while increasing reliability and performance.
In summary, it provides a built-in ISE for distributed storage systems that include multiple rotatable spindles. Each spindle supports a storage medium in close proximity to an independently moving actuator, which stores and retrieves data with and from the storage medium. ISE also includes an ISP that maps virtual storage volumes to multiple media for use by remote devices in distributed storage systems.
In certain embodiments, the ISE has multiple spindles and media housed in a common sealed housing. Preferably, the ISP allocates memory in the virtual storage volume to store the data in a fault-tolerant manner such as RAID. In addition, the ISP can implement controlled reliability schemes during the data storage process, such as voluntarily initiating a deterministic preventive recovery step upon detection of a predicted memory failure. Preferably, the ISE is formed by a plurality of data storage devices, each of which is formed from a data storage medium of two or more disks and has a disk stack.
In another embodiment, the ISE comprises a plurality of built-in discrete data storage devices and an ISP that communicates with the data storage device, retrieves commands received from the remote device, and associates the associated memory with the distributed storage. Used for the system. Preferably, the ISP maps the virtual storage volume to multiple data storage devices for use by one or more remote devices in a distributed storage system. As before, multiple data storage devices and media may be housed in a common sealed housing. Preferably, the ISP allocates memory in the virtual storage volume in order to store the data in a fault-tolerant manner such as RAID. In addition, the ISP will voluntarily initiate a deterministic preventive recovery step within the data storage when it detects a predicted memory failure.
Another embodiment provides a distributed storage system that includes a host, a rear storage subsystem that communicates with the host over a network, and means for virtualizing internal storage capacity independently of the host.
The virtualization means may feature a plurality of discrete and individually accessible data storage units. The virtualization means may be characterized by mapping virtual blocks of storage capacity associated with a plurality of data storage units. The virtualization means may be characterized by hermetically confining a plurality of data storage units and associated controls. The means for virtualization may be characterized in that data is stored by a method such as a RAID method that can withstand a failure, although the means is not limited. The means of virtualization may be characterized by spontaneously initiating a deterministic preventive recovery step upon detection of a predicted memory failure. The means of virtualization may feature a multi-spindle data storage array.
For the purposes here, the term "virtualization means" does not explicitly consider previously attempted solutions that include system intelligence to map the data storage space somewhere other than their respective data storage subsystems. .. For example, "virtualization means" do not consider using a storage manager to control the functionality of the data storage subsystem, nor do they consider placing a manager or switch within a SAN structure or host.
Alternatively, this embodiment features a data storage system comprising a virtual engine connected to the remote device via a network for passing access commands between the remote device and the storage space. The data storage system also has multiple intelligent storage elements (ISEs) that the virtual engine can uniquely address to pass access commands. Independent of the access commands being passed between the virtual engine and the first ISE, the ISE moves data from the first ISE to the second ISE at the same time.
In certain embodiments, each ISE has multiple rotatable spindles, each spindle supporting a storage medium in close proximity to an independently moving actuator, which stores data to and from the storage medium. Search. Multiple spindles and media may be housed in a common sealed housing.
Each ISE has a processor for mapping and managing virtual storage volumes to multiple media. Each ISE processor allocates memory within virtual storage capacity to store data in a fault-tolerant manner, such as preferably one of a redundant array (RAID) scheme of different independent drives.
Each ISE processor may voluntarily take a deterministic preventive recovery step when it detects a memory failure. When doing this, each ISE processor may allocate a second virtual storage capacity when it detects a memory failure. In some embodiments, each ISE processor allocates a second virtual storage capacity within a different ISE.
This embodiment is further characterized as a method for processing access commands between a virtual engine and an intelligent storage element while simultaneously transferring data from the intelligent storage element to another storage space.
The steps to be processed may be characterized in that an intelligent storage element maps and manages virtual storage volumes to internal physical storage. Preferably, the transfer step is characterized in that the intelligent memory element spontaneously initiates a deterministic preventive recovery step upon detection of a memory failure.
The transfer step may be characterized in that the intelligent storage element allocates a second virtual storage volume when it detects a memory failure. In certain embodiments, the transferring step is characterized in that the virtual engine allocates a second virtual storage volume for physical storage with different addresses in the processing step. For example, the transfer step may be characterized by allocating a second virtual storage volume inside an intelligent storage element. Alternatively, the transfer step may be characterized by allocating a second virtual storage volume outside the intelligent storage element. That is, the transfer step may be characterized by allocating a second virtual storage volume within the second intelligent storage element.
The processing step may be characterized by allocating memory and storing data in a fault-tolerant manner. The processing step may also be characterized by moving the data transfer element and the storage medium with each other while transferring the data within a common sealed housing.
Alternatively, this embodiment features a data storage system that includes a plurality of intelligent storage elements that can be individually addressed by the virtual engine and means for transferring data between the intelligent storage elements. For the purposes of this description and claims, the term "means of transfer" with respect to the structures and their equivalents described herein provides data without interrupting normal I / O command processing associated with host access commands. It means moving from one logical unit to another.
As will be appreciated, many features and advantages of the various embodiments of the invention have been described in the previous description, along with details of the structure and function of the various embodiments of the invention. Is merely an example, and changes may be made in detail to the extent indicated in the broad general sense of the terms expressing the claims, especially with respect to the structure and arrangement of parts within the principles of the invention. For example, certain elements may be modified according to a particular processing environment as long as they do not deviate from the spirit and scope of the present invention.
Further, although the embodiments described herein relate to data storage arrays, as those skilled in the art will recognize, the claimed subject matter is not limited thereto and does not deviate from the spirit and scope of the present invention. Various other processing systems may be used as long as they are used.
This application is a partial continuation of US Application No. 11 / 145,403 filed on June 3, 2005 and transferred to the assignee of this application.
<figref num="1">FIG. 1 is a diagram of a computer system for which embodiments of the present invention are useful.</figref><figref num="2">FIG. 2 is a simple diagram of the computer system of FIG.</figref><figref num="3">FIG. 3 is an assembled / disassembled isometric view of an intelligent storage element constructed according to an embodiment of the present invention.</figref><figref num="4">FIG. 4 is a partially assembled and disassembled isometric view of the multiple disk array of intelligent storage elements of FIG.</figref><figref num="5">FIG. 5 is an exemplary data storage device used in the multiple disk array of FIG.</figref><figref num="6">FIG. 6 is a functional block diagram of the intelligent storage element of FIG.</figref><figref num="7">FIG. 7 is a functional block diagram of the intelligent storage processor circuit board of the intelligent storage element of FIG.</figref><figref num="8">FIG. 8 is a functional block diagram of the intelligent storage processor of the intelligent storage element of FIG.</figref><figref num="9">FIG. 9 is a functional block diagram representation of the command retrieval and associated memory mapping services performed by the intelligent storage elements of Figure 3.</figref><figref num="10">FIG. 10 is a functional block diagram of another exemplary data service performed by the intelligent storage element of FIG.</figref><figref num="11">FIG. 11 is a diagram showing a method of wide area reserve according to the embodiment of the present invention.</figref><figref num="12">FIG. 12 is a diagram showing a method of wide area reserve according to the embodiment of the present invention.</figref><figref num="13">FIG. 13 is a diagram showing a method of wide area reserve according to the embodiment of the present invention.</figref><figref num="14">FIG. 14 is a flow chart of steps for executing the method of wide area reserve according to the embodiment of the present invention.</figref><figref num="15">Figure 15 shows<u style="single">4</u>It is the same as the above, but is the figure which shows the thing which puts the data storage device and a circuit board in a closed container.</figref>
Code description
102 Remote device (host) 108 Intelligent Memory Element (ISE) 109 storage space 200 virtual engine
15 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| JP200699763A | Cites | Japan |
| JP2005266933A | Cites | Japan |
| JP200518193A | Cites | Japan |
| JP200648676A | Cites | Japan |
| JP2005115506A | Cites | Japan |
| US20060069864A1 | Cites | United States of America |
| JP200612156A | Cites | Japan |
| JP2005525619A | Cites | Japan |
| JP2005157739A | Cites | Japan |
| JP5314674A | Cites | Japan |
16 members in 2 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 11478028 | United States of America | – | |
| 47802806 | United States of America | A | |
| 47802806 | United States of America | A | |
| 2006478028 | – | – | – |
| US20060478028 | – | – | – |
Members16
| Document | Office | Kind | |
|---|---|---|---|
| US2006277380A1 | United States of America | A1 | |
| JP2006344218A | Japan | A | |
| US2006288155A1 | United States of America | A1 | |
| US2007011417A1 | United States of America | A1 | |
| US2007011425A1 | United States of America | A1 | |
| JP2008016028A | Japan | A | |
| JP2008027437A | Japan | A | |
| JP2008033921A | Japan | A | |
| US2008281830A1 | United States of America | A1 | |
| US7644228B2 | United States of America | B2 | |
| JP2010157257A | Japan | A | |
| JP4530372B2This record | Japan | B2 | |
| US7913038B2 | United States of America | B2 | |
| US7966449B2 | United States of America | B2 | |
| US7984258B2 | United States of America | B2 | |
| JP5150947B2 | Japan | B2 |
21 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Notification of appointment of power of attorneyJAPANESE INTERMEDIATE CODE: A7423RD03 | RD03 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Transfer to examiner for re-examination before appeal (zenchi)AppealJAPANESE INTERMEDIATE CODE: A911A911 | A911 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Decision of refusalJAPANESE INTERMEDIATE CODE: A02A02 | A02 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written permission of extension of timeJAPANESE INTERMEDIATE CODE: A602A602 | A602 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Written permission of extension of timeJAPANESE INTERMEDIATE CODE: A602A602 | A602 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Written permission of extension of timeJAPANESE INTERMEDIATE CODE: A602A602 | A602 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 |
Numbers
- Publication
- 4530372
- Publication, DOCDB
- 4530372
- Publication, EPODOC
- JP4530372B
- Application
- 173121
- Application, DOCDB
- 2007173121
- Application, EPODOC
- JP20070173121
Titles2
- Japanese
- 広域予備化した分散記憶システム
- English
- Wide area reserved distributed storage system
Classification
- IPC, 2
- G06F3 06
- G06F13 10
