Storage cluster
13 claims: 4 independent, 9 dependent
- 1単一シャーシ内の複数のストレージノードにおいて、 前記複数のストレージノードは、ストレージクラスターとして一緒に通信するように構成され、 前記複数のストレージノードの各々は、ユーザデータ記憶のための不揮発性ソリッドステートメモリを有し、及び 前記複数のストレージノードは、前記複数のストレージノードの2つが失われても、前記複数のストレージノードが、消去コードを使用してユーザデータを読み取る能力を維持するように、前記複数のストレージノード全体にわたりユーザデータ及びユーザデータに関連したメタデータを配布するように構成 され 、 前記単一シャーシは、前記複数のストレージノードを結合する配電及び内部通信バスを伴うエンクロージャであり、かつ、 前記単一シャーシは、前記複数のストレージノードを外部通信バスに結合 し、 前記複数のストレージノードは、2つの消去コードスキームに従いユーザデータを配布するように構成され、及び前記2つの消去コードスキームは、前記複数のストレージノードに共存し、 前記複数のストレージノードの2つが失われた後に一方の前記消去コードスキームの前記消去コードを使用してユーザデータを回復し、かつ、前記回復されたユーザデータ及び前記2つの消去コードスキームとは異なる消去コードスキームの消去コードを、残りの前記複数のストレージノードに書き込むように構成される 、 複数のストレージノード。
- 2前記不揮発性ソリッドステートメモリは、フラッシュメモリのアレイを含む、請求項1に記載の複数のストレージノード。
- 3前記不揮発性ソリッドステートメモリは、 コントローラ、 前記コントローラに結合された揮発性メモリ、及び 前記揮発性メモリに結合されたエネルギー貯蔵器、を備え、前記不揮発性ソリッドステートメモリは、停電が検出されると、前記揮発性メモリ内のデータを前記フラッシュメモリのアレイに転送する、請求項2に記載の複数のストレージノード。
- 4前記複数のストレージノードの各々は、質問を監視し及びそれに応答するように構成され、質問に対する応答がないことは、前記複数のストレージノードの1つが欠陥であることを示す、請求項1に記載の複数のストレージノード。
- 5前記単一シャーシは、複数のスロットを含み、それらの複数のスロットのうちの各スロットは、前記複数のストレージノードのうちの1つのストレージノードを収容するように構成される、請求項1に記載の複数のストレージノード。
- 6単一シャーシ内に複数のストレージノードを備え、 前記複数のストレージノードの各々は、ユーザデータ記憶のための不揮発性ソリッドステートメモリを有し、及び 前記複数のストレージノードは、前記複数のストレージノードの2つが欠陥である状態で、前記複数のストレージノードが、消去コードを経て、ユーザデータにアクセスできるように、前記複数のストレージノード全体にわたりユーザデータ及びユーザデータに関連したメタデータを配布するように構成され、 前記単一シャーシは、前記複数のストレージノードを包囲し、かつ、 前記単一シャーシは、前記複数のストレージノード間に通信を与える通信バス及び前記複数のストレージノードに電力を供給する配電システムを備え、 前記複数のストレージノードは、前記複数のストレージノードにわたり第1の消去コードスキームに従って書かれたデータを読み取るように構成され、 前記複数のストレージノードは、前記複数のストレージノードにわたり第2の消去コードスキームに従って書かれたデータを読み取るように構成され、前記第1の消去コードスキームに従って書かれたデータは、前記複数のストレージノードにおいて、前記第2の消去コードスキームに従って書かれたデータと共存し、 前記複数のストレージノードの2つが失われた後に前記第1の消去コードスキームの前記消去コードを使用してユーザデータを回復し、かつ、前記回復されたユーザデータ及び前記第1及び前記第2の消去コードスキームとは異なる消去コードスキームの消去コードを、残りの前記複数のストレージノードに書き込むように構成される、 ストレージクラスター。
- 7前記複数のストレージノードのうちの各ストレージノードは、フラッシュメモリアレイに結合されたプロセッサを有するプリント回路板を含む、請求項6に記載のストレージクラスター。
- 8前記単一シャーシは、複数のスロットを含み、それらの複数のスロットのうちの各スロットは、前記複数のストレージノードのうちの1つのストレージノードを収容するように構成される、請求項6に記載のストレージクラスター。
- 9前記複数のストレージノードの各々は、ユーザデータを読み取る試みとは独立して、前記複数のストレージノードの1つの欠陥を決定するように構成される、請求項6に記載のストレージクラスター。
- 10前記複数のストレージノードは、ユーザデータの異なる部分にアクセスするために異なる形式の消去コードを適用するように構成される、請求項6に記載のストレージクラスター。
- 11不揮発性ソリッドステートメモリを有する複数のストレージノードのユーザデータにアクセスするための方法において、 消去コードを通して複数のストレージノード全体にわたってユーザデータを配布し、前記複数のストレージノードは、それらストレージノードをクラスターとして結合する単一シャーシ内に収容され、 前記複数のストレージノードのうちの2つが到達不能であることを決定し、及び 前記複数のストレージノードの残りから、消去コードを経て、ユーザデータにアクセスし、プロセッサが少なくとも1つの方法オペレーションを遂行する、ことを含み、 前記単一シャーシは、前記複数のストレージノード間に通信を与える通信相互接続を含み、 前記単一シャーシは、前記複数のストレージノードを結合する配電バスを含み、 第1の形式の消去コードを使用して前記複数のストレージノードの残りにわたりユーザデータを読み取り、及び 第2の形式の消去コードを使用して前記複数のストレージノードの残りにわたりユーザデータを書き込む、 ことを更に含み、 前記消去コードは、前記複数のストレージノードに共存する2つの異なる消去コードスキームを含み、 前記複数のストレージノードの2つが失われた後に一方の前記消去コードスキームの前記消去コードを使用してユーザデータを回復し、かつ、前記回復されたユーザデータ及び前記2つの消去コードスキームとは異なる消去コードスキームの消去コードを、残りの前記複数のストレージノードに書き込むこと含む、 方法。
- 12前記複数のストレージノードのうちの2つが到達不能であることを決定するのは、心臓鼓動の欠如、質問に対する応答の欠如、又は時間切れの1つに基づく、請求項11に記載の方法。
- 13複数の消去コードスキームのどれをユーザデータの配布に適用するか、前記複数のストレージノードにわたって協働的に決定することを更に含む、請求項11に記載の方法。
Independent claims13
52 paragraphs, as filed
Augment or replace traditional hard disk drives (HDDs), writable CDs (compact discs), writable DVD (digital diversity discs) drives, and tape drives, commonly known as spin media. Solid-state drives (SSDs), such as flash, are currently used to store large amounts of data. Flash and other solid-state memories have different characteristics than SpinMedia. Also, for compatibility reasons, many solid-state drives are designed to meet the standards of hard disk drives, which gives an improvement in the characteristics of flash and other solid-state memories, or from a unique perspective. It makes it difficult to take advantage of it.
The embodiment is generated in such a situation.
In one embodiment, multiple storage nodes in a single chassis are provided. Multiple storage nodes in a single chassis are configured to communicate together as a storage cluster. Each of the plurality of storage nodes includes a non-volatile solid state memory as user data storage. Multiple storage nodes distribute user data and metadata related to user data across multiple storage nodes, and multiple storage nodes use the erase code even if two of the multiple storage nodes are lost. It is configured to maintain the ability to read user data. The chassis includes the ability to install distribution buses, high-speed communication buses, and one or more storage nodes that use those distribution and communication buses. Further, a method for accessing user data in a plurality of storage nodes having non-volatile solid state memory is also provided.
Other aspects and effects of the embodiments described herein will become apparent from the following detailed description with reference to the accompanying drawings showing the principles of those embodiments as an example.
The embodiments described herein and their effects will be best understood by reference to the following description relating to the accompanying drawings. These drawings do not limit any modification to the embodiments described herein that will be made by one of ordinary skill in the art without departing from the spirit and scope of the embodiments described herein.
<figref num="1">In one embodiment, it is a perspective view of a storage cluster with a plurality of storage nodes and internal storage coupled to each storage node to form network-mounted storage.</figref><figref num="2">In one embodiment, it is a system diagram of an enterprise computing system in which one or more of the storage clusters in FIG. 1 can be used as storage resources.</figref><figref num="3">FIG. 5 illustrates a plurality of storage nodes suitable for use in the storage cluster of FIG. 1 and non-volatile solid state storage with different capabilities, according to one embodiment.</figref><figref num="4">It is a block diagram which shows the interconnection switch which connects a plurality of storage nodes by a certain embodiment.</figref><figref num="5">In one embodiment, it is a multi-level block diagram showing the contents of a storage node and the contents of one of the non-volatile solid state storage units.</figref><figref num="6">It is a flowchart of the operation method of the storage cluster which can be carried out by an embodiment of a storage cluster, a storage node and / or a non-volatile solid state storage by an embodiment.</figref><figref num="7">It is a figure which shows the normative computing device which embodies the embodiment described here.</figref>
The following embodiments describe storage clusters that store user data, such as user data generated from one or more users or client systems or other external sources. This storage cluster uses erase codes and redundant copies of metadata to distribute user data across storage nodes housed within the chassis. Erase code refers to a method of data protection or reconstruction in which data is stored across a set of different locations, such as disks, storage nodes, or geographic locations. Flash memory is a form of solid-state memory that is integrated with an embodiment, but embodiments can be extended to other forms of solid-state memory or other storage media, including non-solid-state memory. .. Storage location and workload control is distributed across storage locations in a clustered peer-to-peer system. All tasks such as arbitrating communication between different storage nodes, detecting when a storage node becomes unavailable, and balancing I / O (inputs and outputs) across different storage nodes are all distributed-based. It is handled in. In some embodiments, the data is spread or distributed across multiple storage nodes with data fragments or stripes that support data recovery. Data ownership can be redesignated within the cluster independently of input and output patterns. This architecture, detailed below, allows storage nodes in the cluster to fail while the system is running. This is because the data can be reconstructed from other storage nodes and is therefore kept available for input and output operations. In various embodiments, the storage node is also referred to as a cluster node, blade or server.
The storage cluster is housed in a chassis, an enclosure that houses one or more storage nodes. A mechanism for supplying power to each storage node, such as a distribution bus, and a communication mechanism, for example, a communication bus that enables communication between storage nodes, are included within the chassis. The storage cluster can operate as an independent system in one location, according to certain embodiments. In one embodiment, the chassis houses at least two instances of both independently enabled or disabled power distribution and communication buses. The internal communication bus is an Ethernet® table bus, but other technologies such as Peripheral Component Interconnect (PCI) Express, InfiniBand, etc. are equally suitable. The chassis forms a port of an external communication bus to allow direct communication between multiple chassis or communication through a switch, and communication with a client system. External communication uses technologies such as Ethernet, InfiniBand, and Fiber Channel. In certain embodiments, the external communication bus uses different communication bus technologies for chassis-to-chassis and client communication. When a switch is deployed within or between chassis, the switch acts as a translator between multiple protocols or technologies. When multiple chassis are connected to form a storage cluster, the storage cluster is a proprietary or standard interface, such as Network File System (NFS), Common Internet File System (CIFS), Small Computer System Interface. Accessed by clients using (SCSI) or Hypertext Transfer Protocol (HTTP). Conversion from the client protocol takes place within the switch, chassis external communication bus, or each storage node.
Each storage node is one or more storage servers, and each storage server is connected to one or more non-volatile solid-state memory units, also referred to as storage units. Some embodiments include, but are not limited to, a single storage server within each storage node and between 1 to 8 non-volatile solid-state memory units. The storage server includes a processor for internal communication buses, dynamic random access memory (DRAM) and interfaces, and distribution means for each power bus. Inside a storage node, interfaces and storage units share a communication bus, for example, in some embodiments, PCI Express. The non-volatile solid-state memory unit either directly accesses the internal communication bus interface through the storage node communication bus or requires the storage node to access the bus interface. The non-volatile solid state memory unit accommodates an embedded central processing unit (CPU), a solid state storage controller, and, in certain embodiments, a large solid state mass storage of, for example, 2-32 terabytes (TB). Non-volatile solid-state memory units include embedded volatile storage media such as DRAM and energy storage. In certain embodiments, the energy storage is a capacitor, supercapacitor or battery that allows a subset of DRAM content to be transferred to a stable storage medium in the event of a power outage. In one embodiment, the non-volatile solid state memory unit is composed of storage class memory, eg, a phase change or magnetoresistive random access memory (MRAM) that replaces DRAM and enables a power saving maintainer. Will be done.
One of the many features of storage nodes and non-volatile solid-state storage is the ability to proactively reconstruct data in storage clusters. When a storage node and non-volatile solid-state storage cannot reach a storage node or non-volatile solid-state storage in a storage cluster, independent of any attempt to read data about that storage node or non-volatile solid-state storage. Can be decided. Storage nodes and non-volatile solid-state storage then work together to recover and reconstruct data at least partially in new locations. This constitutes a forward-looking rebuild in that the system rebuilds the data without waiting for the data to be needed for read access initiated by the client system using the storage cluster. These and further details of the storage memory and its operation are described below.
FIG. 1 is a perspective view of a storage cluster 160 with multiple storage nodes 150 and internal solid state memory coupled to each storage node to form a network-mounted storage or storage area network, according to an embodiment. is there. Network-mounted storage, storage area networks, storage clusters, or other storage memory is one or more storage in a flexible and reconfigurable configuration of both physical components and the large amount of storage memory provided by them. It can contain one or more storage clusters 160, each with a node 150. The storage cluster 160 is designed to fit in a rack, and one or more racks can be set up and populated as described for storage memory. The storage cluster 160 has a chassis 138 with a plurality of slots 142. Chassis 138 is also referred to as a housing, enclosure or rack unit. In one embodiment, the chassis 138 has 14 slots 142, but other numbers of slots are readily conceivable. For example, one embodiment has 4 slots, 8 slots, 16 slots, 32 slots, or any other suitable number of slots. Each slot 142 can accommodate one storage node 150 in certain embodiments. Chassis 138 includes flaps 148 that can be used to mount chassis 138 in the rack. Fan 144 provides air circulation to cool the storage node 150 and its components, but other cooling components can also be used, or embodiments without cooling components are devised. The switch fabric 146 joins the storage nodes 150 in chassis 138 together and also joins the network for communication to memory. In the embodiment shown in FIG. 1, the switch fabric 146 and the fan The slot 142 on the left side of the switch 144 is shown occupied by the storage node 150, while the slot 142 on the right side of the switch fabric 146 and the fan 144 is empty and the storage node 150 is inserted for illustrative purposes. Can be used for. This configuration is only an example, and in yet various other configurations, one or more storage nodes 150 can occupy slot 142. The array of storage nodes does not have to be sequential or adjacent in some embodiments. The storage node 150 is hot pluggable, which means that the storage node 150 can be inserted into or removed from slot 142 of chassis 138 without stopping or powering down the system. When inserting or removing the storage node 150 into or from slot 142, the system automatically reconfigures to recognize and adapt to changes. Reconstruction, in certain embodiments, includes restoring redundancy and / or rebalancing data or loads.
Each storage node 150 has a plurality of components. In the embodiment shown here, the storage node 150 comprises a CPU 156, a printed circuit board 158 popularized by a processor, a memory 154 coupled to the CPU 156, and a non-volatile solid state storage 152 coupled to the CPU 156. In yet another embodiment, other attachments and / or components can be used. Memory 154 has instructions executed by CPU 156 and / or data manipulated by CPU 156. As further described below, the non-volatile solid state storage 152 includes a flash, or in yet another embodiment, a solid state memory of another form.
FIG. 2 is a system diagram of an enterprise computing system 102 that can use one or more of the storage nodes, storage clusters, and / or non-volatile solid-state storage of FIG. 1 as storage resources 108. For example, the flash storage 128 of FIG. 2 integrates the storage node, storage cluster and / or non-volatile solid state storage of FIG. 1 in one embodiment. The enterprise computing system 102 has a processing resource 104, a network resource 106, and a storage resource 108 including a flash storage 128. The flash storage 128 includes a flash controller 130 and a flash memory 132. In various embodiments, the flash storage 128 can include one or more storage nodes or storage clusters together with a flash controller 130 including a CPU and a flash memory 132 including non-volatile solid state storage of the storage nodes. In certain embodiments, the flash memory 132 includes a different type of flash memory, or the same type of flash memory. The enterprise computing system 102 presents an environment suitable for deploying flash storage 128, which is used for other larger or smaller computing systems or devices, or for fewer or additional. It can be used for various enterprise computing systems 102 with resources. The enterprise computing system 102 is coupled to a network 140, such as the Internet, to provide or use services. For example, the enterprise computing system 102 can provide cloud services, physical computing services, or virtual computing services.
In the enterprise computing system 102, various resources are arranged and managed by various controllers. The processing controller 110 manages a processing resource 104 including a processor 116 and a random access memory (RAM) 118. Network controller 112 manages network resource 106, including router 120, switch 122, and server 124. The storage controller 114 manages a storage resource 108 including a hard drive 126 and a flash storage 128. This embodiment may include other types of processing resources, network resources, and storage resources. In one embodiment, the flash storage 128 completely replaces the hard drive 126. The enterprise computing system 102 can provide or allocate various resources as physical computing resources or, in a variant thereof, as virtual computing resources supported by the physical computing resources. For example, various resources can be embodied using one or more servers running software. The storage resource 108 stores files or data objects or other forms of data.
In various embodiments, the enterprise computing system 102 comprises multiple racks populated by storage clusters, which are located in a single physical location, such as a cluster or server farm. In other embodiments, the racks can be located in multiple physical locations such as different cities, states or countries and connected by a network. Each rack, each storage cluster, each storage node, and each non-volatile solid-state storage is individually configured with each amount of storage space, which storage space can then be configured independently of the others. Therefore, storage capacity can be flexibly added, upgraded, deducted, recovered, and / or reconfigured in each of the non-volatile solid state storages. As mentioned above, each storage node can embody one or more servers in certain embodiments.
FIG. 3 is a block diagram showing a plurality of storage nodes 150 suitable for use in the chassis of FIG. 1 and a non-volatile solid state storage 152 with different capabilities. Each storage node 150 can have one or more units of non-volatile solid state storage 152. Each non-volatile solid state storage 152 includes, in certain embodiments, a different capacity than the other non-volatile solid state storage 152 at the storage node 150 or another storage node 150. Alternatively, all the non-volatile solid state storage 152 in the storage node or the plurality of storage nodes may have the same capacity, or may have the same and / or a combination of different capacities. This flexibility is shown in FIG. 3, which shows a storage node 150 with a mixed non-volatile solid state storage 152 with TB capacities of 4, 8 and 32; a non-volatile solid state storage 152 with 32 TB capacities each. An example is shown of another storage node 150 with; and yet another storage node with each 8 TB capacity non-volatile solid state storage 152; According to the teachings described herein, yet various other combinations and capacities can be easily devised. In the situation of clustering, for example, clustering storage to form a storage cluster, the storage node may be non-volatile solid state storage 152 or may include it. As further described below, the non-volatile solid state storage 152 is a convenient clustering point. This is because the non-volatile solid state storage 152 includes a non-volatile random access memory (NVRAM) component.
With reference to FIGS. 1 and 3 , the storage cluster 160 is extensible, which means that storage capacity of non-uniform storage size is easily added, as described above. One or more storage nodes 150 can be plugged in and out of each chassis, and in some embodiments the storage cluster self-configures. The plug-in storage node 150 can be of different sizes regardless of whether it is installed in the chassis at the time of delivery or added later. For example, in some embodiments, the storage node 150 is a multiple of 4TB, such as 8TB, 12TB, 16TB, 32TB, and so on. In yet another embodiment, the storage node 150 is another storage amount or multiple of capacity. The storage capacity of each storage node 150 is broadcast and influences the determination of how to strip the data. For maximum storage efficiency, in one embodiment, in stripes, subject to the predetermined requirement that one or two non-volatile solid-state storage units 152 or storage nodes 150 in the chassis continue to operate. Can be self-configuring as widely as possible.
FIG. 4 is a block diagram showing a communication interconnection unit 170 and a distribution bus 172 that connect a plurality of storage nodes 150. Returning to FIG. 1, the communication interconnect 170 is, in certain embodiments, embodied or included in the switch fabric 146. When a plurality of storage clusters 160 occupy a rack, the communication interconnect 170 is, in certain embodiments, embodied or included in the top of the rack switch. As shown in Figure 4, the storage cluster 160 is enclosed within a single chassis 138. The external port 176 is coupled to the storage node 150 through the communication interconnect 170, while the external port 174 is directly connected to the storage node. External power port 178 is coupled to distribution bus 172. Storage node 150 includes non-volatile solid state storage 152 in varying amounts and different capacities, as described with reference to FIG. Further, the one or more storage nodes 150 may be a calculation-only storage node as shown in FIG. The authority 168 is embodied in the non-volatile solid state storage 152, for example, as a list or other data structure stored in memory. In certain embodiments, the authority is stored in the non-volatile solid state storage 152 and is supported by software running on the controller or other processor of the non-volatile solid state storage 152. In yet another embodiment, the authority 168 is embodied in storage node 150, for example as a list or other data structure stored in memory 154, and is supported by software running on CPU 156 of storage node 150. Ru. Authority 168, in certain embodiments, controls where and how data is stored in the non-volatile solid state storage 152. This control determines which form of erasure code scheme is applied to the data and which storage node 150 Helps determine which part of the data you have. Each authority 168 is designated as non-volatile solid state storage 152. Each authority controls the range of inode numbers, segment numbers, or other data identifiers specified in the data by the file system, storage node 150 or non-volatile solid state storage 152 in various embodiments.
Each piece of data and each piece of metadata, in certain embodiments, has redundancy in the system. In addition, each piece of data and each piece of metadata has an owner, also referred to as the authority. If that authority is unreachable, for example due to a defect in the storage node, there is an inheritance plan on how to find that data or its metadata. In various embodiments, there are redundant copies of authority 168. Authority 168, in certain embodiments, has a relationship with storage node 150 and non-volatile solid state storage 152. Each authority 168 covering a range of data segment numbers or other data identifiers is designated for a particular non-volatile solid state storage 152. In certain embodiments, the authority 168 for all such ranges is distributed across the non-volatile solid state storage 152 of the storage cluster. Each storage node 150 has a network port that provides access to the non-volatile solid state storage 152 of that storage node 150. The data is stored in the segment associated with the segment number, which in some embodiments is an indirection to the configuration of a RAID (redundant array of independent disks) stripes. Therefore, the designation and use of authority 168 establishes an indirect reference to the data. Indirect references, according to certain embodiments, are also referred to in this case as the ability to indirectly reference data via authority 168. The segment identifies a set of non-volatile solid-state storage 152, and the local identifier is for the set of non-volatile solid-state storage 152 that contains the data. In some embodiments, the local identifier is an offset to the device and is sequentially reused by multiple segments. In other embodiments, the local identifier is in a particular segment. On the other hand, it is unique and never reused. The offset of the non-volatile solid state storage 152 is applied to the position data for writing to or reading from the non-volatile solid state storage 152 (in the form of a RAID stripe). Data is striped across multiple units of non-volatile solid-state storage 152, which may include or be different from non-volatile solid-state storage 152 having authority 168 for a particular data segment. ..
For example, if there is a change in the location of a particular segment of data during data movement or data reconstruction, that data segment in the non-volatile solid-state storage 152 or storage node 150 with that authority 168. Must be discussed with Authority 168. In this embodiment, a hash value of a data segment is calculated or an inode number or data segment number is applied to search for a particular data fragment. The output of this operation points to the non-volatile solid state storage 152 having authority 168 for that particular piece of data. In one embodiment, there are two stages for this operation. The first stage maps an entity identifier (ID), such as a segment number, inode number, or directory number, to an authority identifier. This mapping involves calculations such as hashes or bitmasks. The second stage is to map the authority identifier to a particular non-volatile solid state storage 152, which is done through explicit mapping. This operation can be repeated to ensure that when the calculation is done, the calculation result repeatedly and reliably points to a particular non-volatile solid state storage 152 with authority 168. This behavior includes a set of reachable storage nodes as input. As the set of reachable non-volatile solid-state storage units changes, so does the optimal set. In some embodiments, the persistence value is the current specified value (always true), and the calculated value is the targeted value where the cluster attempts to reconfigure. This calculation is used to determine the best non-volatile solid-state storage 152 for authority in the presence of a set of non-volatile solid-state storage 152 that is reachable and constitutes the same cluster. This calculation also authorizes even if the specified non-volatile solid state storage is unreachable. It also determines an ordered set of peer non-volatile solid-state storage 152 that records its authority in the mapping of non-volatile solid-state storage so that it is determined. In certain embodiments, a copy or substitute authority 168 is consulted when a particular authority 168 is not available.
With reference to FIGS. 1-4, two of the many tasks of CPU 156 at storage node 150 are decomposing write data and reassembling read data. When the system decides that data should be written, authority 168 is searched for that data as described above. When the segment ID of the data has already been determined, the write request is forwarded from that segment to the non-volatile solid state storage 152 currently determined to be the host of the determined authority 168. The host CPU 156 of the storage node 150 in which the non-volatile solid-state storage 152 and the corresponding authority 168 are present then decomposes or fragmentes the data and sends the data to various non-volatile solid-state storage 152. The transmitted data is written as a data stripe according to the erasure code scheme. In some embodiments it is required to pull the data, and in other embodiments the data is pushed. Conversely, when reading data, the authority 168 for the segment ID containing the data is searched for as described above. The host CPU 156 of the storage node 150 in which the non-volatile solid state storage 152 and the corresponding authority 168 are present requests data from the non-volatile solid state storage pointed out by the authority and the corresponding storage node. In certain embodiments, the data is read from the flash storage as data stripes. Host CPU 156 on storage node 150 then reassembles the read data, corrects any errors (if any) based on the appropriate erase code scheme, and transfers the reassembled data to the network. In yet another embodiment, some or all of these tasks are handled in the non-volatile solid state storage 152. In some embodiments, the segment host is a stray.
In some systems, such as UNIX-style file systems, data is handled by index nodes or inodes, which identify the data structures that represent objects in the file system. The object is, for example, a file or directory. Metadata accompanies objects as attributes, such as authorization data and generation timestamps, among other attributes. Segment numbers are specified for all or part of such objects in the file system. In other systems, the data segment is treated with a segment number specified somewhere. Descriptionally, the unit of distribution is an entity, and an entity is a file, directory, or segment. That is, an entity is a unit of data or metadata stored by a storage system. Entities are grouped into sets called authorities. Each authority has an authority owner, which is a storage node that has the exclusive right to update an entity in the authority. In other words, the storage node contains the authority, and then the authority contains the entity.
A segment is, according to some embodiments, a logical container for data. Also, the segment is the address space between the media address space and the physical flash location, i.e., this address space has the data segment number. The segment also contains metadata, which can restore data redundancy (rewrite to a different flash location or device) without involving high-level software. In some embodiments, the internal format of the segment includes client data and medium mapping to determine the location of that data. Each data segment is protected from, for example, memory and other defects by dividing the segment into a large number of data and parity shards (if applicable). The data and parity shards are distributed, or striped, across the non-volatile solid state storage 152 coupled to the host CPU 156 (Figure 5) by the erase code scheme. The use of a period segment, in some embodiments, refers to a container and its location in the address space of the segment. The use of period stripes refers to the same shard set as a segment and, according to certain embodiments, includes how shards are distributed with redundancy or parity information.
A series of address space translations is performed across the storage system. At the top is a directory entity (filename) that links to the inode. Inode refers to the medium address space where the data is logically stored. Media addresses are mapped through a set of indirect media to distribute the load on large files or embody data services such as copy exclusion or snapshots. Media addresses are mapped through a set of indirect media to distribute the load on large files or embody data services such as copy exclusion or snapshots. The segment address is then translated into a physical flash position. The physical flash location, in certain embodiments, has an address range limited by the amount of flash in the system. The media address and segment address are logical containers, and in some embodiments, 128-bit or higher identifiers are used to make them virtually infinite, and the likelihood of reuse is calculated to be longer than the expected life of the system. Will be done. Addresses from the logical container are, in some embodiments, assigned in a hierarchy form. First, each non-volatile solid-state storage 152 is given a range of address spaces. Within this specified range, the non-volatile solid state storage 152 can be assigned an address without synchronizing with other non-volatile solid state storage 152.
Data and metadata are stored by a set of basic storage layouts optimized for changing workload patterns and storage devices. These layouts combine multiple redundancy schemes, compression formats and indexing algorithms. Some of these layouts store information about the authority and authority master, while others store file metadata and file data. Redundancy schemes allow error correction codes that allow decay bits in a single storage device (such as NAND flash chips), erase codes that allow defects in multiple storage nodes, and data center or area defects. Includes copying scheme. In some embodiments, a low density parity check (LDPC) code is used within a single storage unit. In some embodiments, Reed-Solomon encoding is used within the storage cluster, and mirroring is used within the storage grid. The metadata is an ordered log-structured index (for example, Log Structured Merge). It is stored using Tree), and the log-structured layout does not store large amounts of data.
To maintain consistency across multiple copies of an entity, storage nodes implicitly agree on two things through computation. (1) an authority that contains an entity, and (2) a storage node that contains an authority. Specifying an entity to an authority is done by assigning the entity to the authority in a pseudo-random manner, dividing the entity into ranges based on externally generated keys, or placing a single entity in each authority. be able to. Pseudo-random schemes are, for example, linear hashes, and the replication under-scalable hash (RUSH) family of hashes, which includes controlled replication under-scalable hashes (CRUSH). In some embodiments, the pseudo-random designation is only used to assign an authority to a node because the set of nodes can change. The set of authorities does not change, so in these embodiments subjective functions apply. Some placement schemes automatically place the authority on the storage node, while other placement schemes rely on the explicit mapping of the authority to the storage node. In some embodiments, a pseudo-random scheme is used to map from each authority to a set of candidate authority owners. The pseudo-random data distribution feature associated with CRUSH specifies the authority to the storage node and produces a list of where the authority is specified. Each storage node has a copy of the pseudo-random data distribution function and arrives at the same calculation to distribute and later discover or explore the authority. Each of the pseudo-random schemes, in certain embodiments, requires a reachable set of storage nodes as input to end at the same target node. When an entity is placed in an authority, it is physically placed so that expected defects do not lead to unexpected data loss. Stored in the device. In one embodiment, the rebalancing algorithm attempts to store copies of all entities in the authority on the same machine set and in the same layout.
Possible defects include, for example, equipment defects, machine thefts, data center fires, and regional catastrophes, such as nuclear accidents and geological events. Different defects result in different levels of tolerable data loss. In some embodiments, storage node theft does not affect the security or reliability of the system, but depending on the configuration of the system, local events are no loss of data, lost for seconds or minutes. It may lead to update or complete data loss.
In these embodiments, the placement of data for storage redundancy is independent of the placement of authorities for data consistency. In some embodiments, the storage node containing the authority does not include persistent storage. Rather, the storage node is connected to non-volatile solid-state storage that does not contain authority. The communication interconnect between the storage node and the non-volatile solid state storage unit consists of multiple communication technologies and has non-uniform performance and defect tolerance characteristics. In one embodiment, as described above, the non-volatile solid state storage unit is connected to the storage node via PCI Express, and the storage nodes are connected together in a single chassis using an Ethernet backplane. The chassis are then connected together to form a storage cluster. In some embodiments, the storage cluster is connected to the client using Ethernet or Fiber Channel. If multiple storage clusters are configured on a storage grid, the multiple storage clusters may use the Internet or other long-distance network links, such as "metroscale" or private links that do not cross the Internet. Be connected.
The authority owner has the exclusive right to modify an entity, move the entity from one non-volatile solid state storage unit to another non-volatile solid state storage unit, and add and remove copies of the entity. This allows the maintenance of basic data redundancy. When the authority owner fails, is retired, or is overloaded, the authority is migrated to a new storage node. Transient defects are not trivial to ensure that all non-defective machines agree on a new authority position. Ambiguities caused by transient flaws can be generated by consensus protocols, such as Paxos, hot warm failover schemes, or through manual intervention by remote system administrators, or by local hardware administrators (eg, from clusters). It is automatically cleared (by physically removing the defective machine or pressing a button on the defective machine). In some embodiments, a consensus protocol is used, and failover is automatic. If too many defects or copying events occur in too short a period of time, in one embodiment the system goes into self-preservation mode and suspends copying and data movement activities until the intervention of an administrator.
Authorities are transferred between storage nodes, and the authority owns update entities to those authorities, so the system transfers messages between the storage nodes and the non-volatile solid-state storage unit. With respect to persistent messages, messages of different purpose are of different form. Based on the format of the message, the system maintains a different order and durability guarantee. When a persistent message is processed, the message is temporarily stored in multiple durable and non-durable storage hardware technologies. In certain embodiments, messages are stored in RAM, NVRAM and NAND flash devices, and various protocols are used to make efficient use of each storage medium. Latency-sensitive client demands are persisted in copy NVRAM and then NAND, while background rebalancing behavior is persisted directly into NAND.
Persistent messages are persistently stored until they are copied. This allows the system to continue to serve client requests despite defects and component replacements. Many hardware components include system administrators, manufacturers, hardware supply chains, and unique identifiers that appear to the ongoing monitoring quality control infrastructure, but applications running at the top of the infrastructure address have the address. Virtualize. These virtualized addresses do not change over the life of the storage system, regardless of component defects and replacements. This allows each component of the storage system to be replaced over time without restructuring or interrupting client request processing.
In some embodiments, the virtualized address is stored with sufficient redundancy. The continuous monitoring system correlates the hardware and software state with the hardware identifier. This allows the detection and prediction of defects by defective components and manufacturing details. Surveillance systems also, in certain embodiments, also allow foresight transfer of authorities and entities from affected devices until defects occur by removing components from critical paths.
FIG. 5 is a multi-level block diagram showing the contents of the storage node 150 and the contents of the non-volatile solid state storage 152 of the storage node 150. In some embodiments, data is communicated to and from storage node 150 by network interface controller (NIC) 202. As mentioned above, each storage node 150 has a CPU 156 and one or more non-volatile solid state storage 152. Moving down one level in FIG. 5, each non-volatile solid state storage 152 has a relatively fast non-volatile solid state memory, such as a non-volatile random access memory (NVRAM) 204, and a flash memory 206. In certain embodiments, NVRAM204 is a component (DRAM, MRAM, PCM) that does not require a program / erase cycle and is a memory that can support writes much more frequently than memory reads. Moving down another level in FIG. 5, NVRAM 204 is embodied as a fast volatile memory, such as Dynamic Random Access Memory (DRAM) 216, backed up by energy storage 218 in certain embodiments. The energy storage 218 provides sufficient power to keep the DRAM 216 energized long enough to transfer content to the flash memory 206 in the event of a power outage. In certain embodiments, the energy storage 218 is a capacitor, supercapacitor, battery, or other device that provides sufficient energy supply to transfer the contents of the DRAM 216 to a stable storage medium in the event of a power outage. The flash memory 206 is embodied as a plurality of flash dies 222, which is also referred to as a package of flash dies 222 or an array of flash dies 222. The flash die 222 includes one die per package, multiple dies per package (ie, a multi-chip package), and a hybrid package. It is clear that it can be packaged in many ways, including bare dies, encapsulation dies, etc. on di, printed circuit boards or other substrates. In the embodiments shown herein, the non-volatile solid state storage 152 has a controller 212 or processor and input / output (I / O) ports 210 coupled to the controller 212. The I / O port 210 is coupled to the CPU 156 and / or network interface controller 202 of the flash storage node 150. The flash input / output (I / O) port 220 is coupled to the flash die 222, and the direct memory access unit (DMA) 214 is coupled to the controller 212, DRAM 216 and flash die 222. In the embodiments shown herein, the I / O port 210, the controller 212, the DMA unit 214 and the flash I / O port 220 are embodied in a programmable logic device (PLD) 208, such as a field programmable gate array (FPGA). .. In this embodiment, each flash die 222 has a page organized as 16 kB (kilobytes) page 224 and a register 226 from which data is written to or read from the flash die 222. In yet another embodiment, other forms of solid state memory are used in place of or in addition to the flash memory shown within the flash die 222. , And the direct memory access unit (DMA) 214 is coupled to the controller 212, DRAM 216 and flash die 222. In the embodiments shown herein, the I / O port 210, the controller 212, the DMA unit 214 and the flash I / O port 220 are embodied in a programmable logic device (PLD) 208, such as a field programmable gate array (FPGA). .. In this embodiment, each flash die 222 has a page organized as 16 kB (kilobytes) page 224 and a register 226 from which data is written to or read from the flash die 222. In yet another embodiment, other forms of solid state memory are used in place of or in addition to the flash memory shown within the flash die 222. , And the direct memory access unit (DMA) 214 is coupled to the controller 212, DRAM 216 and flash die 222. In the embodiments shown herein, the I / O port 210, the controller 212, the DMA unit 214 and the flash I / O port 220 are embodied in a programmable logic device (PLD) 208, such as a field programmable gate array (FPGA). .. In this embodiment, each flash die 222 has a page organized as 16 kB (kilobytes) page 224 and a register 226 from which data is written to or read from the flash die 222. In yet another embodiment, other forms of solid state memory are used in place of or in addition to the flash memory shown within the flash die 222.
FIG. 6 is a flowchart of how to operate the storage cluster. This method can be implemented in various embodiments of storage clusters and storage nodes described herein, or by various embodiments. The various steps of this method can be performed by a processor such as a processor in a storage cluster or a processor in a storage node. This method can be carried out in part or in whole with software, hardware, firmware, or a combination thereof. This method begins with action 602, which distributes the user data with the erase code. For example, user data can be distributed across storage nodes in a storage cluster using one or more erasure code schemes. Two (or more in some embodiments) erasure code schemes can coexist on a storage node in a storage cluster. In certain embodiments, each storage node can determine which of the plurality of erasure code schemes to apply when writing data, and which erasure code scheme to apply when reading data. These may be the same erasure code scheme or different erasure code schemes.
This method goes to action 604 and the storage nodes in the storage cluster are checked. In one embodiment, storage nodes are checked for heartbeats, and each storage node periodically generates a message that acts as a heartbeat. In another embodiment, the check is in the form of asking the storage node, and the lack of response to the question indicates that the storage node is defective. In decision action 606, it is determined whether the two storage nodes are unreachable. For example, if two of the storage nodes no longer generate a heartbeat, or if two of the storage nodes do not answer the question, or are a combination of them, or are other instructions, then of the other storage node. One can determine that two of the storage nodes are unreachable. If this is not the case, the flow returns to action 602 and continues to distribute user data, for example, writing user data to the storage node for storage when user data arrives. If two of the storage nodes are determined to be unreachable, the flow continues to action 608.
In decision action 608, the erase code is used to access user data on the remaining storage nodes. It is clear that user data, in certain embodiments, refers to data originating from one or more users or client systems or other sources outside the storage cluster. In one embodiment, the erase code format includes double redundancy, in this case with two defective storage nodes, the remaining storage nodes have readable user data. The erase code format contains an error correction code that allows the loss of 2 bits of the codeword, and the data is distributed across the storage nodes so that the data can be recovered even if two of the storage nodes are lost. Judgment action 610 determines whether the data should be reconstructed. If the data should not be rebuilt, the flow returns to action 602 and continues to distribute the user data with the erase code. If the data should be rebuilt, the flow branches to action 612. In one embodiment, the decision to rebuild the data is made after the two storage nodes are unreachable, while in other embodiments, the decision to rebuild the data is made unreachable by one storage node. It may be done after it becomes. The various mechanisms considered in the decision to reconstruct the data include error correction counts, error correction rates, read defects, write defects, heartbeat loss, response defects to questions, and the like. Appropriate changes to the method of FIG. 6 will be readily understood for these and yet other embodiments.
In action 612, the erase code is used to recover the data. This is due to the erase code example described above for Action 608. More specifically, the data is either recovered from the remaining storage nodes, for example using an error correction code, or read from the remaining storage nodes. If more than one form of erasure code coexists on the storage node, you can use more than one form of erasure code to recover the data. In decision action 614, it is determined whether the data must be read in parallel. In some embodiments, there are two or more data paths (eg, double redundancy of the data), and the data can be read in parallel across the two paths. If the data should not be read in parallel, the flow branches to action 618. If the data should be read in parallel, the flow diverges to action 616, resulting in competition for results. The winner of the competition is then used as the recovered data.
In action 618, the erase code scheme is determined for reconstruction. For example, in one embodiment, each storage node can determine which of two or more erase code schemes to apply when writing data across storage units. In some embodiments, the storage nodes work together to determine the erase code scheme. This can be done by determining which storage node plays a role for the erase code scheme for a particular data segment, or by specifying a storage node to play that role. In certain embodiments, various mechanisms such as testimony, voting, or judgment logic are used to accomplish this action. Non-volatile solid-state storage acts as a witness (in some embodiments) or voter (in some embodiments), and the remaining functionality and authority of non-volatile solid-state storage in the event that an authoritative copy becomes defective. Allow the remaining copy of to determine the content of the defective authority. In action 620, the recovered data is written over the remaining storage nodes along with the erase code. For example, the erasure code scheme determined for reconstruction is different from the erasure code scheme applied when recovering data, i.e. reading data. More specifically, the loss of two storage nodes means that one erase code scheme can no longer be applied to the remaining storage nodes, and the storage node can be switched to an erase code scheme that can be applied to the remaining storage nodes. means.
It is clear that the methods described here are carried out in digital processing systems such as conventional general purpose computer systems. Alternatively, a special purpose computer designed or programmed to perform only one function may be used. FIG. 7 is a diagram showing a normative computing device that embodies the embodiments described herein. The computing device of FIG. 7 is used, in certain embodiments, to perform an embodiment of a function for a storage node or non-volatile solid state storage. This computing device comprises a central processing unit (CPU) 701, which is coupled to memory 703 and mass storage device 707 through bus 705. Mass storage device 707, in certain embodiments, represents a persistent data storage device such as a disk drive that is local or remote. The mass storage device 707 embodies backup storage in certain embodiments. The memory 703 includes read-only memory, random access memory, and the like. Applications present in a computing device are, in certain embodiments, stored or accessed on a computer-readable medium such as memory 703 or mass storage device 707. The application may also be in the form of a modulated electronic signal accessed via the computing device's network modem or other network interface. It is clear that the CPU701, in certain embodiments, is implemented on a general purpose processor, a special purpose processor, or a specially programmed logic device.
The display 711 communicates with the CPU 701, the memory 703, and the mass storage device 707 via the bus 705. The display 711 is configured to display a visual tool or report associated with the system described herein. The input / output device 709 is coupled to the bus 505 to communicate command selection information to the CPU 701. It is clear that data to and from the external device is communicated via the input / output device 709. CPU701 is defined to perform the functions described herein to enable the functions described with reference to FIGS. 1-6. In some embodiments, the code that performs this function is stored in memory 703 or mass storage device 707 for execution by a processor such as CPU 701. The operating system of the computing device is MS-WINDOWS<sup>TM</sup>, UNIX<sup>TM</sup>, LINUX<sup>TM</sup>, IOS<sup>TM</sup>, CentOS<sup>TM</sup>, Android<sup>TM</sup>, Redhat LINUX<sup>TM</sup>, Z / OS<sup>TM</sup>, Or any other known operating system. Further, it is clear that the embodiments described here can be integrated with a virtual computing system.
Here, embodiments for detailed description are disclosed. However, the particular functional details disclosed herein are merely exemplary to illustrate embodiments. However, those embodiments may be implemented in a number of other embodiments and should not be construed as being limited to the embodiments described herein.
The terms first, second, etc. are used herein to describe the various steps or calculations, but it should be understood that these steps or calculations should not be limited by these terms. .. These terms are only used to distinguish one step or calculation from another. For example, without departing from the scope of this disclosure, the first calculation may be referred to as the second calculation, and similarly, the second step may be referred to as the first step. As used herein, the word "and / or" and the symbol "/" include any and all combinations of one or more of the related items listed.
The singular forms "a", "an" and "the" used herein shall also include the plural form unless the context clearly indicates otherwise. In addition, the terms comprises, comprising, includes and / or including, as used herein, are the features, integers, steps, described above. It identifies the existence of operations, elements, and / or components, but does not preclude the existence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Therefore, the terms used herein are merely for the purpose of describing a particular embodiment and are not limited thereto.
Also note that in another embodiment, the indicated functions / actions may be performed out of the order shown in the figure. For example, two drawings shown in sequence may actually be performed substantially simultaneously, or sometimes in reverse order, based on the functions / actions involved.
With the above embodiments in mind, it should be understood that those embodiments use various computer-implemented operations involving data stored in a computer system. Those operations require physical manipulation of physical quantities. Usually, but not necessarily, their quantities take the form of electrical or magnetic signals that can be stored, transferred, synthesized, compared, or otherwise manipulated. Moreover, the operations performed are often referred to by terms such as occurrence, identification, determination, or comparison. All of the operations described herein that form part of an embodiment are useful machine operations. The embodiments also relate to devices or devices for performing their operations. The device may be specifically configured for the requested purpose, or the device may be a general purpose computer selectively activated or configured by a computer program stored in the computer. In particular, various general purpose machines can be used in computer programs written according to the teachings herein, or it is convenient to configure special equipment to perform the requested operation.
Modules, applications, layers, agents, or other methods-operable entities are embodied as hardware, firmware, or processor execution software, or a combination thereof. If a software-based embodiment is disclosed herein, it will be clear that the software can be implemented on a physical machine such as a controller. For example, the controller can include a first module and a second module. The controller can be configured to perform various actions of, for example, a method, application, layer or agent.
The embodiments can also be embodied as computer-readable code on a non-temporary computer-readable medium. A computer-readable medium is a data storage device that can store data that is later read by a computer system. Computer-readable media include, for example, hard drives, network-attached storage (NAS), read-only memory, random access memory, CD-ROMs, CD-Rs, CD-RWs, magnetic tapes, and other optical and non-optical tapes. Includes optical data storage equipment. Computer-readable media are also distributed across network-coupled computer systems so that computer-readable code is stored and executed in a distributed form. The embodiments described herein include various computer systems including handheld devices, tablets, microprocessor systems, microprocessor-based or programmable consumer electronics, consumer electronics, minicomputers, mainframe computers, and the like. It is embodied in composition. These embodiments can also be embodied in a distributed computing environment in which tasks are performed by remote processing devices linked through wire-based or wireless networks.
The operations of the method have been described in a particular order, but other operations may be performed between the operations described, and the operations described may be adjusted so that they occur at slightly different times. It should be understood that well or described operations may be distributed in a system capable of performing processing operations at various processing-related intervals.
In various embodiments, one or more parts of the methods and mechanisms described herein may form part of a cloud computing environment. In such an embodiment, resources are provided via the Internet as a service according to one or more different models. Such a model includes infrastructure as a service (IaaS), platform as a service (PaaS), and software as a service (SaaS). In IaaS, computer infrastructure is delivered as a service. In such cases, the computing device is generally owned and operated by the service provider. In the PaaS model, software tools and their underlying equipment used by developers to develop software solutions are provided as services and hosted by service providers. SaaS typically includes service provider licensed software as an on-demand service. The service provider hosts the software or deploys the software to the customer for a given period of time. Many combinations of said models are conceivable and intended.
Various units, circuits or other components are described or claimed as "configured" to perform one or more tasks. In this regard, the phrase "configured" implies a structure by indicating that the unit / circuit / component contains a structure (eg, a circuit) that performs one or more tasks during operation. Used for. Therefore, a unit / circuit / component can be said to be configured to perform a task even when the specified unit / circuit / component is not currently operating (eg, not on). Units / circuits / components used with a "configured" language include hardware, such as circuits, memory that stores program instructions that can be executed to perform operations, and the like. Representing a unit / circuit / component as "configured" to perform one or more tasks is 35 for that unit / circuit / component. It is clearly intended not to cite USC112. In addition, "configured" is an inclusive operation operated by operating the software and / or firmware (eg, an FPGA, or a general purpose processor running the software) in a manner capable of performing the task in question (s). Includes a structural structure (eg, an inclusive circuit). Also, "configuring" includes adapting a manufacturing process (eg, a semiconductor manufacturing facility) to manufacture a device (eg, an integrated circuit) for performing or performing one or more tasks.
The above description has been described with reference to a specific embodiment for the sake of explanation. However, the above discussion is not exhaustive, nor does it limit the invention to the exact form disclosed herein. Many changes and modifications can be considered in view of the above teachings. The embodiments have been selected and described to best explain the principles of those embodiments and their practical applications, and therefore those skilled in the art will intend to make those embodiments and various modifications. It will be best utilized to suit the particular application to be made. Accordingly, the embodiments shown herein are considered exemplary and are not limited thereto, and the present invention is not limited to the details set forth herein and is within the scope of the claims and its equivalents. It can be changed.
102: Enterprise Computing System 104: Processing Resources 106: Network Resources 108: Storage Resources 110: Processing Controllers 112: Network Controllers 114: Storage Controllers 116: Processors 118: Random Access Memory (RAM) 120: Routers 122: Switches 124: Server 126: Hard drive 128: Flash storage 130: Flash controller 132: Flash memory 138: Chassis 140: Network 142: Slot 144: Fan 146: Switch fabric 148: Flap 150: Storage node 152: Non-volatile solid state storage 154: Memory 156: CPU 158: Printed circuit board 160: Storage cluster 701: CPU 703: Memory 705: Bus 707: Mass storage 709: Input / output device 711: Display
7 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| JP2013544386A | Cites | Japan |
| JP2007265314A | Cites | Japan |
| JP2007242018A | Cites | Japan |
| WO2014025821A1 | Cites | World Intellectual Property Organization (WIPO) |
| JP2010128886A | Cites | Japan |
85 members in 7 offices
Members85
| Document | Office | Kind | |
|---|---|---|---|
| US8850108B1 | United States of America | B1 | |
| US9201600B1 | United States of America | B1 | |
| US2015355848A1 | United States of America | A1 | |
| US2015355969A1 | United States of America | A1 | |
| WO2015187218A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US9213485B1 | United States of America | B1 | |
| AU2015268889A1 | Australia | A1 | |
| US2016085628A1 | United States of America | A1 | |
| KR20160039265A | Republic of Korea | A | |
| US9357010B1 | United States of America | B1 | |
| CN105706065A | China | A | |
| EP3036639A1 | European Patent Office (EPO) | A1 | |
| EP3036639A4 | European Patent Office (EPO) | A4 | |
| WO2016130301A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US2016246528A1 | United States of America | A1 | |
| US2016277503A1 | United States of America | A1 | |
| US9525738B2 | United States of America | B2 | |
| US9563506B2 | United States of America | B2 | |
| AU2017200947A1 | Australia | A1 | |
| US2017093980A1 | United States of America | A1 | |
| JP2017519320A | Japan | A | |
| KR101770547B1 | Republic of Korea | B1 | |
| AU2016218381A1 | Australia | A1 | |
| KR20170097791A | Republic of Korea | A | |
| WO2017192917A1 | World Intellectual Property Organization (WIPO) | A1 | |
| CN107408074A | China | A | |
| US9836234B2 | United States of America | B2 | |
| EP3256939A1 | European Patent Office (EPO) | A1 | |
| JP2018506123A | Japan | A | |
| US9934089B2 | United States of America | B2 | |
| US2018101321A1 | United States of America | A1 | |
| AU2018202365A1 | Australia | A1 | |
| US9967342B2 | United States of America | B2 | |
| US2018225174A1 | United States of America | A1 | |
| EP3256939A4 | European Patent Office (EPO) | A4 | |
| EP3436923A1 | European Patent Office (EPO) | A1 | |
| CN109416620A | China | A | |
| JP2019517063A | Japan | A | |
| US10379763B2 | United States of America | B2 | |
| US2019356736A1 | United States of America | A1 | |
| US2019369885A1 | United States of America | A1 | |
| EP3436923A4 | European Patent Office (EPO) | A4 | |
| US2020045111A1 | United States of America | A1 | |
| US10574754B1 | United States of America | B1 | |
| AU2018202365B2 | Australia | B2 | |
| EP3036639B1 | European Patent Office (EPO) | B1 | |
| US10671480B2 | United States of America | B2 | |
| KR102118306B1 | Republic of Korea | B1 | |
| US2020257591A1 | United States of America | A1 | |
| US10838633B2 | United States of America | B2 | |
| JP6796589B2 | Japan | B2 | |
| AU2016218381B2 | Australia | B2 | |
| US2021081119A1 | United States of America | A1 | |
| US2021160318A1 | United States of America | A1 | |
| JP2021099814A | Japan | A | |
| US11057468B1 | United States of America | B1 | |
| JP6903005B2This record | Japan | B2 | |
| CN107408074B | China | B | |
| US2021314404A1 | United States of America | A1 | |
| US2021329072A1 | United States of America | A1 | |
| US11310317B1 | United States of America | B1 | |
| CN109416620B | China | B | |
| US2022217206A1 | United States of America | A1 | |
| US2022232075A1 | United States of America | A1 | |
| US11399063B2 | United States of America | B2 | |
| JP7135129B2 | Japan | B2 | |
| US11500552B2 | United States of America | B2 | |
| US2023058369A1 | United States of America | A1 | |
| US11652884B2 | United States of America | B2 | |
| US11671496B2 | United States of America | B2 | |
| US11677825B2 | United States of America | B2 | |
| EP3436923B1 | European Patent Office (EPO) | B1 | |
| EP3436923C0 | European Patent Office (EPO) | C0 | |
| US11714715B2 | United States of America | B2 | |
| US2023275965A1 | United States of America | A1 | |
| US2023308512A1 | United States of America | A1 | |
| US2023376379A1 | United States of America | A1 | |
| US12101379B2 | United States of America | B2 | |
| US12137140B2 | United States of America | B2 | |
| US12141449B2 | United States of America | B2 | |
| US2024380815A1 | United States of America | A1 | |
| US12212624B2 | United States of America | B2 | |
| US12341848B2 | United States of America | B2 | |
| US2025317491A1 | United States of America | A1 | |
| US12549632B2 | United States of America | B2 |
21 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Transfer to examiner for re-examination before appeal (zenchi)AppealJAPANESE INTERMEDIATE CODE: A911A911 | A911 | |
| Notice of transfer of a case for reconsideration by examiners before appeal proceedingsAppealJAPANESE INTERMEDIATE CODE: C21C21 | C21 | |
| Written invitation by the commissioner to file amendmentsJAPANESE INTERMEDIATE CODE: C11C11 | C11 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Trial request (containing other claim documents, opposition documents)OppositionJAPANESE INTERMEDIATE CODE: C60C60 | C60 | |
| Decision of refusalJAPANESE INTERMEDIATE CODE: A02A02 | A02 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Notification of resignation of power of attorneyJAPANESE INTERMEDIATE CODE: A7424RD04 | RD04 | |
| Notification of appointment of power of attorneyJAPANESE INTERMEDIATE CODE: A7423RD03 | RD03 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 6903005
- Application
- 2017516635
Titles2
- Japanese
- ストレージクラスター
- English
- Storage cluster
Classification
- CPC, 18
- G06F12/0246
- G06F11/1068
- G06F11/1076
- G06F11/108
- G06F3/0607
- G06F3/0619
- G06F3/0632
- G06F3/065
- G06F3/067
- G06F2212/7207
- G06F3/06
- G06F3/0655
- G06F3/0688
- G06F11/1092
- G06F2201/845
- G06F2212/7206
- G06F3/0613
- H03M13/154
- IPC, 5
- G06F11 10
- G06F3 06
- G06F3 08
- G06F11 20
- G06F13 10
