Information processing system, i/o switch, and rotation processing method of i/o path
12 claims: 3 independent, 9 dependent
- 1プロセッサとメモリを備えた物理サーバと、 前記物理サーバの計算機資源を仮想化して仮想サーバを実行する仮想化部と、 前記仮想サーバを実行する物理サーバを複数備え、これら物理サーバと1つ以上のI/Oデバイスを接続するI/Oスイッチと、を備えて、前記仮想化部が前記仮想サーバのマイグレーションを行う情報処理システムであって、 前記I/Oスイッチは、 前記I/Oデバイスから仮想サーバへのトランザクション発行の抑止を指示する抑止指示情報を格納するレジスタと、 前記レジスタに前記抑止指示情報が格納されたときに、前記I/Oデバイスからマイグレーションを行う仮想サーバへのトランザクションの発行を抑止し、前記トランザクションの抑止前に前記I/Oデバイスから発行されたトランザクションの完了を保証するトランザクション抑止制御部と、 前記仮想サーバのメモリアドレスと、当該仮想サーバを実行する物理サーバのメモリ上のアドレスとの対応関係を保持するアドレス変換部を備えて、前記仮想サーバのメモリアドレスを前記物理サーバのメモリ上のアドレスに変換する仮想化アシスト部と、 当該I/Oスイッチの構成を管理するスイッチ管理部と、 前記I/Oデバイスから該物理サーバへのトランザクションを保持する第1のバッファと、 前記物理サーバから前記I/Oデバイスへのトランザクションを保持する第2のバッファと、を備え、 前記仮想化部は、 前記物理サーバ上の仮想サーバと他の仮想サーバに割り付けられたI/Oデバイスの対応を管理するI/Oデバイス管理部と、 前記I/Oデバイス管理部からマイグレーション対象の仮想サーバに割り付けられたI/Oデバイスを特定し、当該I/Oデバイスに対応する前記I/Oスイッチの前記レジスタに対して、前記抑止指示情報を設定および解除するトランザクション指示部と、 前記スイッチ管理部に対して当該I/Oスイッチの構成変更を指示し、前記仮想化アシスト部で管理されるアドレス変換部に対して物理サーバの前記アドレスの変更を指令する構成変更指示部と、を備え、 前記トランザクション指示部は、マイグレーションの開始時に前記抑止指示情報を前記レジスタに設定して、前記I/Oデバイスからマイグレーション対象の仮想サーバへのトランザクション発行を抑止し、前記マイグレーションが完了すると前記レジスタに設定した抑止指示情報を解除して前記レジスタに設定して前記I/Oデバイスからマイグレーションが完了した前記仮想サーバへのトランザクション発行を許可し、 前記トランザクション抑止制御部は、 前記第1のバッファからのトランザクションの発行を抑止するトランザクション抑止部と、 レスポンス付きトランザクションを生成して物理サーバに発行するレスポンス付きトランザクション発行部と、 前記第2のバッファを監視し、前記レスポンス付きトランザクションの完了を確認する応答確認部と、 前記応答確認部が前記レスポンス付きトランザクションの完了を確認したときに、トランザクション抑止の完了を前記仮想化部に通知する完了通知部と、を含むことを特徴とする情報処理システム。
- 2前記情報処理システムは、 前記スイッチ管理部を制御するI/Oマネージャをさらに有し、 前記I/Oマネージャは、前記仮想化部と通信を行う設定インタフェースを備え、 前記仮想化部が当該設定インタフェースを介して前記I/Oマネージャに対して構成変更を指示すると、前記I/Oマネージャが前記スイッチ管理部を制御してI/Oスイッチの構成を変更することを特徴とする請求項1に記載の情報処理システム。
- 3前記レジスタは、 前記抑止指示情報を格納するフィールドと、アドレス情報を格納するフィールドと、を有することを特徴とする請求項1に記載の情報処理システム。
- 4前記アドレス情報を格納するフィールドは、前記マイグレーション対象の仮想サーバのメモリのアドレスを格納することを特徴とする請求項3に記載の情報処理システム。
- 5前記情報処理システムは、 前記物理サーバを管理するサーバマネージャをさらに有し、 前記サーバマネージャは、 前記仮想化部に対して、前記仮想サーバのマイグレーションの開始を指示する開始指示部を備えたことを特徴とする請求項1記載の情報処理システム。
- 6前記レスポンス付きトランザクションは、先行するトランザクションを追い抜かないことを特徴とする請求項1に記載の情報処理システム。
- 71つ以上の物理サーバと1枚以上のI/Oデバイスを接続するI/Oスイッチにおいて、 前記I/Oスイッチは、 前記物理サーバを接続する1つ以上の第1のポートと、前記I/Oデバイスを接続する1つ以上の第2のポートを備え、 前記第2のポートは、 前記I/Oデバイスからのトランザクション発行の抑止を指示する抑止指示情報を保持するレジスタと、 当該第2のポートに接続する前記I/Oデバイスから物理サーバへのトランザクションを保持する第1のバッファと、 前記物理サーバから該第2のポートに接続するI/Oデバイスへのトランザクションを保持する第2のバッファと、 前記レジスタに前記抑止指示情報が設定されたことを契機に前記第1のバッファからのトランザクション発行を抑止するトランザクション抑止部と、 該トランザクション抑止部によるトランザクション発行の抑止後に、前記抑止指示情報に基づきレスポンス付きトランザクション要求を発行し、前記レスポンス付きトランザクションの完了を確認する滞留トランザクション完了確認部と、を備え、 前記滞留トランザクション完了確認部は、 前記レスポンス付きトランザクション要求を生成して前記物理サーバに発行するレスポンス付きトランザクション発行部と、 前記第2のバッファを監視し、前記レスポンス付きトランザクションの完了を確認する応答確認部と、を含んで、前記I/Oデバイスから前記物理サーバへ発行されたトランザクションの完了を保証することを特徴とするI/Oスイッチ。
- 8前記レジスタは、 前記抑止指示情報を格納するフィールドと、 アドレス情報を保持するフィールドと、を有することを特徴とする請求項7に記載のI/Oスイッチ。
- 9前記物理サーバは、 該物理サーバ上で稼動する1つ以上の仮想サーバと、 該仮想サーバに割当てられたI/Oデバイスの対応を管理する仮想化部と、を備え、 該仮想化部が、前記レジスタに対して前記抑止指示情報を設定することを特徴とする請求項7に記載のI/Oスイッチ。
- 10前記第2のポートは、 前記レジスタの前記抑止指示情報が解除されたのを契機に、前記I/Oデバイスから該物理サーバへのトランザクションを保持する第1のバッファからのトランザクション発行を再開するトランザクション再開部を含むことを特徴とする請求項7に記載のI/Oスイッチ。
- 11前記レスポンス付きトランザクションは、先行するトランザクションを追い抜かないことを特徴とする請求項7に記載のI/Oスイッチ。
- 121つ以上のサーバと1枚以上のI/Oデバイスを1つ以上のI/Oスイッチで接続する情報処理システムにおいて、前記I/Oスイッチに設定したI/Oパスを切り替えるI/Oパスの交替処理方法であって、 前記I/Oスイッチは、 前記I/Oデバイスからのトランザクション発行の抑止を指示する抑止指示情報を保持するレジスタと、 前記I/Oデバイスから前記サーバへのトランザクションを抑止し、前記トランザクションの抑止前に前記I/Oデバイスから発行されたトランザクションの完了を通知するトランザクション抑止制御部と、 前記I/Oスイッチの構成制御を管理するスイッチ管理部と、を備え、 前記サーバを停止するステップと、 前記レジスタに対して前記抑止指示情報を設定するステップと、 前記I/Oデバイスから該物理サーバへのトランザクションを保持す る第 1のバッファからのトランザクションの発行を抑止するステップと、 レスポンス付きトランザクションを生成して物理サーバに発行するステップと、 前記物理サーバから前記I/Oデバイスへのトランザクションを保持す る第 2のバッファを監視して、前記レスポンス付きトランザクションの完了を確認するステップと、 前記レスポンス付きトランザクションの完了を確認したときに、トランザクション抑止の完了を通知するステップと、 前記完了通知を受け、前記サーバに割り付けられたI/Oデバイスの構成変更を行うステップと、 前記レジスタに設定された該抑止指示情報を解除するステップと、 前記サーバの停止を解除するステップと、を含むことを特徴とするI/Oパスの交替処理方法。
Independent claims12
130 paragraphs, as filed
The present invention relates to an information processing system in which a server computer and an I / O device are connected by a PCI switch, and more particularly to a technique of virtualizing a server computer and moving it to another physical computer.
In recent years, with the increase in operation management costs due to the increase in the number of servers that make up IT systems and the high performance of physical servers (physical computers) due to CPU multi-core processing, etc., one physical server is logically divided and virtualized. Server consolidation that reduces the number of physical servers by using server virtualization technology that operates as a server is attracting attention.
There are several implementation methods for the above server virtualization technology, and among them, the virtualization method using a hypervisor is known to be characterized by a small overhead associated with virtualization (for example, Non-Patent Document 1). ). The server virtualization method shown in Non-Patent Document 1 includes a hypervisor implemented as firmware of a physical server and a virtualization assist function implemented as hardware. The hypervisor controls each virtual server (virtual computer) on the physical server, and the virtualization assist function translates the address of memory access from the I / O device. In the server virtualization method shown in Non-Patent Document 1, since the hypervisor directly assigns the I / O device to the virtual server, the same OS environment can be used between the physical server and the virtual server. In addition, since the hardware-based virtualization assist function performs the address translation processing of the I / O device, the address translation processing overhead associated with the I / O processing can be reduced and the I / O throughput can be improved.
On the other hand, as the performance of physical servers increases due to CPU multi-core processing, I / O throughput according to the performance of physical servers is required. As a technology for improving the I / O throughput according to the performance of the physical server, a PCI switch capable of connecting a plurality of I / O devices to a plurality of physical servers is promising.
Furthermore, standardization of I / O virtualization technology, which combines the above server virtualization technology and PCI switch to make the correspondence between virtual servers and I / O devices flexible, is being promoted (SR / MR-IoV (non-patented). Reference 2)).
Against this background, information processing systems that combine server virtualization technology and PCI switches are expected to become mainstream in the future.
By the way, live migration is one of the functions to increase the flexibility and availability of computer systems to which server virtualization technology is applied (Non-Patent Document 3).
Live migration is a function that moves a running virtual server to another physical server. With live migration, you can change the layout of running virtual servers according to the load on the physical server, and save the running virtual servers from the physical servers to be maintained. As a result, the flexibility and availability of the system can be increased.
When realizing live migration in an information processing system that combines server virtualization technology and PCI switches as described above, it is necessary to maintain and take over the three states of the virtual server in operation. The three states are (a) CPU operating state, (b) memory contents, and (c) I / O device state. The CPU operating state of (a) can be maintained by stopping the operation of the virtual server by the hypervisor.
However, since a normal hypervisor cannot stop memory access such as DMA from the I / O device assigned to the virtual server and transaction processing originating from an external I / O device, (b) memory contents and (c) I / O The device state cannot be maintained. Therefore, when live migration is realized in the server virtualization technology as described above, the influence of the operating I / O device is eliminated and (b) the memory contents and (c) the I / O device state are retained. There is a need.
There are several techniques for maintaining these states.
For example, Patent Document 1 describes an I / O device during migration in order to realize migration of a virtual I / O device in a system in which a physical server and an I / O device are connected by a PCI switch and managed by a PCI manager. A method of stopping transaction processing from is disclosed.
In Patent Document 1, a migration bit is provided in the configuration space of the I / O device, and the I / O device refers to the migration bit to determine whether or not the virtual I / O device is migrating.
The I / O device stops its own processing if it is migrating, and suppresses the issuance of transactions from the I / O device to the memory. In addition, the PCI manager saves the transaction from the I / O device to the memory in the PCI switch to a specific storage area during migration, and restores the transaction saved in the storage area after the migration is completed.
This prevents transactions originating from the I / O device during migration from rewriting the memory contents and changing the I / O device state.
Further, in Patent Document 2, a method in which a host OS provides a device emulator emulating an I / O device to a virtual server and the guest OS of the virtual server indirectly accesses the I / O device using the device emulator. Is disclosed.
The host OS recognizes that the migration is in progress, and if the migration is in progress, the memory contents and the I / O device state of the virtual server being migrated can be retained by stopping the processing of the emulation device.<patcit num="1"><text>U.S. Patent Application Publication No. 2007/0186025</text></patcit><patcit num="2"><text>U.S. Pat. No. 6,496,847</text></patcit><nplcit num="1"><text>Co-authored by Hitoshi Ueno et al., "Virtage", a server virtualization mechanism for "Blade Symphony" that improves the operational efficiency of information systems, July 2007, Internet <http://www.hitachihyoron.com/2007/07/pdf/ 07a10.pdf></text></nplcit><nplcit num="2"><text>"I / O Virtualization and Sharing", co-authored by Michael Krause et al., November 2006, http://www.pcisig.com/developers/main/training_materials/</text></nplcit><nplcit num="3"><text>"Live Migration of Virtual Machines", co-authored by Christopher Clark et al., May 2005, NSDI (Networked Systems Design and Implementation) '05</text></nplcit>
<p> However, the method using the above-mentioned Patent Document 1 as in the above-mentioned conventional example can be applied only when the I / O device has a function of determining whether or not the virtual server is being migrated. Therefore, there is a problem that general-purpose I / O cards that are widely distributed for PC applications cannot be targeted.</p><p> Further, in the method using the above-mentioned Patent Document 2, since the virtual server needs to switch between the guest OS and the host OS in order to access the I / O device, this switching becomes an overhead and the processing performance is improved. There is a problem that it decreases. Furthermore, when allocating an I / O device directly to a virtual server, there is also a problem that the device emulator method cannot be applied.</p><p> In order to solve the above problem, in an information system in which a physical server and an I / O device are connected via a PCI switch, processing is performed even if a general-purpose I / O device is directly assigned to the virtual server. It is necessary to retain the memory contents and I / O device state of the virtual server during migration while reducing the overhead of.</p><p> The subject of the present invention is to reduce the processing overhead and the memory contents and I / O device state of the virtual server being migrated even when a general-purpose I / O device is directly assigned to the virtual server. It is to provide a holding mechanism.</p>
<p> According to one embodiment of the present invention, in an information processing system having a PCI switch for connecting a plurality of physical servers and one or more I / O devices, the physical server is a virtual server and I assigned to the virtual server. It has a virtualization unit that manages the correspondence of / O devices.</p><p> The I / O switch is a register that indicates a request to suppress transaction issuance from the I / O device to the virtual server, and an I / O device that suppresses transactions from the I / O device to the virtual server and is issued before suppression. It is equipped with a transaction suppression control unit that guarantees the completion of transactions from, a virtualization assist unit that converts the virtual server address to an address on the memory of the physical server, and a switch management unit that manages the configuration of the I / O switch.</p><p> The virtualization unit manages the switch by receiving a transaction suppression request (suppression instruction information) and a transaction instruction unit that sets the memory address of the virtual server for the register of the I / O switch, and a completion notification from the I / O switch. It is provided with a configuration change instruction unit that instructs the unit to change the configuration and an instruction to change the address conversion unit to the virtualization assist unit.</p><p> The transaction indicator suppresses transactions from the I / O device to the virtual server and performs transaction completion guarantee processing from the I / O device issued before the suppression, and the transaction from the I / O device is in the memory state of the virtual server. Do not rewrite. In addition, the configuration change instruction unit updates the address translation unit of the virtualization assist unit so that the memory address such as DMA held by the I / O device remains valid even after the virtual server is moved, and the status of the I / O device is changed. maintain.</p>
<p> Therefore, the present invention determines the memory contents and I / O device state of the virtual server being migrated while reducing the processing overhead even when a general-purpose I / O device is directly assigned to the virtual server. Can be retained. This makes it possible to smoothly move the virtual computer to another physical computer in a state where a general-purpose I / O device is assigned to the virtual computer via an I / O switch such as a PCI switch.</p>
Hereinafter, an embodiment of the present invention will be described with reference to the accompanying drawings.
Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings.
FIG. 1 is a block diagram showing an example of a configuration of an information processing system 100 provided with a mechanism for holding a state of a virtual server according to the first embodiment of the present invention.
The information processing system 100 includes one or more physical servers 110a to 110b, one or more I / O devices 120a to 120b, a server manager 140 that controls a virtual server on the physical server, and an I / O device 120a ~. It includes a PCI switch 150 that connects 120b and physical servers 110a to 110b, and a PCI manager 130 that manages the PCI switch 150. The physical servers 110a to 110b, the PCI manager 130, the server manager 140, and the I / O devices 120a to 120b are connected via the PCI switch 150. The physical servers 110a to 110b, PCI manager 130, and server manager 140 are Ethernet (registered trademark) or I2C (Inter-Integrated). It is connected by a management network 102 such as Circuit). Alternatively, a mechanism for inbound access via the PCI switch 150 may be provided. Any connection method may be used as long as information can be exchanged between the physical servers 110a to 110b, the PCI manager 130, and the server manager 140. Further, in the first embodiment of the present invention, an example in which two physical servers and two I / O devices are connected to a PCI switch is shown, but the number is not limited to this.
The physical server 110a includes hardware 116a and a hypervisor 111a, and the virtual server 115a operates on the physical server 110a.
The hardware 116a includes a CPU (processor), a chipset, a memory, and the like, which are hardware resources on a physical server. The physical connection configuration of hardware 116a will be described later (see Fig. 3). The PCI manager 130 and the server manager 140 are also computers composed of a CPU, a chipset, a memory, and the like, like the hardware 116a of the physical server 110a.
The hypervisor 111a is a firmware or application implemented on the physical server 110a, and manages the virtual server 115a on the physical server 110a. Resources such as the CPU and memory of the hardware 116a managed by the hypervisor 111a are allocated to the virtual server 115a. In order to allocate CPU resources to the virtual server 115a, the hypervisor 111a holds a CPU scheduler (not shown) that manages the allocation of CPU resources to the virtual server 115a. A well-known or known technique may be used for the CPU scheduler, and detailed description thereof will be omitted in the embodiment of the present invention.
In the first embodiment of the present invention, only the virtual server 115a is shown, but the number of virtual servers is not limited to one. The hypervisor 111a creates a virtual server as needed.
The hypervisor 111a includes an I / O device management table 117a, an I / O departure Tx suppression instruction unit 112a, an I / O configuration change instruction unit 113a, an I / O departure Tx restart instruction unit 114a, and a setting interface 101a. ..
I / O device management Table 117a manages the correspondence between the virtual server 115a and the I / O devices 120a and 120b assigned to the virtual server 115a. The configuration of I / O device management table 117a will be described later (see Fig. 2).
The I / O departure Tx (transaction) suppression instruction unit 112a instructs the PCI switch 150 to suppress transactions issued by the I / O devices 120a and 120b assigned to the virtual server 115a. The Tx suppression indicator 112a from I / O issues a write transaction to the setting register 161 provided in the configuration register 158 of the PCI switch 150, for example.
The I / O configuration change instruction unit 113a instructs the PCI manager 130 to change the configuration of the PCI switch 150. Further, the I / O configuration change instruction unit 113a instructs the virtualization assist unit 153 of the PCI switch 150 to change the address conversion table 152.
The I / O departure Tx restart instruction unit 114a instructs the PCI switch 150 to issue a transaction (I / O departure transaction) from the I / O devices 120a and 120b assigned to the virtual server 115a. Specifically, a write transaction is issued to the setting register 161 provided in the PCI switch 150.
The setting interface 101a is an interface for exchanging setting information between the server manager 140 and the PCI manager 130 and the hypervisor 111a. The setting information may be exchanged using a normal network, or an information sharing register may be prepared in the hypervisor 111a and information may be exchanged via register access. The information to be exchanged is, for example, information that triggers an I / O configuration change.
The physical server 110a is configured as described above, and the physical server 110b also includes a hypervisor 111b, a virtual server 115b, and hardware 116b configured in the same manner as the physical server 110a.
The I / O devices 120a to 120b are general-purpose I / O interface cards such as NICs and Fiber Channel HBAs.
The PCI manager 130 is a calculator that includes a management program that manages the PCI switch 150. The PCI manager 130 is implemented on hardware equipped with a CPU, memory, and the like. The PCI manager 130 includes an I / O configuration change unit 131 for changing the configuration of the PCI switch 150, and a setting interface 101d for exchanging setting information between the hypervisors 111a to 111b and the server manager 140. Further, in the first embodiment of the present invention, only a single PCI manager 130 is shown, but a plurality of PCI managers may be provided for improving reliability. In that case, information is controlled so that consistency can be obtained among a plurality of PCI managers.
The server manager 140 is a computer including a program that manages the entire information processing system 100, and is implemented on hardware including a CPU and memory. The server manager 140 includes a processing start instruction unit 141 that instructs the start of migration processing of the virtual server 115a (115b), and a setting interface 101e that exchanges setting information between the hypervisors 111a to 111b and the PCI manager 130. .. In the first embodiment of the present invention, only a single server manager 140 is shown, but a plurality of server managers may be provided for improving reliability. In this case, the information is controlled so that it is consistent among multiple server managers.
The PCI switch 150 is a switch fabric that connects a plurality of physical servers 110a to 110b and I / O devices 120a to 120b, and sends and receives transactions between the physical servers 110a and 110b and the I / O devices 120a and 120b. .. The details of the transaction will be described later (see Fig. 10).
The information processing system 100 shown in the first embodiment of the present invention shows only a single PCI switch 150, but may include a plurality of PCI switches. Each PCI switch may be managed by a separate PCI manager, or may be managed by a single PCI manager. Further, in the first embodiment of the present invention, the PCI switch 150 is a switch fabric for the PCI-Express protocol. However, the PCI switch 150 may be a switch fabric (I / O switch) for other protocols such as the PCI protocol and the PCI-X protocol.
The PCI switch 150 has a configuration that stores one or more upstream ports 151a to 151c, one or more downstream ports 160a to 160b, a switching unit 157, a PCI switch management unit 154, a setting register 161 and a routing table. A switch register 158 is provided, and the switching unit 157 includes a virtualization assist unit 153. In the first embodiment of the present invention, three Upstream ports and two Downstream ports are shown, but the number of ports is not limited to this number.
Upstream ports 151a to 151c are ports for connecting the physical servers 110a and 110b and the PCI manager 130.
Downstream ports 160a to 160b are ports for connecting I / O devices 120a and 120b. The Downstream port 160a includes a Tx suppression control unit 162 and a setting register 161. The Tx suppression control unit 162 suppresses and restarts transactions from the I / O devices 120a and 120b assigned to the virtual servers 115a and 115b according to the information set in the setting register 161. The setting register 161 holds the deterrence instruction request issued from the Tx deterrence instruction unit 112a from the I / O of the hypervisor 111a. Details of Downstream port 160a will be described later (see Fig. 5).
The PCI switch management unit 154 is a functional element that controls the switching unit 157 of the PCI switch 150, and operates in cooperation with the PCI manager 130. The PCI switch management unit 154 manages the PCI switch 150 by dividing it into one or a plurality of PCI trees. A PCI tree consists of a pair of Upstream ports and one or more Downstream ports. The PCI switch management unit 154 upstream prevents physical servers 110a and 110b and I / O devices 120a and 120b connected to a port of a PCI tree from accessing physical servers and I / O devices that are not connected to the PCI tree. Routes transactions between ports and Downstream ports. When a plurality of PCI trees exist, the PCI switch management unit 154 uniquely identifies the PCI tree by using the PCI tree identifier. In addition, the PCI switch management unit 154 rewrites the configuration register 158 of the PCI switch 150 according to the instruction from the I / O configuration change unit 131 of the PCI manager 130, and changes the settings of the Upstream port and the Downstream port belonging to the PCI tree. I do.
The switching unit 157 transfers transactions between the Upstream port and the Downstream port according to the tree structure managed by the PCI switch management unit 154.
The virtualization assist unit 153 is a control circuit located on the paths 155 and 156 (switching unit 157) connecting the upstream ports 151a to 151c and the downstream ports 160a to 160b, and includes the address translation table 152. The address translation table 152 is a table that manages the correspondence between the addresses (virtual addresses) of the virtual servers 115a and 115b and the addresses (physical addresses) of the physical servers 110a and 110b. Details of the address translation table 152 will be described later (see FIG. 4). The virtualization assist unit 153 converts the address of the transaction issued from the I / O devices 120a and 120b to the virtual servers 110a and 110b by referring to the address translation table 152, and transfers the transaction to the memory of the physical servers 110a and 110b. Is issued. Further, the virtualization assist unit 153 operates in cooperation with the hypervisors 111a and 111b, sets the correspondence between the virtual address and the physical address set in the address translation table 152, and changes the correspondence as necessary. In the first embodiment of the present invention, the virtualization assist unit 153 is shown as a single control element, but it may be distributed to each upstream port or downstream port.
FIG. 3 is a diagram illustrating the configuration of the hardware 116a of the physical server 110a. Hardware 116a includes CPU 301, chipset 302, memory 303, and Root port 304. CPU301 is a processor that executes a program. The CPU 301 accesses the register of the PCI switch 150 via MMI / O (memory-mapped I / O) in which the register of the PCI switch 150 is mapped to a part of the memory address area. In addition, the CPU 301 and the chipset 302 support processing of a CPU-generated Tx suppression request (Quiescence request) and a CPU-based Tx suppression release request (Dequiesce request), and can suppress and release a CPU 301-based transaction. Further, the chipset 302 connects the CPU 301, the memory 303, and the root port 304. The memory 303 is the main storage area of the physical server 110a, and a part of the area is used as the memory area of the virtual server 115a. The Root port 304 connects to the Upstream port 151b of the PCI switch 150 and serves as the root of the connecting PCI tree.
FIG. 10 shows the configuration of transactions sent and received between the physical servers 110a and 110b and the I / O devices 120a and 120b. A transaction is a packet consisting of header 1201 and payload 1202.
Header 1201 is information required for the PCI switch 150 to route a transaction, and includes a transaction source identifier 1203 (Requester ID), a destination address 1204, and a traffic class 1205. If there are multiple PCI trees, the PCI tree identifier is added to the header 1201. Since the header information is determined by the standardized specifications, detailed description thereof will be omitted in the embodiment of the present invention. Here, only the header information related to the present invention will be described.
The transaction source identifier 1203 is an identifier composed of a bus number, a device number, and a function number in the PCI tree of the source I / O device or root port, and is a downstream port or root connected to the source I / O device. The Upstream port that connects to the port can be uniquely identified.
The destination address 1204 is the address of the memory of the transaction destination. The traffic class 1205 is information for uniquely identifying the Virtual Channel in the PCI switch 150 through which the transaction passes. In the embodiment of the present invention, the traffic class can be set to a value from 0 to 7.
Payload 1202 stores the data held by the transaction. For example, in the case of a memory write transaction, the data to be written to the memory is stored.
FIG. 2 is a diagram showing the configuration of I / O device management table 117a. The I / O device management table 117b of the physical server 110b is also configured in the same manner.
The I / O device management table 117a is managed by the hypervisor 111a and includes a virtual server identifier 201 and an I / O device identifier 202. The virtual server identifier 201 is a number that uniquely identifies the virtual server 115a in the physical server 110a that holds the I / O device management table 117a. The I / O device identifier 202 is an identifier that uniquely identifies the I / O device assigned to the physical server 110a. In the first embodiment of the present invention, the transaction source identifier 1203 shown in FIG. 10 is used as the I / O device identifier 202. The transaction source identifier is included in the transaction header 1201 and can uniquely identify the Downstream port to which the I / O device connects. However, other identifiers may be used as long as they are included in the transaction and can uniquely identify the Downstream port to which the I / O device connects.
FIG. 4 is a diagram showing the details of the address translation table 152. The address translation table 152 is managed for each PCI tree and holds a set consisting of the I / O device identifier 401 and the translation offset 402. The I / O device identifier 401 uniquely identifies the I / O device in the PCI tree by using the transaction source identifier 1203 included in the header of the transaction originating from the I / O. The conversion offset 402 indicates the start address of the memory area of the virtual server 115a to which the I / O device is assigned in the memory area of the physical server 110a.
The virtualization assist unit 153 obtains a conversion offset corresponding to the transaction source identifier 1203 included in the transaction from I / O with reference to the address conversion table 152, and uses the conversion offset to send the destination of the transaction from I / O. Translate address 1204 from a virtual address to a physical address. In the first embodiment of the present invention, the address translation table 152 is divided and managed for each PCI tree, but a plurality of PCI trees may be collectively managed by one address translation table.
Figure 5 shows the block configuration of Downstream port 160a. The Downstream port 160a is connected to the setting register 161, the Tx suppression control unit 162, the receive buffer 507 that sends the transaction from the connected I / O device to the Upstream port, and the transaction from the Upstream port to this Downstream port. It is equipped with a transmission buffer 508 for transmitting to the I / O device, and sends and receives transactions flowing between the I / O devices 120a and 120b and the physical server 110a.
As shown in FIG. 11, the setting register 161 is provided in the configuration register 158 of the PCI switch 150 and is provided for each of the downstream ports 160a and 160b. The setting register 161 corresponds to the downstream port 160a and is set. Register 161b corresponds to Downstream port 160b. The setting registers 161, 161b are registers that can be accessed from the physical servers 110a and 110b and the PCI manager 130 via MMI / O, and include fields for storing the suppression bit 509 and the address 510 for each Downstream port. The suppression bit 509 is set and released according to the instructions from the I / O departure Tx suppression instruction unit 112a and the I / O departure Tx restart instruction unit 114a of the hypervisor 111a. In the first embodiment of the present invention, a transaction for writing to the PCI configuration space is used. For example, the I / O outgoing Tx suppression indicator 112 sets the suppression bit 501 to 1 by a write transaction to the configuration space, and the I / O outgoing Tx restart indicator 114a uses a write transaction to the configuration space. Then, the value of the suppression bit 509 is cleared to 0. The virtual address of the virtual server to be migrated is set in the address 510 of the setting register 161 according to the instruction from the Tx suppression instruction unit 112a from I / O. In addition, the Tx restart instruction unit 112a from I / O clears the value of address 510 by using a write transaction to the configuration space.
The Tx suppression control unit 162 includes a Tx suppression unit 501, a retention Tx completion confirmation unit 502, and a Tx restart unit 503.
The Tx suppression unit 501 controls the transactions originating from I / O in the receive buffer 507 so as not to be issued to the upstream port 507. For example, the Tx suppression unit 501 does not return an ACK (response) to the transaction issued by the receive buffer 507 by using the flow control mechanism performed between the path 156 (switching unit 157) in the PCI switch 150 and the receive buffer 507. By controlling in this way, the issuance of transactions in the receive buffer 507 is suppressed.
The stagnant Tx completion confirmation unit 502 includes a Tx issue unit 504 with a response, a Tx response confirmation unit 505, and a Tx completion notification unit 506, and guarantees the completion of the I / O issue transaction issued before suppression.
The response-attached Tx issuer 504 generates a memory read transaction with the address 510 of the setting register 161 as the destination address, and issues the generated memory read transaction. A memory read transaction is a type of Tx with a response.
The Tx response confirmation unit 505 confirms the reception of the response of the memory read transaction issued by the Tx issuer 504 with a response. In the first embodiment of the present invention, the Tx response confirmation unit 505 monitors the transmission buffer 508, and confirms that the transmission buffer 508 has received the response of the memory read transaction issued by the Tx issuer 504 with a response.
Upon receiving the confirmation of the response by the Tx response confirmation unit 505, the Tx completion notification unit 506 notifies the hypervisor 111a of the physical server assigned to the I / O device connected to the Downstream port 160a of the completion of the Tx suppression control. To do.
When the Tx restart unit 503 detects that the suppression bit 509 of the setting register 161 is cleared, the Tx restart unit 503 restarts the issuance of the I / O transaction in the receive buffer 507. For example, when the Tx restart unit 503 releases the suppression of returning the ACK for the transaction issued by the receive buffer 507, the issuance of the transaction in the receive buffer 507 is restarted.
The Downstream port 160b is also configured in the same manner as the Downstream port 160a, and includes a setting register 161b and a Tx suppression control unit 162b. Further, the Upstream ports 151a to 151c connected to the computer need only have at least a transmission buffer and a reception buffer, and the setting register and the Tx suppression control unit may not be set.
FIG. 6 is a flowchart showing an example of the processing performed by the Tx suppression control unit 162 (or 162b) described with reference to FIG. Further, FIG. 7 is a diagram showing a transaction flow related to the processing of the Tx suppression control unit 162. Hereinafter, with reference to FIGS. 6 and 7 Tx suppression system describing the processing flow of the control section. In the following description, an example of processing related to the Downstream port 160a is shown, but the same processing can be performed for the Downstream port 160b.
This process in FIGS. 6 and 7 is started when the I / O device 120a is connected to the downstream port 160a (S600).
First, the Tx suppression unit 401 of the Tx suppression control unit 162 monitors the suppression bit 509 of the setting register 161 and checks whether the suppression bit is set (S601). Specifically, the Tx suppression control unit 162 repeatedly reads the value of the suppression bit 509 from the setting register 161 until it detects the transition of the suppression bit 509 of the setting register 161 from 0 to 1 (701 in FIG. 7). .. If the suppression bit 509 is set to 1, perform the next step S602. If the suppression bit 509 is 0, step S601 is repeated.
Next, the issuance of transactions in the receive buffer 507 is suppressed (S602). Specifically, the Tx suppression unit 501 suppresses the return of an ACK (response) to the transaction issued by the receive buffer 507 toward the upstream ports 151a to 151c (702 in FIG. 7).
Next, in order to guarantee the completion of the I / O-issued transaction issued before the transaction issuance is suppressed by S602, a Tx with a response is generated and issued for the address 510 held by the setting register 161 (S603). .. If the path between the I / O device and the physical server (I / O path) is divided into multiple paths, issue a transaction with a response for all I / O paths. Specifically, the Tx issuing unit 504 with a response of the staying Tx completion confirmation unit 502 acquires the address 510 from the setting register 161 (703 in FIG. 7), and is one of the Tx with a response to the address. Generate a read transaction (704 in Figure 7) and issue it to the Upstream port (705 in Figure 7). Also, if the PCI switch 150 has multiple Virtual Channels, all available Virutals of the PCI switch 150 To guarantee the completion of I / O outgoing transactions for Channel, issue 8 types of memory read transactions with headers of traffic class 0 to 7. By issuing memory read transactions for all traffic classes, I / O outgoing transactions can be guaranteed for all available Virtual Channels.
Next, confirm the completion of the transaction with response issued in S603 (S604). According to the PCI-Express ordering rule, the transaction with response does not overtake the preceding memory write transaction, so when the issued transaction with response is completed, the memory write transaction from I / O issued before suppression is completed. Guaranteed to be. That is, after confirming the completion of the transaction with response issued by S603, the memory contents of the virtual server 115a are retained. Step S604 is repeated until the Tx response confirmation unit 505 of the retention Tx completion confirmation unit 502 confirms all the responses of the memory read transaction (706 in FIG. 7).
Next, the Tx suppression control unit 162 sends a transaction suppression completion notification to the hypervisor 111a (S605). Specifically, the Tx completion notification unit 506 of the retention Tx completion confirmation unit 502 issues the processing completion notification 708 to the hypervisor 111a. The means for issuing the processing completion notification 708 may be an interrupt to the physical server 110a or a write to the memory 303 of the physical server 110a.
Next, wait until the suppression instruction information of the transaction originating from I / O is released (S606). Specifically, the Tx restart unit 503 checks whether the suppression bit 509 of the setting register 161 is cleared to 0 (701 in FIG. 7), and receives the transition of the suppression bit 509 to 0 and performs the next processing. Proceed to.
Finally, the suppression of transactions originating from I / O in the receive buffer 507 is released (S607). Specifically, the Tx restart unit 503 releases the suppression of the return of the ACK for the transaction issued by the receive buffer 507, thereby restarting the issuance of the transaction from the receive buffer 507 to the upstream ports 151a to 151c (Fig.). 7 709).
With the above series of processing, the processing of the Tx suppression control unit 162 is completed (S608). As described above, the Tx suppression control unit 162 can suppress the I / O outgoing transaction and guarantee the completion, so that the memory contents of the physical server 110a can be protected from the I / O outgoing transaction. That is, after starting the suppression of the transaction originating from I / O, the transaction with response is issued to the address of the migration source virtual server set in the address 510 of the setting register 161 and the completion of the issued transaction with response is confirmed. Therefore, it is possible to guarantee that the transaction issued before the start of suppression has been completed.
FIG. 8 is a flowchart of a process for realizing live migration of a virtual server in the information processing system 100. Further, FIG. 9 is a block diagram showing a main part of the information processing system 100 showing the transaction flow related to the live migration process.
The flowchart of FIG. 8 is described on the assumption that the virtual server 115a to which the I / O device 120a is assigned is migrated from the physical server 110a to the physical server 110b.
The virtual server 115a that is the target of migration is called the virtual server that is the target of migration. The physical server 110a in which the migration target virtual server exists is called the migration source physical server, and the hypervisor 111a on the migration source physical server 110a is called the migration source hypervisor. The migration destination physical server 110b is called a migration destination physical server, and the hypervisor 111b on the migration destination physical server 110b is called a migration destination hypervisor.
In the flowchart of FIG. 8, the processing start instruction unit 141 of the server manager 140 issues a migration start request for the migration target virtual server 115a to the migration source hypervisor 111a, and then the migration destination hypervisor 111b completes the migration to the server manager 140. The processing flow until notification is shown.
The live migration process is started when the migration source hypervisor 111a receives the migration start request 901 by the process start instruction unit 141 of the server manager 140 (S801).
The migration start request 901 in FIG. 9 includes the virtual server identifier of the migration source virtual server 115a and the identifier of the migration destination physical server 110b. The migration source hypervisor 111a receives the migration start request 901 and performs the following processing.
First, the migration source hypervisor 111a stops the migration target virtual server 115a (S802). Known means will be used for realization. In the first embodiment of the present invention, first, the migration source hypervisor 111a requests the CPU 301 and the chipset 302 of the hardware 116a to suppress the Tx from the CPU. Next, the hypervisor 111a changes the setting of the CPU scheduler so that the CPU resource is not allocated to the virtual server 115a, and stops the operation of the virtual server 115a. Finally, make a request to release Tx suppression from the CPU. By the above processing, the stop of the virtual server 115a and the completion of the transaction from the virtual server 115a are guaranteed.
Next, the migration source hypervisor 111a instructs the PCI switch 150 to suppress transactions from the I / O device 120a connected to the migration target virtual server 115a (S803). Specifically, the I / O departure Tx suppression instruction unit 112a of the migration source hypervisor 111a extracts the I / O device 120a assigned to the virtual server 115a by referring to the I / O device management table 117a. Next, the Tx suppression indicator 112a from I / O sends the setting register 161 of the downstream port 160a connected to the I / O device 120a to the configuration register 158 via the MMI / O of the PCI switch 150. Write (902 in Figure 9). Specifically, 1 is set in the suppression bit 509 of the setting register 161 and a part of the memory address used by the virtual server 115a is set in the address 510. As a result, the Tx suppression control unit 162 of the Downstream port 160a starts the transaction suppression processing from the I / O device 120a assigned to the virtual server 115a. The transaction suppression processing flow from this I / O device is as explained in Fig. 6.
Next, the migration source hypervisor 111a waits for the suppression completion notification of the transaction originating from the I / O device 120a from the Tx suppression control unit 162 of the PCI switch 150 (S804). When the migration source hypervisor 111a receives the suppression completion notification 708, it can be confirmed that all transactions addressed to the migration target virtual server 115a have been completed, and the memory contents of the virtual server 115a are not rewritten by the transaction originating from I / O. I can guarantee that.
Next, the migration source hypervisor 111a moves the migration target virtual server 115a from the migration source physical server 110a to the migration destination physical server 110b (S805). Known means will be used for realization. Specifically, the hypervisor 111a copies the OS and application images of the virtual server 115a to the migration destination physical server 110b. In the first embodiment of the present invention, the hypervisor 111a migrates the memory contents of the virtual server 115a and the configuration information of the virtual server 115a held by the hypervisor 111a by using outbound communication using the management network 102. Copy to the destination physical server 110b. However, the virtual server may be moved by inbound communication via the PCI switch 150.
Next, the migration source hypervisor 111a instructs to take over the I / O device 120a assigned to the migration target virtual server 115a (S806). Specifically, the I / O configuration change instruction unit 113a of the hypervisor 111a changes the setting of the PCI switch management unit 154 and changes the PCI tree to which the I / O device 120a connects. That is, the allocation of the I / O device 120a is switched from the physical server 110a to the virtual server 115a of the physical server 110b. Also, access the virtualization assist unit 153 and change the correspondence between the physical address and the virtual address in the address translation table 152. This change corresponds to the translation offset 402 of the entry where the I / O device identifier 401 in address translation table 152 matches the I / O device 120a, to the physical address of the memory of the physical server 110b to which the virtual server 115a is moved. Update to the offset value. That is, the configuration change instruction unit 113a of the hypervisor 111a instructs the PCI switch management unit 154 of the I / O device identifier 401 and the conversion offset 402 of the new physical server 110b. After that, the hypervisor 111a deletes the information about the I / O device 120a from the I / O device management table 117a because the virtual server 115a has been deleted from the physical server 110a.
A known means is used to change the setting of the PCI switch management unit 154. Specifically, the I / O configuration change instruction unit 113a of the hypervisor 111a issues an I / O configuration change request 906 to the I / O configuration change unit 131 of the PCI manager 130. The I / O configuration change request 906 is sent to the PCI manager 130 via the PCI switch 150 or the configuration interfaces 101a and 101d. The I / O configuration change request 906 includes the identifier of the PCI switch and the I / O device identifier of the I / O device to be taken over. When a plurality of PCI trees are included in the PCI switch 150, the I / O configuration change request 906 includes the PCI tree identifier of the takeover source PCI tree and the PCI tree identifier of the takeover destination PCI tree. The I / O configuration change unit 131 receives the I / O configuration change request 906 and issues the setting change request 907 to the PCI switch management unit 154.
Further, in order to change the setting of the address translation table 152, the I / O configuration change instruction unit 113a issues an address translation table update request 904 to the virtualization assist unit 153. The address translation table update request 904 contains information on the I / O device identifier and translation offset. If the PCI switch 150 contains multiple PCI trees, the address translation table update request 904 contains the PCI tree identifier. The virtualization assist unit 153 updates the address translation table 152 in response to the address translation table update request 904.
By changing the settings of the PCI switch management unit 154 and the settings of the virtualization assist unit 153, the transactions existing in the receive buffer 507 of the downstream port 160a can be allocated to the virtual server 115a on the migration destination physical server 110b. Written in.
Next, the migration destination hypervisor 111b waits for the completion notification of the I / O configuration change (S807). The completion notification may be explicitly received by the hypervisor 111b from the I / O configuration change section 131 of the PCI manager 130, or the hypervisor 111b shares a register that the I / O device 120a has been added to the migration destination physical server 110b. It may be detected via such as. Any method may be used as long as the hypervisor 111b can recognize that the I / O configuration change has been completed. The hypervisor 111b receives the completion notification of the I / O configuration change and adds information about the I / O device 120a to the I / O device management table 117b.
Next, the migration destination hypervisor 111b restarts the processing of the migration target virtual server 115a (S808). Known means will be used for realization. For example, the hypervisor 111b changes the CPU scheduler setting of the hypervisor 111b so as to allocate CPU resources to the virtual server 115a, and restarts the operation of the virtual server 115a.
Next, the migration destination hypervisor 111b instructs the PCI switch 150 to resume the transaction originating from the I / O device (S809). Specifically, the I / O departure Tx restart instruction unit 114b of the hypervisor 111b refers to the I / O device management table 117b and is connected to the I / O device 120a assigned to the migration target virtual server 115a. Extract port 160a. Next, register access is performed to the configuration register 161 corresponding to the Downstream port 160a by using a write transaction to the configuration space (908 in FIG. 9). In register access 908, the suppression bit 509 and address 510 of the setting register 161 are cleared to 0. The Tx restart unit 503 of the Downstream port 160a detects that the various setting information of the setting register 161 is cleared to 0, and restarts the transaction transmission from the I / O device 120a.
Finally, in step 810 (S810), the migration destination hypervisor 111b notifies the server manager 140 that the migration is complete (909 in FIG. 9).
The live migration is completed with the above series of processes (S811). As described above, by performing the live migration process, it is possible to prevent the transaction from the I / O device 120a assigned to the migration target virtual server 115a from being written to the memory area of the virtual server being migrated. Therefore, the memory state and the I / O device state of the live migration target virtual server 115a can be maintained.
<Modification example 1> The mechanism for holding the state of the virtual server described in the first embodiment of the present invention can be applied not only to the live migration use of the virtual server but also to the I / O path switching function. The I / O path switching function prepares multiple routes (I / O paths) between the physical server and the I / O devices assigned to the physical servers, the active system and the standby system, and ports on the I / O path. This is a function that fails over the I / O path from the active system to the standby system when a failure occurs. The I / O path switching function can prevent the information processing system from stopping due to a port failure of the PCI switch, improving the availability of the information processing system.
FIG. 12 is a block diagram showing a configuration of an information processing system 1000 having an I / O path switching function according to the first modification of the present invention.
The information processing system 1000 includes one or more physical servers 1010, one or more I / O devices 120a, a PCI manager 1030, a server manager 1040, and one or more PCI switches 150a to 150b. The physical server 1010, the PCI manager 1030, the server manager 1040, and the I / O device 120a are connected via PCI switches 150a to 150b. The PCI switch 150a and the PCI switch 150b are connected by two routes: Upstream port 151a and Downstream port 160c, and Upstream port 151b and Downstream port 160d. The physical server 1010, the PCI manager 1030, and the server manager 1040 are connected by the management network 102. The PCI switches 150a and 150b include a configuration register and a switching unit as in the first embodiment, but are not shown in FIG. Further, the Downstream ports 160a to 160d are provided with the Tx suppression control unit and the setting register 161 set in the configuration register, but the setting register and the like are omitted for the Downstream ports 160b to 160d.
An I / O device 120a is assigned to the physical server 1010, and the upstream I / O path between the physical server 1010 and the I / O device 120a is Upstream port 151d, Downstream port 160d, Upstream port 151b, Downstream. An I / O path is set up via port 160a. Further, as a standby I / O path, an I / O path via Upstream port 151d, Downstream port 160c, Upstream port 151a, and Downstream port 160a is set.
Since the configuration of the components constituting the information processing system 1000 is similar to the configuration of the components constituting the information processing system 100 shown in the first embodiment of the present invention, the information processing system 1000 and the information processing system 100 will be described below. The difference will be described.
The physical server 1010 includes hardware 116 including a CPU, a chipset, and memory. On the physical server 1010, OS1015 runs, and on OS1015, application 1016 runs.
OS1015 includes a driver module 1017, an I / O failure detection unit 1011, a server-generated Tx suppression unit 1012, an I / O path replacement instruction unit 1013, and a server-generated Tx restart unit 1014. The driver module 1017 is a driver for the I / O device 120a. Since the I / O failure detection unit 1011, the server-initiated Tx suppression unit 1012, the I / O path change instruction unit 1013, and the server-initiated Tx restart unit 1014 are implemented as one function of the OS 1015, the application 1016 can be used. I / O path replacement processing can be performed without being aware of it.
The I / O failure detection unit 1011 detects the failure notification from the I / O device or PCI switch used by OS1015, and if the failure is related to the I / O path, starts the I / O path replacement process. Specifically, the I / O failure detection unit 1011 analyzes the I / O failure notification received by, for example, the Advanced Error Reporting function of the PCI-Express switch, and if the failure is related to the I / O path, the physical server 1010 Starts the I / O path replacement process without resetting.
The server-initiated Tx suppression unit 1012 suppresses the issuance of transactions (server-initiated transactions) to I / O devices that use the failed I / O path. Known means will be used for realization. For example, OS 1015 on physical server 1010 has a Hot Plug function for I / O devices, and uses the Hot Plug mechanism to disconnect the I / O devices assigned to physical server 1010.
The I / O path replacement instruction unit 1013 instructs the PCI manager 1030 to replace the failed I / O path. Specifically, the I / O path replacement instruction unit 1013 issues an I / O path replacement request to the PCI manager 1030. The I / O path replacement request includes the identifier of the failed PCI switch and the identifier of the failed port. The I / O path change request may be notified from the physical server 1010 to the PCI manager 1030 via the PCI switch 150b, or from the physical server 1010 via the management network 102 to the PCI manager 1030 via the server manager 1040. May be notified to.
The server-initiated Tx restart unit 1014 resumes the issuance of transactions to the I / O device suppressed by the server-initiated Tx suppression unit 1012. Known means will be used for realization. For example, OS1015 has a Hot Plug function for I / O devices, and uses the Hot Plug mechanism to connect the I / O devices assigned to the physical server 1010.
The server-initiated Tx suppression unit 1012 and the server-initiated Tx restart unit 1014 are not limited to the Hot Plug function as long as they are mechanisms for controlling server-initiated transactions. For example, when the physical server 1010 is provided with a hypervisor, the hypervisor may control the transaction originating from the server by using the method described in S802 or S808. Further, if the PCI switch 150b connected to the physical server 1010 has a mechanism for controlling transactions from the physical server 1010, the physical server 1010 does not have to have a mechanism for controlling transactions originating from the server.
The PCI manager 1030 completes the I / O path replacement with the I / O configuration change unit 131, the I / O departure Tx suppression instruction unit 112d, the I / O configuration change instruction unit 113d, and the I / O departure Tx restart instruction unit 114d. A notification unit 1031 is provided. The I / O departure Tx suppression instruction unit 112d, the I / O configuration change instruction unit 113d, and the I / O departure Tx restart instruction unit 114d are included in the physical server 110a shown in the first embodiment of the present invention. It is the same as the / O departure Tx suppression instruction unit 112a, the I / O configuration change instruction unit 113a, and the I / O departure Tx restart instruction unit 114a.
The I / O path replacement completion notification unit 1031 notifies the physical server 1010 that has instructed the I / O path replacement of the completion of the I / O path replacement. The physical server 1010 receives the notification of the completion of the I / O path change and restarts the transaction originating from the server.
FIG. 13 is a flowchart of a process for realizing I / O path change on the PCI switches 150a and 150b. Below, it is assumed that the upstream port 151b of the I / O path between the physical server 1010 and the I / O device 120a assigned to the physical server 1010 fails and the I / O path fails from the active system to the standby system. Then, the I / O path change processing shown in FIG. 12 will be described.
The I / O path replacement process is started when a failure occurs in the path between the PCI switches (S1100).
First, the physical server 1010 detects that a failure has occurred in the I / O path (S1101). Specifically, the I / O failure detection unit 1011 detects that a failure has occurred in the upstream port 151b of the PCI switch 150a by, for example, the Advanced Error Reporting function of the PCI-Express switch.
Next, the physical server 1010 suppresses the issuance of transactions (transactions originating from the server) to the I / O device 120a that uses the failed I / O path, and instructs the PCI manager 1030 to replace the I / O path. (S1102). Specifically, the server-generated Tx suppression unit 1012 disconnects the I / O device 120a using the Hot Plug mechanism. As a result, transactions from the physical server 1010 to the I / O path for which the I / O path is to be replaced are suppressed. In addition, the I / O path replacement instruction unit 1013 issues an I / O path replacement request to the PCI manager 1030. The I / O path replacement request includes the identifier of the failed PCI switch 150a and the identifier of the failed port 151b.
When the PCI manager 1030 receives the I / O path change request, it first suppresses the transaction from the I / O device 120a connected to the failed I / O path (S1103). Specifically, the I / O departure Tx suppression indicator 112b of the PCI manager 1030 configures the setting register 161 of the downstream port 160a to which the I / O device 120a is connected via the MMI / O of the PCI switch. Write to the operation register. As a result, the Tx suppression control unit 162 of the Downstream port 160a starts the transaction suppression processing from the I / O device 120a as described above. The transaction suppression processing flow from the I / O device is as explained in Fig. 6.
Next, the PCI manager 1030 waits for the suppression completion notification from the Tx suppression control unit 162 (S1104). The PCI manager 1030 can guarantee that the transaction from the I / O device 120a has been suppressed by receiving the suppression completion notification.
Next, the PCI manager 1030 instructs the PCI switches 150a to 150b to change the I / O configuration (S1105). Specifically, the I / O configuration change instruction unit 113a of the PCI manager 1030 generates I / O path replacement information related to the PCI switches 150a to 150b based on the standby system I / O path information that avoids the failure port 151b. To do. The I / O path replacement information related to the PCI switch 150a includes the pre-replacement path information "Upstream port 151b, Downstream port 160a" and the post-replacement path information "Upstream port 151a, Downstream port 160a". Further, the I / O path replacement information regarding the PCI switch 150b includes the pre-replacement path information "Upstream port 151b, Downstream port 160a" and the post-replacement path information "Upstream port 151a, Downstream port 160a". Next, the I / O configuration change instruction unit 113a issues an I / O configuration change request to the I / O configuration change unit 131. The I / O configuration change request includes the I / O path replacement information related to the above-mentioned PCI switches 150a to 150b.
The I / O configuration change unit 131 receives the I / O configuration change request and issues the setting change request to the PCI switch management units 154a to 154b. The setting change request to the PCI switch management unit 154a includes the above-mentioned I / O path replacement information regarding the PCI switch 150a. Further, the setting change request to the PCI switch management unit 154b includes the above-mentioned I / O path replacement information regarding the PCI switch 150b. The PCI switch management units 154a to 154b change the configuration of the port belonging to the corresponding PCI tree according to the setting change request. Since the I / O path replacement process does not involve the movement of the virtual server, the setting of the address translation table is not changed.
Next, the PCI manager 1030 instructs the PCI switch 150a to resume the transaction from the I / O device 120a whose I / O path is to be replaced (S1106). Specifically, the I / O departure Tx restart indicator 114b of the PCI manager 1030 goes through the MMI / O of the PCI switch 150a to the setting register 160 of the downstream port 160a to which the I / O device 120a is connected. Write to the configuration register. As a result, the Tx suppression control unit 162 of the Downstream port 160a resumes the transaction from the I / O device 120a.
Next, the PCI manager 1030 notifies the physical server 1010 of the completion of the I / O path change (S1107). Specifically, the I / O path replacement completion notification unit 1031 of the server manager 1040 notifies the physical server 1010 of the completion of the I / O path replacement via the management network 102.
The physical server 1010 receives the notification of the completion of the I / O path change and restarts the transaction originating from the server (S1108). Specifically, the server-originated Tx restart unit 1014 receives the notification of the completion of the I / O path replacement and connects the I / O device 120a using the Hot Plug mechanism. As a result, the issuance of transactions from the physical server 1010 to the I / O device 120a is resumed.
The I / O path replacement process is completed with the above series of processes (S1109). As described above, by performing the I / O path replacement process, I / O path replacement can be realized in a state where it is guaranteed that no transaction exists on the I / O path to be replaced, and I / O path replacement can be realized. It is possible to avoid loss of transactions due to processing.
<Transformation example 2> In the second modification, the driver module 1017 of the I / O device 120a has the I / O failure detection unit 1011 of the modification 1 shown in FIG. 12, the server-generated Tx suppression unit 1012, and the I / O device 120a, as compared with the above modification 1. The difference is that it is equipped with a / O path replacement instruction unit 1013 and a server-generated Tx restart unit 1014, and is further equipped with a mechanism for realizing I / O path replacement. The mechanism that realizes the I / O path alternation provided in the driver module 1017 manages multiple I / O paths related to the I / O device 120a by linking with the I / O device 120a, and arbitrarily selects those I / O paths. It is a mechanism to change to.
The I / O path replacement instruction unit 1013 of the driver module 1017 instructs the PCI manager 1030 to suppress transactions originating from I / O, and uses the mechanism provided in the driver module 1017 to realize the I / O path replacement. Change the I / O path.
Since the driver module 1017 is equipped with an I / O path replacement processing mechanism, I / O path replacement can be performed even if the OS 1015 or application 1016 does not have an I / O path replacement processing mechanism. On the other hand, the I / O device 120a and the driver module 1017 cannot be applied to a general-purpose I / O device because they need to have a mechanism for I / O path alternation processing.
<Modification example 3> The modified example 3 is different from the modified examples 1 and 2 in that the PCI manager 1030 is provided with the I / O failure detection unit 1011.
In this modification 3, the PCI manager 1030 has the I / O failure detection unit 1011 of the modification 1 shown in FIG. 12, and when the I / O failure detection unit 1011 detects an I / O path failure, it is a physical server. Notify 1010 of I / O path failure. The server-initiated Tx suppression unit 1012 of the physical server 1010 receives the notification of the I / O path failure and suppresses the issuance of the server-initiated transaction.
As described above, in the present invention, a computer system including an I / O switch that dynamically changes the connection between a computer and an I / O device, or a computer system that dynamically changes the route in the I / O switch. Can be applied.
<figref num="1">It is a block diagram which shows the 1st Embodiment and shows an example of the structure of the information processing system which executes a virtual server.</figref><figref num="2">It is explanatory drawing which shows the 1st Embodiment and shows an example of the I / O device management table.</figref><figref num="3">It is a block diagram which shows the 1st Embodiment and shows the hardware configuration of a physical server.</figref><figref num="4">Explanatory drawing which shows the 1st Embodiment and shows an example of the address translation table.</figref><figref num="5">It is a block diagram which shows the 1st Embodiment and shows the structure of the Downstream port.</figref><figref num="6">It is a flowchart which shows the 1st Embodiment and shows an example of the process performed in the Tx suppression control part.</figref><figref num="7">It is a block diagram which shows the 1st Embodiment and shows the transaction flow concerning the processing of the Tx suppression control part.</figref><figref num="8">It is a flowchart which shows the 1st Embodiment and shows an example of the live migration process of a virtual server performed in an information processing system.</figref><figref num="9">It is a block diagram which shows the main part of the information processing system which showed the 1st Embodiment, and showed the transaction flow about the process of live migration.</figref><figref num="10">It is a block diagram which shows the 1st Embodiment and shows the structure of a transaction.</figref><figref num="11">The block diagram which shows the 1st Embodiment and shows the set setting register set in the configuration register.</figref><figref num="12">It is a block diagram which shows the modification 1 and shows the structure of the information processing system which has the I / O path change function.</figref><figref num="13">It is a flowchart of I / O path change processing of PCI switch 150a, 150b performed in a computer system.</figref>
Code description
100, 1000 information processing system 110a, 110b, 110c physical server 111a, 111b, 111c hypervisor 112a, 112b, 112d I / O departure Tx deterrence indicator 113a, 113b, 113d I / O configuration change indicator 114a, 114b, 114d I / O departure Tx restart instruction section 115a, 115b virtual server 120a, 120b I / O device 130 PCI Manager 140 Server Manager 150 PCI switch 153 Virtualization Assist Department 154 PCI Switch Management Department 160a, 160b Downstream port 161 Setting register 152 Tx Suppression Control
16 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| US20070186025A1 | Cites | United States of America |
| JP06290067A | Cites | Japan |
| JP2004032224A | Cites | Japan |
4 members in 2 offices
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 2008020923 | Japan | A | |
| JP20080020923 | – | – | – |
Members4
| Document | Office | Kind | |
|---|---|---|---|
| US2009198862A1 | United States of America | A1 | |
| JP2009181418A | Japan | A | |
| US8078764B2 | United States of America | B2 | |
| JP5116497B2This record | Japan | B2 |
14 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 5116497
- Publication, DOCDB
- 5116497
- Publication, EPODOC
- JP5116497B
- Application
- 20923
- Application, DOCDB
- 2008020923
- Application, EPODOC
- JP20080020923
Titles2
- Japanese
- 情報処理システム、I/Oスイッチ及びI/Oパスの交替処理方法
- English
- Information processing system, I / O switch and I / O path replacement processing method
Classification
- CPC, 4
- G06F9/45558
- G06F9/5088
- G06F2009/4557
- G06F2009/45579
- IPC, 2
- G06F13 14
- G06F13 10
