US7533288B2

Method of achieving high reliability of network boot computer system

Summary by NHIP

Network Fault Recovery Method

The method detects a fault in a server, network, or external disk device within a multi-server system. It then identifies an alternative server to access a replacement disk containing identical data and transmits a boot instruction to that server via the management network.

Claim Score by NHIP

Read claim 13, the broadest

Abstract

In a network computer system, recovery may be impossible from a fault when the fault occurs in a network switch in a network or a device such as an external disk device. Provided is a computer system that includes a plurality of servers, a plurality of networks, a plurality of external disk devices, and a management computer, in which the management computer detects a fault which is occurred, retrieves an application stop server inaccessible to the used disk due to the fault, retrieves the disk for storing the same contents as contents stored in the disk used by the retrieved application stop server and the external disk device including the disk, retrieves an application resuming server capable of accessing the retrieved external disk device, and transmits an instruction to boot by using the retrieved disk to the retrieved application resuming server.

US7533288B2, drawing sheet 1
Sheet 1 of 34

Term

Projected expiry 10 November 2027.

  1. Priority
  2. Filed
  3. Granted
  4. Today
  5. Projected expiry

14 claims: 3 independent, 11 dependent

  1. 1
    A method of controlling a computer system including a plurality of servers, a plurality of networks connected to the plurality of servers, a plurality of external disk devices connected to the plurality of networks, and a management computer connected to the plurality of servers, the plurality of networks, and the plurality of external disk devices via a management network, each of the external disk devices including at least one disk for storing data, and the management computer including an interface connected to the management network, a first processor connected to the interface, and a first memory connected to the first processor, the method comprising:detecting, by the first processor, a fault which is occurred in any of the server, the network, and the external disk device;retrieving, by the first processor, an application stop server inaccessible to the used disk due to the fault among the plurality of servers;retrieving, by the first processor, a disk for storing the same contents as contents stored in the disk used by the retrieved application stop server among the plurality of disks, and the external disk device including the retrieved disk;retrieving, by the first processor, an application resuming server capable of accessing the retrieved external disk device via the network in which the fault is not occurred among the plurality of servers;and transmitting, by the first processor, an instruction to boot by using the retrieved disk to the retrieved application resuming server via the management network.
  2. 11
    A program for controlling a management computer in a computer system including a plurality of servers, a plurality of networks connected to the plurality of servers, a plurality of external disk devices connected to the plurality of networks, and a management computer connected to the plurality of servers, the plurality of networks, and the plurality of external disk devices via a management network, each of the external disk devices including at least one disk for storing data, and the management computer including an interface connected to the management network, a processor connected to the interface, and a memory connected to the processor to store the program, the program causing the processor to execute:a first step of detecting a fault which is occurred in any of the server, the network, and the external disk device;a second step of retrieving an application stop server inaccessible to the used disk due to the fault among the plurality of servers;a third step of retrieving a disk for storing the same contents as contents stored in the disk used by the retrieved application stop server among the plurality of disks, and the external disk device including the retrieved disk;a fourth step of retrieving an application resuming server capable of accessing the retrieved external disk device via the network in which the fault is not occurred among the plurality of servers;and a fifth step of transmitting an instruction to boot by using the retrieved disk to the retrieved application resuming server via the management network.
  3. 13
    Broadest claimClaim Score 46, average(NHIP)A computer system, comprising:a plurality of servers;a plurality of networks connected to the plurality of servers;a plurality of external disk devices connected to the plurality of networks;and a management computer connected to the plurality of servers, the plurality of networks, and the plurality of external disk devices via a management network, wherein: each of the external disk devices includes at least one disk for storing data;the management computer includes an interface connected to the management network, a processor connected to the interface, and a memory connected to the processor;the processor detects a fault which is occurred in any of the server, the network, and the external disk device;the processor retrieves an application stop server inaccessible to the used disk due to the fault among the plurality of servers;the processor retrieves a disk for storing the same contents as contents stored in the disk used by the retrieved application stop server among the plurality of disks, and the external disk device including the retrieved disk;the processor retrieves an application resuming server capable of accessing the retrieved external disk device via the network in which the fault is not occurred among the plurality of servers;and the processor transmits an instruction to boot by using the retrieved disk to the retrieved application resuming server via the management network.