EP2561444B1

Automated recovery and escalation in complex distributed applications

Abstract

This record has no abstract on file.

EP2561444B1, drawing sheet 1
Sheet 1 of 6

Term

4.5 yearsleft in the term

Expires 30 March 2031.

  1. Priority and filed
  2. Granted
  3. Today
  4. Expires

14 claims: 3 independent, 11 dependent

  1. 1
    A method to be executed at least in part in a computing device for automated recovery and escalation of alerts in distributed systems, the method comprising:detecting, by a monitoring engine (103), a problem associated with at least one of a device and a software application within a distributed system;transmitting, by the monitoring engine (103), an alert (113) based on the detected problem to an automation engine (102);receiving the alert associated with a detected problem from the monitoring engine (103) at the automation engine (102);the method further comprising the following steps performed by the automation engine (102);collecting diagnostic information associated with the detected problem;attempt to map the alert to a recovery action, wherein the automation engine (102) performs a wild card search of a troubleshoot database to determine multiple recovery actions in response to the received alert, the troubleshoot database storing profiles of alerts matched to recovery actions further classified by device or software programs;if the alert is mapped to a recovery action, perform the recovery action, wherein, if the alert matches several wild card mappings, only the most specific mapping is applied;else escalate (111) the alert to a designee (101) together with the collected diagnostic information;and update, by employing the collected diagnostic information, records associated with alert - recovery action mapping;the method further comprising the monitoring engine (103) receiving a feedback response from a device or software program after execution of the recovery action and passing the response on to the automation engine (102);in response, the automation engine (102) updating the troubleshoot database.
  2. 7
    A system for automated recovery and escalation of alerts in distributed systems, the system comprising:a server executing a monitoring engine (103) and an automation engine (102), wherein the monitoring engine is configured to: detect a problem associated with at least one of a device and a software application within a distributed system;and transmit an alert (113) based on the detected problem;and the automation engine (102) is configured to: receive the alert (113);collect diagnostic information associated with the detected problem;attempt to map the alert (113) to a recovery action, wherein the automation engine (102) performs a wild card search of a troubleshoot database to determine multiple recovery actions in response to the received alert, the troubleshoot database storing profiles of alerts matched to recovery actions further classified by device or software programs;if the alert is mapped to a recovery action, perform the recovery action, wherein, if the alert matches several wild card mappings, only the most specific mapping is applied;else escalate (111) the alert to a designee (101) along with the collected diagnostic information;and update records in the troubleshoot database by employing the collected diagnostic information;the monitoring engine (103) further configured for receiving a feedback response from a device or software program after execution of the recovery action and passing the response on to the automation (engine (102);the automation engine further configured for updating the troubleshoot database.
  3. 12
    A computer-readable storage medium with instructions stored thereon for automated recovery and escalation of alerts in distributed systems, that, when executed by a computing system, cause the computing system to perform a method comprising:detecting, by a monitoring engine (103), a problem associated with at least one of a device and a software application within a distributed system;transmitting, by the monitoring engine (103), an alert (113) based on the detected problem to an automation engine (102);receiving the alert (113) associated with a detected problem from the monitoring engine (103) at the automation (engine (102);the method further comprising the following steps performed by the automation engine (102);collecting diagnostic information associated with the detected problem;attempt to map the alert (113) to a recovery action, wherein the automation engine performs a wild card search of a troubleshoot database to determine multiple recovery actions in response to the received alert, the troubleshoot database storing profiles of alerts matched to recovery actions further classified by device or software programs;if the alert is mapped to a recovery action, perform the recovery action, wherein, if the alert matches several wild card mappings, only the most specific mapping is applied;else escalate (111) the alert to a designee (101) together with the collected diagnostic information;and update, by employing the collected diagnostic information, records associated with alert - recovery action mapping;the method further comprising the monitoring engine (103) receiving a feedback response from a device or software program after execution of the recovery action and passing the response on to the automation engine (102);in response, the automation engine (102) updating the troubleshoot database.