Systems and methods for automated computer support
Abstract
An automated computer support method comprising: i) receiving (602) a snapshot from a computer (116a, b); ii) compare (608) the snapshot with a database of computer states; and, iii) identify an anomaly based on the result of the comparison; characterized by: iv) receiving (602) snapshots from a plurality of computers (116a, b) within a population of computers, where the individual snapshots include data indicating a status of a respective computer; v) store (604) snapshots in a data store; vi) automatically create (606) an adaptive reference model (206, 402) based at least in part on snapshots and comprising a set of rules customized to the characteristics of the computer population, the set of rules being developed identifying patterns between the snapshots of the plurality of computers so that the adaptive reference model is indicative of normal states in the computers within the population; vii) compare (608) snapshots of at least one of the plurality of computers with the adoptive reference model; and viii) determine (610), based on the result of the comparison, if an anomaly (720) is present in the state of the at least one of the computers.
Term
Term ended
Projected expiry passed 11 August 2024, 2.1 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
15 claims: 11 independent, 4 dependent
- 1ES 2 640 191 T3 REIVINDICACIONES 1. Un método de soporte informático automatizado que comprende:i) recibir (602) una instantánea desde un ordenador (116a, b);ii) comparar (608) la instantánea con una base de datos de estados de ordenador;y, iii) identificar una anomalía basada en el resultado de la comparación;caracterizado por: iv) recibir (602) instantáneas desde una pluralidad de ordenadores (116a, b) dentro de una población de ordenadores, en donde las instantáneas individuales incluyen datos que indican un estado de un ordenador respectivo;v) almacenar (604) las instantáneas en un almacén de datos;vi) crear (606) automáticamente un modelo (206, 402) de referencia adaptativo basado al menos en parte en las instantáneas y que comprende un conjunto de reglas personalizado a las características de la población de ordenadores, estando desarrollado el conjunto de reglas identificando patrones entre las instantáneas de la pluralidad de ordenadores de manera que el modelo de referencia adaptativo es indicativo de estados normales en los ordenadores dentro de la población;vii) comparar (608) instantáneas de al menos uno de la pluralidad de ordenadores con el modelo de referencia adoptivo;y viii) determinar (610), basado en el resultado de la comparación, si está presente una anomalía (720) en el estado del al menos uno de los ordenadores.
- 2El método de la reivindicación 1, que comprende además hacer coincidir (612) al menos una anomalía (720) con al menos un filtro (216) de reconocimiento para diagnosticar una condición en al menos uno de la pluralidad de ordenadores (116a, b).
- 3El método de la reivindicación 2, en donde el filtro (216) de reconocimiento comprende un patrón particular de anomalías que indica la presencia de una condición de causa raíz particular o una clase genérica de condiciones.
- 4El método de la reivindicación 3, que comprende responder (614) a la condición mediante al menos uno de:- i) generar una notificación;ii) enviar un tique de problema a un sistema de gestión de problemas;iii) solicitar permiso para tomar una acción;y, iv) eliminar la condición de al menos uno de la pluralidad de ordenadores (116a, b).
- 5El método de la reivindicación 4, en donde la eliminación de la condición comprende hacer que un programa de reparación sea ejecutado en al menos uno de la pluralidad de ordenadores (116a, b) afectados por la condición.
- 6El método de cualquiera de las reivindicaciones 3 a 5, que comprende además:i) determinar cuáles de la pluralidad de ordenadores (116a, b) están afectados por la condición;y, ii) causar una respuesta (614) a la condición a ser ejecutada en nombre de cada uno de la pluralidad de ordenadores afectados por la condición.
- 7El método de cualquiera de las reivindicaciones 3 a 6, en donde el diagnóstico (612) de la condición comprende identificar una causa raíz de al menos una anomalía (720).
- 8El método de cualquiera de las reivindicaciones 2 a 7, en donde al menos un filtro (216) de reconocimiento está asociado con una respuesta (214) automatizada para la condición.
- 9El método de cualquiera de las reivindicaciones 2 a 8, en donde la condición es una clase que comprende un grupo de condiciones.
- 10El método de cualquiera de las reivindicaciones 2 a 9, que además comprende determinar una calidad de una coincidencia (612) entre al menos un filtro (216) de reconocimiento y al menos una anomalía (720).
- 11El método de cualquiera de las reivindicaciones 1 a 10, en donde el modelo (402) de referencia adaptativo comprende una pluralidad de activos cada uno asociado con un tipo de activo que comprende uno de:- i) un archivo, ii) una clave de registro, ES 2 640 191 T3 iii) una medida de rendimiento, iv) un servicio, v) un componente de hardware, vi) un proceso en ejecución, vii) un registro, y viii) un puerto de comunicación.
- 12El método de cualquiera de las reivindicaciones 1 a 11, en donde comparar (912) al menos una de la pluralidad de instantáneas con el modelo (402) de referencia adaptativo comprende:- i) generar un resultado (914);y, ii) proporcionar (920) el resultado a un usuario.
- 13El método de cualquiera de las reivindicaciones 1 a 12, en donde la pluralidad de instantáneas se crea mediante un agente de software que reside en cada uno de la pluralidad de ordenadores (116a, b).
- 14Un sistema (102) de soporte informático automatizado que comprende:i) un componente (108) Colector;y, ii) un componente (110) analítico en comunicación con el componente colector;caracterizado por que: iii) el sistema (102) de soporte automatizado está organizado y dispuesto para efectuar las reivindicaciones del método en cualquiera de las reivindicaciones 1 a 13.
- 15Un medio legible por ordenador en el cual se codifica un código de programa que comprende instrucciones ejecutables por ordenador para efectuar el método de soporte informático automatizado de cualquiera de las reivindicaciones 1 a 13.
Independent claims15
175 paragraphs in 13 sections, as filed
IS 2 640 191 T3
DESCRIPTION
Systems and methods for automated IT support
FIELD OF THE INVENTION
The present invention relates generally to systems and methods for automated computer support.
BACKGROUND
As information technology continues to increase in complexity, problem management costs will escalate as the frequency of support incidents rises and the skill set requirements of human analysts become more demanding. Conventional problem management tools are designed to reduce costs by increasing the efficiency of the human beings who perform these support tasks. This is typically accomplished by at least partially automating the capture of trouble ticket information and facilitating access to knowledge bases. Although useful, this type of automation has reached the point of diminishing returns, insofar as it does not address the fundamental weakness in the support model itself, its reliance on humans.
Table 1 illustrates the distribution of labor costs associated with incident resolution in the conventional human-based support model. Data displayed is provided by Motive Communications, Inc. of Austin, Texas (www.motive.com), a leading provider of technical support service software. The highest cost items are those associated with tasks that require analysis and / or human interaction (eg, Diagnostics, Investigation, Resolution).
Table 1
<td>Support Tasks</td><td>% Labor Cost</td>
<td>Simple and Repetitive Problems (30%)</td><td></td>
<td>Desktop Settings (User Inflicted)</td><td> 4%</td>
<td>Desktop Environment (Software Malfunction)</td><td> 9%</td>
<td>Interconnection and Connectivity</td><td> 7%</td>
<td>How (questions)</td><td> 10%</td>
<td>Complex and Dynamic Problems (70%)</td><td></td>
<td>Classification (Identify user and support rights)</td><td> 7%</td>
<td>Diagnostics (Analyze machine status)</td><td> 11%</td>
<td>Investigation (Finding the source of the problem)</td><td> 35%</td>
<td>Resolution and Repair (Walk with the user through the repair)</td><td> 18%</td>
Conventional software solutions for automated problem management strive to lower these costs and add value across a wide range of service levels. Forrester Research, Inc. of Cambridge, MA (www.forrester.com) provides a helpful characterization of these service levels. Forrester Research divides conventional automated IT support solutions into five levels of service, including: (1) Mass Healing - resolving incidents before they occur; (2) Self-healing - resolving incidents when they occur; (3) Self-service - resolve incidents before a user calls; (4) Assisted Service - resolve incidents when a user calls; and (5) Visit from Technical Support Side - resolve incidents when all else fails. According to Forrester, the cost per incident using a conventional self-healing service is less than a dollar. However, the cost escalates rapidly, reaching more than $ 300 per incident if a technical support visit is finally required.
The goal of Mass Heal is to resolve incidents before they occur. In conventional systems, this goal is achieved by making all PC configurations the same or, at the very least, by ensuring that a problem found on one PC cannot be replicated on any other PCs. Conventional products typically associated with this level of service consist of software distribution tools and configuration management tools. Security products such as antivirus scanners, intrusion detection systems, and data integrity checkers are also considered part of this tier, as they focus on preventing incidents from occurring.
Conventional products that attempt to address this level of service operate by limiting the managed population to a small number of known good configurations and detecting and removing a relatively small number of known bad configurations (eg, virus signatures). The problem with this approach is that it assumes that: (1) all good and bad configurations can be known ahead of time; and (2) once they are known to be relatively stable. As the complexity of computing and interconnection systems increases, the stability of any particular node on the network tends to decrease. Both hardware and software on any particular node are likely to change frequently. For example, many software products are capable of automatically updating themselves using software patches.
ES 2 640 191 T3 accessed over an internal network or the Internet. Since there are an infinite number of good and bad settings and since they are constantly changing, these conventional self-healing products can never be more than partially effective.
Furthermore, virus authors continue to develop increasingly intelligent viruses. Conventional virus detection and eradication software relies on the ability to identify a known pattern to detect and eradicate a virus. However, as the number and complexity of viruses increases, the resources required to maintain a database of known viruses and fixes for those viruses combined with the resources required to distribute the fixes to the population of nodes on a network they become overwhelming. In addition, a conventional PC running a Microsoft Windows operating system includes more than 7,000 system files and more than 100,000 registry keys, all of which are multi-valued. Therefore, for all practical purposes, there can be an infinite number of good states and an infinite number of bad states, making the task of identifying bad states more complicated.
The goal of the Self-Heal level is to automatically detect and correct problems before they result in a call to technical support, ideally before the user is even aware that a problem exists. Conventional Self-Healing tools and utilities have been around since the late 1980s when Peter Norton introduced a PC repair and diagnostic toolkit (www.Symantec.com). These tools also include tools that allow a user to restore a PC to a restore point set prior to installing a new product. However, none of the conventional tools work well under real world conditions.
A fundamental problem with these conventional tools is the difficulty in creating a reference model with sufficient scope, granularity and flexibility to allow "normal" to be reliably distinguished from "abnormal". Compounding the problem is the fact that the definition of "normal" must constantly change as new software updates and applications are rolled out. This is a formidable technical challenge and one that has yet to be conquered by any of the conventional tools.
The goal of the Self-Service tier is to reduce the volume of technical support calls by providing a collection of automated tools and knowledge bases that enable end users to help themselves. Conventional Self-Service products consist of “how-to” knowledge bases and collections of software solutions that automate low-risk, repetitive support functions such as resetting forgotten passwords. These conventional solutions have a significant disadvantage in that they increase the probability of self-inflicted damage. For this reason they are limited to specific types of problems and applications.
The goal of the Assisted Service level is to improve human efficiency by providing an automated infrastructure to manage a service request and by providing capabilities to remotely control a personal computer and interact with end users. Conventional Assisted Service products include technical support service software, online reference materials, and remote control software.
Although products at this service level are perhaps the most mature of the conventional products and solutions described herein, they do not yet fully meet the requirements of users and organizations. Specifically, the ability of these products to automatically diagnose problems is severely limited both in terms of the types of problems that can be correctly identified, as well as the accuracy of the diagnosis (often multiple-choice).
A Visit from the Technical Support Side becomes necessary when all else fails. This level of service includes any "shortcut" activities that may be necessary to restore a computer that cannot be remotely diagnosed / repaired. It also includes the monitoring and management of these activities to ensure a timely resolution. Of all service levels, this level will most likely require significant time from highly trained and therefore expensive human resources.
Conventional products at this level consist of specialized software products and diagnostic tools that track and resolve customer issues over time and potentially through multiple customer service representatives.
Thus, what is needed is a paradigm shift, which is necessary to significantly reduce support costs. This change will be characterized by the appearance of a new support model in which machines will serve as primary agents to make decisions and initiate actions.
COMPENDIUM OF PREVIOUS TECHNIQUE
US 2003/0028825 A1 (Hines) - February 6, 2003 refers to a debugging and error correction method and describes:
ES 2 640 191 T3 “the service guru system works automatically to process an image from a computer system to identify, if any, which problem preconditions are satisfied (i.e. the proactive case) and then identify particular problems of this smaller set that match an accurate description of the problem symptom (ie, the reactive case) "
The method described is limited to correcting a single computer system.
US 2002/0194550 A1 (LOPKE) - December 19, 2002 refers to an end-user diagnostic system and describes:
"As shown in FIG. 1, the inspector 40 is linked to the system register 26. The inspector 40 obtains configuration data from register 26. The configuration data refers to information required to configure the software and hardware components that define the computer system 20. Preferably, the inspector 40 obtains configuration data that reflects real-time configuration information. "
Document EP 1 172 732 A1 (HITACHI) - January 16, 2002 refers to "[0014] an object of the present invention is to provide a computer system in which a computer can acquire fault information even in the case where it occurs in the computer a fault that disables an OS from executing the fault processing. "
Document US 2003/0110248 A1 (RITCHE) - June 12, 2003 refers to "a computer system and an associated method, with a tool for processing error alerts issued during the distribution of application packages to network client devices" [resume]
COMPENDIUM OF THE INVENTION
According to the present invention, an automated computer support method comprises:
i) receive a snapshot from a plurality of computers:
ii) comparing the snapshot with a database of computer states; and, iii) identify an anomaly, the method also comprising:
iv) receiving snapshots from a plurality of computers within a population of computers, wherein the individual snapshots include data indicating a status of a respective computer;
v) storing the plurality of snapshots in a data store;
vi) automatically create an adaptive reference model based at least in part on snapshots and comprising a set of rules customized to the characteristics of the population of computers, the set of rules that is developed by identifying patterns between snapshots of the plurality of computers so that the adaptive reference model is indicative of normal states in computers within the population;
vii) comparing snapshots of at least one of the plurality of computers with the adaptive reference model; and viii) determining, the method further, if at least one anomaly is present in the state of the at least one of the computers.
Additional embodiments of the method, an organized and arranged system for performing the method, and a computer-readable medium on which the program code for performing the method of the present invention is encoded, are set forth in appended claims 2 to 15.
Embodiments of the present invention provide systems and methods for automated computer support. A method according to an embodiment of the present invention comprises receiving a plurality of snapshots from a plurality of computers, storing the plurality of snapshots in a data store, and creating an adaptive reference model based at least in part on the plurality of snapshots. The method further comprises comparing at least one of the plurality of snapshots with the adaptive reference model, and identifying at least one anomaly based on the comparison. In another embodiment, a computer-readable medium (such as, for example, a random access memory or a computer disk) comprises a code to carry out such a method.
These embodiments are mentioned not to limit or define the invention, but to provide exemplary embodiments of the invention to aid in an understanding thereof. Illustrative embodiments are discussed in the Detailed Description, and a further description of the invention is provided there. The advantages offered by the various embodiments of the present invention can be further understood by examining this specification.
BRIEF DESCRIPTION OF THE FIGURES
These and other characteristics, aspects and advantages of the present invention are better understood when the following Detailed Description is read with reference to the accompanying drawings, in which:
Figure 1 illustrates an exemplary environment for the implementation of an embodiment of the present invention;
Figure 2 is a block diagram illustrating a flow of information and actions in an embodiment of the present invention;
Figure 3 is a flow chart illustrating an overall anomaly detection process in one embodiment of the present invention; and Figure 4 is a block diagram illustrating components of an adaptive reference model in one embodiment of the present invention;
Figure 5 is a flow chart illustrating a registration information normalization process in an agent in an embodiment of the present invention;
Figure 6 is a flow chart illustrating a method for identifying and responding to an anomaly in one embodiment of the present invention;
Figure 7 is a flow chart illustrating a process for identifying certain types of anomalies in one embodiment of the present invention;
Figure 8 is a flow chart illustrating a process for generating an adaptive reference model in one embodiment of the present invention;
Figure 9 is a flow chart, illustrating a process for proactive anomaly detection in one embodiment of the present invention;
Figure 10 is a flow chart illustrating a reactive process for anomaly detection in one embodiment of the present invention;
Figure 11 is a screenshot of a user interface for creating an adaptive reference model in one embodiment of the present invention;
Figure 12 is a screenshot of a user interface for managing an adaptive reference model in an embodiment of the present invention;
Figure 13 is a screenshot of a user interface for selecting a snapshot to use for creating a recognition filter in one embodiment of the present invention;
Figure 14 is a screenshot of a user interface for managing a recognition filter in one embodiment of the present invention;
Figure 15 is a screen shot illustrating a user interface for selecting a "great system" for use in a policy template in one embodiment of the present invention; and Figure 16 is a screenshot of a user interface for selecting policy template assets in one embodiment of the present invention.
DETAILED DESCRIPTION
Embodiments of the present invention provide systems and a method for automated computer support. Referring now to the drawings in which like numbers indicate like elements throughout the various figures, Figure 1 is a block diagram illustrating an exemplary environment for implementing an embodiment of the present invention. The embodiment shown includes an automated support facility 102. Although the automated support facility 102 is shown as a single facility in Figure 1, it may comprise multiple facilities or be incorporated into the site where the managed population resides. The automated support facility includes a firewall 104 in communication with a network 106 to provide security for data stored within the automated support facility 102. The automated support facility 102 also includes a Collector component 108. The Collector component 108 provides, among other features, a mechanism for transferring data in and out of the automated support facility 102. The transfer routine can use a standard protocol such as File Transfer Protocol (FTP) or Hypertext Transfer Protocol (HTTP) or it can use a proprietary protocol. The Collector component also provides the necessary processing logic to download, decompress, and parse incoming snapshots.
The automated support facility 102 shown also includes an Analytical component 110 in communication with the Collector component 108. The Analytical component 110 includes hardware and software to implement the adaptive reference model described herein and store the adaptive reference model in a Database component 112. The Analytical component 110 extracts snapshots and adaptive reference models from a Database component 112, analyzes the snapshot in the context of the reference model, identifies and filters out any anomalies, and broadcasts the responding agent (s) when appropriate. The Analytical component 110 also provides the user interface for the system.
The embodiment shown also includes a Database component 112 in communication with the Collector component 108 and the Analytical component 110. Database component 112 provides a means for storing data for agents and for processes performed by one embodiment of the present invention. A primary function of the Database component can be to store snapshots and adaptive reference models. It includes a set of database tables, as well as the necessary processing logic to automatically manage those tables. The embodiment shown includes only a Database component 112 and an Analytical component 110. Other embodiments include many Database and Analytics components 112, 110. One embodiment includes a Database component and multiple Analytics components, allowing multiple support personnel to share a single database while performing analytical tasks in parallel.
IS 2 640 191 T3
One embodiment of the present invention provides automated support to a managed population 114 which may comprise a plurality of client computers 116a, b. The managed population provides data to the automated support facility 102 over the network 106.
In the embodiment shown in Figure 1, an Agent component 202 is deployed within each monitored machine 116a, b. Agent component 202 gathers data from customer 116. At scheduled intervals (for example, once a day) or in response to a command from Analytical component 110, Agent component 202 takes a detailed snapshot of the state of the machine in the residing. This snapshot includes a detailed examination of all system files, designated application files, the registry, performance counters, processes, services, communication ports, hardware configuration, and log files. The results of each scan are then compressed and transmitted as a Snapshot to a Collector component 108.
Each of the servers, computers, and network components shown in Figure 1 comprise processors and computer-readable media. As is well known to those skilled in the art, an embodiment of the present invention can be configured in a number of ways by combining multiple functions on a single computer or, alternatively, using multiple computers to perform a single task.
Processors used by an embodiment of the present invention may include, for example, digital logic processors capable of processing input, executing algorithms, and generating output as necessary in support of processes according to the present invention. Such processors can include a microprocessor, an ASIC, and state machines. Such processors include, or may be in communication with, means, for example computer-readable media, that store instructions that, when executed by the processor, cause the processor to perform the steps described herein.
Computer-readable media embodiments include, but are not limited to, an electronic, optical, magnetic, or other transmission or storage device capable of providing a processor, such as the processor in communication with a touch-sensitive input device, with computer readable instructions. Other examples of suitable media include, but are not limited to, a floppy disk, CD-ROM, magnetic disk, memory chip, ROM, RAM, an ASIC, a configured processor, all optical media, all magnetic tapes, or other media. magnetic, or any other medium from which a computer processor can read instructions. Also, various other forms of computer-readable media can transmit or transport instructions to a computer, including a router, a private or public network, or other device or transmission channel, both wired and wireless. The instructions can comprise code from any computer programming language, including, for example, C, C #, C ++, Visual Basic, Java, and JavaScript.
Figure 2 is a block diagram illustrating a flow of information and actions in an embodiment of the present invention. The embodiment shown comprises an Agent component 202. Agent component 202 is the part of the system that is deployed within each monitored machine. It can perform three main functions. First of all, you may be responsible for gathering data. The Agent component 202 may perform an extensive scan of the client machine 116a, b at scheduled intervals, in response to a command from the Analytical component 110, or in response to events of interest detected by the Agent component 202. This scan may include a detailed examination of all system files, designated application files, the registry, performance counters, hardware configuration, logs, running tasks, services, network connections, and other relevant data. The results of each scan are compressed and transmitted over the network 106 in the form of a "snapshot" to the Collector component 108.
In one embodiment, Agent component 202 reads each byte of files to be scanned and creates a digital signature or random check for each file. The digital signature identifies the exact content of each file rather than simply providing metadata, such as size and creation date. Some conventional viruses change the information in the file header in an attempt to fool systems that rely on metadata for detection. Such an embodiment is capable of successfully detecting such viruses.
Scanning the client by Agent component 202 can be resource intensive. In one embodiment, a full scan is performed periodically, eg, daily, during a time when the user is not using the client machine. In another embodiment, Agent component 202 performs a delta scan of the client machine, only logging changes since the last scan. In another embodiment, the scans by Agent component 202 are run on demand, providing a valuable tool for a technician or support person attempting to remedy a failure on the client machine.
The second main function performed by agent 202 is behavior blocking. Agent 202 constantly (or substantially constantly) monitors access to key system resources, such as system files and the registry. It is capable of selectively blocking access to these resources in real time to prevent malicious software damage. While behavior monitoring occurs on an ongoing basis, the behavior lock is enabled as part of a repair action. For example, if the Analytical 110 component suspects the presence of a virus, it can download a repair action
ES 2 640 191 T3 to make the client block the virus from accessing key information resources within the managed system. Client component 202 provides monitoring process information as part of the snapshot.
The third main function performed by Agent component 202 is to provide a runtime environment for responsive agents. Response agents are mobile software components that implement automated procedures to address various types of problem conditions. For example, if the Analytical component 110 suspects the presence of a virus, it can download a response agent to cause the Agent component 202 to remove suspicious assets from the managed system. Agent component 202 may run as a service or other background process on the computer being monitored. Due to the scope and granularity of the information provided by an embodiment of the present invention, the repair can be performed with more precision than with conventional systems. Although described in terms of a client, the managed population 114 may comprise PCs, workstations, servers, or any other type of computer.
The embodiment shown also includes an adaptive reference model component 206. A difficult technical challenge in building an automated support product is creating a reference model that can be used to distinguish between normal and abnormal states of the system. The state of a modern computer system is determined by many multivalued variables, and consequently there are a virtually infinite number of normal and abnormal states. To make matters worse, these variables change frequently as new software updates are rolled out and as end users communicate. The adaptive reference model 206 in the embodiment shown analyzes snapshots from many computers and identifies statistically significant patterns using a generic data extraction algorithm or a proprietary data extraction algorithm specifically designed for this purpose. The resulting set of rules is extremely rich (hundreds of thousands of rules) and is customized to the unique characteristics of the managed population. In the embodiment shown, the process of building a new reference model is fully automatic and can be run periodically to allow the model to adapt to desirable changes such as the planned deployment of a software update.
Since the adaptive reference model 206 is used for the analysis of statistically significant patterns from a population of machines, in one embodiment, a minimum number of machines are analyzed to ensure the accuracy of the statistical measurements. In one embodiment, a minimum population of about 50 machines is tested to achieve consistently relevant standards for machine analysis. Once a baseline is established, samples can be used to determine if something abnormal is occurring within the entire population or any member of the population.
In another embodiment, the Analytical component 110 calculates a set of maturity metrics that allow the user to determine when a sufficient number of samples has accumulated to provide an accurate analysis. These maturity metrics indicate the percentage of available relationships at each level of the model that have met predefined criteria for various confidence levels (for example, High, Medium, and Low). In such an embodiment, the user monitors the metrics and ensures that enough snapshots have been assimilated to create a mature model. In another such embodiment, the Analytical component 110 assimilates samples until it reaches a predefined maturity goal set by the user. In any such embodiment, it is not necessary to assimilate a certain number of samples (eg, 50).
The embodiment shown in Figure 2 also comprises a Policy Template component 208. Policy Template component 208 allows the service provider to manually insert rules in the form of "policies" into the adaptive reference model. Policies are combinations of attributes (files, registry keys, etc.) and values that when applied to a model, override some of the statistically generated information in the model. This mechanism can be used to automate a variety of common maintenance activities such as checking for compliance with security policies and checking to ensure that the appropriate software updates have been installed.
When something goes wrong with a computer, it often impacts a number of different information assets (files, registry keys, etc.). For example, a "Trojan" could install malicious files, add certain registry keys to ensure those files are running, and open ports for communication. The embodiment shown in Figure 2 detects these undesirable changes as anomalies by comparing the snapshot of the infected machine with the standard embedded in the adaptive reference model. An anomaly is defined as an asset that is unexpectedly present, an asset that is unexpectedly absent, or an asset that has an unknown value. The anomalies are compared against a Library of Recognition Filters 216. An Acknowledgment Filter 216 comprises a particular pattern of anomalies that indicates the presence of a particular root cause condition or a generic class of conditions. Acknowledgment Filters 216 also associate conditions with an indication of severity, a textual description, and a link to a responding agent. In another embodiment, a Recognition Filter 216 can be used to identify and interpret benign abnormalities. For example, if a user adds a new application that the administrator trusts will not cause any problems, the system according to the present invention will continue to report the new application as a set of anomalies. If the
ES 2 640 191 T3 application is new, so the notification of the assets added as anomalies is correct. However, the administrator can use an Acknowledgment Filter 216 to interpret the anomalies produced by adding the application as benign.
In one embodiment of the present invention, certain attributes refer to continuous processes. For example, performance data is made up of multiple counters. These counters measure the occurrence of various events during a particular period of time. To determine whether the value of such a counter is normal across a population, an embodiment of the present invention calculates a mean and a standard deviation. An anomaly is declared if the counter value falls more than a certain number of standard deviations away from the mean.
In another embodiment, a mechanism handles the case where the adaptive reference model 206 assimilates a snapshot containing an anomaly. Once a model reaches the desired level of maturity, it undergoes a process that eliminates any anomalies that may have been assimilated. These anomalies are visible in a mature model as isolated exceptions to strong relationships. For example, if file A appears in conjunction with file B on 999 machines, but file A is present on 1 machine, but file B is absent, the process will assume that the posterior relationship is abnormal and will be removed from the model. When the model is later used for testing, any machine that contains file A, but not file B, will be marked as failing.
The embodiment of the invention shown in Figure 2 also includes a response agent library 212. Response agent library 212 enables the service provider to authorize and store automated responses for specific problem conditions. These automated responses are built from a collection of scripts that can be dispatched to a managed machine to perform actions such as replacing a file or changing a registry value. Once a problem condition has been analyzed and a response agent defined, any subsequent occurrences of the same problem condition should be automatically corrected.
Figure 3 is a flow chart illustrating an overall anomaly detection process in one embodiment of the present invention. In the embodiment shown, Agent component 202 takes a snapshot on a periodic basis, for example once per day 302. This snapshot involves collecting a massive amount of data and can take anywhere from a few minutes to hours. to run, depending on the client configuration. When the scan is complete, the results are compressed, formatted, and transmitted as a snapshot to a secure server known as the 304 Collector component. The Collector component acts as a central repository for all snapshots that are submitted from the managed population. Each snapshot is then decompressed, parsed, and stored in various tables in the database by the Collector component.
The detection function (218) uses the data stored in the adaptive reference model component (206) to check the contents of the snapshot against hundreds of thousands of statistically relevant relationships known to be normal for that managed population 308. If no anomaly is found 310, the process terminates 324.
If an anomaly is 310 found, the Acknowledgment Filters 210 are queried to determine if the anomaly matches any known conditions 312. If the answer is yes, then the abnormality is reported based on the condition that has been diagnosed 314. Otherwise, the abnormality is reported as an unrecognized abnormality 316. The Acknowledgment Filter (216) also indicates whether or not an automated response has been authorized for that particular type of condition 318.
In one embodiment, the Acknowledgment Filters (216) can recognize and consolidate multiple anomalies. The Failure Awareness Filter matching process is performed after the entire snapshot has been analyzed and all anomalies associated with that snapshot have been detected. If a match is found between a subset of anomalies and an Acknowledgment Filter, the name of the Acknowledgment Filter will be associated with the subset of anomalies in the output stream. For example, the presence of a virus could generate a set of file failures, process failures, and registry failures. A Recognition Filter could be used to consolidate these anomalies, so that the user would simply see a descriptive name relating all anomalies to a probable common cause, that is, a virus.
If the automated response has been authorized, then the response agent library (212) downloads the appropriate response agents to the affected machine 320. Agent component 202 on the affected machine then runs the sequence of scripts necessary to correct the problematic 322 condition. The process shown then terminates 324.
Embodiments of the present invention substantially reduce the cost of maintaining a population of personal computers and servers. One embodiment achieves this goal by automatically detecting and correcting problem conditions before they escalate to helpdesk and providing information.
ES 2 640 191 T3 to shorten the time required for a support analyst to resolve any issues not addressed automatically.
Anything that reduces the frequency with which incidents occur has a significant positive impact on the cost of IT support. An embodiment of the present invention monitors and adjusts the state of a managed machine so that it is more resistant to threats. Using Policy Templates, service providers can routinely monitor the security posture of each managed system, automatically adjusting security settings and installing software updates to eliminate known vulnerabilities.
In a human-based support model, problem conditions are detected by end users, reported to a helpdesk, and diagnosed by human experts. This process accumulates costs in a number of ways. First, there are costs associated with lost productivity while the end user waits for resolution. Also, there is the cost of data collection, usually done by technical support staff. In addition, there is the cost of diagnosis, which requires the services of a trained (expensive) support analyst. In contrast, a machine-based support model implemented in accordance with the present invention automatically detects, reports, and diagnoses many troublesome software-related conditions. Adaptive Reference Model technology enables the detection of anomalous conditions in the presence of extreme diversity and changes with sensitivity and precision previously impossible.
In one embodiment of the present invention, to avoid false positives, the system can be configured to operate at various levels of confidence, and anomalies that are known to be benign can be filtered out using Recognition Filters. Acknowledgment Filters can also be used to alert the service provider to the presence of specific types of unwanted or malicious software.
In conventional systems, computer incidents are typically resolved by humans through the application of a series of trial and error repair actions. These remedial actions tend to be of the "mace" variety, meaning solutions that affect much more than the problem conditions they were intended to correct. Multiple choice repair procedures and hub solutions are a consequence of an inadequate understanding of the problem and a source of unnecessary costs. Because a system according to the present invention has the data to fully characterize the problem, it can reduce repair cost in two ways. First, you can automatically resolve the incident if an Acknowledgment Filter has been defined that specifies the required automated response. Second, if automatic repair is not possible, the system's diagnostic capabilities take the guesswork out of the human-based repair process, reducing execution time and allowing for greater accuracy.
Figure 4 is a block diagram illustrating components of an adaptive reference model in one embodiment of the present invention. Figure 4 is merely exemplary.
The embodiment shown in Figure 4 illustrates a single silo, multi-layer adaptive reference model 402. In the embodiment shown, bin 404 comprises three layers: value layer 406, grouping layer 408, and profile layer 410.
The value layer 406 tracks the active / value pair values provided by the Agent component (202) described herein through the managed population (114) of Figure 1. When comparing a snapshot to model 402 of adaptive reference, the value layer 406 of the adaptive reference model 402 evaluates the value part of each active / value pair contained therein. This evaluation consists of determining whether an asset value in the snapshot violates a statistically significant pattern of asset values within the managed population as represented by the adaptive reference model 402.
For example, an Agent (116b) transfers a snapshot that includes a digital signature for a particular system file. During the assimilation process (when the adaptive reference model is being built), the model records the values it finds for each asset name and the number of times that value is found. In this way, for each asset name, the model knows the “legal” values that it has seen in the population. When the model is used for testing, the value layer 406 determines whether the value of each attribute in the snapshot matches one of the "legal" values in the model. For example, in the case of a file, a series of “legal” values is possible because different versions of the file could exist in the managed population. An anomaly would be declared if the model contained one or more file values that were statistically consistent and the snapshot contained a file value that did not match any of the file values in the model. The model can also detect situations where there is no "legal" value for an attribute. For example, log files have no legal value since they change frequently. If there is no "legal" value, then the attribute value in the snapshot will be ignored during the check.
In one embodiment, the adaptive reference model 402 implements criteria to ensure that an anomaly is truly an anomaly and not just a new file variant. The criteria may include a level of
ES 2 640 191 T3 confidence. Confidence levels do not stop a single file from being reported as a failure. Confidence levels limit the relationships used in the model during the testing process to those relationships that meet certain criteria. The criteria associated with each level are designed to achieve a certain statistical probability. For example, in one embodiment, the criteria for the high confidence level are designed to achieve a statistical probability greater than 90%. If a lower confidence level is specified, then additional relationships that are not as statistically reliable are included in the checking process. The process of considering viable but less likely relationships is similar to the human process of speculation when we need to make a decision without all the information that would allow us to be sure. In a continuously changing environment, the administrator may want to filter out anomalies associated with low confidence levels, that is, the administrator may want to eliminate as many false positives as possible.
In an embodiment that implements the trust level, if a user reports that something is wrong with a machine, but the administrator is unable to see any anomalies in the default trust level, the administrator can lower the trust level, allowing the analysis process considers relationships that have less statistical significance and are ignored at higher confidence levels. By lowering the confidence level, the administrator allows the adaptive reference model 402 to include patterns that may not have enough samples to be statistically significant, but could provide clues as to what the problem is. In other words, the administrator is allowing the machine to speculate.
In another embodiment, the value layer 406 automatically removes the asset values from the adaptive reference model 402 if, after assimilating a specified number of snapshots, the asset values have not exhibited any stable pattern. For example, many applications generate log files. The values in the log files are constantly changing and are rarely the same from machine to machine. In one embodiment, these file values are initially evaluated and then, after a specified number of evaluations, they are removed from the adaptive reference model 402. By removing these types of file values from the model 402, the system eliminates unnecessary comparisons during the discovery process 218 and reduces database storage requirements by reducing low-value information.
An embodiment of the present invention is not limited to removing asset values from the adaptive reference model 402. In one embodiment, the process also applies to asset names. Certain asset names are "unique in nature," that is, they are unique to a particular machine, but they are a by-product of normal operation. In one embodiment, a separate process handles unstable asset names. This process in such an embodiment identifies asset names that are unique in nature and allows them to remain in the model so that they are not reported as anomalies.
The second layer shown in Figure 4 is grouping layer 408. The grouping layer 408 tracks the relationships between the asset names. An asset name can apply to a variety of entities, including a file name, a registry key name, a port number, a process name, a service name, a performance counter name, or a performance characteristic. hardware. When a particular set of asset names is generally present in tandem on machines in a managed population (114), the grouping layer 408 is capable of flagging a failure when a member of the asset name set is absent.
For example, many applications on a computer running a Microsoft Windows operating system require a multitude of dynamic link libraries (DLLs). Each DLL will often depend on one or more DLLs. If the first DLL is present, then the other DLLs must also be present. The bundling layer 408 tracks this dependency and if one of the DLLs is missing or corrupted, the bundling layer 408 alerts the administrator that a failure has occurred.
The third layer in the adaptive reference model 402 shown in Figure 4 is the profile layer 410. Profile layer 410 in the embodiment shown detects anomalies based on grouping relationship violations. There are two types of relationships, associative (groupings appear together) and exclusive (groupings never appear together). The profile layer 410 allows the adaptive reference model to detect missing assets not detected by the grouping layer, as well as conflicts between assets. Profile layer 410 determines which groupings have strong associative and exclusive relationships with each other. In such an embodiment, if a particular cluster is not detected in a snapshot where it would normally be expected by virtue of the presence of other clusters with which it has strong associative relationships, then the profile layer 410 detects the absence of that cluster as an anomaly. Similarly, if a cluster is detected in a snapshot where it would not normally be expected due to the presence of other clusters with which it has strong mutually exclusive relationships, then the profile layer 410 detects the presence of the first cluster as an anomaly. The profile layer 410 allows the adaptive reference model 402 to detect anomalies that would not be detectable at lower levels of the bin 404.
The adaptive reference model 402 shown in Figure 4 can be implemented in a number of ways that are well known to those of skill in the art. By optimizing the processing of the adaptive reference model 402 and providing sufficient processing and storage resources, a realization of the
The present invention is capable of supporting an unlimited number of managed populations and individual clients. Both assimilating a new model and using the model in testing involve comparing hundreds of thousands of attribute names and values. Performing these comparisons using the text strings for the names and values is a very demanding processing task. In one embodiment of the present invention, each unique string in an incoming snapshot is assigned an integer identifier. Comparisons are then made using the integer identifiers rather than the strings. Because computers can compare integers much faster than the long strings associated with file names or registry key names, processing efficiency is greatly improved.
The adaptive reference model 402 depends on data from the Agent component (202). The functionality of the Agent component (202) has been described above, which is a functional summary of the user interface and the Agent component (202) in one embodiment of the present invention.
An embodiment of the present invention is capable of comparing log entries across client machines in a managed population. One difficulty in comparing registry keys across different machines running a Microsoft Windows operating system stems from the use of a Globally Unique Identifier ("GUID"). A GUID for a particular item on one machine may differ from the GUID for the same item on a second machine. Accordingly, one embodiment of the present system provides a mechanism for normalizing GUIDs for comparison purposes.
Figure 5 is a flow chart illustrating a customer registration information normalization process in an embodiment of the present invention. In the embodiment shown, GUIDs are first grouped into two groups 502. The first group is for GUIDs that are not unique (duplicates) across machines in the managed population. The second group includes GUIDs that are unique across machines, that is, the same key has a different GUID on different machines within the managed population. The keys for the second group are classified 504 below. In this way, the relationship between two or more keys within the same machine can be identified. The intention is to normalize such relationships in a way that will allow them to be compared across multiple machines.
The embodiment shown below creates a random check for the values in the keys 506. This creates a unique signature for all names, path names, and other values contained in the key. The random check is then replaced by GUID 508. In this way, the singularity is maintained within the machine, but the same random check appears on each machine so that the relationship can be identified. The relationship enables the adaptive reference model to identify anomalies within the managed population.
For example, conventional viruses often change registry keys so that the infected machine will run the executable that spreads the virus. An embodiment of the present invention is capable of identifying registry changes on one or more machines in the population due to its ability to normalize registry keys.
Figure 6 is a flow chart illustrating a method for identifying and responding to an anomaly in one embodiment of the present invention. In the embodiment shown, a processor, such as the Collector component 108, receives a plurality of snapshots from a plurality of computers 602. Although the following discussion describes the process shown in Figure 6 as being performed by the Analytical component (110), any suitable processor can perform the process shown. The plurality of snapshots can comprise as few as two snapshots from two computers. Alternatively, the plurality of snapshots may comprise thousands of snapshots. Snapshots comprise data about computers in a population to be examined. For example, the plurality of snapshots can be received from each of the computers in communication with a local area network of the organization. Each snapshot comprises a collection of asset / value pairs that represent the state of a computer at a particular point in time.
As the Collector component 108 receives the snapshots, it stores them 604. Storing the snapshots may comprise storing them in a data store, such as in database 112 or in memory (not shown). Snapshots can be stored temporarily or permanently. Also, in one embodiment of the present invention, the entire snapshot is stored in a data store. In another embodiment, only the parts of the snapshot that have changed from a previous version (that is, a delta snapshot) are stored.
The Analytical component (110) uses the data in the plurality of snapshots to create an adaptive reference model 606. Each of the snapshots comprises a plurality of assets, comprising a plurality of pairs of asset names and asset values. An asset is an attribute of a computer, such as a file name, a registry key name, a performance parameter, or a communication port. Assets reflect a state of a computer, real or virtual, within the population of computers analyzed. An asset value is the state of an asset at a particular point in time. For example, for a file, the value
ES 2 640 191 T3 may comprise an MD5 random check representing the contents of the file; for a registry key, the value can comprise a text string representing the data assigned to the key.
The adaptive reference model also comprises a plurality of assets. The assets of the adaptive reference model can be compared to the resources of a snapshot to identify anomalies and for other purposes. In one embodiment, the adaptive reference model comprises a collection of data about various relationships between assets that characterize one or more normal computers at a particular point in time.
In one embodiment, the Analytical component (110) identifies a pool of asset names. A grouping comprises one or more non-overlapping groups of asset names that appear together. The Analytical component (110) may also attempt to identify relationships between the clusters. For example, the Analytical component (110) can compute a probability matrix that predicts, given the existence of a particular cluster in a snapshot, the probability of the existence of any other cluster in the snapshot. Probabilities that are based on a large number of snapshots and that are either very high (for example, greater than 95%) or very low (for example, less than 5%) can be used by the model to detect anomalies. Probabilities that are based on a small number of snapshots (that is, a number that is not statistically significant) or that are neither very high nor very low are not used to detect anomalies.
The adaptive reference model can comprise a confidence criterion to determine when a relationship can be used to test a snapshot. For example, the confidence criterion can comprise a minimum threshold for a number of snapshots contained in the adaptive reference model. If the threshold is not exceeded, the relationship will not be used. The adaptive reference may comprise, also or instead, a minimum threshold for a number of snapshots contained in the adaptive reference model that include the ratio, using the ratio only if the threshold is exceeded. In one embodiment, the adaptive reference model comprises a maximum threshold for a ratio of the number of different asset values to the number of snapshots that contain the asset values. The adaptive reference model can comprise one or more minimum and maximum thresholds associated with numerical asset values.
Each of the plurality of assets in the adaptive reference model or in a snapshot can be associated with an asset type. The asset type can comprise, for example, a file, a registry key, a performance measure, a service, a hardware component, a running process, a registry, and a communication port. Other types of actives can also be used by embodiments of the present invention. In order to conserve space, asset names and asset values can be compressed. For example, in one embodiment of the present invention, the Collector component 108 identifies the first occurrence of an asset name or asset value in one of the plurality of snapshots and generates an identifier associated with that first occurrence. Subsequently, if the Collector component (108) identifies a second occurrence of the asset name or asset value, the Collector component (108) associates the identifier with the second asset name and asset value. The asset identifier and name or asset value can then be stored in an index, while only the identifier is stored with the data in the adaptive or snapshot reference model. This minimizes the space required to store frequently repeated asset names or values.
The adaptive reference model can be generated automatically. In one embodiment, the adaptive reference model is generated automatically and then manually reviewed for the knowledge of technical support personnel or others. Figure 11 is a screenshot of a user interface for creating an adaptive reference model in one embodiment of the present invention. In the embodiment shown, a user selects the snapshots to be included in the model by moving them from the Machine Selection Menu window 1102 to the Job Machines window 1104. When the user completes the selection process and presses the button 1106 Finish, an automatic task is created that causes the model to be generated. Once the model has been created, the user can use another interface screen to manage it. Figure 12 is a screenshot of a user interface for managing an adaptive reference model in one embodiment of the present invention.
Referring again to Figure 6, once the adaptive reference model has been created, the Analytical component (110) compares at least one of the plurality of snapshots to the adaptive reference model 608. For example, the Collector component 108 can receive and store one hundred snapshots in the Database component 112. Component (110) Analytical uses the hundred snapshots to create an adaptive reference model. The Analytical component (110) then begins comparing each of the snapshots in the plurality of snapshots with the adaptive reference model. At some point later the Collector component 108 may receive 100 new snapshots from the Agent components, which can then be used by the Analytical component to generate a revised version of the adaptive reference model and to identify anomalies.
IS 2 640 191 T3
In one embodiment of the present invention, comparing one or more snapshots with an adaptive reference model comprises examining relationships between asset names. For example, the probability of existence of a first asset name can be high when a second asset name is present. In one embodiment, the comparison comprises determining whether all the asset names in a snapshot exist within the adaptive reference model and are consistent with a plurality of high probability relationships between asset names.
Still referring to Figure 6, in one embodiment, the Analytical component (110) compares the snapshot to the adaptive reference model in order to identify any anomalies that may be present in a computer 610. An anomaly is an indication that some part of a snapshot deviates from normal as defined by the adaptive reference model. For example, an asset name or value may deviate from the normal asset name and expected asset value in a particular situation as defined by an adaptive reference model. The anomaly may or may not signal that a known or new problem or problem condition exists on or in relation to the computer with which the snapshot is associated. A condition is a group of anomalies that are related. For example, a group of anomalies can be related because they arise from a single root cause. For example, an anomaly may indicate the presence of a particular application on a computer when the application is not generally present on the other computers within a given population. Anomaly recognition can also be used for functions such as capacity balancing. For example, by evaluating the performance metrics of multiple servers, the Analytical component (110) is able to determine when to trigger the automatic deployment and configuration of a new server to address changing demands.
A condition comprises a group of related abnormalities. For example, a group of anomalies may be related because they arise from a single root cause, such as the installation of an application program or the presence of a “worm”. A condition can comprise a kind of condition. The condition class allows several conditions to be grouped with each other.
In the embodiment shown in Figure 6, if an abnormality is found, the Analytical component (110) attempts to match the abnormality with a recognition filter in order to diagnose a condition 612. The abnormality can be identified as a benign abnormality with in order to eliminate noise during analysis, that is, in order to avoid obscuring real problem conditions due to the presence of anomalies that are the result of normal operating processes. A check is a comparison of a snapshot with an adaptive reference model. A check can be performed automatically. The output of a check can comprise a set of anomalies and conditions that have been detected. In one embodiment, the anomaly is matched to a plurality of acknowledgment filters. A recognition filter comprises a signature of a condition or a class of conditions. For example, a recognition filter may comprise a collection of asset name and value pairs that, when taken together, represent the signature of a condition that it is desirable to recognize, such as the presence of a worm. A generic recognition filter can provide a template for creating more specific filters. For example, a recognition filter that is tailored to search for worms in general can be tailored to search for a specific worm.
In one embodiment of the present invention, a recognition filter comprises at least one of: an asset name associated with the condition, an asset value associated with the condition, a combination of asset name and asset value associated with the condition , a maximum threshold associated with an asset value and with the condition, and a minimum threshold associated with an asset value and with the condition. The asset name / value pairs in a snapshot can be compared to the recognition filter name / value pairs to find a match and diagnose a condition. The name / value match can be exact or the recognition filter can comprise a wildcard, which allows a partial value to be entered into the recognition filter and then matched to the snapshot. A particular asset name and / or value can be matched with a plurality of recognition filters in order to diagnose a condition.
You can create a recognition filter in a number of ways. For example, in one embodiment of the present invention, a user copies anomalies from a machine where the condition of interest is present. Anomalies can be presented in an anomaly summary from which they can be selected and copied to the filter. In another embodiment, a user enters a wildcard character in a filter definition. For example, a piece of spyware called Gator generates thousands of registry keys beginning with the string "hklm \ software \ gator \". An embodiment of the present invention can provide a wild card mechanism to effectively deal with this situation. The wildcard character can be, for example, the percent sign (%), and can be used before a text string, after a text string, or in the middle of a text string. Continuing with the Gator example, if the user enters the string "hklm \ software \ gator \%" in the filter body, then any key beginning with "hkml \ software \ gator" will be recognized by the filter. The user may wish to build a filter for a condition that has not yet been experienced in the managed population. For example, a filter for a virus based on information publicly available on the Internet rather than an actual case of the virus within the managed population. To address this situation, the user enters the relevant information directly into a filter.
IS 2 640 191 T3
Figure 13 is a screenshot of a user interface for selecting a snapshot to use for creating a recognition filter in one embodiment of the present invention. A user accesses the screenshot shown to select the snapshots to be used to create the recognition filter. Figure 14 is a screenshot of a user interface for creating or editing a recognition filter in one embodiment of the present invention. In the embodiment shown, the assets of the snapshot selected in the interface illustrated in Figure 13 are displayed in the Data Source window 1402. The user selects these assets and copies them to the 1404 Source window to create the recognition filter.
In one embodiment, the match between a recognition filter and a set of anomalies is associated with a quality measure. For example, an exact match of all asset names and asset values in the recognition filter with asset names and asset values in the anomaly set may be associated with a higher quality measure than a subset match. of the asset names and asset values in the recognition filter with asset names and asset values in the anomaly set.
The recognition filter may also comprise other attributes. For example, in one embodiment, the recognition filter comprises a check mark to determine whether to include the asset name and asset value in the adaptive reference model. In another embodiment, the recognition filter comprises one or more textual descriptions associated with one or more conditions. In yet another embodiment, the recognition filter comprises a severity indicator that indicates the severity of a condition in terms of, for example, how much damage it can cause, how difficult it can be to remove, or some other suitable measure.
The recognition filter may comprise fields that are administrative in nature. For example, in one embodiment, the recognition filter comprises a recognition filter identifier, a creator name, and an update date-time.
Still referring to Figure 6, Analytical component (110) below responds to condition 614. Responding to the condition may include, for example, generating a notification, such as an email to a support technician, submitting a trouble ticket to a problem management system, requesting permission to take an action, for example, requesting the confirmation of a support technician to install a patch, and clear the condition of at least one of the plurality of computers. Eliminating the condition may comprise, for example, causing a response agent to run on any of the plurality of computers affected by the condition. The condition may be associated with an automatic response. Steps for diagnosis 612 and response to conditions 614 can be repeated for each condition. Also, the process of finding anomalies 610 can be repeated for each individual snapshot.
In the embodiment shown in Figure 6, the Analytical component (110) then determines whether an additional 616 snapshots are to be analyzed. If so, the steps of comparing the snapshot to the adaptive reference 608, finding anomalies 610, matching the anomalies with a recognition filter to diagnose a 612 condition, and responding to condition 614 are repeated for each snapshot. Once all snapshots have been analyzed, the process ends 618.
In one embodiment of the present invention, once the Analytical component 110 has identified a condition, the Analytical component 110 attempts to determine which of the plurality of computers within a population are affected by the condition. For example, the Analytical component (110) may examine the snapshots to identify a particular set of anomalies. The Analytical component (110) can then cause a response to the condition to be executed on behalf of each of the affected computers. For example, in one embodiment, an Agent component (202) resides on each of the plurality of computers. The Agent component (202) generates the snapshot that is evaluated by the Analytical component (110). In one such embodiment, the Analytical component (110) uses the Agent component (202) to execute a response program if the Analytical component (110) identifies a condition on one of the computers. In diagnosing a condition, the Analytical component (110) may or may not be able to identify a root cause of a condition.
Figure 7 is a flow chart illustrating a process for identifying certain types of anomalies in one embodiment of the present invention. In the embodiment shown, the Analytical component (110) evaluates snapshots for a plurality of computers 702. These snapshots can be base snapshots comprising the full state of the computer or delta snapshots comprising changes in the state of the computer since the last Base Snapshot. . The Analytical component (110) uses the snapshots to create an adaptive reference model 704. Note that when delta snapshots are used for this purpose, the Analytical component must first reconstitute the equivalent of a base snapshot by applying the changes described in the delta snapshot to the most recent base snapshot. The Analytical component (110) subsequently receives a second snapshot (base or delta) for at least one of the plurality of computers 706. The snapshot can be created based on various events, such as the passage of a predetermined amount of time, the installation of a new program, or some other suitable event.
IS 2 640 191 T3
The Analytical component (110) compares the second snapshot with the adaptive reference model to try and detect anomalies. There can be various types of anomalies in a computer. In the embodiment shown, the Analytical component (110) first attempts to identify asset names that are unexpectedly absent 710. For example, all or substantially all computers within a population may include a particular file. The existence of the file is signaled in the adaptive reference model by the presence of an asset name. If the file is unexpectedly absent from one of the computers within the population, that is, the asset name cannot be found, some condition may be affecting the computer where the file is missing. If the asset name is unexpectedly absent, the absence is identified as a 712 failure. For example, an entry that identifies the computer, the date, and an unexpectedly missing asset can be entered into a data warehouse.
The Analytical component (110) then attempts to identify asset names that are unexpectedly present 714. The presence of an unexpected asset name, such as a file name or registry entry, may indicate the presence of a problematic condition, such as a computer worm. An asset name is unexpectedly present if it has never been seen before or if it has never been seen before in the context in which it is located. If the asset name is present unexpectedly, the presence is identified as a 720 anomaly.
The Analytical component (110) then attempts to identify an unexpected asset value 718. For example, in one embodiment, the Analytical component (110) attempts to identify a string asset value that is unknown to the asset name associated with it. In another embodiment, the Analytical component (110) compares a numeric asset to the minimum or maximum thresholds associated with the corresponding asset name. In embodiments of the present invention, thresholds can be adjusted automatically based on the mean and standard deviation for asset values within a population. According to the embodiment shown, if an unexpected asset value is detected, it is identified as an anomaly 720. The process then terminates 722.
Although the process in Figure 7 is shown as a serial process, comparing a snapshot to the adaptive reference model and identifying anomalies can occur in parallel. Furthermore, each of the steps represented can be repeated numerous times. Additionally, either the delta snapshots or the base snapshots can be compared to the adaptive reference model during each cycle.
Once the analysis is complete, the Analytical component (110) can generate a result, such as an anomaly report. This report can also be provided to a user. For example, the Analytical component (110) may generate a web page comprising the results of a comparison of a snapshot with an adaptive reference model. Embodiments of the present invention may provide a means for performing automated security audits, file and registry integrity checking, anomaly-based virus detection, and automated repair.
Figure 8 is a flow chart illustrating a process for generating an adaptive reference model in one embodiment of the present invention. In the embodiment shown, the Analytical component (110) accesses a plurality of snapshots from a plurality of computers through the Database component. Each of the snapshots comprises a plurality of pairs of asset names and asset values. The Analytical component (110) automatically creates an adaptive reference model that is based, at least in part, on snapshots.
The adaptive reference model can comprise any of a number of attributes, relationships, and measures of the various asset names and values. In the embodiment shown in Figure 8, the Analytical component (110) first finds one or more unique asset names and then determines the number of times that each unique asset name occurs within the plurality of snapshots 804. For example, a file for a basic operating system driver can occur on substantially all computers within a population. The file name is a unique asset name; it will appear only once within a snapshot, but will likely occur in substantially all snapshots.
In the embodiment shown, the Analytical component (110) then determines the unique asset values associated with each asset name 806. For example, the file name asset for the controller described in connection with step 804 will likely have the same value for each occurrence of the filename asset. In contrast, the file value for a log file will likely have as many different values as there are occurrences, that is, a log file on any particular computer will contain a different number of entries from all the other computers in a population.
Since the population can be very large, in the embodiment shown in Figure 8, if the number of unique values associated with an asset name exceeds a threshold 808, the determination stops 810. In other words, in the log file example described above, whether or not the computer is in a normal state does not depend on a log file having a consistent value. The log file contents are expected to vary on each computer. Note, however, that the presence or absence of the log file can be stored in the adaptive reference model as an indication of normality or an abnormality.
IS 2 640 191 T3
In the embodiment shown in Figure 8, the Analytical component (110) then determines the unique string asset values associated with each asset name 812. For example, in one embodiment, there are only two types of asset values, strings and numbers. File and registry key random check values are examples of strings; a performance counter value is an example of a number.
The Analytical component (110) then determines a statistical measure associated with unique numerical values associated with an asset name 814. For example, in one embodiment, the Analytical component (110) captures a performance measure, such as a memory search. . If a computer in a population often searches memory, it may be an indication that a malicious program is running in the background and requiring substantial memory resources. However, if all or a significant number of computers in a population frequently search memory, it may indicate that the computers generally lack memory resources. In one embodiment, the Analytical component (110) determines a mean and standard deviation for numerical values associated with a unique asset name. In the memory example, if the memory search measure for a computer falls far outside the statistical mean for the population, an anomaly can be identified.
In one embodiment of the present invention, the adaptive reference model can be modified by applying a policy template. A policy template is a collection of asset / value pairs that are identified and applied to an adaptive reference model to establish a rule that reflects a specific policy. For example, the policy template may comprise a plurality of asset name and asset value pairs that are expected to be present on a normal computer. In one embodiment, applying the policy template comprises modifying the adaptive reference model so that the pairs of asset names and asset values present in the policy template appear to have been present in each of the plurality of snapshots, that is, they seem to be the normal state of a computer in the population.
Figure 15 is a screen shot illustrating a user interface for selecting an "excellent system" for use in a policy template in one embodiment of the present invention. As described above, the user first selects the excellent system on which the policy template is to be based. Figure 16 is a screenshot of a user interface for selecting policy template assets in one embodiment of the present invention. As with the user interface to create recognition filters. The user selects assets from a Data Source window 1602 and copies them to a content window, Template content window 1604.
Figure 9 is a flow chart, illustrating a process for proactive anomaly detection in one embodiment of the present invention. In the embodiment shown, when an analysis occurs, the Analytical component (110) establishes a connection with the database (112) that stores the snapshots to be analyzed 902. In the embodiment shown, only one database is used. However, in other embodiments, data from multiple databases can be analyzed.
Before the diagnostic checks are run, one or more reference models are 904 created. The reference models are periodically updated, for example once a week, to ensure that the information they contain remains current. One embodiment of the present invention provides a task scheduler that allows model creation to be configured as a fully automated procedure.
Once a reference model has been created, it can be processed in a variety of ways to allow for different types of analysis. For example, it is possible to define a policy template 906 as described above. For example, a policy template might require that all machines in a managed population have antivirus software installed and operational. Once a policy template has been applied to a model, the diagnostic checks against that model will include a policy compliance test. Policy templates can be used in a variety of applications including automated security audits, performance threshold checking, and Windows update management. A policy template comprises the set of assets and values that will be forced into the model as a rule. In one embodiment, the template editing process is based on a "great system" approach. An excellent system is one that presents the assets and values that a user wishes to incorporate into the template. The user locates the snapshot that corresponds to the excellent system and then selects each asset / value pair that the user wants to include in the template.
In the process shown in Figure 9, the policy template is then applied to a model to modify its definition of normal 908. This allows the model to be formed in ways that allow it to check compliance against user-defined policies. as described herein.
A model can also be converted to 910. The conversion process alters a reference model. For example, in one embodiment, the conversion process removes from the model any information assets that are unique, that is, any assets that occur in one and only one snapshot. When running a
ES 2 640 191 T3 check against a converted model, all unique information assets will be reported as anomalies. This type of check is useful in emerging previously unknown problem conditions that exist at the time the Agent components are first installed. Converted models are useful in establishing an initial baseline since they expose unique characteristics. For this reason the converted models are sometimes called baseline models in embodiments of the present invention.
In another embodiment, the model building process removes any information assets that match a recognition filter from the model, ensuring that known problem conditions are not incorporated into the model. When the system is first installed, the managed population quite often contains a number of known problem conditions that have not yet been reported. It is important to discover these conditions and remove them from the model, since they will otherwise be incorporated into the adaptive reference model as part of the normal state of a machine.
Agent component 202 takes a snapshot of the status of each managed machine on a scheduled basis 924. The snapshot is transmitted and entered into the database as a snapshot. Snapshots can also be generated on demand or in response to a specific event, such as the installation of an application.
In the proactive problem management process shown, a periodic 912 check of the latest snapshots is performed against an updated reference model. The output of a periodic check is a set of anomalies, which are displayed to a user as results 914. The results also include any conditions that are identified as a result of matching the anomalies with the recognition filters. The recognition filters can be defined as described 916 above. Anomalies are passed through the recognition filters for interpretation resulting in a set of conditions. Conditions can range in severity from something as benign as a Windows update to something as serious as a Trojan.
The troublesome conditions that can occur in a computer change as the hardware and software components that make up that computer evolve. Consequently, there is a continuing need to define and share new recognition filters as new combinations of anomalies are discovered. Recognition filters can be thought of as a very detailed and structured way of documenting problem conditions and as such represent an important mechanism for facilitating collaboration. The embodiment shown comprises a mechanism for exporting recognition filters to an XML file and importing recognition filters from an XML file.
Once conditions are identified, 920 reports are generated that document the results of a proactive check. The reports may comprise, for example, a summary description of all detected conditions or a detailed description of a particular condition.
Figure 10 is a flow chart, illustrating a reactive process in one embodiment of the present invention. In the process shown in Figure 10, it is assumed that an adaptive reference model has already been created. The process shown begins when a user calls technical support to report a 1002 problem. In the traditional helpdesk paradigm, the next step would be to verbally gather information about the symptoms being experienced by the user. In contrast, in the embodiment of the present invention shown, the next step is to run a diagnostic check of the suspect machine against the most recent snapshot 1003. If this does not produce an immediate diagnosis of a problem condition, there may be three possibilities: (1) The condition has occurred since the last snapshot was taken; (2) the condition is new and is not being recognized by your filters; or (3) the condition is beyond the scope of the analysis, for example, a hardware problem.
If the problem condition is suspected to have occurred since the last snapshot was taken, then the user can have the Agent component (202) on the client machine take another snapshot 1006. Once the resulting snapshot is available, it can be run a new 1004 diagnostic check.
If the problem condition is suspected to be new, the analyst can run a compare function that provides a breakdown of changes in the state of a machine over a specific time window such as new applications that may have been installed 1008. The user You can also see a detailed representation of the state of a machine at various points in 1010 time. If the analyst identifies a new problem condition, the user can identify the asset set as a recognition filter for subsequent 1012 analyzes.
Although conventional products have focused on improving the efficiency of the human-based model of support, embodiments of the present invention are designed around a different paradigm, a machine-based model of support. This fundamental difference in approach manifests itself most profoundly in the areas of data collection and analysis. Since much of the analysis of the collected data will be performed by a machine rather than a human being, the collected data can be voluminous. For
For example, in one embodiment, data collected from a single machine, known as a "health check" or snapshot for the machine, includes values for hundreds of thousands of attributes. The ability to collect a large volume of data provides embodiments of the present invention with a significant advantage over conventional systems in terms of the number and variety of conditions that can be detected.
Another embodiment of the present invention provides powerful analytical capabilities. The basis for high-value analysis in such an embodiment is the ability to accurately distinguish between normal and abnormal conditions. For example, a system according to the present invention synthesizes its reference model automatically by extracting statistically significant relationships from the snapshot data it collects from its customers. The resulting "adaptive" reference model defines what is normal for that particular managed population at that particular point in time.
One embodiment of the present invention combines the data collection and adaptive analysis features described above. In such an embodiment, the superior data collection capabilities combined with the analytical power of the adaptive reference model translate into a number of significant competitive advantages, including the ability to provide automatic protection against security threats by conducting daily security audits and checking software updates to eliminate vulnerabilities. Such an embodiment may also be capable of proactively scanning all managed systems on a routine basis for problems before they result in loss of productivity or calls for technical support.
An embodiment of the present invention that implements the capabilities of the adaptive reference model is also capable of detecting previously unknown problem conditions. Furthermore, such an embodiment is synthesized and maintained automatically, requiring little or no vendor updates to be effective. Such an embodiment is automatically customized for a particular managed population, allowing you to detect failure modes unique to that population.
A further advantage of an embodiment of the present invention is that in the event that a problem condition cannot be resolved automatically, such an embodiment can provide a massive amount of structured technical information to facilitate the work of the support analyst.
An embodiment of the present invention provides the ability to automatically repair an identified problem. Such an embodiment, when combined with the adaptive reference model of the previously described embodiment, is uniquely capable of automated repair due to its ability to identify all aspects of a problem condition.
Embodiments of the present invention also provide many advantages over conventional systems and methods in terms of the levels of service described herein. For example, in terms of the Mass Healing service level, it is considerably less expensive to prevent an incident than to resolve an incident after damage has occurred. Embodiments of the present invention substantially increase the percentage of incidents that can be detected / prevented without the need for human intervention and in a manner that encompasses the diverse and dynamic nature of computers in real world environments.
Furthermore, an embodiment of the present invention is capable of directing the Self-Healing service level by automatically detecting and repairing both known and unknown anomalies. An embodiment that implements the adaptive reference model described herein is uniquely suited to automatic detection and repair. Automated service and repair also help eliminate or at least minimize the need for Self-Service and Visits from the Technical Support side.
Embodiments of the present invention provide benefits at the Assisted Service level by providing superior diagnostic capabilities and extensive information resources. One embodiment collects and analyzes massive amounts of end-user data, facilitating a variety of needs associated with the human-based support model including: security audits, configuration audits, inventory management, performance analysis, and problem diagnosis.
Contents13
5 priority claims, no other members on record
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 494225P | United States of America | – | |
| 49422503 | United States of America | P | |
| 2004026186 | United States of America | W | |
| 916800 | United States of America | – | |
| 91680004 | United States of America | A |
Numbers
- Publication
- 2640191
- Application
- 4786501
Titles2
- Spanish
- Sistemas y métodos para soporte informático automatizado
- English
- Systems and methods for automated computer support
Classification
- CPC, 6
- G06F11/0793
- G06F11/0748
- G06F11/079
- G06F21/55
- G06F21/577
- G06F11/3452
- IPC, 7
- G06F17 30
- G06F11 07
- G06F
- G06F9 44
- G06F11 34
- G06F17 00
- G06F21 00