EP0942384A2

Method and apparatus to model the variables of a data set

Abstract

The present invention relates to modelling the variables of a data set by means of a probabilistic network including data nodes and causal links. The term 'probabilistic networks' includes Bayesian networks, belief networks, causal networks and knowledge maps. The variables of an input data set are registered and a population of genomes is generated each of which individually models the input data set. Each genome has a chromosome to represent the data nodes in a probabilistic network and a chromosome to represent the causal links between the data nodes. A crossover operation is performed between the chromosome data of parent genomes in the population to generate offspring genomes. The offspring genomes are then added to the genome population. A scoring operation is performed on genomes in the said population to derive scores representing the correspondence between the genomes and the input data. Genomes are selected from the population according to their scores and the crossover, scoring, addition and selecting operations for a plurality of generations of the genomes. Finally a genome is selected from the last generation according to the best score. A mutation operation may be performed on the genomes. The mutation may consist of the addition or deletion of a data node and the addition or deletion of a causal link.

EP0942384A2, drawing sheet 1
Sheet 1 of 4

Term

Term ended

Projected expiry passed 10 February 2019, 7.6 years ago.

  1. Priority
  2. Filed
  3. Published
  4. Projected expiry
  5. Today

22 claims: 6 independent, 16 dependent

  1. 1
    A method of modelling the variables in an input data set by means of a probabilistic network including data nodes and causal links, the method comprising the steps of;registering the input data set, generating a population of genomes each individually modelling the input data set by means of chromosome data to represent the data nodes in a probabilistic network and the causal links between the data nodes, performing a crossover operation between the chromosome data of parent genomes in the population to generate offspring genomes, performing an addition operation to add the offspring genomes to the said population, performing a scoring operation on genomes in the said population to derive scores representing the correspondence between the genomes and the input data set, performing a selecting operation to select genomes from the population according to the scores, repeating the crossover, scoring, addition and selecting operations for a plurality of generations of the genomes, and selecting, as an output model, a genome from the last generation.
  2. 7
    A method as claimed in any one of the preceding claims, wherein the chromosome data of each genome includes a node chromosome in the form of a linear array of node data and a causal link chromosome in the form of a matrix of causal links.
  3. 9
    A method as claimed in any one of the preceding claims, comprising the further step of culling genomes which fail to meet predetermined structural constraints.
  4. 12
    Apparatus for modelling the variables in an input data set by means of a probabilistic network including data nodes and causal links, the apparatus comprising;data register means to register the input data set, generating means for generating a population of genomes each individually modelling the input data set by means of chromosome data to represent the data nodes in a probabilistic network and the causal links between the data nodes, crossover means for performing a crossover operation between the chromosome data of parent genomes in the population to generate offspring genomes, adding means to perform an addition operation to add the offspring genomes to the said population, scoring means for performing a scoring operation on genomes in the said population to derive scores representing the correspondence between the genomes and the input data set, selecting means for performing a selecting operation to select genomes from the population according to the scores, control means to control the crossover, scoring, addition and selecting means to repeat their operations for a plurality of generations of the genomes, and output means to select, as an output model, a genome from the last generation.
  5. 18
    Apparatus as claimed in any one of claims 12 to 17, wherein the generating means is adapted to generate genomes each of which includes a node chromosome in the form of a linear array of node data and a causal link chromosome in the form of a matrix of causal links.
  6. 20
    Apparatus as claimed in any one of claims 12 to 19, further comprising culling means for culling genomes which fail to meet predetermined structural constraints.