EP0805434B1

Method and system for speech recognition using continuous density hidden Markov models

Abstract

This record has no abstract on file.

EP0805434B1, drawing sheet 1
Sheet 1 of 21

Term

Term ended

Expired 29 April 2017, 9.4 years ago.

  1. Priority
  2. Filed
  3. Granted
  4. Expired
  5. Today

18 claims: 3 independent, 15 dependent

  1. 1
    A method for use in a computer system for matching an input speech utterance to a linguistic expression, the method comprising the steps of:for each of a plurality of phonetic units of speech, providing (124) a plurality of more-detailed acoustic models and a less-detailed acoustic model to represent the phonetic unit, each acoustic model having a plurality of states followed by a plurality of transitions, each state representing a portion of a speech utterance occurring in the phonetic unit at a certain point in time and having an output probability indicating a likelihood of a portion of an input speech utterance occurring in the phonetic unit at a certain point in time;selecting sequences of more-delaited acoustic models representing a plurality of linguistic expressions likely to match the input speech utterance;for each of the select sequences of more-detailed acoustic models, determining (126) how close the input speech utterance matches the sequence, the matching further comprising the step of for each state of the select sequence of more-detailed acoustic models, determining (134) an accumulative output probability as a combination of the output probability of the state and a same state of the less-detailed acoustic model representing the same phonetic unit;and determining (136) the sequence which best matches the input speech utterance, the sequence representing the linguistic expression.
  2. 4
    The method according to one of the claims 1-3 wherein the step of providing a plurality of more-detailed acoustic models further comprises the step of training each acoustic model using an amount of training data of speech utterances;and wherein the step of determining the output probability further comprises the step of weighing the less-detailed model and more-detailed model output probabilities relative to the amount of training data used to train each acoustic model.
  3. 12
    A computer system for matching an input speech utterance to a linguistic expression, comprising:a storage device (28) for storing a plurality of more-detailed and less-detailed acoustic models representing respective ones of phonetic units of speech, the plurality of more-detailed acoustic models which represent each phonetic unit having at least one associated less-detailed acoustic model representing the phonetic unit of speech, each acoustic model comprising states having transitions, each state representing a portion of the phonetic unit at a certain point in time and having an output probability indicating a likelihood of a portion of the input speech utterance occurring in the phonetic unit at a certain point in time;a model sequence generator (20) which provides select sequences of more-detailed acoustic models representing a plurality of linguistic expressions likely to match the input speech utterance;a processor (34) for determining how well each of the select sequence of models matches the input speech utterance, the processor matching a portion of the input speech utterance to a state in the sequence by utilizing an accumulative output probability for each state of the sequence, the accumulative output probability including the output probability of each state of the more-detailed acoustic model combined with the output probability of a same state of the associated less-detailed acoustic model;and a comparator (34) to determine the sequence which best matches the input speech utterance, the sequence representing the linguistic expression.