EP0750293A2

State transition model design method and voice recognition method and apparatus using same

Abstract

An object of the invention is to provide a method of generating a state transition model capable of high speed voice recognition and to provide a voice recognition method and apparatus using the state transition model. To this end, a method is provided which generates a state transition model in which a state shared structure of the state transition model is designed, the method including a step of setting the states of a triphone state transition model in an acoustic space as initial clusters, a clustering step of generating a cluster containing the initial clusters by top-down clustering, a step of determining a state shared structure by assigning a short distance cluster among clusters generated by the clustering step, to the state transition model and a step of learning a state shared model by analyzing the states of the triphones in accordance with the determined state shared structure.

EP0750293A2, drawing sheet 1
Sheet 1 of 12

Term

Term ended

Projected expiry passed 18 June 2016, 10.3 years ago.

  1. Priority
  2. Filed
  3. Published
  4. Projected expiry
  5. Today

17 claims: 6 independent, 11 dependent

  1. 1
    A method of processing signals representative of speech samples to generate a state transition model in which a state shared structure of the state transition model is determined, the method comprising:a step of setting the states of a triphone state transition model in an acoustic space as initial clusters;a clustering step of generating a cluster containing said initial clusters by top-down clustering;a step of determining a state shared structure by assigning a short distance cluster among clusters generated by said clustering step, to the state transition model;and a step of learning a state shared model by analyzing the states of the triphones in accordance with the determined state shared structure.
  2. 6
    A voice recognition apparatus using a state transition model, comprising:input means for inputting voice information;analyzing means for analyzing the voice information input from said input means;likelihood calculation means for calculating a likelihood between the voice information analyzed by said analyzing means and the state transition model;and output means for outputting as a recognition result a language series having a largest likelihood determined by said likelihood calculation means, wherein the state transition model is a model obtained by: setting the states of a triphone state transition model in an acoustic space as initial clusters;generating a cluster containing said initial clusters by top-down clustering;determining a state shared structure by assigning a short distance cluster among clusters generated by said clustering step, to the state transition model;and learning a state shared model by analyzing the states of the triphones in accordance with the determined state shared structure.
  3. 11
    A voice recognition method using a state transition model, comprising:an input step of inputting voice information;an analyzing step of analyzing the voice information input from said input means;a likelihood calculation step of calculating a likelihood between the voice information analyzed by said analyzing means and the state transition model;and an output step of outputting as a recognition result a language series having a largest likelihood determined by said likelihood calculation means, wherein the state transition model is a model obtained by: setting the states of a triphone state transition model in an acoustic space as initial clusters;generating a cluster containing said initial clusters by top-down clustering;determining a state shared structure by assigning a short distance cluster among clusters generated by said clustering step, to the state transition model;and learning a state shared model by analyzing the states of the triphones in accordance with the determined state shared structure.
  4. 12
    A method of generating a state transition model in which a state shared structure of the state transition model is determined, the method comprising:the step of arranging the states of a transition model in an acoustic space as an initial cluster;the step of iteratively dividing said states in said acoustic space into a number of sub-clusters;and the step of determining a state shared structure by grouping acoustically similar clusters and assigning them to the state transition model.
  5. 15
    A method of generating a state transition model in which a state shared structure of the state transition model is determined, the method comprising:a step of setting the states of a triphone state transition model in an acoustic space as initial clusters;a clustering step of generating a cluster containing said initial clusters by top-down clustering;a step of determining a state shared structure by assigning a short distance cluster among clusters generated by said clustering step, to the state transition model;and a step of learning a state shared model by analysing the states of the triphones in accordance with the determined state shared structure.
  6. 16
    A data carrier programmed with instructions for carrying out the method according to any of claims 1 to 5 or 11 to 15.
  7. 17
    A data carrier conveying a state transition model as generated by a method according to any one of claims 1 to 5 or 12 to 15.