EP2192575A1

Speech recognition based on a multilingual acoustic model

Abstract

The present invention relates to a method for generating a multilingual speech recognizer comprising a multilingual acoustic model, comprising the steps of providing a first speech recognizer comprising a first codebook consisting of first Gaussians and a first Hidden Markov Model, HMM, comprising first states; providing at least one second speech recognizer comprising a second codebook consisting of second Gaussians and a second Hidden Markov Model, HMM, comprising second states; replacing each of the second Gaussians of the at least one second speech recognizer by the respective closest one of the first Gaussians and/or each of the second states of the second HMM of the at least one second speech recognizer with the respective closest state of the first HMM of the first speech recognizer to obtain at least one modified second speech recognizer and combining the first speech recognizer and the at least one modified second speech recognizer to obtain the multilingual speech recognizer.

EP2192575A1, drawing sheet 1
Sheet 1 of 11

Term

2.2 yearsto projected expiry

Projected expiry 27 November 2028, counted from filing; an application has no term until it is granted.

  1. Priority and filed
  2. Published
  3. Today
  4. Projected expiry

15 claims: 5 independent, 10 dependent

  1. 1
    Method for generating a multilingual speech recognizer comprising a multilingual acoustic model, comprising the steps of providing a first speech recognizer comprising a first codebook consisting of first Gaussians and a first Hidden Markov Model, HMM, comprising first states;providing at least one second speech recognizer comprising a second codebook consisting of second Gaussians and a second Hidden Markov Model, HMM, comprising second states;replacing each of the second Gaussians of the at least one second speech recognizer by the respective closest one of the first Gaussians and/or each of the second states of the second HMM of the at least one second speech recognizer with the respective closest state of the first HMM of the first speech recognizer to obtain at least one modified second speech recognizer;and combining the first speech recognizer and the at least one modified second speech recognizer to obtain the multilingual speech recognizer.
  2. 2
    Method for generating a speech recognizer comprising a multilingual acoustic model, comprising the steps of providing a first speech recognizer comprising a first codebook consisting of first Gaussians and a first Hidden Markov Model, HMM, comprising first states;providing at least one second speech recognizer comprising a second codebook consisting of second Gaussians and a second Hidden Markov Model, HMM, comprising second states;determining mean vectors of states for the first states of the first HMM of the first speech recognizer;determining HMMs of the first speech recognizer based on the determined mean vectors of states;replacing the second HMM of the at least one second speech recognizer by the closest HMM of the first speech recognizer to obtain at least one modified second speech recognizer;and combining the first speech recognizer and the at least one modified second speech recognizer to obtain the multilingual speech recognizer.
  3. 9
    The method according to one of the claims 2, 3 and 5 to 7, wherein the closest HMM of the first speech recognizer is determined based on Euclidean distances between first states of the first HMM of the first speech recognizer and second states of the second HMM of the second speech recognizer.
  4. 10
    The method according to one of the preceding claims, wherein the first speech recognizer is modified by modifying the first codebook before combining it with the at least one modified second speech recognizer to obtain the multilingual speech recognizer, wherein the step of modifying the first codebook comprises adding at least one of the second Gaussians of the second codebook of the at least one second speech recognizer to the first codebook.
  5. 13
    Speech recognition means or speech dialog system or speech control system comprising a multilingual speech recognizer generated by the method according to one of the preceding claims.
  6. 14
    Audio device, in particular, an MP3 or MP4 player, cell phone or a Personal Digital Assistant, or a video device comprising a speech recognition or speech dialog system or speech control system means comprising a multilingual speech recognizer generated according to the method according to one of the claims 1 to 12.
  7. 15
    Computer program product, comprising one or more computer readable media having computer-executable instructions for performing the steps of the method according to one of the claims 1 to 12 when run on a computer.