EP0805433B1

Method and system of runtime acoustic unit selection for speech synthesis

Abstract

This record has no abstract on file.

EP0805433B1, drawing sheet 1
Sheet 1 of 14

Term

Term ended

Expired 29 April 2017, 9.4 years ago.

  1. Priority
  2. Filed
  3. Granted
  4. Expired
  5. Today

19 claims: 2 independent, 17 dependent

  1. 1
    A computer readable medium having stored thereon instructions for performing speech synthesis (36), comprising instructions for generating:a speech unit store (28) according to the steps of: obtaining an estimate of hidden Markov models (HMMs) for a plurality of speech units;receiving training data as a plurality of speech waveforms (64);segmenting (52) the speech waveforms (64) by performing the steps of: obtaining text (62) associated with the speech waveforms (64);and converting the text (62) into a speech unit string (66) formed of a plurality of training speech units (70);re-estimating (54) the HMMs based on the training speech units (70), each HMM having a plurality of states, each state having a corresponding senone (72, 74, 76);and repeating (56) the steps of segmenting (52) and re-estimating (54) until a probability of the parameters of the HMMs generating the plurality of speech waveforms reaches a threshold level;and mapping (58) each waveform to one or more states and corresponding senones of the HMMs to form a plurality of instances corresponding to each training speech unit (70) and storing the plurality of instances in the speech unit store (28);and a speech synthesizer (36) component configured to synthesize an input linguistic expression by performing the steps of: converting (124) the input linguistic expression into a sequence of input speech units;generating (130) a plurality of sequences of instances corresponding to the sequence of input speech units based on the plurality of instances in the speech unit store;and generating (132) speech based on one of the sequences of instances having a lowest dissimilarity between adjacent instances in the sequence of instances.
  2. 11
    A method of performing speech synthesis, comprising:obtaining an estimate of hidden Markov models (HMMs) for a plurality of speech units;receiving training data as a plurality of speech waveforms (64);segmenting (52) the speech waveforms (64) by performing the steps of: obtaining text (62) associated with the speech waveforms (64);and converting the text (62) into a speech unit string (66) formed of a plurality of training speech units (70);re-estimating (54) the HMMs based on the training speech units (70), each HMM having a plurality of states, each state having a corresponding senone (72, 74, 76);repeating (56) the steps of segmenting (52) and re-estimating (54) until a probability of the parameters of the HMMs generating the plurality of speech waveforms reaches a threshold level;mapping (58) each waveform to one or more states and corresponding senones of the HMMs to form a plurality of speech unit instances corresponding to each training speech unit (70), and storing the plurality of speech unit instances;receiving (122) an input linguistic expression;converting (124) the input linguistic expression into a sequence of input speech units;generating (130) a plurality of sequences of instances corresponding to the sequence of input speech units based on the plurality of speech unit instances stored;and generating (132) speech based on one of the sequences of instances having a lowest dissimilarity between adjacent instances in the sequence of instances.