EP0805433A2

Method and system of runtime acoustic unit selection for speech synthesis

Abstract

The present invention pertains to a concatenative speech synthesis system and method which produces a more natural sounding speech. The system provides for multiple instances of each acoustic unit which can be used to generate a speech waveform representing an linguistic expression. The multiple instances are formed during an analysis or training phase of the synthesis process and are limited to a robust representation of the highest probability instances. The provision of multiple instances enables the synthesizer to select the instance which closely resembles the desired instance thereby eliminating the need to alter the stored instance to match the desired instance. This in essence minimizes the spectral distortion between the boundaries of adjacent instances thereby producing more natural sounding speech.

EP0805433A2, drawing sheet 1
Sheet 1 of 15

Term

Term ended

Projected expiry passed 29 April 2017, 9.4 years ago.

  1. Priority
  2. Filed
  3. Published
  4. Projected expiry
  5. Today

20 claims: 5 independent, 15 dependent

  1. 1
    A method in a computer system for producing speech from an input linguistic expression, said method comprising the steps of:converting the input linguistic expression into a plurality of acoustic units of speech;providing a plurality of instances for each acoustic unit, each instance indicating acoustic properties of a speech signal used to generate the speech associated with the acoustic unit;forming a plurality of sequences of instances which correspond to the acoustic units in the linguistic expression;for each sequence, determining the dissimilarity between adjacent instances in the sequence;selecting the best sequence having minimal dissimilarities between adjacent instances;and generating the speech which results from the best sequence.
  2. 2
    In a computer system having a storage device, a method of synthesizing speech, comprising the steps of:providing multiple instances of a first acoustical unit in the storage device;providing multiple instances of a second acoustical unit in the storage device;and synthesizing speech by selecting instances to minimize distortion between selected instances and concatenating one of the instances provided for the first acoustical unit and one of the instances provided for the second acoustical unit.
  3. 8
    In a computer system, a method comprising the steps of:providing a set of instances of an acoustical unit;pruning the set of instances of the acoustical unit to produce a robust set of instances of the acoustical unit;and selecting one of the instances from the robust set of instances of the acoustical unit to synthesize speech.
  4. 14
    In a computer system having a storage device, a method of synthesizing speech, comprising the steps of:processing an input text string into a phoneme string;converting the phoneme string into a diphone string having diphones with boundaries;providing in the storage, multiple instances of each diphone in the diphone string;selecting ones of the instances of the diphones in the diphone string that result in minimal spectral distortion between the boundaries of adjacent diphones;and concatenating the selected ones of the instances of the diphones to synthesize speech.
  5. 16
    A computer system, comprising:a storage device for storing multiple instances of an acoustical unit;a speech synthesizer for synthesizing speech, comprising: a selection unit for selecting one of the instances of the stored multiple instances of the acoustical unit;and a speech output unit for using the selected one of instances of the acoustical unit with at least one other instance of a different acoustical unit to output synthesized speech.