EP1501075A2

Speech synthesis using concatenation of speech waveforms

Abstract

A high quality speech synthesizer in various embodiments concatenates speech waveforms referenced by a large speech database. Speech quality is further improved by speech unit selection and concatenation smoothing.

EP1501075A2, drawing sheet 1
Sheet 1 of 8

Term

Term ended

Projected expiry passed 12 November 2019, 6.9 years ago.

  1. Priority
  2. Filed
  3. Published
  4. Projected expiry
  5. Today

9 claims: 7 independent, 2 dependent

  1. 1
    A speech synthesizer comprising:a. a speech database referencing speech waveforms;b. a speech waveform selector, in communication with the speech database, that selects waveforms referenced by the database using designators that correspond to a phonetic transcription input;and c. a speech waveform concatenator, in communication with the speech database, that concatenates waveforms selected by the speech waveform selector to produce a speech signal output,    wherein, for at least one ordered sequence of a first waveform and a second waveform, the concatenator selects (i) a location of a trailing edge of the first waveform and (ii) a location of a leading edge of the second waveform, each location being selected so as to produce an optimization of a phase match between the first and second waveforms in regions near the locations.
  2. 2
    A speech synthesizer comprising:a. a speech database referencing speech waveforms;b. a speech waveform selector, in communication with the speech database, that selects waveforms referenced by the database using designators that correspond to a phonetic transcription input;and c. a speech waveform concatenator, in communication with the speech database, that concatenates waveforms selected by the speech waveform selector to produce a speech signal output,    wherein, for at least one ordered sequence of a first waveform and a second waveform, the concatenator selects the location of a trailing edge of the first waveform, the location being selected so as to produce an optimization of a phase match between the first and second waveforms in regions near the location and a leading edge of the second waveform.
  3. 3
    A speech synthesizer comprising:a. a speech database referencing speech waveforms;b. a speech waveform selector, in communication with the speech database, that selects waveforms referenced by the database using designators that correspond to a phonetic transcription input;and c. a speech waveform concatenator, in communication with the speech database, that concatenates waveforms selected by the speech waveform selector to produce a speech signal output,    wherein, for at least one ordered sequence of a first waveform and a second waveform, the concatenator selects the location of a leading edge of the second waveform, the location being selected so as to produce an optimization of a phase match between the first and second waveforms in regions near the location and a trailing edge of the first waveform.
  4. 4
    A speech synthesizer according to any of claims 1 through 3, wherein the optimization is determined on the basis of similarity in shape of the first and second waveforms in the regions near the locations.
  5. 5
    A speech synthesizer according to 4, wherein similarity is determined using a cross-correlation technique.
  6. 7
    A speech synthesizer according to any of claims 1 through 3 and 5, wherein the optimization is determined using at least one non-rectangular window.
  7. 8
    A speech synthesizer according to any of claims 1 through 3, and 5, wherein the optimization is determined in a plurality of successive stages in which time resolution associated with the first and second waveforms is made successively finer.