US6785652B2

Method and apparatus for improved duration modeling of phonemes

Summary by NHIP

Root sinusoidal phoneme modeling

The method identifies a non-exponential functional transformation containing an inflection point and incorporates it into a generalized additive model. This transformation uses a root sinusoidal form where parameters A and B define minimum and maximum phoneme durations, while alpha and beta control the slope and inflection location.

Claim Score by NHIP

Read claim 1, the broadest

Abstract

A method and an apparatus for improved duration modeling of phonemes in a speech synthesis system are provided. According to one aspect, text is received into a processor of a speech synthesis system. The received text is processed using a sum-of-products phoneme duration model that is used in either the formant method or the concatenative method of speech generation. The phoneme duration model, which is used along with a phoneme pitch model, is produced by developing a non-exponential functional transformation form for use with a generalized additive model. The non-exponential functional transformation form comprises a root sinusoidal transformation that is controlled in response to a minimum phoneme duration and a maximum phoneme duration. The minimum and maximum phoneme durations are observed in training data. The received text is processed by specifying at least one of a number of contextual factors for the generalized additive model. An inverse of the non-exponential functional transformation is applied to duration observations, or training data. Coefficients are generated for use with the generalized additive model. The generalized additive model comprising the coefficients is applied to at least one phoneme of the received text resulting in the generation of at least one phoneme having a duration. An acoustic sequence is generated comprising speech signals that are representative of the received text.

US6785652B2, drawing sheet 1
Sheet 1 of 19

Term

Term ended

Expired 18 December 2017, 8.8 years ago.

  1. Priority
  2. Filed
  3. Granted
  4. Expired
  5. Today

41 claims: 5 independent, 36 dependent

  1. 1
    Broadest claimClaim Score 87, broad(NHIP)A method comprising:identifying a non-exponential functional transformation that defines a shape containing an inflection point, wherein the functional transformation comprises a root sinusoidal transformation;and incorporating the functional transformation into a generalized additive model for modeling phoneme durations.
  2. 10
    A computer-readable medium having executable instructions to cause a processor to perform a method comprising:identifying a non-exponential functional transformation that defines a shape containing an inflection point, wherein the functional transformation comprises a root sinusoidal transformation;and incorporating the functional transformation into a generalized additive model for modeling phoneme durations.
  3. 19
    A system comprising:a processor coupled to a memory through a bus;and a process executed from the memory by the processor to cause the processor to identify a non-exponential functional transformation that defines a shape containing an inflection point, and incorporate the functional transformation into a generalized additive model for modeling phoneme durations, wherein the functional transformation comprises a root sinusoidal transformation.
  4. 28
    An apparatus comprising:means for identifying a non-exponential functional transformation that defines a shape containing an inflection point, wherein the functional transformation comprises a root sinusoidal transformation;and means for incorporating the functional transformation into a generalized additive model for modeling phoneme durations.
  5. 37
    An apparatus comprising:means for receiving text signals;means for synthesizing an acoustic sequence from the text signals using a phoneme duration model, the phoneme duration model produced by incorporating a functional transformation form with an inflection point into a generalized additive model that calculates phoneme durations, wherein the functional transformation form comprises a root sinusoidal transformation, the root sinusoidal transformation controlled in response to a minimum phoneme duration and a maximum phoneme duration;and means for providing speech signals representative of the received text.