EP0745971A2

Pitch lag estimation system using linear predictive coding residual

Abstract

A pitch estimation device and method utilizing a multi-resolution approach to estimate a pitch lag value (614) of input speech. The system includes determining the LPC residual of the speech and sampling the LPC residual (602). A discrete Fourier transform is applied (606) and the result is squared (608). A DFT on the squared amplitude is then performed (610) to transform the LPC residual samples into another domain. An initial pitch lag (614) can then be found with lower resolution. After getting the low-resolution pitch lag estimate, a refinement algorithm is applied (618) to get a higher-resolution pitch lag. The refinement algorithm is based on minimizing the prediction error in the time domain. The refined pitch lag then can be used directly in the speech coding.

EP0745971A2, drawing sheet 1
Sheet 1 of 35

Term

Term ended

Projected expiry passed 22 May 2016, 10.3 years ago.

  1. Priority
  2. Filed
  3. Published
  4. Projected expiry
  5. Today

14 claims: 7 independent, 7 dependent

  1. 1
    A system for estimating pitch lag for speech quantization and compression, the speech having a linear predictive coding (LPC) residual signal defined by a plurality of LPC residual samples, wherein the estimation of a current LPC residual sample is determined in the time domain according to a linear combination of past samples, the system comprising:means for applying a first discrete Fourier transform (DFT) (606) to the plurality of LPC residual samples, the first DFT having an associated amplitude;means for squaring the amplitude (608) of the first DFT;means for applying a second DFT over the squared amplitude (610), the second DFT having associated time domain-transformed samples;and means for determining an initial pitch lag value (614) according to the time domain-transformed samples.
  2. 4
    A system operable with a computer for estimating pitch lag for input speech quantization and compression, the speech having a linear predictive coding (LPC) residual signal defined by a plurality of LPC residual samples (602), wherein the estimated pitch lag falls within a predetermined minimum and maximum pitch lag value range, the system comprising:means for selecting a pitch analysis window (604) among the LPC residual samples (602), the pitch analysis window being at least twice as large as the maximum pitch lag value;means for applying a first discrete Fourier transform (DFT) (606) to the windowed plurality of LPC residual samples, the first DFT having an associated amplitude;means for applying a second DFT (610) over the amplitude of the second DFT having associated time domain-transformed samples;means for applying a weighted average to the time domain-transformed samples, wherein at least two samples are combined to produce a single sample;means for searching the time-domain transformed speech samples to find at least one sample having a maximum peak value;and means for estimating an initial pitch lag value (714) according to the sample having the maximum peak value.
  3. 7
    The apparatus of claims 1 or 4, further comprising a low pass filter for filtering high frequency components of the amplitude of the first DFT (709).
  4. 8
    The system of claims 1 or 4, further comprising means for applying a Hamming window (705) to the LPC residual samples before applying the first DFT (606).
  5. 10
    The system of claims 1 or 9, further comprising:speech input means for receiving the input speech;means for determining the LPC residual signal of the input speech;a processor for processing the initial pitch lag value to represent the LPC residual signal as coded speech;and speech output means for outputting the coded speech.
  6. 11
    A method of estimating pitch lag for speech quantization and compression, the speech being represented by a linear predictive coding (LPC) residual which is defined by a plurality of LPC residual samples (602), wherein the estimation of a current LPC residual sample is determined in the time domain according to a linear combination of past samples, the method comprising the steps of:applying a first discrete Fourier transform (DFT (606)) to the LPC residual samples, the first DFT having an associated amplitude;squaring the amplitude of the first DFT (608);applying a second DFT (610) over the squared amplitude of the first DFT to produce time domain-transformed LPC residual samples;determining an initial pitch lag value (614) according to the time domain-transformed LPC residual samples, the initial pitch lag value having an associated prediction error;refining the initial pitch lag value (618), wherein the associated prediction error is minimized;and coding the LPC residual samples according to the refined pitch lag value.
  7. 13
    A speech coding method for reproducing and coding input speech, the speech coding apparatus operable with a linear predictive coding (LPC) excitation signal defining the decoded LPC residual of the input speech, LPC parameters, and an innovation codebook representing pseudo-random signals which form a plurality of vectors which are referenced to excite speech reproduction to generate speech, the speech coding method comprising the steps of:receiving and processing the input speech;processing the input speech, wherein the step of processing includes: determining the LPC residual of the input speech, determining a coding frame (810) within the LPC residual (602), subdividing the coding frame (810) into plural pitch subframes (802, 804), defining a pitch analysis window (806) having N LPC residual samples (602), the pitch analysis window extending across the pitch subframes, roughly estimating an initial pitch lag value (614) for each pitch subframe, dividing each pitch subframe into multiple coding subframes (808), such that the initial pitch lag estimate for each pitch subframe represents the lag estimate for the last coding subframe of each pitch subframe, and interpolating the estimated pitch lag values (720) between the pitch subframes for determining a pitch lag estimate for each coding subframe, and refining the linearly interpolated lag values (722);and outputting speech reproduced according to the refined pitch lag values.