EP0762386A2

Method and apparatus for CELP coding an audio signal while distinguishing speech periods and non-speech periods

Abstract

For the CELP (Code Excited Linear Prediction) coding of an input audio signal (S), an autocorrelation matrix (R), a speech/noise decision signal (v) and a vocal tract prediction coefficient (a) are fed to an adjusting section (111). In response, the adjusting section (222) computes a new autocorrelation matrix (Ra) based on the combination of the autocorrelation matrix of the current frame and that of a past period determined to be noise. The new autocorrelation matrix (Ra) is fed to an LPC (Linear Prediction Coding) analyzing section (103). The analyzing section computes a vocal tract prediction coefficient (a) based on the autocorrelation matrix (R) and delivers it to a prediction gain computing section (112). At the same time, in response to the above new autocorrelation matrix (Ra), the analyzing section (103) computes an optimal vocal tract prediction coefficient (aa) by correcting the vocal tract prediction coefficient (a). The optimal vocal tract prediction coefficient (aa) is fed to a synthesis filter (104).

EP0762386A2, drawing sheet 1
Sheet 1 of 11

Term

Term ended

Projected expiry passed 22 August 2016, 10.1 years ago.

  1. Priority
  2. Filed
  3. Published
  4. Projected expiry
  5. Today

18 claims: 6 independent, 12 dependent

  1. 1
    A method of CELP coding an input audio signal (S), comprising the steps of:(a) classifying the input acoustic signal (S) into a speech period and a noise period frame by frame;(b) computing a new autocorrelation matrix (Ra) based on a combination of an autocorrelation matrix (R) of a current noise period frame and an autocorrelation matrix of a previous noise period frame;(c) performing LPC analysis with said new autocorrelation matrix (Ra);(d) determining a synthesis filter coefficient (aa) based on a result (a) of the LPC analysis, quantizing said synthesis filter coefficient (aa), and sending a resulting quantized synthesis filter coefficient;and (e) searching for an optimal codebook vector based on said quantized synthesis filter coefficient.
  2. 2
    A method in accordance with claim 1, wherein step (d) comprises:(f) transforming a synthesis filter coefficient of a noise period to an LSP coefficient (l);(g) determining a spectrum characteristic of a synthesis filter, and comparing said spectrum characteristic with a past spectrum characteristic of said synthesis filter occurred in a past noise period to thereby produce a new LSP coefficient (la) having reduced spectrum fluctuation;and (h) transforming said new LSP coefficient to said synthesis filter coefficient (aa).
  3. 3
    A method in accordance with claim 1, wherein step (d) comprises (i) interpolating the synthesis filter coefficient of a noise period with the synthesis filter coefficient of a past noise period to thereby directly compute said new synthesis filter coefficient (aa) of the current noise period.
  4. 4
    A method of CELP coding an input audio signal (S), comprising the steps of:(a) determining whether the input audio signal (S) is a speech or noise subframe by subframe;(b) computing an autocorrelation matrix (R) of a noise period;(c) performing LPC analysis with said autocorrelation matrix (R);(d) determining a synthesis filter coefficient (aa) based on a result (a) of the LPC analysis, quantizing said synthesis filter coefficient (aa), and sending a resulting quantized synthesis filter coefficient (aa);(e) selecting an amount of noise reduction and a noise reducing method on the basis of a speech/noise decision performed in step (a);(f) computing a target signal vector (t) with the noise reducing method selected;and (g) searching for an optimal codebook vector by using said target signal vector (t).
  5. 5
    An apparatus for CELP coding an input audio signal, including autocorrelation analyzing means (102) for producing autocorrelation information (R) from the input audio signal (S), and vocal tract prediction coefficient analyzing means (103) for computing a vocal tract prediction coefficient (a) from a result of analysis (R) output from said autocorrelation analyzing means (102), CHARACTERIZED BY comprising:prediction gain coefficient analyzing means (112) for computing a prediction gain coefficient (pg) from said vocal tract prediction coefficient (a);autocorrelation adjusting means (110, 111) for detecting a non-speech signal period on the basis of the input audio signal (S), said vocal tract prediction coefficient (a) and said prediction gain coefficient (pg), and adjusting said autocorrelation information (R) in the non-speech signal period;vocal tract prediction coefficient correcting means (103) for producing from adjusted autocorrelation information (Ra) a corrected vocal tract prediction coefficient (aa) having said vocal tract prediction coefficient (a) of the non-speed signal period corrected;and coding means (104-109, 113-117, 130) for CELP coding the input audio signal (S) by using said corrected vocal tract prediction coefficient and an adaptive excitation signal (ex).
  6. 6
    An apparatus in accordance with claim 5, CHARACTERIZED IN THAT said vocal tract prediction coefficient analyzing means (103) and said vocal tract prediction coefficient correcting means (103) perform LPC analysis with said autocorrelation information (R, Ra) to thereby output said vocal tract prediction coefficient (a, aa).
  7. 7
    An apparatus in accordance with claim 5, CHARACTERIZED IN THAT said coding means (104-109, 113-117, 130) includes an IIR degital filter (104) for filtering said adaptive excitation signal (ex) by using said corrected vocal tract prediction coefficient (aa) as a filter coefficient.
  8. 8
    An apparatus for CELP coding an input audio signal, including autocorrelation analyzing means (102) for producing autocorrelation information (R) from the input audio signal (S), vocal tract prediction coefficient analyzing means (103A) for computing a vocal tract prediction coefficient (a) from a result of analysis (R) output from said autocorrelation analyzing means (102), CHARACTERIZED BY comprising:prediction gain coefficient analyzing means (112) for computing a prediction gain coefficient (pg) from said vocal tract prediction coefficient (a);LSP coefficient adjusting means (119, 110, 121) for computing an LSP coefficient (l) from said vocal tract prediction coefficient (a), detecting a non-speech signal period of the input audio signal (S) from the input audio signal (S), said vocal tract prediction coefficient (a) and said prediction gain coefficient (pg), and adjusting said LSP coefficient (l) of the non-speech signal period;vocal tract prediction coefficient correcting means (120) for producing from adjusted LSP coefficient (la) a corrected vocal tract prediction coefficient (aa) having said vocal tract prediction coefficient (a) of the non-speech signal period corrected;and coding means for CELP coding the input audio signal (S) by using said corrected vocal tract coefficient (aa) and an adaptive excitation signal (ex).
  9. 9
    An apparatus in accordance with claim 8, CHARACTERIZED IN THAT said vocal tract prediction coefficient analyzing means (103A) performs LPC analysis with said autocorrelation information (R) to thereby output said vocal tract prediction coefficient (a).
  10. 10
    An apparatus in accordance with claim 8, CHARACTERIZED IN THAT said coding means (104-109, 113-117, 130) includes an IIR digital filter (104) for filtering said adaptive excitation signal (ex) by using said corrected vocal tract prediction coefficient (aa) as a filter coefficient.
  11. 11
    An apparatus for CELP coding an input audio signal, including autocorrelation analyzing means (102) for producing autocorrelation information (R) from the input audio signal (S), and vocal tract prediction coefficient analyzing means (103A) for computing a vocal tract prediction coefficient (a) from a result of analysis (R) output from said autocorrelation analyzing means (102), CHARACTERIZED BY comprising:prediction gain coefficient analyzing means (112) for computing a prediction gain coefficient (pg) from said vocal tract prediction coefficient (a);vocal tract coefficient adjusting means for detecting a non-speech signal period on the basis of the input audio signal (S), said vocal tract prediction coefficient (a) and said prediction gain coefficient (pg), and adjusting said vocal tract prediction coefficient (a) to thereby output an adjusted vocal tract prediction coefficient (aa);coding means for CELP coding the input audio signal (S) by using said adjusted vocal tract prediction coefficient and an adaptive excitation signal (ex).
  12. 12
    An apparatus in accordance with claim 11, CHARACTERIZED IN THAT said vocal tract prediction coefficient analyzing means (103A) performs LPC analysis with said autocorrelation information (R) to thereby output said vocal tract prediction coefficient(a).
  13. 13
    An apparatus in accordance with claim 11, CHARACTERIZED IN THAT said coding means (104-109, 113-117, 130) includes an IIR digital filter (104) for filtering said adaptive excitation signal (ex) by using said corrected vocal tract prediction coefficient (aa) as a filter coefficient.
  14. 14
    An apparatus for CELP coding an input audio signal, including autocorrelation analyzing means (102) for producing autocorrelation information (R) from the input audio signal (S), and vocal tract prediction coefficient analyzing means (103A) for computing a vocal tract prediction coefficient (a) from a result of analysis (R) output from said autocorrelation analyzing mans (102), CHARACTERIZED BY comprising:prediction gain coefficient analyzing means (112) for computing a prediction gain coefficient (pg) from said vocal tract prediction coefficient (a);noise cancelling means (124, 110B, 125, 122) for detecting a non-speech signal period on the basis of bandpass signals (Sbpl-SbpN) produced by bandpass filtering the input audio signal (S) and said prediction gain coefficient (pg), performing signal analysis on the non-speech signal period to thereby generate a filter coefficient (nc) for noise cancellation, and performing noise cancellation with the input audio signal (S) by using said filter coefficient (nc) to thereby generate a target signal (t) for the generation of a synthetic speech signal (Sw);synthetic speech generating means (104) for generating said synthetic speech signal (Sw) by using said vocal tract prediction coefficient (a);and coding means (104-109, 113-117, 130) for CELP coding the input audio signal by using said vocal tract prediction coefficient (a) and said target signal (t).
  15. 15
    An apparatus in accordance with claim 14, CHARACTERIZED IN THAT said vocal tract prediction coefficient analyzing means (103A) performs LPC analysis with said autocorrelation information (R) to thereby output said vocal tract prediction coefficient (a).
  16. 16
    An apparatus in accordance with claim 14, CHARACTERIZED IN THAT said coding means (104-109, 113-117, 130) includes an IIR digital filter (104) for filtering said adaptive excitation signal (ex) by using said corrected vocal tract prediction coefficient (aa) as a filter coefficient.
  17. 17
    An apparatus in accordance with claim 14, CHARACTERIZED IN THAT said noise cancelling means (124, 110B, 125, 122) includes a plurality of bandpass filters (124) each having a particular passband for filtering the input audio signal (S).
  18. 18
    An apparatus in accordance with claim 17, CHARACTERIZED IN THAT said noise cancelling mean (124, 110B, 125, 122) includes an IIR filter (122) for cancelling noise of the input audio signal (S) in accordance with said filter coefficient (nc) to thereby generate said target signal (t).
Independent claims18