CA2249792C

Audio signal compression method, audio signal compression apparatus, speech signal compression method, speech signal compression apparatus, speech recognition method, and speech recognition apparatus

Abstract

An audio signal compression apparatus for compressively coding an input audio signal comprises a time-to-frequency transformation unit for transforming the input audio signal to a frequency domain signal; a spectrum envelope calculation unit for calculating a spectrum envelope having different resolutions for different frequencies, from the input audio signal, using a weighting function on frequency based on human auditory characteristics; a normalization unit for normalizing the frequency domain signal using the spectrum envelope to obtain a residual signal; a power normalization unit for normalizing the residual signal by the power; an auditory weighting calculation unit for calculating weighting coefficients on frequency, based on the spectrum of the input audio signal and human auditory characteristics; and a multi-stage quantization means having plural stages of vector quantizers connected in series, to which the normalized residual signal is input, and at least one of the vector quantizers quantizing the residual signal using the weighting coefficients. Therefore, a low frequency band, which is auditively important, can be analyzed with a higher frequency resolution as compared with a high frequency band, whereby efficient signal compression utilizing human auditory characteristics is realized.

CA2249792C, drawing sheet 1
Sheet 1 of 13

Term

Term ended

Expired 2 October 2018, 8 years ago.

  1. Priority
  2. Filed
  3. Granted
  4. Expired
  5. Today

20 claims: 10 independent, 10 dependent

  1. 1
    CA 02249792 2008-04-30 The embodiments of the invention in which an exclusive property or privilege is claimed are defined as follows:1. An audio signal compression method for coding an inputted audio signal, and compressing an amount of data thereof, including the steps of: obtaining an autocorrelation function on a melfrequency axis by using the inputted audio signal, and an audio signal that is obtained by subjecting the inputted audio signal to warping on a frequency axis corresponding to human auditory characteristics;obtaining mel-linear predictive coefficients from the autocorrelation function on the mel-frequency axis;setting the mel-linear predictive coefficients themselves as a spectrum envelope, or obtaining a spectrum envelope from the mel-linear predictive coefficients;and flattening the inputted audio signal for each frame using the spectrum envelope.
  2. 2
    An audio signal compression method for coding an inputted audio signal, and compressing an amount of data thereof, including the steps of:cutting out a predetermined time length of audio signal from the inputted audio signal, and filtering the predetermined time length of audio signal through multiple stages of all-pass filters to obtain filter output signals from the respective filter stages;obtaining an autocorrelation function on a melfrequency axis which is subjected to warping on a frequency axis corresponding to human auditory characteristics, by performing a product-sum operation (formula (1)) between the inputted audio signal and the CA 02249792 2008-04-30 filter output signal outputted from each filter stage, which product-sum operation is performed by a finite number of times, wherein φ (i, j) is the autocorrelation function, x[n] is the inputs signal, and y(±-j) [n] is the filter output signal from each filter stage;obtaining mel-linear predictive coefficients from the autocorrelation function on the mel-frequency axis;setting the mel-linear predictive coefficients themselves as a spectrum envelope, or obtaining a spectrum envelope from the mel-linear predictive coefficients;and flattening the inputted audio signal for each frame using the spectrum envelope.
  3. 3
    An audio signal compression method as defined in Claim 2 wherein said all-pass filters are first order all-pass filters .
  4. 4
    An audio signal compression method as defined in Claim 2 or 3 wherein filter coefficients of the all-pass filters are subjected to weighting on a frequency corresponding to human auditory characteristics, using a bark scale or a mel scale. 5.. An audio signal compression apparatus for coding an inputted audio signal, and compressing an amount of data thereof, comprising:a time-to-frequency transformation means for transforming the inputted audio signal to a frequency CA 02249792 2008-04-30 domain signal, and outputting the frequency domain signal;a spectrum envelope calculation means for obtaining an autocorrelation function on a mel-frequency axis by using the inputted audio signal, and an audio signal that is obtained by subjecting the inputted audio signal to warping on a frequency axis corresponding to human auditory characteristics, and setting, as a spectrum envelope, mel-linear predictive coefficients obtained from the autocorrelation function on the mel-frequency axis, or obtaining a spectrum envelope from the mellinear predictive coefficients;a normalization means for normalizing the frequency domain signal using the spectrum envelope to obtain a residual signal;a power normalization means for normalizing the residual signal on the basis of a maximum value or an average value of power to obtain a normalized residual signal;a vector quantization means for vector-quantizing the normalized residual signal according to a residual code book to transform the residual signal into residual codes ;an auditory weighting calculation means for subjecting said spectrum envelope to weighting on a frequency corresponding to human auditory characteristics to output auditory weighting coefficients;and said vector quantization means performing quantization of the normalized residual signal by using the auditory weighting coefficients.
  5. 5
    6. An audio signal compression apparatus as defined in Claim 5 wherein:CA 02249792 2008-04-30 said vector quantization means is a multiple quantization means comprising plural vector quantization means connected to plural vertical lines;and at least one of the vector quantization means constituting the multiple quantization means performs quantization of the residual signal using the weighting coefficients.
  6. 6
    7. An audio signal compression apparatus as defined in Claim 5 or 6 wherein:said spectrum envelope calculation means cuts out a predetermined time length of audio signal from the inputted audio signal, and filters the predetermined time length of audio signal through multiple stages of allpass filters to obtain filter output signals from the respective filter stages;obtains an autocorrelation function on a melfrequency axis which is subjected to warping on a frequency axis corresponding to human auditory characteristics, by performing a product-sum operation (formula (2)) between the inputted audio signal and the filter output signal outputted from each filter stage, which product-sum operation is performed by a finite number of times, N-1 -..(2) wherein φ (i, j) is the autocorrelation function, x[n] is the input signal, and y(i-j) [n] is the filter output signal from each filter stage;obtains mel-linear predictive coefficients from the autocorrelation function on the mel-frequency axis;and CA 02249792 2008-04-30 sets the mel-linear predictive coefficients themselves as a spectrum envelope, or obtains a spectrum envelope from the mel-linear predictive coefficients.
  7. 7
    8. An audio signal compression apparatus as defined in Claim 7 wherein said all-pass filters are first order all-pass filters .
  8. 8
    9. An audio signal compression apparatus as defined in Claim 7 or 8 wherein filter coefficients of the all-pass filters are subjected to weighting on a frequency corresponding to human auditory characteristics, using a bark scale or a mel scale.
  9. 9
    10. A speech signal compression method for coding an inputted speech signal, and compressing an amount of data thereof, including the steps of:obtaining an autocorrelation function on a melfrequency axis by using the inputted speech signal, and a speech signal that is obtained by subjecting the inputted speech signal to warping on a frequency axis corresponding to human auditory characteristics;obtaining mel-linear predictive coefficients from the autocorrelation function on the mel-frequency axis;setting the mel-linear predictive coefficients themselves as a spectrum envelope, or obtaining a spectrum envelope from the mel-linear predictive coefficients;and flattening the inputted speech signal using the spectrum envelope. CA 02249792 2008-04-30
  10. 10
    11. A speech signal compression method for coding an inputted speech signal, and compressing an amount of data thereof, including the steps of:cutting out a predetermined time length of speech signal from the inputted speech signal, and filtering the predetermined time length of speech signal through multiple stages of all-pass filters to obtain filter output signals from the respective filter stages;obtaining an autocorrelation function on a melfrequency axis which is subjected to warping on a frequency axis corresponding to human auditory characteristics, by performing a product-sum operation (formula (3)) between the inputted speech signal and the filter output signal outputted from each filter stage, which product-sum operation is performed by a finite number of times, Ύ-1 ^>7')=^Ψ]'Ζ-λΗ ...(3) wherein φ (i, j) is the autocorrelation function, x[n] is the input signal, and y(i-j> [n] is the filter output signal from each filter stage;obtaining mel-linear predictive coefficients from the autocorrelation function on the mel-frequency axis;setting the mel-linear predictive coefficients themselves as a spectrum envelope, or obtaining a spectrum envelope from the mel-linear predictive coefficients;and flattening the inputted speech signal using the spectrum envelope.
  11. 11
    12. A speech signal compression method as defined in Claim 11 wherein said all-pass filters are first order allpass filters. CA 02249792 2008-04-30
  12. 12
    13. A speech signal compression method as defined in Claim 11 or 12 wherein filter coefficients of the all-pass filters are subjected to weighting on a frequency corresponding to human auditory characteristics, using a bark scale or a mel scale.
  13. 13
    14. A speech signal compression apparatus for coding an inputted speech signal, and compressing an amount of data thereof, comprising:a time-to-frequency transformation means for transforming the inputted speech signal to a frequency domain signal, and outputting the frequency domain signal;a parameter transformation means for obtaining an autocorrelation function on a mel-frequency axis by using the inputted speech signal, and a speech signal that is obtained by subjecting the inputted speech signal to warping on a frequency axis corresponding to human auditory characteristics, and transforming mel-linear predictive coefficients obtained from the autocorrelation function on the mel-frequency axis to parameters representing a spectrum envelope;an envelope normalization means for reversely filtering the inputted speech signal with the parameters to normalize the speech signal, thereby obtaining a residual signal;a power normalization means for normalizing the residual signal on the basis of a maximum value or an average value of power to obtain a normalized residual signal;a vector quantization means for vector-quantizing the normalized residual signal according to a residual CA 02249792 2008-04-30 code book to transform the residual signal to residual codes;an auditory weighting calculation means for subjecting said spectrum envelope to weighting on a frequency corresponding to human auditory characteristics to output auditory weighting coefficients;and said vector quantization means performing quantization of the normalized residual signal by using the auditory weighting coefficients.
  14. 14
    15. A speech signal compression apparatus as defined in Claim 14 wherein:said parameter calculation means cuts out a predetermined time length of speech signal from the inputted speech signal, and filters the predetermined time length of speech signal through multiple stages of all-pass filters to obtain filter output signals from the respective filter stages;obtains an autocorrelation function on a melfrequency axis which is subjected to warping on a frequency axis corresponding to human auditory characteristics, by performing a product-sum operation (formula (4)) between the inputted speech signal and the filter output signal outputted from each filter stage, which product-sum operation is performed by a finite number of times, N-l /) = V Ψ]· /(,-,)H ---(4) wherein φ (i, j) is the autocorrelation function, x[n] is the input signal, and y(±-j) [n] is the filter output signal from each filter stage;obtains mel-linear predictive coefficients from the autocorrelation function on the mel-frequency axis;and CA 02249792 2008-04-30 transforms the mel-linear predictive coefficients to parameters representing a spectrum envelope.
  15. 15
    16. A speech signal compression apparatus as defined in Claim 15 wherein said all-pass filters are first order allpass filters.
  16. 16
    17. A speech signal compression apparatus as defined in Claim 15 or 16 wherein filter coefficients of the all-pass filters are subjected to weighting on a frequency corresponding to human auditory characteristics, using a bark scale or a mel scale.
  17. 17
    18. A speech recognition method for recognizing speech from an inputted speech signal, comprising:obtaining an autocorrelation function on a melfrequency axis by using the inputted speech signal, and a speech signal that is obtained by subjecting the inputted speech signal to warping on a frequency axis corresponding to human auditory characteristics;obtaining mel-linear predictive coefficients from the autocorrelation function on the mel-frequency axis;and obtaining parameters representing a spectrum envelope from the mel-linear predictive coefficients.
  18. 18
    19. A speech recognition method for recognizing speech from an inputted speech signal, comprising:cutting out a predetermined time length of speech signal from the inputted speech signal, and filtering the predetermined time length of speech signal through multiple stages of all-pass filters to obtain filter output signals from the respective filter stages;CA 02249792 2008-04-30 obtaining an autocorrelation function on a melfrequency axis which is subjected to warping on a frequency axis corresponding to human auditory characteristics, by performing a product-sum operation (formula (5)) between the inputted speech signal and the filter output signal outputted from each filter stage, which product-sum operation is performed by a finite number of times, JV-J wherein φ (i, j) is the autocorrelation function, x[n] is the input signal, and y(±-j) [n] is the filter output signal from each filter stage;obtaining mel-linear predictive coefficients from the autocorrelation function on the mel-frequency axis;and obtaining parameters representing a spectrum envelope from the mel-linear predictive coefficients.
  19. 19
    20. A speech recognition method as defined in Claim 19 wherein said all-pass filters are first order all-pass filters.
  20. 20
    21. A speech recognition method as defined in Claim 19 or 20 wherein filter coefficients of the all-pass filters are subjected to weighting on a frequency corresponding to human auditory characteristics, using a bark scale or a mel scale.
Independent claims20