CA2898677A1

Low-frequency emphasis for lpc-based coding in frequency domain

Abstract

The invention provides an audio encoder and method for encoding a non-speech audio signal so as to produce therefrom a bitstream, the audio encoder comprising: a combination (2, 3) of a linear predictive coding filter (2) having a plurality of linear predictive coding coefficients (LC) and a time-frequency converter (3), wherein the combination (2, 3) is configured to filter and to convert a frame (Fl) of the audio signal (AS) into a frequency domain in order to output a spectrum (SP) based on the frame (Fl) and on the linear predictive coding coefficients (LC); a low frequency emphasizer (4) configured to calculate a processed spectrum (PS) based on the spectrum (SP), wherein spectral lines (SL) of the processed spectrum (PS) representing a lower frequency than a reference spectral line (RSL) are emphasized; and a control device (5) configured to control the calculation of the processed spectrum (PS) by the low frequency emphasizer (4) depending on the linear predictive coding coefficients (LC) of the linear predictive coding filter (2). Furthermore, the invention provides a corresponding audio decoder, a system, a method for decoding a bitstream containing quantized spectrums and a plurality of linear predictive coding coefficients and a corresponding computer program.

CA2898677A1, drawing sheet 1
Sheet 1 of 9

Term

7.3 yearsto projected expiry

Projected expiry 28 January 2034, counted from filing; an application has no term until it is granted.

  1. Priority
  2. Filed
  3. Published
  4. Today
  5. Projected expiry

2 claims: 2 independent, 0 dependent

  1. 1
    CA 02898677 2015-07-20 Printed:15/12/2014 CLMSPAMD EP2014051585 FH140112PCT 11.11.2014 marked-up version Claims 1. Audio encoder for encoding a non-speech audio signal (AS) so as to produce therefrom a bitstream (BS), the audio encoder (1) comprising: a combination (2, 3) of a linear predictive coding filter (2) having a plurality of linear predictive coding coefficients (LC) and a time-frequency converter (3), wherein the combination (2, 3) is configured to filter and to convert a frame (FI) of the audio signal (AS) into a frequency domain in order to io output a spectrum (SP) based on the frame (FI) and on the linear predictive coding coefficients (LC);a low frequency emphasizer (4) configured to calculate a processed spectrum (PS) based on the spectrum (SP), wherein spectral lines (SL) of the 15 processed spectrum (PS) representing a lower frequency than a refer| ence spectral line (RSL) are emphasized: and a control device (5) configured to control the calculation of the processed spectrum (PS) by the low frequency emphasizer (4) depending on the lin20 ear predictive coding coefficients (LC) of the linear predictive coding filter (2);a quantization device (6) configured to produce a quantized spectrum (QS) based on the processed spectrum (PS): and a bitstream producer (7) configured to embed the quantized spectrum (QS) and the linear predictive coding coefficients (LC) into the bitstream (BS). 3o 2. Audio encoder according to the preceding claim, wherein the frame (FI) of the audio signal (AS) is input to the linear predictive coding filter (2), wherein a filtered frame (FF) is output by the linear predictive coding filter 28/11/2014 CA 02898677 2015-07-20 CLMSPAMD Printed: 15/12/2014 EP2014051585 FH140112PCT 11.11.2014 (2) and wherein the time-frequency converter (3) is configured to estimate the spectrum (SP) based on the filtered frame (FF), 3. Audio encoder according to claim 1, wherein the frame (FI) of the audio signal (AS) is input to the time-frequency converter (3), wherein a converted frame (FC) is output by the time-frequency converter (3) and wherein the linear predictive coding filter (2) is configured to estimate the spectrum (SP) based on the converted frame (FC). 4- Audio-encoder according to one of the preceding claims;wherein the audio encoder-(-1)-Gomprises a quantization-device (6) configured to produce a quantized spectrum (QS) based on the processed speGtrum (PS) and a 5r4._Audio encoder according to one of the preceding claims, wherein the control device (5) comprises a spectral analyzer (8) configured to estimate a spectral representation (SR) of the linear predictive coding coefficients (LC), a minimum-maximum analyzer (9) configured to estimate a minimum (Ml) of the spectral representation (SR) and a maximum (MA) of the spectral representation (SR) below a further reference spectral line and an emphasis factor calculator (10, 11) configured to calculate spectral line emphasis factors (SEF) for calculating the spectral lines (SL) of the processed spectrum (PS) representing a lower frequency than the reference spectral line (RSL) based on the minimum (Ml) and on the maximum (MA), wherein the spectral lines (SL) of the processed spectrum (PS) are emphasized by applying the spectral line emphasis factors (SEF) to spectral lines of the spectrum of the filtered frame. 30 | fe-5. Audio encoder according to the preceding claim, wherein the emphasis factor calculator (10, 11) is configured in such way that the spectral line emphasis factors (SEF) increase in a direction from the reference 28/11/2014 CA 02898677 2015-07-20 CLMSPAMD Printed: 15/12/2014 EP2014051585 FH140112PCT 11.11,2014 spectral line (RSL) to the spectral line (SL) representing the lowest frequency of the spectrum (SP). | +6. Audio encoder according to claim 54 or 55, wherein the emphasis fac5 tor calculator (10,11) comprises a first stage (10) configured to calculate a basis emphasis factor (BEF) according to a first formula γ = (a min I max) p , wherein a is a first preset value, with a > 1, β is a second preset value, with 0 < β s 1, min is the minimum (Ml) of the of the spectral representation (SR), max is the maximum (MA) of the spectral representation io (SR) and γ is the basis emphasis factor (BEF), and wherein the emphasis factor calculator (10,11) comprises a second stage (11) configured to calculate spectral line emphasis factors (SEF) according to a second formula Ci = γ'wherein i’ is a number of the spectral lines (SL) to be emphasized, i is an index of the respective spectral line (SL), the index in15 creases with the frequencies of the spectral lines, with i = 0 to i’-1, y is the basis emphasis factor (BEF) and q is the spectral line emphasis factor (SEF) with index i. j _Audio encoder according to the preceding claim, wherein the first pre20 set value is smaller than 42 and larger than 22, in particular smaller than 38 and larger than 26, more particular smaller 34 and larger than 30. j θ-8._Audio encoder according to claim 76 or 87, wherein the second preset value is determined according to the formula β = 1 / (θ · i'), wherein i’ is 25 the number of the spectral lines being emphasized, Θ is a factor between 3 and 5, in particular between 3,4 and 4,6, more particular between 3,8 and 4,2. J 4£r9. Audio encoder according to one of the preceding claims, wherein the 30 reference spectral line (RSL) represents a frequency between 600 Hz and 1000Hz, in particular between 700 Hz and 900 Hz, more particular between 750 Hz and 850 Hz. 28/11/2014 CA 02898677 2015-07-20 CLMSPAMD Printed: 15/12/2014 EP2014051585 FH140112PCT 11.11.2014 j 44-10. Audio encoder according to one of the claims 54 to 409, wherein the further reference spectral line represents the same or a higher frequency than the reference spectral line (RSL). j 42t1 1. Audio encoder according to one of the preceding claims, wherein the control device (5) is configured in such way that the spectral lines (SL) of the processed spectrum (PS) representing a lower frequency than the reference spectral line (RSL) are emphasized only if the maximum (MA) is io less than the minimum (Ml) multiplied with the first preset value. J 4-3-12. Audio decoder for decoding a bitstream (BS) based on a non-speech audio signal (AS) so as to produce from the bitstream (BS) a non-speech audio output signal (OS), in particular for decoding a bitstream (BS) pro15 duced by an audio encoder (1) according to claims 1 to 12, the bitstream (BS) containing quantized spectrums (QS) and a plurality of linear predictive coding coefficients (LC), the audio decoder (12) comprising: a bitstream receiver (13) configured to extract the quantized spectrum 20 (QS) and the linear predictive coding coefficients (LC) from the bitstream (BS);a de-quantization device (14) configured to produce a de-quantized spectrum (DQ) based on the quantized spectrum (QS);a low frequency de-emphasizer (15) configured to calculate a reverse processed spectrum (RS) based on the de-quantized spectrum (DQ), wherein spectral lines (SLD) of the reverse processed spectrum (RS) representing a lower frequency than a reference spectral line (RSLD) are 30 deemphasized;and a control device (16) configured to control the calculation of the reverse 28/11/2014 CA 02898677 2015-07-20 CLMSPAMD Printed: 15/12/2014 EP2014051585 FH140112PCT 11.11.2014 processed spectrum (RS) by the low frequency de-emphasizer (15) depending on the linear predictive coding coefficients (LC) contained in the bitstream (BS). 5 J 4443. Audio decoder according to the preceding claim, wherein the audio decoder (12) comprises combination (17, 18) of a frequency-time converter (17) and an inverse linear predictive coding filter (18) receiving the plurality of linear predictive coding coefficients (LC) contained in the bitstream (BS), wherein the combination (17, 18) is configured to inverse10 filter and to convert the reverse processed spectrum (RS) into a time domain in order to output the output signal (OS) based on the reverse processed spectrum (RS) and on the linear predictive coding coefficients (LC). 15 J 45r14. Audio decoder according to the preceding claim, wherein the frequency-time converter (17) is configured to estimate a time signal (TS) based on the reverse processed spectrum (RS) and wherein the inverse linear predictive coding filter (18) is configured to output the output signal (OS) based on the time signal (TS). J 4645, Audio decoder according to claim 4413, wherein the inverse linear predictive coding filter (18) is configured to estimate an inverse filtered signal (IFS) based on the reverse processed spectrum (RS) and wherein the frequency-time converter (17) is configured to output the output signal 25 (OS) based on the inverse filtered signal (IFS). j 47416. Audio decoder according to one of the claims 43-12 to 4315. wherein the control device (16) comprises a spectral analyzer (19) configured to estimate a spectral representation (SR) of the linear predictive coding co30 efficients (LC), a minimum-maximum analyzer (20) configured to estimate a minimum (Ml) of the spectral representation (SR) and a maximum (MA) of the spectral representation (SR) below a further reference spectral line 28/11/2014 CA 02898677 2015-07-20 CLMSPAMD Printed: 15/12/2014 EP2014051585 FH140112PCT 11.11.2014 and a de-emphasis factor calculator (21, 22) configured to calculate spectral line de-emphasis factors (SDF) for calculating the spectral lines (SLD) of the reverse processed spectrum (RS) representing a lower frequency than the reference spectral line (RSLD) based on the minimum (Ml) and 5 on the maximum (MA), wherein the spectral lines (SLD) of the reverse processed spectrum (RS) are de-emphasized by applying the spectral line de-emphasis factors (SDF) to spectral lines of the spectrum of the dequantized spectrum (DQ), io j 44^17. Audio decoder according to the preceding claim, wherein the deemphasis factor calculator (21, 22) is configured in such way that the spectral line de-emphasis factors (SDF) decrease in a direction from the reference spectral line (RSLD) to the spectral line (SL) representing the lowest frequency of the reverse process spectrum (RS). J 48r18. Audio decoder according to claim 47-16 or 48Î7, wherein the deemphasis factor calculator (21, 22) comprises a first stage (21) configured to calculate a basis de-emphasis factor (BDF) according to a first formula δ = (a · min / max)' p , wherein a is a first preset value, with a > 1, β is a 20 second preset value, with 0 < β s 1, min is the minimum (Ml) of the of the spectral representation (SR), max is the maximum (MA) of the spectral representation (SR) and δ is the basis de-emphasis factor (BDF), and wherein the de-emphasis factor calculator (21,22) comprises a second stage (22) configured to calculate spectral line de-emphasis factors (SDF) 25 according to a second formula ζ, = δ* ' 1 , wherein i’ is a number of the spectral lines (SLD) to be de-emphasized, i is an index of the respective spectral fine (SLD), the index increases with the frequencies of the spectral lines, with i = 0 to i’-1, δ is the basis de-emphasis factor (BDF) and ζ is the spectral line de-emphasis factor (SDF) with index i. 28/11/2014 CA 02898677 2015-07-20 CLMSPAMD Printed: 15/12/2014 EP2014051585 FH140112PCT 11.11.2014 | 26:19, Audio decoder according to the preceding claim, wherein the first preset value is smaller than 42 and larger than 22, in particular smaller than 38 and larger than 26, more particular smaller 34 and larger than 30. 5 | 2-420. Audio decoder according to claim 4-9-18 or 2019, wherein the second preset value is determined according to the formula β = 1 / (Θ i’), wherein i’ is the number of the spectral lines (SLD) being de-emphasized, Θ is a factor between 3 and 5, in particular between 3,4 and 4,6, more particular between 3,8 and 4,2. J 22:21. Audio decoder according to one of the claims 43-12 to 2420, wherein the reference spectral line (RSLD) represents a frequency between 600 Hz and 1000Hz, in particular between 700 Hz and 900 Hz, more particular between 750 Hz and 850 Hz. J 23:22. Audio decoder according to one of the claims 47-16 to 2221, wherein the further reference spectral line represents the same or a higher frequency than the reference spectral line (RSLD). 20 J 24:23. Audio decoder according to one of the claims 43-12 to 2322, wherein the control device (16) is configured in such way that the spectral lines (SLD) of the reverse processed spectrum (RS) representing a lower frequency than the reference spectral line (RSLD) are de-emphasized only if the maximum (MA) is less than the minimum (Ml) multiplied with the first 25 preset value. 24424. A system comprising a decoder (1) and an encoder (12), wherein the encoder (1) is designed according to one of the claims 1 to 42-11 and/or the decoder is designed according to one of the claims 43-12 to 2423. J 26:25. Method for encoding a non-speech audio signal (AS) so as to produce therefrom a bitstream (BS), the method comprising the steps: 28/11/2014 CA 02898677 2015-07-20 CLMSPAMD Printed: 15/12/2014 EP2014051585 FH140112PCT 11.11.2014 filtering with a linear predictive coding filter (2) having a plurality of linear predictive coding coefficients (LC) and converting a frame (FI) ofthe audio signal (AS) into a frequency domain in order to output a spectrum 5 (SP) based on the frame (FI) and on the linear predictive coding coefficients (LC);calculating a processed spectrum (PS) based on the spectrum (SP), wherein spectral lines (SL) of the processed spectrum (PS) representing 10 a lower frequency than a reference spectral line (RSL) are emphasized;and controlling the calculation ofthe processed spectrum (PS) depending on the linear predictive coding coefficients (LC) of the linear predictive coding 15 filter (2),. producing a quantized spectrum (QS) based on the processed spectrum (PS);and 20 embedding the quantized spectrum (QS) and the linear predictive coding coefficients (LC) into the bitstream (BS). | 2/726. Method for decoding a bitstream (BS) based on a non-speech audio signal (AS) so as to produce from the bitstream (BS) a non-speech audio 25 output signal (OS), in particular for decoding a bitstream (BS) produced by the method according to the preceding claim, the bitstream (BS) containing quantized spectrums (QS) and a plurality of linear predictive coding coefficients (LC), the method comprising the steps: 30 extracting the quantized spectrum (QS) and the linear predictive coding coefficients (LC) from the bitstream (BS);28/11/2014 CA 02898677 2015-07-20 GLMSWAMU tK2U14Ui)10«b FH140112PCT 11.11.2014 κππϊβσ: io/iz/2ui4 producing a de-quantized spectrum (DQ) based on the quantized spectrum (QS);calculating a reverse processed spectrum (RS) based on the de5 quantized spectrum (DQ), wherein spectral tines (SLD) of the reverse processed spectrum (RS) representing a lower frequency than a reference spectral line (RSLD) are deemphasized;and controlling the calculation of the reverse processed spectrum (RS) deio pending on the linear predictive coding coefficients (LC) contained in the bitstream (BS).
  2. 2
    2£r27. Computer program for performing, when running on a computer or a processor, the method of claim 2§-25 or 2726. 28/11/2014