EP2015293A1

Method and apparatus for encoding and decoding an audio signal using adaptively switched temporal resolution in the spectral domain

Abstract

Perceptual audio codecs make use of filter banks and MDCT in order to achieve a compact representation of the audio signal, by removing redundancy and irrelevancy from the original audio signal. During quasi-stationary parts of the audio signal a high frequency resolution of the filter bank is advantageous in order to achieve a high coding gain, but this high frequency resolution is coupled to a coarse temporal resolution that becomes a problem during transient signal parts by producing audible pre-echo effects. The invention achieves improved coding/decoding quality by applying on top of the output of a first filter bank a second non-uniform filter bank, i.e. a cascaded MDCT. The inventive codec uses switching to an additional extension filter bank (or multi-resolution filter bank) in order to re-group the time-frequency representation during transient or fast changing audio signal sections. By applying a corresponding switching control, pre-echo effects are avoided and a high coding gain and a low coding delay are achieved.

EP2015293A1, drawing sheet 1
Sheet 1 of 14

Term

Projected expiry 14 June 2027.

  1. Priority and filed
  2. Published
  3. Today
  4. Projected expiry

12 claims: 5 independent, 7 dependent

  1. 1
    Method for encoding an input signal (CIS), e.g. an audio signal, using a first transform (MDCT-1) into the frequency domain being applied to first-length (N L ) sections of said input signal, and using adaptive switching of the temporal resolution, followed by quantisation and entropy encoding (QUCOD) of the values of the resulting frequency domain bins, wherein control (PSYM, FBCTL) of said switching, quantisation and/or entropy encoding is derived from a psycho-acoustic analysis of said input signal, characterised by the step of:- adaptively controlling (SW1, SW2, SWI) said temporal resolution by performing a second transform (MDCT-2) following said first transform (MDCT-1) and being applied to second-length (N short ) sections of said transformed first-length sections, wherein said second length is smaller than said first length (N L ) and either the output values of said first transform or the output values of said second transform are processed in said quantisation and entropy encoding (QUCOD);- attaching (STRPCK) to the encoding output signal (COS) corresponding temporal resolution control information (SWI) as side information.
  2. 2
    Apparatus for encoding an input signal (CIS), e.g. an audio signal, said apparatus including:- first transform means (MDCT-1) being adapted for transforming first-length (N L ) sections of said input signal into the frequency domain;- second transform means (MDCT-2) being adapted for transforming second-length (N short ) sections of said transformed first-length sections, wherein said second length is smaller than said first length (N L );- means (QUCOD) being adapted for quantising and entropy encoding the output values of said first transform means or the output values of said second transform means;- means (PSYM, FBCTL) being adapted for controlling said quantisation and/or entropy encoding and for controlling adaptively whether said output values of said first transform means or the output values of said second transform means are processed in said quantising and entropy encoding means, wherein said controlling is derived from a psycho-acoustic analysis of said input signal;- means (STRPCK) being adapted for attaching to the encoding apparatus output signal (COS) corresponding temporal resolution control information (SWI) as side information.
  3. 3
    Method for decoding an encoded signal (DIS), e.g. an audio signal, that was encoded using a first transform (MDCT-1) into the frequency domain being applied to first-length (N L ) sections of said input signal, wherein the temporal resolution was adaptively switched (SW1, SW2) by performing a second transform (MDCT-2) following said first transform (MDCT-1) and being applied to second-length (N short ) sections of said transformed first-length sections, wherein said second length is smaller than said first length (N L ) and either the output values of said first transform or the output values of said second transform were processed in a quantisation and entropy encoding (QUCOD), and wherein control (PSYM, FBCTL) of said switching, quantisation and/or entropy encoding was derived from a psycho-acoustic analysis of said input signal and corresponding temporal resolution control information (SWI) was attached (STRPCK) to the encoding output signal (COS) as side information, said decoding method including the steps of:- providing (DPCRQU) from said encoded signal (DIS) said side information (SWI);- inversely quantising and entropy decoding (DPCRQU) said encoded signal (DIS);- corresponding to said side information, either (SW3, SW4) performing a first inverse transform (iMDCT-1) into the time domain, said first inverse transform operating on first-length (N L ) signal sections of said inversely quantised and entropy decoded signal and said first inverse transform providing the decoded signal (DOS), or processing second-length (N short ) sections of said inversely quantised and entropy decoded signal in a second inverse transform (iMDCT-2) before performing said first inverse transform (iMDCT-1).
  4. 4
    Apparatus for decoding an encoded signal (DIS), e.g. an audio signal, that was encoded using a first transform (MDCT-1) into the frequency domain being applied to first-length (N L ) sections of said input signal, wherein the temporal resolution was adaptively switched (SW1, SW2) by performing a second transform (MDCT-2) following said first transform (MDCT-1) and being applied to second-length (N short ) sections of said transformed first-length sections, wherein said second length is smaller than said first length (N L ) and either the output values of said first transform or the output values of said second transform were processed in a quantisation and entropy encoding (QUCOD), and wherein control (PSYM, FBCTL) of said switching, quantisation and/or entropy encoding was derived from a psycho-acoustic analysis of said input signal and corresponding temporal resolution control information (SWI) was attached (STRPCK) to the encoding output signal (COS) as side information, said apparatus including:- means (DPCRQU) being adapted for providing from said encoded signal (DIS) said side information (SWI) and for inversely quantising and entropy decoding said encoded signal;- means (iMDCT-1, iMDCT-2, SW3, SW4) being adapted for, corresponding to said side information, either performing a first inverse transform into the time domain, said first inverse transform operating on first-length (N L ) signal sections of said inversely quantised and entropy decoded signal and said first inverse transform providing the decoded signal (DOS), or processing second-length (N short ) sections of said inversely quantised and entropy decoded signal in a second inverse transform before performing said first inverse transform.
  5. 8
    Method according to one of claims 1, 3 and 5 to 7, or apparatus according to one of claims 2 and 4 to 7, wherein in case more than one different second length is used successively, the lengths increase starting from frequency bins representing low frequency lines.
  6. 10
    Digital video signal that is encoded according to the method of one of claims 1 and 5 to 9.
  7. 12
    Use of the method according to one of claims 1 and 5 to 9 in a watermark embedder.