EP2947653A1

Multi-channel audio coding using complex prediction and window shape information

Abstract

An audio encoder and an audio decoder are based on a combination of two audio channels (201, 202) to obtain a first combination signal (204) as a mid signal and a residual signal (205) which can be derived using a predicted side signal derived from the mid signal. The first combination signal and the prediction residual signal are encoded (209) and written (212) into a data stream (213) together with the prediction information (206) derived by an optimizer (207) based on an optimization target (208). A decoder uses the prediction residual signal, the first combination signal and the prediction information to derive a decoded first channel signal and a decoded second channel signal. In an encoder example or in a decoder example, a real-to-imaginary transform can be applied for estimating the imaginary part of the spectrum of the first combination signal. For calculating the prediction signal used in the derivation of the prediction residual signal, the real-valued first combination signal is multiplied by a real portion of the complex prediction information and the estimated imaginary part of the first combination signal is multiplied by an imaginary portion of the complex prediction information.

EP2947653A1, drawing sheet 1
Sheet 1 of 18

Term

4.5 yearsto projected expiry

Projected expiry 23 March 2031, counted from filing; an application has no term until it is granted.

  1. Priority
  2. Filed
  3. Published
  4. Today
  5. Projected expiry

15 claims: 11 independent, 4 dependent

  1. 1
    Audio decoder for decoding an encoded multi-channel audio signal (100), the encoded multi-channel audio signal comprising an encoded first combination signal generated based on a combination rule for combining a first channel audio signal and a second channel audio signal of a multi-channel audio signal, an encoded prediction residual signal and prediction information, comprising:a signal decoder (110) for decoding the encoded first combination signal (104) to obtain a decoded first combination signal (112), and for decoding the encoded residual signal (106) to obtain a decoded residual signal (114);and a decoder calculator (116) for calculating a decoded multi-channel signal having a decoded first channel signal (117), and a decoded second channel signal (118) using the decoded residual signal (114), the prediction information (108) and the decoded first combination signal (112), so that the decoded first channel signal (117) and the decoded second channel signal (118) are at least approximations of the first channel signal and the second channel signal of the multi-channel signal, wherein the prediction information (108) comprises a real-valued portion different from zero and/or an imaginary portion different from zero, wherein the decoder calculator (116) comprises: a predictor (1160) for applying the prediction information (108) to the decoded first combination signal (112) or to a signal (601) derived from the decoded first combination signal to obtain a prediction signal (1163);a combination signal calculator (1161) for calculating a second combination signal (1165) by combining the decoded residual signal (114) and the prediction signal (1163);and a combiner (1162) for combining the decoded first combination signal (112) and the second combination signal (1165) to obtain a decoded multi-channel audio signal having the decoded first channel signal (117) and the decoded second channel signal (118), and wherein the predictor (1160) is configured for receiving window shape information (109) and for using different filter coefficients for calculating an imaginary spectrum, where the different filter coefficients depend on different window shapes indicated by the window shape information (109).
  2. 2
    Audio decoder in accordance with claim 1, in which the encoded first combination signal (104) and the encoded residual signal (106) have been generated using an aliasing generating time-spectral conversion, wherein the decoder further comprises:a spectral-time converter (52, 53) for generating a time-domain first channel signal and a time-domain second channel signal using a spectral-time conversion algorithm matched to the time-spectral conversion algorithm;an overlap/add processor (522) for conducting an overlap-add processing for the time-domain first channel signal and for the time-domain second channel signal to obtain an aliasing-free first time-domain signal and an aliasing-free second time-domain signal.
  3. 3
    Audio decoder in accordance with one of the preceding claims, in which the prediction information (108) comprises a real factor different from zero, in which the predictor (1160) is configured for multiplying the decoded first combination signal by the real factor to obtain a first part of the prediction signal, and in which the combination signal calculator is configured for linearly combining the decoded residual signal and the first part of the prediction signal.
  4. 4
    Audio decoder in accordance with one of the preceding claims, in which the prediction information (108) comprises an imaginary factor different from zero, and in which the predictor (1160) is configured for estimating (1160a) an imaginary part of the decoded first combination signal (112) using a real part of the decoded first combination signal (112), in which the predictor (1160) is configured for multiplying the imaginary part (601) of the decoded first combination signal by the imaginary factor of the prediction information (108) to obtain a second part of the prediction signal;and in which the combination signal calculator (1161) is configured for linearly combining the first part of the prediction signal and the second part of the prediction signal and the decoded residual signal to obtain a second combination signal (1165).
  5. 5
    Audio decoder in accordance with one of the preceding claims, in which the encoded or decoded first combination signal (104) and the encoded or decoded prediction residual signal (106) each comprises a first plurality of subband signals, wherein the prediction information comprises a second plurality of prediction information parameters, the second plurality being smaller than the first plurality, wherein the predictor (1160) is configured for applying the same prediction parameter to at least two different subband signals of the decoded first combination signal, wherein the decoder calculator (116) or the combination signal calculator (1161) or the combiner (1162) are configured for performing a subband-wise processing;and wherein the audio decoder further comprises a synthesis filterbank (52, 53) for combining subband signals of the decoded first combination signal and the decoded second combination signal to obtain a time-domain first decoded signal and a time-domain second decoded signal.
  6. 6
    Audio decoder in accordance with claim 1, in which the decoded first combination signal comprises a sequence of real-valued signal frames, and in which the predictor (1160) is configured for estimating (1160a) an imaginary part of the current signal frame using only the current real-valued signal frame or using the current real-valued signal frame and either only one or more preceding or only one or more following real-valued signal frames or using the current real-valued signal frame and one or more preceding real-valued signal frames and one or more following real-valued signal frames.
  7. 7
    Audio decoder in accordance with one of claims 1 to 6, in which the encoded multi-channel signal comprises, as side information, a real indicator indicating that all prediction coefficients for a frame of the encoded multi-channel signal are real valued, wherein the audio decoder is configured for extracting the real indicator from the encoded multi-channel audio signal (100), and wherein the decoder calculator (116) is configured for not calculating an imaginary signal for a frame, for which the real indicator is indicating only real-valued prediction coefficients.
  8. 8
    Audio encoder for encoding a multi-channel audio signal having two or more channel signals, comprising:an encoder calculator (203) for calculating a first combination signal (204) and a prediction residual signal (205) using a first channel signal (201) and a second channel signal (202) and prediction information (206), so that a prediction residual signal, when combined with a prediction signal derived from the first combination signal or a signal derived from the first combination signal and the prediction information (206) results in a second combination signal (2032), the first combination signal (204) and the second combination signal (2032) being derivable from the first channel signal (201) and the second channel signal (202) using a combination rule;an optimizer (207) for calculating the prediction information (206) so that the prediction residual signal (205) fulfills an optimization target (208);a signal encoder (209) for encoding the first combination signal (204) and the prediction residual signal (205) to obtain an encoded first combination signal (210) and an encoded residual signal (211);and an output interface (212) for combining the encoded first combination signal (210), the encoded prediction residual signal (211) and the prediction information (206) to obtain an encoded multi-channel audio signal, wherein the encoder calculator (203) comprises: a combiner (2031) for combining the first channel signal (201) and the second channel signal (202) in two different ways to obtain the first combination signal (204) and the second combination signal (2032);a predictor (2033) for applying the prediction information (206) to the first combination signal (204) or a signal (600) derived from the first combination signal (204) to obtain a prediction signal (2035);and a residual signal calculator (2034) for calculating the prediction residual signal (205) by combining the prediction signal (2035) and the second combination signal (2032), wherein the predictor (2033) is configured for multiplying the first combination signal (204) by a real part of the prediction information (2073) to obtain a first part of the prediction signal;for estimating (2070) an imaginary part (600) of the first combination signal using the first combination signal (204);and for multiplying the imaginary part of the first combined signal by an imaginary part of the prediction information (2074) to obtain a second part of the prediction signal;wherein the residual calculator (2034) is configured for linearly combining the first part signal of the prediction signal or the second part signal of the prediction signal and the second combination signal to obtain the prediction residual signal (205), and wherein the predictor (2033) is configured for receiving window shape information (109) and for using different filter coefficients for calculating an imaginary spectrum, where the different filter coefficients depend on different window shapes indicated by the window shape information (109).
  9. 9
    Audio encoder in accordance with claim 8, in which the predictor (2033) comprises a quantizer for quantizing the first channel signal, the second channel signal, the first combination signal or the second combination signal to obtain one or more quantized signals, and wherein the predictor (2033) is configured for calculating the residual signal using quantized signals.
  10. 10
    Audio encoder in accordance with one of claims 8 to 9, in which the first channel signal is a spectral representation of a block of samples;in which the second channel signal is a spectral representation of a block of samples, wherein the spectral representations are either pure real spectral representations or pure imaginary spectral representations, in which the optimizer (207) is configured for calculating the prediction information (206) as a real-valued factor different from zero and/or as an imaginary factor different from zero, and in which the encoder calculator (203) is configured to calculate the first combination signal and the prediction residual signal so that the prediction signal is derived from the pure real spectral representation or the pure imaginary spectral representation using the real-valued factor.
  11. 11
    Audio encoder in accordance with one of claims 8 to 10, in which the first channel signal is a spectral representation of a block of samples;in which the second channel signal is a spectral representation of a block of samples, wherein the spectral representations are either pure real spectral representations or pure imaginary spectral representations, in which the optimizer (207) is configured for calculating the prediction information (206) as a real-valued factor different from zero and/or as an imaginary factor different from zero, and in which the predictor of the encoder calculator (203) comprises a real-to-imaginary transformer (2070) or an imaginary-to-real transformer for deriving a transform spectral representation from the first combination signal, and in which the encoder calculator (203) is configured to calculate the first combined signal (204) and the first residual signal (2032) so that the prediction signal is derived from the transformed spectrum using the imaginary factor.
  12. 12
    Method of decoding an encoded multi-channel audio signal (100), the encoded multi-channel audio signal comprising an encoded first combination signal generated based on a combination rule for combining a first channel audio signal and a second channel audio signal of a multi-channel audio signal, an encoded prediction residual signal and prediction information, comprising:decoding (110) the encoded first combination signal (104) to obtain a decoded first combination signal (112), and decoding the encoded residual signal (106) to obtain a decoded residual signal (114);and calculating (116) a decoded multi-channel signal having a decoded first channel signal (117), and a decoded second channel signal (118) using the decoded residual signal (114), the prediction information (108) and the decoded first combination signal (112), so that the decoded first channel signal (117) and the decoded second channel signal (118) are at least approximations of the first channel signal and the second channel signal of the multi-channel signal, wherein the prediction information (108) comprises a real-valued portion different from zero and/or an imaginary portion different from zero, wherein the calculating the decoded multi-channel signal (116) comprises: applying the prediction information (108) to the decoded first combination signal (112) or to a signal (601) derived from the decoded first combination signal to obtain a prediction signal (1163);calculating a second combination signal (1165) by combining the decoded residual signal (114) and the prediction signal (1163);and combining the decoded first combination signal (112) and the second combination signal (1165) to obtain a decoded multi-channel audio signal having the decoded first channel signal (117) and the decoded second channel signal (118), and wherein the applying the prediction information comprises receiving window shape information (109) and using different filter coefficients for calculating an imaginary spectrum, where the different filter coefficients depend on different window shapes indicated by the window shape information (109).
  13. 13
    Method of encoding a multi-channel audio signal having two or more channel signals, comprising:calculating (203) a first combination signal (204) and a prediction residual signal (205) using a first channel signal (201) and a second channel signal (202) and prediction information (206), so that a prediction residual signal, when combined with a prediction signal derived from the first combination signal or a signal derived from the first combination signal and the prediction information (206) results in a second combination signal (2032), the first combination signal (204) and the second combination signal (2032) being derivable from the first channel signal (201) and the second channel signal (202) using a combination rule;calculating (207) the prediction information (206) so that the prediction residual signal (205) fulfills an optimization target (208);encoding (209) the first combination signal (204) and the prediction residual signal (205) to obtain an encoded first combination signal (210) and an encoded residual signal (211);and combining (212) the encoded first combination signal (210), the encoded prediction residual signal (211) and the prediction information (206) to obtain an encoded multi-channel audio signal, wherein the calculating (203) comprises: combining the first channel signal (201) and the second channel signal (202) in two different ways to obtain the first combination signal (204) and the second combination signal (2032);applying the prediction information (206) to the first combination signal (204) or a signal (600) derived from the first combination signal (204) to obtain a prediction signal (2035);and calculating the prediction residual signal (205) by combining the prediction signal (2035) and the second combination signal (2032), wherein the applying the prediction information comprises: multiplying the first combination signal (204) by a real part of the prediction information (2073) to obtain a first part of the prediction signal;estimating (2070) an imaginary part (600) of the first combination signal using the first combination signal (204);and multiplying the imaginary part of the first combined signal by an imaginary part of the prediction information (2074) to obtain a second part of the prediction signal;wherein the calculating the residual signal (2034) comprises linearly combining the first part signal of the prediction signal or the second part signal of the prediction signal and the second combination signal to obtain the prediction residual signal (205), and wherein the applying the prediction information comprises receiving window shape information (109) and using different filter coefficients for calculating an imaginary spectrum, where the different filter coefficients depend on different window shapes indicated by the window shape information (109).
  14. 15
    Encoded multi-channel audio signal comprising an encoded first combination signal generated based on a combination rule for combining a first channel audio signal and a second channel audio signal of a multi-channel audio signal, an encoded prediction residual signal, prediction information, and window shape information as side information.