EP2947653B1

Multi-channel audio coding using complex prediction and window shape information

Abstract

This record has no abstract on file.

EP2947653B1, drawing sheet 1
Sheet 1 of 19

Term

4.5 yearsleft in the term

Expires 23 March 2031.

  1. Priority
  2. Filed
  3. Granted
  4. Today
  5. Expires

19 claims: 9 independent, 10 dependent

  1. 1
    Audio decoder for decoding an encoded multi-channel audio signal (100), the encoded multi-channel audio signal (100) comprising an encoded first combination signal (104) generated based on a combination rule for combining a first channel audio signal and a second channel audio signal of a multi-channel audio signal, an encoded prediction residual signal (106) and prediction information (108), comprising:a signal decoder (110) for decoding the encoded first combination signal (104) to obtain a decoded first combination signal (112), and for decoding the encoded prediction residual signal (106) to obtain a decoded residual signal (114);and a decoder calculator (116) for calculating a decoded multi-channel audio signal having a decoded first channel signal (117), and a decoded second channel signal (118) using the decoded residual signal (114), the prediction information (108) and the decoded first combination signal (112), so that the decoded first channel signal (117) and the decoded second channel signal (118) are at least approximations of the first channel audio signal and the second channel audio signal of the multi-channel audio signal, wherein the prediction information (108) comprises an imaginary portion different from zero, wherein the decoder calculator (116) comprises: a predictor (1160) for applying the prediction information (108) to the decoded first combination signal (112) or to a signal (601) derived from the decoded first combination signal (112) to obtain a prediction signal (1163);a combination signal calculator (1161) for calculating a second combination signal (1165) by combining the decoded residual signal (114) and the prediction signal (1163);and a combiner (1162) for combining the decoded first combination signal (112) and the second combination signal (1165) to obtain the decoded multi-channel audio signal having the decoded first channel signal (117) and the decoded second channel signal (118), wherein the predictor (1160) comprises a real-to-imaginary converter (1160) for estimating (1160a) an imaginary spectrum of the decoded first combination signal (112) using a real part of the decoded first combination signal (112) directly in the frequency domain using two-dimensional filtering, the real part of the decoded first combination signal (112) being subject to window switching, wherein the predictor (1160) is configured for multiplying an imaginary part (601) of the decoded first combination signal (112) by the imaginary part of the prediction information (108) to obtain at least a part of the prediction signal (1163), and wherein the predictor (1160) is configured for receiving window shape information (109) and for using different filter coefficients by the real-to-imaginary converter (1160) for calculating the imaginary spectrum of the decoded first combination signal (112), wherein the different filter coefficients depend on different window shapes indicated by the window shape information (109), wherein the filter coefficients used by the predictor (1160) depend on a complete window, and wherein a set of filter coefficients is required for every window type and for every window transition.
  2. 2
    Audio decoder in accordance with claim 1, in which the encoded first combination signal (104) and the encoded prediction residual signal (106) have been generated using an aliasing generating time-spectral conversion, wherein the decoder further comprises:a spectral-time converter (52, 53) for generating a time-domain first channel signal and a time-domain second channel signal using a spectral-time conversion algorithm matched to the time-spectral conversion algorithm;an overlap/add processor (522) for conducting an overlap-add processing for the time-domain first channel signal and for the time-domain second channel signal to obtain an aliasing-free first time-domain signal and an aliasing-free second time-domain signal.
  3. 3
    Audio decoder in accordance with one of the preceding claims, in which the prediction information (108) further comprises a real factor different from zero, in which the predictor (1160) is configured for multiplying the decoded first combination signal (112) by the real factor to obtain a first part of the prediction signal (1163), and in which the combination signal calculator (1161) is configured for linearly combining the decoded residual signal (114) and the first part of the prediction signal (1163) and the at least a part of the prediction residual signal.
  4. 4
    Audio decoder in accordance with one of the preceding claims, in which the encoded first combination signal (104) or the decoded first combination signal (112) and the encoded prediction residual signal (106) or the decoded residual signal (114) each comprises a first plurality of subband signals, wherein the prediction information (108) comprises a second plurality of prediction information parameters, the second plurality being smaller than the first plurality, wherein the predictor (1160) is configured for applying the same prediction parameter to at least two different subband signals of the decoded first combination signal (112), wherein the decoder calculator (116) or the combination signal calculator (1161) or the combiner (1162) are configured for performing a subband-wise processing;and wherein the audio decoder further comprises a synthesis filterbank (52, 53) for combining subband signals of the decoded first combination signal (112) and the decoded second combination signal (1165) to obtain a time-domain first decoded signal and a time-domain second decoded signal.
  5. 5
    Audio decoder in accordance with claim 1, in which the decoded first combination signal (112) comprises a sequence of real-valued signal frames, and in which the predictor (1160) is configured for estimating (1160a), as the imaginary spectrum of the decoded first combination signal, an imaginary part of the current signal frame using only the current real-valued signal frame or using the current real-valued signal frame and either only one or more preceding or only one or more following real-valued signal frames or using the current real-valued signal frame and one or more preceding real-valued signal frames and one or more following real-valued signal frames.
  6. 6
    Audio decoder in accordance with one of claims 1 to 5, in which the encoded multi-channel audio signal (100) comprises, as side information, a real indicator indicating that all prediction coefficients for a frame of the encoded multi-channel audio signal (100) are real valued, wherein the audio decoder is configured for extracting the real indicator from the encoded multi-channel audio signal (100), and wherein the decoder calculator (116) is configured for not calculating an imaginary signal for a frame, for which the real indicator is indicating only real-valued prediction coefficients.
  7. 7
    Audio decoder in accordance with claim 1, wherein a previous frame's spectrum or a next frame's spectrum is an MDCT spectrum, wherein the filter coefficients applied to the previous frame's spectrum or applied to the next frame's spectrum depend only on the window half overlapping with a current frame, wherein a set of coefficients is required only for each window type.
  8. 8
    Audio decoder in accordance with claim 1, wherein a window type is either a sine window or a Kaiser Bessel Derived window and, subject to a given window sequence configuration, the window type can be a long window, a start window, a stop window, a stop-start window, or a short window.
  9. 9
    Audio decoder in accordance with claim 1, wherein a window configuration can be a window configuration of long window, short window, start window, stop window or stop start window.
  10. 10
    Audio decoder in accordance with claim 1, wherein the predictor (1160) is configured to calculate an MDST spectrum as the imaginary spectrum using an MDCT spectrum of a current frame as the real part of the decoded first combination signal (112), wherein the filter coefficients used by the predictor (1160) are MDST filter coefficients and depend on the window shape of a left half of the current window and the right half of the current window, wherein:either the left half is a sine shape and the right half is a sine shape, or the left half is a Kaiser Bessel Derived shape and the right half is a Kaiser Bessel Derived shape, or the left half is a sine shape and the right half is a Kaiser Bessel Derived shape, or the left half is a Kaiser Bessel Derived shape and the right half is a sine shape.
  11. 11
    Audio decoder in accordance with claim 1, wherein the predictor (1160) is configured to calculate an MDST spectrum as the imaginary spectrum using an MDCT spectrum of a current frame using the following MDST filter coefficients selected for a left half and a corresponding right half of a current window and a corresponding current window sequence in accordance with the following Table A:Table A - MDST Filter Coefficients for Current Window Current Window Sequence Left Half: Sine Shape Right Half: Sine Shape Left Half: KBD Shape Right Half: KBD Shape ONLY_LONG_SEQUENCE, EIGHT_SHORT_SEQUENCE [0.000000, 0,000000, 0.500000, 0.000000, -0.500000, 0.000000, 0.000000] [0.091497, 0.000000, 0.581427, 0.000000, -0.581427, 0.000000, -0.091497] LONG_START_SEQUENCE [0.102658, 0.103791, 0.567149, 0.000000, -0.567149, -0.103791, -0.102658] [0.150512, 0.047969, 0.608574, 0.000000, -0.608574, -0.047969, -0.150512] LONG_STOP_SEQUENCE [0.102658, -0.103791, 0.567149, 0.000000, -0.567149, 0.103791, -0.102658] [0.150512, -0.047969, 0.608574, 0.000000, -0.608574, 0.047969, -0.150512] STOP_START_SEQUENCE [0.205316, 0.000000, 0.634298, 0.000000, -0.634298, 0.000000, -0.205316] [0.209526, 0.000000, 0.635722, 0.000000, -0.635722, 0.000000, -0.209526] Current Window Sequence Left Half: Sine Shape Right Half: KBD Shape Left Half: KBD Shape Right Half: Sine Shape ONLY_LONG_SEQUENCE EIGHT_SHORT_SEQUENCE [0.045748, 0.057238, 0.540714, 0.000000, -0.540714, -0.057238, -0.045748] [0.045748, -0.057238, 0.540714, 0.000000, -0.540714, 0.057238, -0.045748] LONG_START_SEQUENCE [0.104763, 0.105207, 0.567861, 0.000000, -0.567861, -0.105207, -0.104763] [0.148406, 0.046553, 0.607863, 0.000000, -0.607863, -0.046553, -0.148406] LONG_STOP_SEQUENCE [0.148406, -0.046553, 0.607863, 0.000000, -0.607863, 0.046553, -0.148406] [0.104763, -0.105207, 0.567861, 0.000000, -0.567861, 0.105207, -0.104763] STOP_START_SEQUENCE [0.207421, 0.001416, 0.635010, 0.000000, -0.635010, -0.001416, -0.207421] [0.207421, -0.001416, 0.635010, 0.000000, -0.635010, 0.001416, -0.207421]
  12. 12
    Audio decoder in accordance with claim 11, wherein the predictor (1160) is configured to calculate an MDST spectrum as the imaginary spectrum using, additionally, an MDCT spectrum of a previous frame and using the following MDST filter coefficients selected for a left half of a current window and a corresponding current window sequence in accordance with the following Table B:Table B - MDST Filter Coefficients for Previous Window Current Window Sequence Left Half of Current Window: Sine Shape Left Half of Current Window: KBD Shape ONLY_LONG_SEQUENCE, LONG_START_SEQUENCE, EIGHT_SHORT_SEQUENCE [ 0.000000, 0.106103, 0.250000, 0.318310, 0.250000, 0.106103, 0.000000] [ 0.059509, 0.123714, 0.186579, 0.213077, 0.186579, 0.123714, 0.059509] LONG_STOP_SEQUENCE, STOP_START_SEQUENCE [0.038498, 0.039212, 0.039645, 0.039790, 0.039645, 0.039212, 0.038498] [0.026142, 0.026413, 0.026577, 0.026631, 0.026577, 0.026413, 0.026142 ]
  13. 13
    Audio encoder for encoding a multi-channel audio signal having two or more channel signals, comprising:an encoder calculator (203) for calculating a first combination signal (204) and a prediction residual signal (205) using a first channel signal (201) and a second channel signal (202) and prediction information (206), so that a prediction residual signal (205), when combined with a prediction signal (2035) derived from the first combination signal (204) or a signal derived from the first combination signal (204) and the prediction information (206) results in a second combination signal (2032), the first combination signal (204) and the second combination signal (2032) being derivable from the first channel signal (201) and the second channel signal (202) using a combination rule;an optimizer (207) for calculating the prediction information (206), so that the prediction residual signal (205) fulfills an optimization target (208);a signal encoder (209) for encoding the first combination signal (204) and the prediction residual signal (205) to obtain an encoded first combination signal (210) and an encoded prediction residual signal (211);and an output interface (212) for combining the encoded first combination signal (210), the encoded prediction residual signal (211) and the prediction information (206) to obtain an encoded multi-channel audio signal, wherein the encoder calculator (203) comprises: a combiner (2031) for combining the first channel signal (201) and the second channel signal (202) in two different ways to obtain the first combination signal (204) and the second combination signal (2032);a predictor (2033) for applying the prediction information (206) to the first combination signal (204) or a signal (600) derived from the first combination signal (204) to obtain the prediction signal (2035);and a residual signal calculator (2034) for calculating the prediction residual signal (205) by combining the prediction signal (2035) and the second combination signal (2032), wherein the predictor (2033) is configured for multiplying the first combination signal (204) by a real part (2073) of the prediction information (206) to obtain a first part of the prediction signal (2035);for estimating (2070) an imaginary part (600) of the first combination signal using the first combination signal (204), wherein the predictor (2033) comprises a real-to-imaginary converter (2070) for estimating, directly in the frequency domain, an imaginary spectrum of the first combination signal as the imaginary part (600) of the first combination signal using the first combination signal (204) using two-dimensional filtering, the first combination signal (204) being subject to window switching;and for multiplying the imaginary part (600) of the first combination signal by an imaginary part (2074) of the prediction information (206) to obtain a second part of the prediction signal (2035);wherein the residual calculator (2034) is configured for linearly combining the first part of the prediction signal (2035) or the second part of the prediction signal (2035) and the second combination signal (2032) to obtain the prediction residual signal (205), and wherein the predictor (2033) is configured for receiving window shape information (109) and for using different filter coefficients for calculating, using the real-to-imaginary converter (2070), the imaginary spectrum of the first combination signal, wherein the different filter coefficients depend on different window shapes indicated by the window shape information (109), wherein the filter coefficients used by the predictor (2033) depend on a complete window, and wherein a set of filter coefficients is required for every window type and for every window transition.
  14. 14
    Audio encoder in accordance with claim 13, in which the predictor (2033) comprises a quantizer for quantizing the first channel signal, the second channel signal, the first combination signal (204), or the second combination signal (2023) to obtain one or more quantized signals, and wherein the predictor (2033) is configured for calculating the prediction residual signal (205) using quantized signals.
  15. 15
    Audio encoder in accordance with one of claims 13 to 14, in which the first channel signal is a spectral representation of a block of samples;in which the second channel signal is a spectral representation of a block of samples, wherein the spectral representations are either pure real spectral representations or pure imaginary spectral representations, in which the optimizer (207) is configured for calculating the prediction information (206) as a real-valued factor different from zero and/or as an imaginary factor different from zero, and in which the encoder calculator (203) is configured to calculate the first combination signal (204) and the prediction residual signal (205), so that the prediction signal (2035) is derived from the pure real spectral representation or the pure imaginary spectral representation using the real-valued factor.
  16. 16
    Audio encoder in accordance with one of claims 13 to 15, in which the first channel signal is a spectral representation of a block of samples;in which the second channel signal is a spectral representation of a block of samples, wherein the spectral representations are either pure real spectral representations or pure imaginary spectral representations, in which the optimizer (207) is configured for calculating the prediction information (206) as a real-valued factor different from zero and/or as an imaginary factor different from zero, and in which the predictor (2033) of the encoder calculator (203) comprises the real-to-imaginary converter (2070) or an imaginary-to-real transformer for deriving a transformed spectral representation from the first combination signal (204), and in which the encoder calculator (203) is configured to calculate the first combination signal (204) and the prediction residual signal (205), , so that the prediction residual signal (205) is derived from the transformed spectral representation using the imaginary factor.
  17. 17
    Method of decoding an encoded multi-channel audio signal (100), the encoded multi-channel audio signal (100) comprising an encoded first combination signal (104) generated based on a combination rule for combining a first channel audio signal and a second channel audio signal of a multi-channel audio signal, an encoded prediction residual signal (106) and prediction information (108), comprising:decoding (110) the encoded first combination signal (104) to obtain a decoded first combination signal (112), and decoding the encoded prediction residual signal (106) to obtain a decoded residual signal (114);and calculating (116) a decoded multi-channel audio signal having a decoded first channel signal (117), and a decoded second channel signal (118) using the decoded residual signal (114), the prediction information (108) and the decoded first combination signal (112), so that the decoded first channel signal (117) and the decoded second channel signal (118) are at least approximations of the first channel audio signal and the second channel audio signal of the multi-channel audio signal, wherein the prediction information (108) comprises an imaginary portion different from zero, wherein the calculating the decoded multi-channel audio signal (116) comprises: applying the prediction information (108) to the decoded first combination signal (112) or to a signal (601) derived from the decoded first combination signal (112) to obtain a prediction signal (1163);calculating a second combination signal (1165) by combining the decoded residual signal (114) and the prediction signal (1163);and combining the decoded first combination signal (112) and the second combination signal (1165) to obtain the decoded multi-channel audio signal having the decoded first channel signal (117) and the decoded second channel signal (118), wherein the applying the prediction information (108) comprises estimating (1160a) an imaginary spectrum of the decoded first combination signal (112) using a real part of the decoded first combination signal (112) directly in the frequency domain using two-dimensional filtering in a real-to-imaginary converter (1160), the real part of the decoded first combination signal (112) being subject to window switching, wherein the applying the prediction information (108) comprises multiplying an imaginary part (601) of the decoded first combination signal (112) by the imaginary part of the prediction information (108) to obtain at least a part of the prediction signal (1163), and wherein the applying the prediction information (108) comprises receiving window shape information (109) and using different filter coefficients by the real-to-imaginary converter (1160) for calculating the imaginary spectrum of the decoded first combination signal (112), wherein the different filter coefficients depend on different window shapes indicated by the window shape information (109), wherein the filter coefficients depend on a complete window, and wherein a set of filter coefficients is required for every window type and for every window transition.
  18. 18
    Method of encoding a multi-channel audio signal having two or more channel signals, comprising:calculating (203) a first combination signal (204) and a prediction residual signal (205) using a first channel signal (201) and a second channel signal (202) and prediction information (206), so that a prediction residual signal, when combined with a prediction signal (2035) derived from the first combination signal (204) or a signal derived from the first combination signal (204) and the prediction information (206) results in a second combination signal (2032), the first combination signal (204) and the second combination signal (2032) being derivable from the first channel signal (201) and the second channel signal (202) using a combination rule;calculating (207) the prediction information (206), so that the prediction residual signal (205) fulfills an optimization target (208);encoding (209) the first combination signal (204) and the prediction residual signal (205) to obtain an encoded first combination signal (210) and an encoded prediction residual signal (211);and combining (212) the encoded first combination signal (210), the encoded prediction residual signal (211) and the prediction information (206) to obtain an encoded multi-channel audio signal, wherein the calculating (203) comprises: combining the first channel signal (201) and the second channel signal (202) in two different ways to obtain the first combination signal (204) and the second combination signal (2032);applying the prediction information (206) to the first combination signal (204) or a signal (600) derived from the first combination signal (204) to obtain a prediction signal (2035);and calculating the prediction residual signal (205) by combining the prediction signal (2035) and the second combination signal (2032), wherein the applying the prediction information (206) comprises: multiplying the first combination signal (204) by a real part (2073) of the prediction information (206) to obtain a first part of the prediction signal (2035);estimating (2070) an imaginary part (600) of the first combination signal using the first combination signal (204) wherein the applying the prediction information (206) comprises estimating, by a real-to-imaginary converter (2070), directly in the frequency domain, an imaginary spectrum of the first combination signal as the imaginary part (600) of the first combination signal using the first combination signal (204) by means of two-dimensional filtering, the first combination signal (204) being subject to window switching;and multiplying the imaginary part (600) of the first combination signal (204) by an imaginary part (2074) of the prediction information (206) to obtain a second part of the prediction signal (2035);wherein the calculating the residual signal (2034) comprises linearly combining the first part of the prediction signal (2035) or the second part of the prediction signal (2035) and the second combination signal (2023) to obtain the prediction residual signal (205), and wherein the applying the prediction information (206) comprises receiving window shape information (109) and using, by the real-to-imaginary converter (2070), different filter coefficients for calculating the imaginary spectrum of the first combination signal, wherein the different filter coefficients depend on different window shapes indicated by the window shape information (109), wherein the filter coefficients depend on a complete window, and wherein a set of filter coefficients is required for every window type and for every window transition.
Independent claims18