Audio encoder, audio decoder, methods and computer program using jointly encoded residual signals
1 claim: 1 independent, 0 dependent
- 1Zastrzeżenia patentowe 1. Dekoder audio (200; 300; 600; 1300; 1600; 2000) do dostarczania co najmniej czterech sygnałów kanału audio (220, 222, 224, 226, 320, 322, 324, 326, 620, 622, 624, 626 1320, 1322, 1324, 1326) na podstawie kodowanej reprezentacji (210; 310, 360; 610, 682; 1310, 1312; 1610), przy czym dekoder audio jest przystosowany do dostarczania pierwszego sygnału resztkowego (232; 332; 684; 1362) i drugiego sygnału resztkowego (234; 334; 686; 1364) na podstawie wspólnie kodowanej reprezentacji (210; 310; 682; 1312) pierwszego sygnału resztkowego i drugiego sygnału resztkowego z wykorzystaniem dekodowania wielokanałowego (230; 330; 680; 1360), który wykorzystuje podobieństwa i/lub zależności między sygnałami resztkowymi; przy czym dekoder audio jest przystosowany do dostarczania pierwszego sygnału kanału audio (220; 320; 642; 1372) i drugiego sygnału kanału audio (222; 322; 644; 1374) na podstawie pierwszego sygnału downmixu (212; 312; 632; 1342) i pierwszy sygnał resztkowy z wykorzystaniem dekodowania wielokanałowego (240; 340; 640; 1370) wspomaganego sygnałem resztkowym; i przy czym dekoder audio jest przystosowany do dostarczania trzeciego sygnału kanału audio (224; 324; 656; 1382) i czwartego sygnału kanału audio (226; 326; 658; 1384) na podstawie drugiego sygnału downmixu (214; 314; 634; 1344) i drugiego sygnału resztkowego z wykorzystaniem dekodowania wielokanałowego (250; 350, 650; 1380) wspomaganego sygnałem resztkowym. 2. Dekoder audio według zastrz. 1, przy czym dekoder audio jest przystosowany do dostarczania pierwszego sygnału downmixu (212; 312; 632; 1342) i drugiego sygnału downmixu (214; 314; 634; 1344) na podstawie wspólnie kodowanej reprezentacji (360; 610; 1310) pierwszego sygnału downmixu i drugiego sygnału downmixu z wykorzystaniem dekodowania wielokanałowego (370; 630; 1340). 3. Dekoder audio według zastrz. 1 albo 2, przy czym dekoder audio jest przystosowany do dostarczania pierwszego sygnału resztkowego i drugiego sygnału resztkowego na podstawie wspólnie kodowanej reprezentacji pierwszego sygnału resztkowego i drugiego sygnału resztkowego z wykorzystaniem dekodowania wielokanałowego opartego na predykcji. 4. Dekoder audio według jednego z zastrz. 1 do 3, przy czym dekoder audio jest przystosowany do dostarczania pierwszego sygnału resztkowego i drugiego sygnału resztkowego na podstawie wspólnie kodowanej reprezentacji pierwszego sygnału resztkowego i drugiego sygnału resztkowego z wykorzystaniem dekodowania wielokanałowego wspomaganego sygnałem resztkowym. 5. Dekoder audio według zastrz. 3, przy czym dekodowanie wielokanałowe oparte na predykcji jest przystosowane do oceniania parametru predykcji opisującego udział elementu składowego sygnału, który pochodzi z elementu składowego sygnału z poprzedniej ramki, do dostarczania sygnałów resztkowych do bieżącej ramki. 6. Dekoder audio według jednego z zastrz. 3, zastrz. 4, o ile zależnie od zastrz. 3, i zastrz. 5, przy czym dekodowanie wielokanałowe oparte na predykcji jest przystosowane do otrzymywania pierwszego sygnału resztkowego i drugiego sygnału resztkowego na podstawie sygnału downmixu pierwszego sygnału resztkowego i drugiego sygnału resztkowego i na podstawie wspólnego sygnału resztkowego pierwszego sygnału resztkowego i drugiego sygnału resztkowego. 7. Dekoder audio według zastrz. 6, przy czym dekodowanie wielokanałowe oparte na predykcji jest przystosowane do zastosowania wspólnego sygnału resztkowego z pierwszym znakiem, do uzyskania pierwszego sygnału resztkowego i do zastosowania wspólnego sygnału resztkowego z drugim znakiem, który jest przeciwny do pierwszego znaku, do uzyskania drugiego sygnału resztkowego. 8. Dekoder audio według jednego z zastrz. 1 do 7, przy czym dekoder audio jest przystosowany do dostarczania pierwszego sygnału resztkowego i drugiego sygnału resztkowego na podstawie wspólnie kodowanej reprezentacji pierwszego sygnału resztkowego i drugiego sygnału resztkowego przy użyciu dekodowania wielokanałowego, które działa w dziedzinie MDCT. 9. Dekoder audio według jednego z zastrz. 1 do 8, przy czym dekoder audio jest przystosowany do dostarczania pierwszego sygnału resztkowego i drugiego sygnału resztkowego na podstawie wspólnie kodowanej reprezentacji pierwszego sygnału resztkowego i drugiego sygnału resztkowego używając USAC Complex Stereo Prediction, gdzie USAC oznacza ujednolicone kodowanie mowy i dźwięku. 10. Dekoder audio według jednego z zastrz. od 1 do 9, przy czym dekoder audio jest przystosowany do dostarczania pierwszego sygnału kanału audio i drugiego sygnału kanału audio na podstawie pierwszego sygnału downmixu i pierwszego sygnału resztkowego z wykorzystaniem dekodowania wielokanałowego wspomaganego sygnałem resztkowym opartego na parametrach; i przy czym dekoder audio jest przystosowany do dostarczania trzeciego sygnału kanału audio i czwartego sygnału kanału audio na podstawie drugiego sygnału downmixu i drugiego sygnału resztkowego z wykorzystaniem dekodowania wielokanałowego wspomaganego sygnałem resztkowym opartego na parametrach. 11. Dekoder audio według zastrz. 10, przy czym oparte na parametrach dekodowanie wielokanałowe wspomagane sygnałem resztkowym jest przystosowane do oceny jednego lub większej liczby parametrów opisujących pożądaną korelację między dwoma kanałami i/lub różnic poziomu między dwoma kanałami w celu dostarczenie dwóch lub więcej sygnałów kanału audio na podstawie odnośnego sygnału downmixu i odpowiadającego sygnału resztkowego. 12. Dekoder audio według jednego z zastrz. 1 do 11, przy czym dekoder audio jest przystosowany do dostarczania pierwszego sygnału kanału audio i drugiego sygnału kanału audio na podstawie pierwszego sygnału downmixu i pierwszego sygnału resztkowego z wykorzystaniem dekodowania wielokanałowego wspomaganego sygnałem resztkowym, które działa w dziedzinie QMF; i przy czym dekoder audio jest przystosowany do dostarczania trzeciego sygnału kanału audio i czwartego sygnału kanału audio na podstawie drugiego sygnału downmixu i drugiego sygnału resztkowego z wykorzystaniem dekodowania wielokanałowego wspomaganego sygnałem resztkowym, które działa w dziedzinie QMF. 13. Dekoder audio według jednego z zastrz. 1 do 12, przy czym dekoder audio jest przystosowany do dostarczania pierwszego sygnału kanału audio i drugiego sygnału kanału audio na podstawie pierwszego sygnału downmixu i pierwszego sygnału resztkowego z wykorzystaniem dekodowania MPEG Surround 2-1-2 lub dekodowania Unified Stereo; i przy czym dekoder audio jest przystosowany do dostarczania trzeciego sygnału kanału audio i czwartego sygnału kanału audio na podstawie drugiego sygnału downmixu i drugiego sygnału resztkowego z wykorzystaniem dekodowania MPEG Surround 2-1-2 lub dekodowania Unified Stereo. 14. Dekoder audio według jednego z zastrz. 1 do 13, przy czym pierwszy sygnał resztkowy i drugi sygnał resztkowy są powiązane z różnymi położeniami poziomymi sceny audio lub z różnymi położeniami azymutu sceny audio. 15. Dekoder audio według jednego z zastrz. 1 do 14, przy czym pierwszy sygnał kanału audio i drugi sygnał kanału audio są powiązane z pionowo sąsiednimi położeniami sceny audio, i przy czym trzeci sygnał kanału audio i czwarty sygnał kanału audio są powiązane z pionowo sąsiednimi położeniami sceny audio. 16. Dekoder audio według jednego z zastrz. 1 do 15, przy czym pierwszy sygnał kanału audio i drugi sygnał kanału audio są powiązane z pierwszym położeniem poziomym lub azymutem sceny audio, i przy czym trzeci sygnał kanału audio i czwarty sygnał kanału audio są powiązane z drugim położeniem poziomym lub położeniem azymutu sceny audio, która różni się od pierwszego położenia poziomego lub pierwszego płożenia azymutu. 17. Dekoder audio według jednego z zastrz. 1 do 16, przy czym pierwszy sygnał resztkowy jest powiązany z lewą stroną sceny audio, i przy czym drugi sygnał resztkowy jest powiązany z prawą stroną sceny audio. 18. Dekoder audio według zastrz. 17, przy czym pierwszy sygnał kanału audio i drugi sygnał kanału audio są powiązane z lewą stroną sceny audio, i przy czym trzeci sygnał kanału audio i czwarty sygnał kanału audio są powiązane z prawą stroną sceny audio. 19. Dekoder audio według zastrz. 18, przy czym pierwszy sygnał kanału audio jest powiązany z dolnym lewym położeniem sceny audio, przy czym drugi sygnał kanału audio jest powiązany z górnym lewym położeniem sceny audio, przy czym trzeci sygnał kanału audio jest powiązany z dolnym prawym położeniem sceny audio i przy czym czwarty sygnał kanału audio jest powiązany z górnym prawym położeniem sceny audio. 20. Dekoder audio według jednego z zastrz. 1 do 19, przy czym dekoder audio jest przystosowany do dostarczania pierwszego sygnału downmixu i drugiego sygnału downmixu na podstawie wspólnie kodowanej reprezentacji pierwszego sygnału downmixu i drugiego sygnału downmixu przy użyciu dekodowania wielokanałowego, przy czym pierwszy sygnał downmixu jest powiązany z lewą stroną sceny audio, a drugi sygnał downmixu jest powiązany z prawą stroną sceny audio. 21. Dekoder audio według jednego z zastrz. 1 do 20, przy czym dekoder audio jest przystosowany do dostarczania pierwszego sygnału downmixu i drugiego sygnału downmixu na podstawie wspólnie kodowanej reprezentacji pierwszego sygnału downmixu i drugiego sygnału downmixu za pomocą dekodowania wielokanałowego opartego na predykcji. 22. Dekoder audio według jednego z zastrz. 1 do 21, przy czym dekoder audio jest przystosowany do dostarczania pierwszego sygnału downmixu i drugiego sygnału downmixu na podstawie wspólnie kodowanej reprezentacji pierwszego sygnału downmixu i drugiego sygnału downmixu z wykorzystaniem dekodowania wielokanałowego wspomaganego sygnałem resztkowym opartego na predykcji. 23. Dekoder audio według jednego z zastrz. 1 do 22, przy czym dekoder audio jest przystosowany do wykonywania pierwszego wielokanałowego rozszerzenia (660; 1390) szerokości pasma na podstawie pierwszego sygnału kanału audio i trzeciego sygnału kanału audio, i przy czym dekoder audio jest przystosowany do wykonywania drugiego wielokanałowego rozszerzenia (670; 1394) szerokości pasma na podstawie drugiego sygnału kanału audio i czwartego sygnału kanału audio. 24. Dekoder audio według zastrz. 23, przy czym dekoder audio jest przystosowany do wykonywania pierwszego wielokanałowego rozszerzenia szerokości pasma w celu uzyskania dwóch lub większej liczby sygnałów (620, 624; 1320, 1324) kanału audio o rozszerzonej szerokości pasma, powiązanych z pierwsza wspólna płaszczyzna poziomą lub pierwsza wspólna wysokość sceny audio na podstawie pierwszego sygnału kanału audio i trzeciego sygnału kanału audio i jednego lub więcej parametrów rozszerzenia szerokości pasma (1338) i, przy czym dekoder audio jest przystosowany do wykonywania drugiego wielokanałowego rozszerzenia szerokości pasma w celu uzyskania dwóch lub większej liczby sygnałów (622, 626:1322, 1326) kanału audio o rozszerzonym paśmie powiązanych z drugą wspólną płaszczyzną poziomą lub drugą wspólną wysokością sceny audio na podstawie drugiego sygnał kanału audio i czwartego sygnału kanału audio i jednego lub więcej parametrów rozszerzenia pasma (1358). 25. Dekoder audio według jednego z zastrz. 1 do 24, przy czym wspólnie zakodowana reprezentacja pierwszego sygnału resztkowego i drugiego sygnału resztkowego zawiera element pary kanałów zawierający sygnał downmixu pierwszego i drugiego sygnału resztkowego i wspólny sygnał audio pierwszego i drugiego sygnału resztkowego. 26. Dekoder audio według jednego z zastrz. 1 do 25, przy czym dekoder audio jest przystosowany do dostarczania pierwszego sygnału downmixu i drugiego sygnału downmixu na podstawie wspólnie kodowanej reprezentacji pierwszego sygnału downmixu i drugiego sygnału downmixu z wykorzystaniem dekodowania wielokanałowego, przy czym wspólnie zakodowana reprezentacja pierwszego sygnału downmixu i drugiego sygnału downmixu zawiera element pary kanałów zawierający sygnał downmixu pierwszego i drugiego sygnału downmixu i wspólny sygnał resztkowy pierwszego i drugiego sygnał downmixu. 27. Koder audio (100;1100;1200;1500;2100) do dostarczania zakodowanej reprezentacji (130;1144, 1154;1220, 1222;2272, 2282) na podstawie co najmniej czterech sygnałów kanału audio (110, 112, 114, 116;1110, 1112, 1114, 1116;1210, 1212, 1214, 1216;2216, 2226, 2218, 2228), przy czym koder audio jest przystosowany do wspólnego kodowania co najmniej pierwszego sygnału kanału audio i drugiego sygnał kanału audio z wykorzystaniem kodowania wielokanałowego wspomaganego sygnałem resztkowym (140;1120;1230;2230), w celu uzyskania pierwszego sygnału downmixu (120;1122;1232;2234) i pierwszego sygnału resztkowego (142;1124;1234;2236);i przy czym koder audio jest przystosowany do wspólnego kodowania co najmniej trzeciego sygnału kanału audio i czwartego sygnału kanału audio z wykorzystaniem kodowania wielokanałowego (150;1130;1240;2240) wspomaganego sygnałem resztkowym w celu uzyskania drugiego sygnału downmixu (122;1132;1242;2244) i drugiego sygnału resztkowego (152;1134;1244;2246);i przy czym koder audio jest przystosowany do wspólnego kodowania pierwszego sygnału resztkowego i drugiego sygnału resztkowego z wykorzystaniem kodowania wielokanałowego (160;1150;1260;2260), który wykorzystuje podobieństwa i/lub zależności między sygnałami resztkowymi, w celu uzyskania wspólnie kodowanej reprezentacja (130;1154;1262;2264) sygnałów resztkowych. 28. Koder audio według zastrz. 27, przy czym koder audio jest przystosowany do wspólnego kodowania pierwszego sygnału downmixu i drugiego sygnału downmixu z wykorzystaniem kodowania wielokanałowego (1140;1250;2250), aby uzyskać wspólnie zakodowaną reprezentację (1144;1252;2254) sygnałów downmixu. 29. Koder audio według zastrz. 28, przy czym koder audio jest przystosowany do wspólnego kodowania pierwszego sygnału resztkowego i drugiego sygnału resztkowego z wykorzystaniem kodowania wielokanałowego opartego na predykcji, i przy czym koder audio jest przystosowany do wspólnego kodowania pierwszego sygnału downmixu i drugiego sygnału downmixu z wykorzystaniem kodowania wielokanałowego opartego na predykcji. 30. Koder audio według jednego z zastrz. 27 do 29, przy czym koder audio jest przystosowany do wspólnego kodowania co najmniej pierwszego sygnału kanału audio i drugiego sygnału kanału audio z wykorzystaniem kodowania wielokanałowego wspieranego sygnałem resztkowym opartego na parametrach, i przy czym koder audio jest przystosowany do wspólnego kodowania co najmniej trzeciego sygnału kanału audio i czwartego sygnału kanału audio z wykorzystaniem opartego na parametrach kodowania wielokanałowego wspomaganego sygnałem resztkowym. 31. Koder audio według jednego z zastrz. 27 do 30, przy czym pierwszy sygnał kanału audio i drugi sygnał kanału audio są powiązane z pionowo sąsiednimi położeniami sceny audio, i przy czym trzeci sygnał kanału audio i czwarty sygnał kanału audio są powiązane z pionowo sąsiadującymi pozycjami sceny audio. 32. Koder audio według jednego z zastrz. 27 do 31, przy czym pierwszy sygnał kanału audio i drugi sygnał kanału audio są powiązane z pierwszym położeniem poziomym lub położeniem azymutu sceny audio, i przy czym trzeci sygnał kanału audio i czwarty sygnał kanału audio jest powiązany z drugim położeniem poziomym lub położeniem azymutu sceny audio, która różni się od pierwszego położenia poziomego lub położenia azymutu. 33. Koder audio według jednego z zastrz. 27 do 32, przy czym pierwszy sygnał resztkowy jest powiązany z lewą stroną sceny audio, i przy czym drugi sygnał resztkowy jest powiązany z prawą stroną sceny audio. 34. Koder audio według zastrz. 33, przy czym pierwszy sygnał kanału audio i drugi sygnał kanału audio są powiązane z lewą stroną sceny audio, i przy czym trzeci sygnał kanału audio i czwarty sygnał kanału audio są powiązane z prawą stroną sceny audio. 35. Koder audio według zastrz. 34, przy czym pierwszy sygnał kanału audio jest powiązany z dolnym lewym położeniem sceny audio, przy czym drugi sygnał kanału audio jest powiązany z górnym lewym położeniem sceny audio, przy czym trzeci sygnał kanału audio jest powiązany z dolnym prawym położeniem sceny audio i przy czym czwarty sygnał kanału audio jest powiązany z górnym prawym położeniem sceny audio. 36. Koder audio według jednego z zastrz. 27 do 35, przy czym koder audio jest przystosowany do wspólnego kodowania pierwszego sygnału downmixu i drugiego sygnału downmixu z wykorzystaniem kodowania wielokanałowego, aby uzyskać wspólnie zakodowaną reprezentację sygnałów downmixu, przy czym pierwszy sygnał downmixu jest powiązany z lewą stroną sceny audio, a drugi sygnał downmixu jest powiązany z prawą stroną sceny audio. 37. Sposób (800) do dostarczania co najmniej czterech sygnałów kanału audio na podstawie kodowanej reprezentacji, który to sposób obejmuje: dostarczanie (810) pierwszego sygnału resztkowego i drugiego sygnału resztkowego na podstawie wspólnie zakodowanej reprezentacji sygnału pierwszego sygnału resztkowego i drugiego sygnału resztkowego z wykorzystaniem dekodowania wielokanałowego, który wykorzystuje podobieństwa i/lub zależności między sygnałami resztkowymi;dostarczanie (820) pierwszego sygnału kanału audio i drugiego sygnału kanału audio na podstawie pierwszego sygnału downmixu i pierwszego sygnału resztkowego z wykorzystaniem dekodowania wielokanałowego wspomaganego sygnałem resztkowym;i dostarczanie (830) trzeciego sygnału kanału audio i czwartego sygnału kanału audio na podstawie drugiego sygnału downmixu i drugiego sygnału resztkowego z wykorzystaniem dekodowania wielokanałowego wspomaganego sygnałem resztkowym. 38. Sposób (700) dostarczania zakodowanej reprezentacji na podstawie co najmniej czterech sygnałów kanału audio, który to sposób obejmuje: wspólne kodowanie (710) co najmniej pierwszego sygnału kanału audio i drugiego sygnału kanału audio z wykorzystaniem kodowania wielokanałowego wspomaganego sygnałem resztkowym, w celu uzyskania pierwszego sygnału downmixu i pierwszego sygnału resztkowego;wspólne kodowanie (720) co najmniej trzeciego sygnału kanału audio i czwartego sygnału kanału audio z wykorzystaniem kodowania wielokanałowego wspomaganego sygnałem resztkowym, w celu uzyskania drugiego sygnału downmixu i drugiego sygnału resztkowego;i wspólne kodowanie (730) pierwszego sygnału resztkowego i drugiego sygnału resztkowego z wykorzystaniem kodowania wielokanałowego, które wykorzystuje podobieństwa i/lub zależności między sygnałami resztkowymi, w celu uzyskania zakodowanej reprezentacji sygnałów resztkowych. 39. Program komputerowy przystosowany do wykonywania sposobu według zastrz. 37 albo 38, gdy program komputerowy działa na komputerze. Fraunhofer-Gesellschaft zur Forderung der angewandten Forschung e.V., Niemcy;Pełnomocnik: EP 3 022 735 B1 Z-16499/17 1/21 FIG1 EP 3 022 735 B1 Z-16499/17 2/21 C\J C\J FIG 2 EP 3 022 735 B1 Z-16499/17 EP 3 022 735 B1 Z-16499/17 4/21 o cz ^4· o |X O LO EP 3 022 735 B1 Z-16499/17 5/21 LO CD LL. EP 3 022 735 B1 Z-16499/17 EP 3 022 735 B1 Z-16499/17 O CN FIG 6B EP 3 022 735 B1 Z-16499/17 8/21 FIG 7 EP 3 022 735 B1 Z-16499/17 FIG 8 EP 3 022 735 B1 Z-16499/17 10/21 FIG 9 EP 3 022 735 B1 Z-16499/17 11/21 1010 1020 1030 1040 FIG 10 1050 EP 3 022 735 B1 Z-16499/17 12/21 O co o fi fi fi £ fi S-H fi, fi s O O fi fi fi Ρ» fi s-o o o C\J Dolny lewy kanał Dolny prawy kanał EP 3 022 735 B1 Z-16499/17 13/21 1200 C\J co co CM co ,^r CM CM CD EP 3 022 735 B1 Z-16499/17 EP 3 022 735 B1 Z-16499/17 15/21 UsacChannelPairElementConfig (sbrRatiolndex) { UsacCoreConfig ();if (sbrRatiolndex 0) { SbrConfig ();stereoCon1iglndex;} else { stereoConfiglndex = 0;} if (stereoConfiglndex 0) { Mps212Conf ig (stereoConf i gl ndex);l· + qcelndex } uimsbf uimsbf FIG 14A FIG14B EP 3 022 735 B1 Z-16499/17 16/21 o Widok z góry kodera audio 3D EP 3 022 735 B1 Z-16499/17 17/21 ca 32-calowe głośniki CNI O co o Q CO '.2 EŚ C3 Uh U O u b on N O Ό co CD O co co EP 3 022 735 B1 Z-16499/17 18/21 1700 1720 Wyjście sygnału miksera Wyjście układu miksera Sygnały głośnika Odtworzenie układu Budowa modułu konwersji formatu FIG 17 EP 3 022 735 B1 Z-16499/17 19/21 FIG 18 2010 2020 2022 2030 2032 2024 2034 FIG 19 EP 3 022 735 B1 Z-16499/17 Dane znaczące SBR EP 3 022 735 B1 Z-16499/17 21/21 Schemat kodera czterokanałowego
218 paragraphs in 1 section, as filed
Technical Field [0001] Embodiments of the invention are associated with an audio decoder for providing at least four audio channel signals based on a coded representation.
[0002] Further embodiments of the invention are associated with an audio encoder for providing a coded representation based on at least four audio channel signals.
[0003] Further embodiments of the invention are associated with a method of providing at least four audio channel signals based on an encoded representation and a method of providing an encoded representation based on at least four audio channel signals.
[0004] Further embodiments of the invention are associated with a computer program for performing one of said methods.
[0005] In general, embodiments of the invention are associated with common coding of n channels.
Background of the invention [0006] The demand for storage and transmission of audio content has been steadily increasing in recent years. Moreover, the quality requirements for the storage and transmission of audio content are also constantly increasing. Therefore, the concepts for encoding and decoding audio content have been improved. For example, so-called "advanced audio coding" ( advanced audio coding; AAC), which is described, for example, in the International Standard ISO / IEC 13818-7: 2003. In addition, several spatial extensions have been created, such as the term "MPEG Surround", which is described, for example, in the international standard ISO / IEC 23003-1: 2007. In addition, additional improvements in the encoding and decoding of spatial information of audio signals are described in the international standard ISO / IEC 23003-2: 2010, which relates to the so-called spatial coding of an audio object (called spatial audio object coding; SAOC).
[0007] Furthermore, a flexible audio coding / decoding concept that provides the ability to encode both general audio and speech signals with good coding efficiency and support for multi-channel audio signals is set out in the international standard ISO / IEC 23003-3: 2012, which describes the so-called unified speech and audio coding concept (USAC).
[0008]
In MPEG USAC [1], the common stereo coding of two channels is performed using complex prediction, MPS 2-1-1 or a unified stereo signal with limited bandwidth or full-band residual signals.
MPEG surround [2] hierarchically combines OTT and TTT frames for joint multi-channel coding with or without residual signal transmission.
[0009] Multichannel audio coding and decoding is for example also disclosed in EP 2194526 A1. However, there is a desire to provide an even more advanced concept for efficient coding and decoding of three-dimensional audio scenes.
Summary of the Invention [0010] An embodiment of the invention creates an audio decoder for providing at least four audio channel signals based on a coded representation. The audio decoder is adapted to provide the first residual signal and the second residual signal based on a jointly coded representation of the first residual signal and the second residual signal using multi-channel decoding that uses the similarities and / or relationships between the residual signals. The audio decoder is also adapted to provide the first audio channel signal and the second audio channel signal based on the first residual signal using multi-channel decoding supported by the residual signal. The audio decoder is also adapted to provide at least a third audio channel signal and a fourth audio channel signal based on a second downmix signal and a second residual signal using multi-channel decoding supported by the residual signal. This embodiment of the invention is based on whether the relationship between four or even more audio channel signals can be used by deriving two residual signals, each of which is used to provide two or more audio channel signals, using multi-channel decode supported by a residual signal from a jointly coded representation of the residual signal. In other words, It found that there are usually some similarities between the residual signal mentioned, such as the bit rate for encoding said residual signals, which contribute to improving sound quality when decoding at least four audio channel signals, can be reduced by deriving two residual signals from the shared coded representation using multi-channel decoding, which uses similarities and / or relationships between residual signals.
[0012] In a preferred embodiment, the audio decoder is adapted to provide the first downmix signal and the second downmix signal based on a jointly coded representation of the first downmix signal and the second downmix signal using multi-channel decoding. Accordingly, the hierarchical structure of the audio decoder is created, wherein both downmix signals and residual signals that are used in multi-channel decoding supported by the residual signal to provide at least four audio channel signals are output using separate multi-channel decoding. This approach is particularly efficient because the two downmix signals typically contain similarities that can be used in multi-channel coding / decoding, and because the two residual signals usually also contain similarities that can be used in multi-channel coding / decoding. Thus, good coding performance is usually obtained using this approach.
[0013] In a preferred embodiment, the audio decoder is adapted to provide the first residual signal and the second residual signal based on a jointly coded representation of the first residual signal and the second residual signal using prediction-based multi-channel decoding. The use of prediction-based multi-channel decoding typically provides comparatively good reconstruction quality for residual signals. This is for example advantageous if the first residual signal represents the left side of the audio stage and the second residual signal represents the right side of the audio stage because human hearing is usually relatively sensitive to differences between the left and right sides of the sound stage.
[0014] In a preferred embodiment, the audio decoder is adapted to provide the first residual signal and the second residual signal based on a jointly coded representation of the first residual signal and the second residual signal using residual signal assisted multi-channel decoding. It has been found that particularly good quality of the first and second residual signals can be obtained if the first residual signal and the second residual signal are provided using multi-channel decoding, which in turn receives the residual signal (and usually also the downmix signal that combines the first residual signal and the second signal) residual). Thus, cascading of decoding steps occurs, in which two residual signals (first residual signal, which is used to provide the first audio channel signal and the second audio channel signal, and a second residual signal, which is used to provide the third audio channel signal and the fourth audio channel signal), are provided based on the downmix input signal and the residual input signal, the latter may also be designated as a common residual signal) of the first residual signal and the second residual signal). Thus, the first residual signal and the second residual signal are actually "intermediate" residual signals that are obtained using multi-channel decoding from the corresponding downmix signal and the corresponding "common" residual signal.
[0015] In a preferred embodiment, the prediction-based multi-channel decoding is adapted to evaluate a prediction parameter describing the contribution of the signal component that is derived using the signal component from the previous frame to provide residual signals (i.e. first residual signal and second residual signal) of the current frame. The use of such prediction-based multi-channel decoding brings particularly good quality of residual signals (first residual signal and second residual signal).
[0016] In a preferred embodiment, prediction-based multi-channel decoding is adapted to receive a first residual signal and a second residual signal based on a (corresponding) downmix signal and a (corresponding) "common" residual signal, in which prediction-based multi-channel decoding is adapted to use a common residual signal with the first character, to obtain the first residual signal and to use the common residual signal with the second character, which is opposite to the first character to obtain the second residual signal. Such prediction-based multi-channel decoding has been found to provide good performance for reconstructing the first residual signal and the second residual signal.
[0017] In a preferred embodiment, the audio decoder is adapted to provide a first residual signal and a second residual signal based on a jointly coded representation of the first residual signal and the second residual signal using multi-channel decoding that operates in the field of modified discrete cosine transformation (MDCT domain) . It has been found that this approach can be implemented efficiently because audio decoding that can be used to provide a jointly coded representation of the first residual signal and the second residual signal preferably works in the MDCT field. Accordingly, intermediate transformations can be avoided by using multi-channel decoding to provide the first residual signal and the second residual signal in the MDCT domain.
[0018] In a preferred embodiment, the audio decoder is adapted to provide the first residual signal and the second residual signal based on a co-coded representation of the first residual signal and the second residual signal using the composite USAC stereo prediction (for example, as mentioned in the USAC standard cited above) ). Such complex stereo USAC prediction has been found to perform well for decoding the first residual signal and the second residual signal. Furthermore, the use of complex stereo USAC prediction to decode the first residual signal and the second residual signal also allows simple implementation of the assumption using decoding blocks that are already available in Unified Speech and Audio Coding (USAC). Accordingly, the unified speech and audio coding decoder can be easily reconfigured to implement the decoding principle described herein.
[0019] In a preferred embodiment, the audio decoder is adapted to provide the first audio channel signal and the second audio channel signal based on the first downmix signal and the first residual signal using parameter-based multichannel assisted decoding. Similarly, the audio decoder is adapted to provide the third audio channel signal and the fourth audio channel signal based on the second downmix signal and the second residual signal using parameter-based multi-channel decoding assisted by the residual signal. It has been found that such multi-channel decoding is well suited for outputting audio channel signals based on the first downmix signal, first residual signal, second downmix signal and second residual signal. In addition, it has been found that such parameter-based, residual-signal multi-channel decoding can be implemented with little effort using processing blocks that are already present in typical multi-channel audio decoders.
[0020] In a preferred embodiment, the parameter-based assisted multi-channel decoding is adapted to evaluate one or more parameters describing the desired correlation between two channels and / or level differences between two channels to provide two or more audio channel signals based on the related the downmix signal and the corresponding corresponding residual signal. It has been found that such parameter-based, residual signal-assisted decoding is well suited for the second stage of cascading multi-channel decoding (wherein, preferably, the first and second downmix signals and the first and second residual signals are provided using prediction-based multi-channel decoding).
[0021] In a preferred embodiment, the audio decoder is adapted to provide the first audio channel signal and the second audio channel signal based on the first downmix signal and the first residual signal using residual signal assisted multi-channel decoding that operates in the QMF domain. Similarly, the audio decoder is preferably adapted to provide a third audio channel signal and a fourth audio channel signal based on a second downmix signal and a second residual signal using residual signal multichannel decoding that operates in the QMF domain. Accordingly, the second step of hierarchical multi-channel decoding operates in the QMF domain, which is well suited to typical post-processing, which is often also performed in the QMF domain, so that intermediate conversions can be avoided.
[0022] In a preferred embodiment, the audio decoder is adapted to provide the first audio channel signal and the second audio channel signal based on the first downmix signal and the first residual signal using MPEG Surround 2-1-2 decoding or unified stereo decoding. Similarly, the audio decoder is preferably adapted to provide the third audio channel signal and the fourth audio channel signal based on the second downmix signal and the second residual signal using MPEG Surround 2-1-2 decoding or unified stereo decoding. It has been found that such decoding concepts are particularly well suited to the second decoding step.
[0023] In a preferred embodiment, the first residual signal and the second residual signal are associated with different horizontal positions (or, equivalently, azimuth positions) of the audio scene. Separation of residual signals that are associated with different horizontal positions (or azimuth positions) has been found to be particularly advantageous in the first stage of hierarchical multi-channel processing because a particularly good audibility impression can be obtained if perceptually valid left / right separation is performed in the first stage hierarchical multi-channel decoding.
[0024] In a preferred embodiment, the first audio channel signal and the second channel signal are associated with vertically adjacent positions of the audio scene (or, equivalently, with adjacent positions of the height of the audio scene). Also, the third audio channel signal and the fourth audio channel signal are preferably associated with vertically adjacent audio scene positions (or, equivalently, with adjacent audio scene height positions). It has been found that good decoding results can be obtained if the separation of upper and lower signals is performed in the second stage of hierarchical audio decoding (which usually has a slightly lower separation accuracy than the first stage) because the human auditory system is less sensitive with respect to the vertical position of the sound source compared to the horizontal position of the sound source.
In a preferred embodiment, the first audio channel signal and the second audio channel signal are associated with the first horizontal position of the audio scene (or equivalent, azimuth position), and the third audio channel signal and the fourth audio channel signal are associated with the second horizontal position of the stage audio (or, equivalently, azimuth position) that differs from the first horizontal position (or equivalent azimuth position).
[0026] Preferably, the first residual signal is associated with the left side of the audio stage and the second residual signal is associated with the right side of the audio stage. Accordingly, left and right separation is performed in the first step of hierarchical audio decoding.
[0027] In a preferred embodiment, the first audio channel signal and the second audio channel signal are associated with the left side of the audio scene, and the third audio channel signal and the fourth audio channel signal are associated with the right side of the audio signal.
[0028] In another preferred embodiment, the first audio channel signal is associated with the lower left side of the audio scene, the second audio channel signal is associated with the upper left side of the audio scene, the third audio channel signal is associated with the lower right side of the audio scene, and the fourth the audio channel signal is associated with the top right of the audio scene. This combination of audio channel signals brings particularly good coding results.
[0029] In a preferred embodiment, the audio decoder is adapted to provide the first downmix signal and the second downmix signal based on a jointly coded representation of the first downmix signal and the second downmix signal using multi-channel decoding, the first downmix signal being associated with the left side of the audio scene , and the second downmix signal is associated with the right side of the audio scene. It has been found that downmix signals can also be encoded with good coding efficiency using multi-channel coding, even if the downmix signals are associated with different sides of the audio scene.
[0030] In a preferred embodiment, the audio decoder is adapted to provide the first downmix signal and the second downmix signal based on a jointly coded representation of the first downmix signal and the second downmix signal using prediction-based multi-channel decoding, and even using residual channel-based decoding on prediction. It has been found that the use of such multi-channel decoding concepts provides a particularly good decoding result. Also existing decoding functions can be reused in some audio decoders.
In a preferred embodiment, the audio decoder is adapted to perform the first multichannel bandwidth extension based on the first audio channel signal and the third audio channel signal. In addition, the audio decoder may be adapted to perform a second (usually separate) extension of the multi-channel bandwidth based on the second audio channel signal and the fourth audio channel signal. It has been found to be advantageous to make possible bandwidth extension based on two audio channel signals that are associated with different sides of the audio scene (with different residual signals usually associated with different sides of the audio scene.
[0032] In a preferred embodiment, the audio decoder is adapted to perform the first multi-channel bandwidth extension to obtain two or more extended band audio channel signals associated with a first common horizontal plane (or, equivalent, with a first common height) of the audio scene at based on the first audio channel signal and the third audio channel signal and one or more bandwidth extension parameters. In addition, the audio decoder is preferably adapted to perform a second multi-channel bandwidth extension to obtain two or more extended band audio channel signals associated with a second common horizontal plane (or equivalent second common height) of the audio stage at the base of the second and fourth channel signals audio channel signal and one or more bandwidth extension parameters. It has been found that such a decoding scheme provides good sound quality, because multi-channel bandwidth extension can take into account stereo characteristics that are important for the listening experience in such a system.
[0033] In a preferred embodiment, the co-coded representation of the first residual signal and the second residual signal comprises a channel pair element comprising a downmix signal of the first and second residual signals and a common residual signal of the first and second residual signals. It has been found that the coding of the downmix signal of the first and second residual signals and the common residual signal of the first and second residual signals using the channel pair element is advantageous because the downmix signal of the first and second residual signals and the joint residual signal of the first and second residual signals usually share many features. Accordingly, the use of an element with a pair of channels typically reduces signaling overhead and consequently allows efficient coding.
[0034] In another preferred embodiment, the audio decoder is adapted to provide the first downmix signal and the second downmix signal based on a jointly coded representation of the first downmix signal and the second downmix signal using multi-channel decoding, wherein the co-coded representation of the first downmix signal and the second downmix signal comprise a channel pair element, the channel pair element comprises a downmix signal of the first and second downmix signals and a common residual signal of the first and second downmix signals. This embodiment is based on the same considerations as the previously described embodiment.
[0035] Another embodiment of the invention creates an audio encoder providing a coded representation based on at least four audio channel signals. The audio encoder is adapted to jointly encode at least the first audio channel signal and the second audio channel signal using the residual signal multi-channel coding to obtain the first downmix signal and the joint coding of the at least the third audio channel signal and the fourth audio channel signal using the assisted multi-channel coding a residual signal to obtain a second downmix signal and a second residual signal. In addition, the audio encoder is adapted to co-encode the first residual signal and the second residual signal using multi-channel coding to obtain a jointly coded representation of the residual signals. This audio encoder is based on the same considerations as the decoder described above.
[0036] Furthermore, the optional patches of this audio encoder, and the preferred audio encoder configurations are substantially parallel to the patches and preferred audio decoder configurations discussed above. Accordingly, refer to the following discussion.
[0037] Another embodiment of the invention creates a method for providing at least four audio channel signals based on an encoded representation that preferably performs the function of an audio encoder as described above and which can be supplemented with the elements and functionality discussed above.
[0038] Another embodiment according to the invention creates a method for coded representation based on at least four audio channel signals that substantially fulfills the functionality of the audio decoder described above.
[0039] Another embodiment of the invention creates a computer program for performing the methods mentioned above.
Brief Description of the Figures [0040] Embodiments of the present invention will then be described with reference to the attached Figures, in which:
Fig. 1 is a basic block diagram of an audio encoder in accordance with an embodiment of the present invention;
Fig. 2 shows a basic block diagram of an audio decoder, according to an embodiment of the present invention;
Fig. 3 is a basic block diagram of an audio decoder in accordance with another embodiment of the present invention;
Fig. 4 shows a basic block diagram of an audio encoder in accordance with an embodiment of the present invention;
Fig. 5 shows a basic block diagram of an audio decoder, according to an embodiment of the present invention;
Fig. 6 shows a basic block diagram of an audio decoder, according to another embodiment of the present invention;
Fig. 7 is a flowchart of a method of providing a coded representation based on at least four audio channel signals in accordance with an embodiment of the present invention;
Fig. 8 is a flowchart of a method for providing at least four audio channel signals based on an encoded representation in accordance with an embodiment of the invention;
Fig. 9 is a flowchart of a method of providing a coded representation based on at least four audio channel signals in accordance with an embodiment of the invention; and
Fig. 10 is a block diagram of a method of providing at least four audio channel signals based on an encoded representation in accordance with an embodiment of the invention;
Fig. 11 shows a basic block diagram of an audio encoder in accordance with an embodiment of the invention;
Fig. 12 is a block diagram of an audio encoder in accordance with another embodiment of the present invention;
Fig. 13 shows a basic block diagram of an audio decoder in accordance with an embodiment of the invention;
Fig. 14a shows a representation of a bit stream syntax that can be used with the audio encoder of Fig. 13;
Fig. 14b shows a tabular representation of the different values of the qceIndex parameter;
Fig. 15 shows a basic block diagram of a 3D audio encoder in which the concepts of the present invention may be used;
Fig. 16 is a basic block diagram of a 3D audio decoder in which the concepts of the present invention may be used; and
Fig. 17 shows the basic block diagram of a format converter.
Fig. 18 is a graphical representation of the topological structure of a Quad Channel Element (QCE) in accordance with an embodiment of the present invention;
Fig. 19 is a block diagram of an audio decoder in accordance with an embodiment of the present invention;
Fig. 20 shows a detailed block diagram of a QCE decoder in accordance with an embodiment of the present invention; and
Fig. 21 shows a detailed block diagram of a four channel element encoder according to an embodiment of the present invention.
A detailed description of the embodiments
1. Audio encoder according to Fig. 1 [0041] Fig. 1 shows the basic block diagram of an audio encoder, which is fully indicated as 100. The audio encoder 100 is adapted to provide a coded representation based on at least four audio channel signals. The audio encoder 100 is adapted to receive a first audio channel signal 110, a second audio channel signal 112, a third audio channel signal 114 and a fourth audio channel signal 116. Furthermore, the audio encoder 100 is adapted to provide an encoded representation of the first downmix signal 120 and the second downmix signal 122, as well as a combined coded representation of the residual signals 130. The audio encoder 100 includes a residual signal multichannel encoder 140 that is adapted to co-encode the first audio channel signal 110 and the second audio channel signal 112 using residual signal multichannel encoding to obtain the first downmix signal 120 and the first residual signal 142. The audio encoder 100 also includes a residual signal multichannel encoder 150 that is adapted to co-encode at least the third audio channel signal 114 and the fourth audio channel signal 116 using the residual signal multichannel encoding to obtain a second downmix signal 122 and a second residual signal 152. The audio decoder 100 also includes a multi-channel encoder 160 that is adapted to jointly encode the first residual signal 142 and the second residual signal 152 using multi-channel coding to obtain a combined coded representation of residual signals 142, 152.
[0042] Regarding the functionality of the audio encoder 100, it should be noted that the audio encoder 100 performs hierarchical coding in which the first audio channel signal 110 and the second audio channel signal 112 are jointly encoded using multi-channel encoding 140 a residual signal provided there is both the first downmix signal 120 and the first residual signal 142. The first residual signal 142 may, for example, describe the differences between the first audio channel signal 110 and the second audio channel signal 112, and / or may describe some or any of the signal features that cannot be represented by the first downmix signal 120 and optional parameters that may be provided by a multi-channel encoder 140 assisted by a residual signal. In other words, the first residual signal 142 may be a residual signal that allows improving the decoding result that can be obtained from the first downmix signal 120 and all possible parameters that can be provided by the residual signal multi-channel encoder 140. For example, the first residual signal 142 may allow at least partial reconstruction of the waveform of the first audio channel signal 110 and the second audio channel signal 112 on the audio decoder side compared to ordinary reconstruction of the high level signal characteristics (e.g., correlation characteristics, covariance characteristics, features level differences and similar). Similarly, the residual signal multi-channel encoder 150 provides both the second downmix signal 122 and the second residual signal 152 based on the third audio channel signal 114 and the fourth audio channel signal 116, such that the second residual signal allows the reconstruction of the third audio channel signal 114 and a fourth audio channel signal 116 on the audio decoder side. The second residual signal 152 may consequently serve the same functionality as the first residual signal 142. However, if the audio channel signals 110, 112, 114, 116 contain some correlation, the first residual signal 142 and the second residual signal 152 are usually also correlated to some extent. Accordingly, the joint coding of the first residual signal 142 and the second residual signal 152 using the multi-channel encoder 160 typically includes high performance, because the multi-channel coding of the correlated signals typically reduces the bit rate by using the relationship. Consequently, the first residual signal 142 and the second residual signal 152 can be encoded with good precision while maintaining a reasonable low rate of the combined coded representation of the residual signals.
[0043] In summary, the embodiment of Fig. 1 provides hierarchical multi-channel coding in which good playback quality can be achieved by using multi-channel encoders 140, 150 assisted by a residual signal, while the required bit rate can be maintained moderately by combining the first and second residual signal coding 142 residual signal 152.
[0044] It is possible to further optionally improve the audio encoder 100. Some of these improvements will be described with reference to the Figures. 4, 11 and 12. However, it should be noted that the audio encoder 100 may also be adapted in parallel with the audio decoder described herein, with the functionality of the audio encoder being usually inverse to the functionality of the audio decoder.
2. Audio decoder according to Fig. 2 [0045] Fig. 2 shows the basic block diagram of an audio decoder, which is indicated in its entirety by 200.
[0046] The audio decoder 200 is adapted to receive an encoded representation which comprises together the encoded representation 210 of the first residual signal and the second residual signal. The audio decoder 200 also receives a representation of the first downmix signal 212 and the second downmix signal 214. The audio decoder 200 is adapted to provide the first audio channel signal 220, the second audio channel signal 222, the third audio channel signal 224 and the fourth audio channel signal 226.
[0047] The audio decoder 200 includes a multi-channel decoder 230 that is adapted to provide the first residual signal 232 and the second residual signal 234 based on a jointly coded representation 210 of the first residual signal 232 and the second residual signal 234. The audio decoder 200 also includes a (first) residual signal supported multi-channel decoder 240 that is adapted to provide the first audio channel signal 220 and the second audio channel signal 222 based on the first downmix signal 212 and the first residual signal 232 using multi-channel decoding. The audio decoder 200 also includes a residual signal assisted multi-channel decoder 250 that is adapted to provide the third audio channel signal 224 and the fourth audio channel signal 226 based on the second downmix signal 214 and the second residual signal 234.
[0048] Regarding the functionality of the audio decoder 200, it should be noted that the audio signal decoder 200 provides a first audio channel signal 220 and a second audio channel signal 222 based on (first) common multi-channel decoding 240 assisted by a residual signal in which the decoding quality for decoding multi-channel is increased by the first residual signal 232 (compared to decoding not supported by the residual signal). In other words, the first downmix signal 212 provides "coarse" information about the first audio channel signal 220 and the second audio channel signal 222, wherein, for example, the differences between the first audio channel signal 220 and the second audio channel signal 222 can be described by ( optional) parameters that can be received by a multi-channel decoder 240 assisted by a residual signal and by the first residual signal 232. Consequently, the first residual signal 232 may, for example, allow partial reconstruction of the waveform of the first audio channel signal 220 and the second audio channel signal 222.
[0049] Similarly (residual signal) second channel decoder 250 provides a third audio channel signal 224 in a fourth audio channel signal 226 based on a second downmix signal 214, wherein the second downmix signal 214 may, for example, "roughly" describe the third signal 224 audio channel and the fourth audio channel signal 226. In addition, the differences between the third audio channel signal 224 and the fourth audio channel signal 226 may be, for example, described by (optional) parameters that can be received by the (second) residual multi-channel decoder 250 and the second residual signal 234. Accordingly, the evaluation of the second residual signal 234 may, for example, allow partial reconstruction of the waveform of the third audio channel signal 224 and the fourth audio channel signal 226. Accordingly, the second residual signal 234 may allow improved reconstruction quality of the third audio channel signal 224 and the fourth audio channel signal 226.
[0050] However, the first residual signal 232 and the second residual signal 234 are derived from a jointly coded representation 210 of the first residual signal and the second residual signal. Such multi-channel decoding, which is performed by the multi-channel decoder 230, allows high decoding performance because the first audio channel signal 220, the second audio channel signal 222, the third audio channel signal 224 and the fourth audio channel signal 226 are usually similar or "correlated". Accordingly, the first residual signal 232 and the second residual signal 234 are also typically similar or "correlated" which can be used by deriving the first residual signal 232 and the second residual signal 234 from the coded representation 210 using multi-channel decoding.
[0051] Consequently, it is possible to obtain high quality decoding with a moderate bit rate by decoding the residual signals 232, 234 based on the jointly coded representation 210, and by using each of the residual signals to decode two or more audio channel signals.
[0052] In summary, the audio decoder 200 allows for high coding efficiency by providing high quality signals 220, 222, 224, 226 of the audio channel.
[0053] It will be appreciated that additional features and functions that may be implemented optionally in the audio decoder 200 will be described later with reference to Figs. 3,5,6 and 13. It should be noted, however, that the audio encoder 200 may include the above-mentioned advantages without any additional modifications.
3. Audio decoder according to Fig. 3 [0054] Fig. 3 shows a basic block diagram of an audio decoder according to another embodiment of the present invention. The audio decoder of Fig. 3 is indicated in its entirety by 300. The audio decoder 300 is similar to the audio decoder 200 according to Fig. 2, so that the above explanations also apply. However, the 300 audio decoder is supplemented with additional features and functionalities compared to the 200 audio decoder, which will be explained below.
[0055] The audio decoder 300 is adapted to receive a jointly coded representation 310 of the first residual signal and the second residual signal. In addition, the audio decoder 300 is adapted to receive a combined coded 360 representation of the first downmix signal and the second downmix signal. In addition, the audio decoder 300 is adapted to provide the first audio channel signal 320, the second audio channel signal 322, the third audio channel signal 324 and the fourth audio channel signal 326. The audio decoder 300 includes a multi-channel decoder 330 that is adapted to receive a jointly coded representation 310 of the first residual signal and the second residual signal and to provide based on this first residual signal 332 and the second residual signal 334. The 300 audio decoder also includes (first) 340 multi-channel decoding assisted by a residual signal, which receives the first residual signal 332 and the first downmix signal 312, and provides the first audio channel signal 320 and the second audio channel signal 322. The audio decoder 300 also includes (second) residual multi-channel decoding 350, which is adapted to receive a second residual signal 334 and a second downmix signal 314, and for providing the third audio channel signal 324 and the fourth audio channel signal 326.
[0056] The audio decoder 300 also includes another multi-channel decoder 370 that is adapted to receive a combined coded representation of the first downmix signal and the second downmix signal, and to provide, based on this first downmix signal 312 and the second downmix signal 314.
[0057] Further specific details of the audio decoder 300 will be described below.
However, it should be noted that the correct audio decoder does not have to implement a combination of all these additional functions and functionalities. Instead, the functions and functions described below can be individually added to the audio decoder 200 (or any other audio decoder) to gradually improve the audio decoder 200 (or any other audio decoder).
[0058] In a preferred embodiment, the audio decoder 300 receives a jointly coded representation 310 of the first residual signal and the second residual signal, wherein the combined coded representation 310 may include a downmix signal of the first residual signal 332 and the second residual signal 334 and a common residual signal of the first residual signal 332 and the second residual signal 334. In addition, the jointly coded representation 310 may, for example, comprise one or more prediction parameters. Accordingly, multi-channel decoder 330 may be a residual signal-based multi-channel decoder based on prediction. For example, 330 multi-channel decoder may be a composite USAC stereo prediction, as described, for example, in the "Complex Stereo Prediction" section of the international ISO / IEC 23003-3: 2012 standard. For example, multi-channel decoder 330 may be adapted to evaluate a prediction parameter describing the contribution of a signal element that is derived from the signal element from the previous frame to provide the first residual signal 332 and the second residual signal 334 for the current frame. In addition, multi-channel decoder 330 may be adapted to use a common residual signal (which is included in the total coded representation 310) with a first character, to obtain the first residual signal 332 and to use a common residual signal (which is included in the total coded representation 310) of a second character, which is opposite to the first character, to obtain the second residual signal 334. Thus, the common residual signal may, at least in part, describe the differences between the first residual signal 332 and the second residual signal 334. However, the multi-channel decoder 330 may evaluate a downmix signal, a common residual signal and one or more prediction parameters that are included in the jointly coded representation 310 to obtain the first residual signal 332 and the second residual signal 334 as described in the abovementioned international ISO standard / IEC 23003-3: 2012. In addition, it should be noted that the first residual signal 332 may be associated with the first horizontal position (or azimuth position), e.g., the left horizontal position, and that the second residual signal 334 may be associated with the second horizontal position (or azimuth position), for example in the right vertical position of the audio scene.
[0059] The combined coded 360 representation of the first downmix signal and the second downmix signal preferably comprises a downmix signal of the first downmix signal and the second downmix signal, a common residual signal of the first downmix signal and the second downmix signal, and one or more prediction parameters. In other words, there is a "common" downmix signal to which the first downmix signal 312 and the second downmix signal 314 are downmixed, and there is a "common" residual signal that can describe, at least in part, the differences between the first downmix signal 312 and the second downmix signal 314 . The multi-channel decoder 370 is preferably a multi-channel residual signal supported decoder based on prediction, e.g. a composite USAC stereo predictive decoder. In other words, the multi-channel decoder 370 that provides the first downmix signal 312 and the second downmix signal 314 may be substantially identical to the multi-channel decoder 330 which provides the first residual signal 332 and the second residual signal 334, such that the above explanations and references also apply. In addition, it should be noted that the first downmix signal 312 is preferably associated with the first horizontal or azimuth position (e.g., left horizontal or azimuth position) of the audio scene, and that the second downmix 314 signal is associated preferably with the second horizontal or azimuth position (e.g., right horizontal or azimuth position) of the audio scene. Accordingly, the first downmix signal 312 and the first residual signal 332 may be associated with the same first horizontal position or azimuth position (e.g., left horizontal position), and the second downmix signal 314 and the second residual signal 334 may be associated with the same, a second horizontal or azimuth position (for example, the right horizontal position). Accordingly, both multi-channel decoder 370 and multi-channel decoder 330 may perform horizontal splitting (or horizontal distribution or horizontal distribution).
[0060] The multi-channel decoder 340 assisted by the residual signal may advantageously be parameter based and may consequently receive one or more parameters 342 describing the desired correlation between two channels (e.g., between the first audio channel signal 320 and the second audio channel signal 322) and / or level differences between the two channels. For example, residual signal-assisted multi-channel decoding 340 may be based on MPEG-Surround coding (as described, for example, in ISO / IEC 23003-1: 2007) with a residual signal extension or a "unified stereo decoding" decoder (as described, for example in ISO / IEC 23003-3, chapter 7.11 (Decoder) and Annex B.21 (Encoder description and definition of the term "Unified Stereo", "Unified Stereo"). Accordingly, the residual signal multi-channel decoder 340 may provide the first audio channel signal 320 and the second audio channel signal 322, wherein the first audio channel signal 320 and the second audio channel signal 322 are associated with vertically adjacent audio scene positions. For example, the first audio channel signal may be associated with the lower left position of the audio scene and the second audio channel signal may be associated with the upper left position of the audio scene (such that the first audio channel signal 320 and the second audio channel signal 322 are e.g. associated with identical horizontal or azimuth positions of the audio scene or with azimuth positions separated by no more than 30 degrees). In other words, a multi-channel decoder 340 assisted by a residual signal can perform a vertical split (or distribution or separation).
[0061] The functionality of the multi-channel decoder 350 supported by the residual signal may be identical to the functionality of the multi-channel decoder 340 supported by the residual signal, wherein the third audio channel signal may for example be associated with the lower right position of the audio scene and in which the fourth audio channel signal may for example be associated with the top right position of the audio scene. In other words, the third audio channel signal and the fourth audio channel signal may be associated with vertically adjacent positions of the audio stage and may be associated with the same horizontal position or azimuth position of the audio stage, with the residual signal multichannel decoder 350 performing vertical division (or chapter or distribution).
[0062] In summary, the audio decoder 300 according to Fig. 3 performs hierarchical audio decoding in which left-right fission is performed in the first stages (multi-channel decoder 330, multi-channel decoder 370), and in which higher-lower fission is performed in the second stage (340, 350 multi-channel decoders supported by the residual signal). In addition, residual signals 332, 334 are also encoded using a jointly coded representation 310 as well as downmix signals 312, 314 (jointly coded representation 360). Thus, correlations between different channels are used both to encode (and decode) downmix signals 312, 314 and to encode (and decode) residual signals 332, 334. Accordingly, high coding efficiency is achieved and the correlations between the signals are well used.
4. Audio encoder according to Fig. 4 [0063] In Fig. 4 an essential block diagram of an audio encoder is shown, according to another embodiment of the present invention. Audio signal encoder according to Fig. 4 it is marked entirely with 400. The audio encoder 400 is adapted to receive four audio channel signals, namely a first audio channel signal 410, a second audio channel signal 412, a third audio channel signal 414 and a fourth audio channel signal 416. Furthermore, the audio signal encoder 400 is adapted to provide a coded representation based on the 410, 412, 414 and 416 audio channel signals, said coded representation including a combined coded representation of two downmix signals 420 as well as an encoded representation of the first set of 422 common width extension parameters bandwidth and the second set of 424 parameters for extending the common bandwidth. The audio encoder 400 includes a first bandwidth extension parameters extractor 430 that is adapted to obtain a first set of 422 common bandwidth extraction parameters based on the first audio channel signal 410 and the third audio channel signal 414. The audio encoder 400 also includes a second bandwidth extension parameter extractor 440 that is adapted to receive a second set of common bandwidth extension parameters 424 based on the second audio channel signal 412 and the fourth audio channel signal 416.
[0064] In addition, the audio signal encoder 400 includes a (first) multi-channel encoder 450 that is adapted to jointly encode at least the first audio channel signal 410 and the second audio channel signal 412 using the multi-channel coding to obtain the first downmix 452 signal. In addition, the audio encoder 400 also includes a (second) multi-channel encoder 460 that is adapted to jointly encode at least a third audio channel signal 414 and a fourth audio channel signal 416 using multi-channel coding to obtain a second downmix 462 signal. In addition, the audio encoder 400 also includes a (third) multi-channel encoder 470 that is adapted to jointly encode the first downmix signal 452 and the second downmix signal 462 using multi-channel coding to obtain a combined coded representation of 420 downmix signals.
[0065] Regarding the functionality of the audio encoder 400, it should be noted that the audio encoder 400 performs hierarchical multi-channel coding in which the first audio channel signal 410 and the second audio channel signal 412 are combined in the first stage, and wherein the third audio channel signal 414 and the fourth signal The audio channel 416 are also combined in a first step to thereby obtain a first downmix signal 452 and a second downmix signal 462. The first downmix signal 452 and the second downmix signal 462 are then encoded together in the second stage. It should be noted, however, that the first bandwidth extension parameters extractor 430 provides the first set of 422 common bandwidth extension parameters based on the audio channel signals 410, 414, which are supported by various multi-channel encoders 450, 460 in the first step of the hierarchical multi-channel coding. Similarly, the second width extension parameter extractor 440 provides a second set of 424 common bandwidth extraction parameters based on different audio channel signals 412, 416, which are supported by different multi-channel encoders 450, 460 in the first processing step. This specific processing order has the advantage that sets 422, 424 of the bandwidth extension parameters are based on channels that are connected only in the second stage of hierarchical coding (i.e. in multi-channel encoder 470). This is advantageous because it is desirable to combine such audio channels in the first stage of hierarchical coding, the relationship of which is not very important in relation to the perception of the positions of the sound source. Rather, it is recommended that the relationship between the first downmix signal and the second downmix signal mainly determines the perception of the location of the audio source, since the relationship between the first downmix 452 signal and the second downmix 462 signal can be better preserved than the relationship between individual signals 410, 412, 414, 416 audio. In other words, turned out, that it is desirable that the first set of 422 common bandwidth extension parameters is based on two audio channels (audio channel signals), which contribute to various 452 downmix signals, 462, and that the second set of common bandwidth extension parameters 424 is provided based on signals 412, 416 audio channel which also contribute to various 452 downmix signals, 462, which is achieved by the processing of audio channel signals described above in hierarchical multi-channel coding. Consequently, the first set of common bandwidth extension parameters 422 is based on a similar channel relation compared to the channel relationship between the first downmix signal 452 and the second downmix signal 462, the latter usually dominating the spatial impression generated on the audio decoder side. Accordingly, providing the first set of bandwidth extension parameters 422, as well as providing the second set of bandwidth extension parameters 424 is well suited to the spatial listening experience that is generated on the side of the audio decoder.
5. Audio decoder according to Fig. 5 [0066] Fig. 5 shows an essential block diagram of an audio decoder, according to another embodiment of the present invention. The audio decoder according to Fig. 5 is marked entirely with 500.
[0067] The audio decoder 500 is adapted to receive a jointly coded representation of a first downmix signal 510 and a second downmix signal. In addition, the audio decoder 500 is adapted to provide the first extended bandwidth channel signal 520, the second extended bandwidth channel signal 522, the extended bandwidth channel signal 524 and the fourth extended bandwidth channel signal 526.
[0068] The audio decoder 500 includes a (first) multi-channel decoder 530 that is adapted to provide a first downmix signal 532 and a second downmix signal 534 based on a jointly coded representation of the first downmix signal 510 and the second downmix signal by multi-channel decoding. The audio decoder 500 also includes a (second) multi-channel decoder 540 that is adapted to provide at least a first audio channel signal 542 and a second audio channel signal 544 based on the first downmix signal 532 using multi-channel decoding. The audio decoder 500 also includes a (third) multi-channel decoder 550, which is adapted to provide at least a third audio channel signal 556 and a fourth audio channel signal 558 based on a second downmix signal 544 using multi-channel decryption. In addition, the 500 audio decoder includes (first) multi-channel 560 bandwidth extension, which is adapted to perform multi-channel bandwidth expansion based on the first audio channel 542 signal and the third audio channel 556 signal, to obtain the first channel signal 520 with extended bandwidth and the third channel signal 524 with extended bandwidth In addition, the audio decoder includes a (second) multi-channel 570 bandwidth extension, which is adapted to perform multi-channel bandwidth extension based on the second audio channel signal 544 and the fourth audio channel signal 558, to obtain a second channel signal 522 with extended bandwidth and a fourth channel signal 526 with extended bandwidth.
[0069] Regarding the functionality of the audio decoder 500, it should be noted that the 500 audio decoder performs hierarchical multi-channel decoding, wherein the separation between the first downmix signal 532 and the second downmix signal 534 is performed in a first stage of the hierarchical decoding process and wherein the first audio channel signal 542 and the second audio channel signal 544 are derived from the first downmix signal 532 in the second hierarchical decoding stage, and wherein the third audio channel signal 556 and the fourth audio channel signal 558 are derived from the second downmix signal 550 in the second hierarchical decoding step. However, both the first multi-channel bandwidth extension 560 and the second multi-channel bandwidth extension 570 receive each one audio channel signal that originates from the first downmix signal 532 and one audio channel signal that originates from the second downmix signal 534. Because better channel separation is usually achieved by (first) 530 multi-channel decoding, which is performed as the first stage of hierarchical multi-channel decoding, compared to the second stage of hierarchical decoding, you will notice that every 560 multi-channel extension, 570 bandwidth receives input signals, which are well separated (because they come from the first downmix 532 signal and the second downmix 534 signal, which are well separated on the channels). Thus, the multi-channel extension 560, 570 of bandwidth may take into account stereo characteristics that are important for the listening experience, and which are well represented by the relationship between the first downmix signal 532 and the second downmix signal 534, and therefore may provide a good hearing impression.
[0070] In other words, the "cross" structure of the audio decoder in which each of the steps 560, 570 of multi-channel bandwidth extension receives input signals from both (second stage) of multi-channel decoders 540, 550 allows for a good multi-channel bandwidth extension that takes into account the relationship stereo between channels.
[0071] It should be noted, however, that the audio decoder 500 can be supplemented with any of the functions and functions described herein with respect to the audio decoders of the Figures. 2, 3, 6 and 13, in which it is possible to introduce individual features to the 500 audio decoder to gradually improve the performance of the audio decoder.
6. Audio decoder according to Fig. 6 [0072] In Fig. 6 shows the basic block diagram of an audio decoder according to another embodiment of the present invention. Audio decoder according to Fig. 6 it is marked entirely with 600. Audio decoder 600 according to Fig. 6 is similar to the 500 audio decoder according to Fig. 5, so that the above explanations also apply. However, the 600 audio decoder has been supplemented with some features and functionalities that can also be added, alone or in combination, to the 500 audio decoder for improvement.
[0073] The audio decoder 600 is adapted to receive a jointly coded representation 610 of the first downmix signal and the second downmix signal and to provide the first signal 620 with extended bandwidth, the second signal 622 with extended bandwidth, the third signal 624 with extended bandwidth and the fourth signal 626 with extended bandwidth. The audio decoder 600 includes a multi-channel decoder 630 that is adapted to receive a jointly coded representation 610 of the first downmix signal and the second downmix signal, and to provide, based on this first downmix signal 632 and the second downmix signal 634. The audio decoder 600 further includes a multi-channel decoder 640 that is adapted to receive a first downmix signal 632 and to provide, based on this first audio channel signal 542 and a second audio channel signal 544. The audio decoder 600 also includes a multi-channel decoder 650 that is adapted to receive a second downmix signal 634 and provide a third audio channel signal 656 and a fourth audio channel signal 658. The audio decoder 600 also includes a (first) multi-channel bandwidth extension 660 that is adapted to receive the first audio channel signal 642 and the third audio channel signal 656 and to provide, based on this first channel signal 620 with extended bandwidth and a third channel signal 624 with extended bandwidth. In addition, the (second) multi-channel bandwidth extension 670 receives the second audio channel signal 644 and the fourth audio channel signal 658 and provides the second channel signal 622 with extended bandwidth and the fourth channel signal 626 with extended bandwidth based thereon.
[0074] The audio decoder 600 also includes another multi-channel decoder 680, which is adapted to receive a jointly coded representation of the first residual signal 682 and the second residual signal, and which based on this provides the first residual signal 684 for use by the multi-channel decoder 640 and the second residual signal 686 for use by the 650 multi-channel decoder.
[0075] The multi-channel decoder 630 is preferably a residual signal-based multi-channel decoder based on prediction. For example, multi-channel decoder 630 may be substantially identical to multi-channel decoder 370 described above. For example, the 630 multi-channel decoder may be a composite USAC stereo predictive decoder as mentioned above and as described in the above USAC standard. Accordingly, the jointly coded representation of the first downmix signal and the second downmix signal may for example comprise the (common) downmix signal of the first downmix signal and the second downmix signal, the (common) residual signal of the first downmix signal and the second downmix signal, and one or more prediction parameters that are rated by a 630 multi-channel decoder.
[0076] It should further be noted that the first downmix signal 632 may for example be associated with the first horizontal position or azimuth position (e.g. left horizontal position) of the audio scene and that the second downmix signal 634 may for example be associated with the second horizontal position or an azimuth position (e.g., right horizontal position) of the audio scene.
[0077] Furthermore, the multi-channel decoder 680 may, for example, be a prediction-based multi-channel decoder associated with the residual signal. The multi-channel decoder 680 may be substantially identical to the multi-channel decoder 330 described above. For example, the 680 multi-channel decoder may be a composite USAC stereo predictive decoder as mentioned above. Consequently, the jointly coded representation of the first residual signal and the second residual signal 682 may include a (common) downmix signal of the first residual signal and the second residual signal, (common) residual signal of the first residual signal and the second residual signal, and one or more prediction parameters that are rated by 680 multi-channel decoder. In addition, it should be noted that the first residual signal 684 can be associated with the first horizontal or azimuth position (e.g., left horizontal position) of the audio scene and that the second residual signal 686 can be associated with the second horizontal residual signal or azimuth position (e.g. right horizontal position) of the audio scene.
[0078] Multi-channel decoder 640 may, for example, be parameter-based multi-channel decoding, such as, for example, MPEG surround multi-channel decoding as described above and in the referenced standard. However, in the presence of the (optional) multi-channel decoder 680 and the (optional) first residual signal 684, the multi-channel decoder 640 may be a residual signal-supported multi-channel decoder based on a decoder parameters such as, for example, a unified stereo decoder. Thus, multi-channel decoder 640 may be substantially identical to multi-channel decoder 340 described above, and multi-channel decoder 640 may, for example, obtain parameters 342 described above.
[0079] Similarly, multi-channel decoder 650 can be substantially identical to multi-channel decoder 640. Accordingly, multi-channel decoder 650 can, for example, be parameter-based and can optionally be assisted by a residual signal (in the presence of the optional multi-channel decoder 680).
[0080] Furthermore, it should be noted that the first audio channel signal 642 and the second audio channel signal 644 are preferably associated with vertically adjacent spatial positions of the audio scene. For example, the first audio channel signal 642 is associated with the lower left position of the audio scene and the second audio channel signal 644 is associated with the upper left position of the audio scene. Accordingly, multi-channel decoder 640 performs vertical splitting (or separating or distributing) of audio content described by the first downmix signal 632 (and, optionally, by the first residual signal 684). Similarly, the third audio channel signal 656 and the fourth audio channel signal 658 are associated with vertically adjacent positions of the audio scene and are preferably associated with the same horizontal position or azimuth of the audio scene. For example, the third audio channel signal 656 is preferably associated with the lower right position of the audio scene, and the fourth audio channel signal 658 is preferably associated with the upper right position of the audio scene. Thus, multi-channel decoder 650 performs the vertical separation (or separation or distribution) of audio content described by the second downmix signal 634 (and, optionally, the second residual signal 686).
[0081] However, the first multi-channel extension 660 receives the first audio channel signal 642 and the third audio channel 656 that are associated with the lower left position and lower right position of the audio scene. Accordingly, the first multi-channel bandwidth extension 660 performs the multi-channel bandwidth extension based on two audio channel signals that are associated with the same horizontal plane (e.g., lower horizontal plane) or the height of the audio scene and different sides (left / right) of the audio scene. Accordingly, multi-channel bandwidth extension may take into account stereo characteristics (e.g., human stereo perception) when performing bandwidth extension. Similarly, the second multi-channel bandwidth extension 670 may also take into account stereo characteristics, because the second multi-channel bandwidth extension works on audio channel signals of the same horizontal plane (e.g., horizontal top plane) or height, but in different horizontal positions (different sides) (left) / right) audio scenes.
[0082] To further state, the hierarchical audio decoder 600 comprises a structure in which the left / right (or chapter or distribution) division is performed in a first step (630, 680 multi-channel decoding), the vertical division (distribution or distribution) being carried out in the second stage (multi-channel 640, 650 decoding), and wherein the multi-channel bandwidth extension operates on the left / right pair of signals (multi-channel bandwidth extension 660, 670). This "intersection" of decoding paths allows that left / right separation, which is particularly important for the listening experience (for example, more important than the upper / lower split), can be performed at the first stage of the hierarchical audio decoder processing and that the multi-channel width extension bands can also be made on a pair of left-right audio signals, which again give a particularly good listening impression. The upper / lower split is performed as an intermediate state between left-right separation and multi-channel bandwidth extension that allows the output of four audio channel signals (or extended bandwidth channel signals) without significantly reducing the listening experience.
7. Method of Fig. 7 [0083] Fig. 7 shows a flowchart of a method 700 for providing a coded representation based on at least four audio channel signals.
[0084] The method 700 comprises encoding 710 at least the first audio channel signal and the second audio channel signal using the residual signal multi-channel coding to obtain the first downmix signal and the first residual signal. The method also includes joint coding
720 at least a third audio channel signal and a fourth audio channel signal using a residual signal assisted multi-channel coding to obtain a second downmix signal and a second residual signal. The method further includes co-coding 730 the first residual signal and the second residual signal using multi-channel coding to obtain an encoded representation of the residual signals. However, it should be noted that method 700 can be complemented by any of the features and functionalities described herein with respect to audio encoders and audio decoders.
8. Method of Fig. 8 [0085] Fig. 8 is a flowchart of a method 800 for providing at least four audio channel signals based on an encoded representation.
[0086] Method 800 includes providing 810 a first residual signal and a second residual signal based on a jointly encoded representation of the first residual signal and the second residual signal using multi-channel decoding. The method 800 also includes providing 820 the first audio channel signal and the second audio channel signal based on the first downmix signal and the first residual signal using residual signal assisted multi-channel decoding. The method also includes providing 830 the third audio channel signal and the fourth audio channel signal based on the second downmix signal and the second residual signal using residual signal assisted multi-channel decoding.
[0087] In addition, it should be noted that the method 800 may be supplemented with any of the features and functionalities described herein with respect to audio decoders and audio encoders.
9. Method of Fig. 9 [0088] Fig. 9 is a flowchart of a method 900 for providing a coded representation based on at least four audio channel signals.
[0089] The method 900 includes obtaining 910 a first set of common bandwidth extension parameters based on the first audio channel signal and the third audio channel signal. The method 900 also includes obtaining 920 a second set of common bandwidth extension parameters based on the second audio channel signal and the fourth audio channel signal. The method also includes co-coding at least a first audio channel signal and a second audio channel signal using multi-channel coding to obtain a first downmix signal and common coding 940 of at least a third audio channel signal and a fourth audio channel signal using multi-channel coding to obtain a second downmix signal. The method also includes jointly encoding 950 the first downmix signal and the second downmix signal using multi-channel coding to obtain an encoded representation of the downmix signals.
[0090] It should be noted that some of the steps of the method 900 that do not contain k correlation can be performed in any order or in parallel. In addition, it should be noted that method 900 may be complemented by any of the features and functionalities described herein with respect to audio encoders and audio decoders.
10. Method of Fig. 10 [0091] Fig. 10 shows a flowchart of a method 1000 for providing at least four audio channel signals based on an encoded representation.
[0092] The method 1000 includes providing 1010 a first downmix signal and a second downmix signal based on a jointly coded representation of the first downmix signal and the second downmix signal using multi-channel decoding, providing 1020 at least a first audio channel signal and a second audio channel signal based on the first downmix signal using multi-channel decoding, providing 1030 at least a third audio channel signal and a fourth audio channel signal based on a second downmix signal using multi-channel decoding, performing 1040 multi-channel bandwidth extension based on the first audio channel signal and the third audio channel signal, to obtain the first channel signal with extended bandwidth and the third signal with extended bandwidth, and performing 1050 multi-channel bandwidth extension based on the second audio channel signal and the fourth audio channel signal, to obtain a second extended bandwidth channel signal and a fourth extended bandwidth channel signal.
[0093] It should be noted that some of the steps of the method 1000 may be performed in parallel or in a different order. In addition, it should be noted that method 1000 can be supplemented with any of the features and functionalities described herein with respect to the audio encoder and audio decoder.
11. Embodiments of Figures 11, 12 and 13 [0094] Some additional embodiments of the present invention and underlying considerations will be described below.
[0095] Fig. 11 shows a schematic block diagram of an audio encoder 1100 in accordance with an embodiment of the invention. The 1100 audio encoder is adapted to receive the lower left channel signal 1110, upper left channel signal 1112, lower right channel signal 1114 and upper right channel signal 1116.
[0096] The audio encoder 1100 includes a first multi-channel audio encoder (or encoding) 1120 which is an MPEG 2-1-2 surround audio encoder (or encoding) or a unified stereo audio encoder (or encoding) and which receives the lower left 1111 signal and upper left channel signal 1112. The first multi-channel 1120 audio encoder provides a left downmix signal 1122 and an optional left residual signal 1124. In addition, the 1100 audio encoder includes a second multi-channel encoder (or encoding) 1130 that is an MPEG-surround 2-1-2 encoder (or encoding) or a unified stereo encoder (or encoding) that receives the lower right channel 1114 signal and the right 1116 signal upper channel. The second multi-channel 1130 audio encoder provides the right 1132 downmix signal and, optionally, the right 1134 residual signal. The 1100 audio encoder also includes a stereo (or encoding) 1140 encoder that receives the left downmix signal 1122 and the right downmix signal 1132. Furthermore, the first 1140 stereo coding, which is a complex predictive stereo coding ( complex prediction stereo coding), receives information about the 1142 psychoacoustic model from the psychoacoustic model. For example, information 1142 regarding the psycho model may describe the psychoacoustic meaning of different frequency bands or frequency subbands, psychoacoustic masking effects and the like. 1140 stereo coding provides a "downmix" channel pair element (CPE) which is denoted by 1144 and which describes the left downmix signal 1122 and the right 1132 downmix signal 1132 in co-coded form. Additionally, the 1100 audio encoder optionally includes a second stereo (or encoding) 1150 encoder that is adapted to receive an optional left residual signal 1124 and an optional right residual signal 1134 as well as psychoacoustic model information 1142. The second stereo coding 1150, which is a composite predictive stereo coding, is adapted to provide a "residual" channel pair (CPE) element that represents left residual signal 1124 and right residual signal 1134 in co-coded form.
[0097] The 1100 encoder (as well as other audio encoders described herein) is based on the assumption that the horizontal and vertical signal dependencies are used by a hierarchical combination of available stereo USAC tools (i.e., coding concepts that are available in USAC coding) . Pairs of vertically adjacent channels are combined using MPEG 2-1-2 surround or unified stereo (designated 1120 and 1130) with limited bandwidth or full band residual signal (designated 1124 and 1134). The output of each pair of vertical channel pair channels is the downmix signal 1122, 1132, and for the unified stereo residual signal 1124, 1134. To meet perceptual requirements for binaural mask removal (ang. binaural unmasking), both downmix signals 1122, 1132 are horizontally combined and jointly coded using a complex prediction (encoder 1140) in the MDCT domain that includes left-right and center-side coding. The same method can be applied to horizontally connected residual signals 1124, 1134. This concept is presented in Fig. 11.
[0098] The hierarchical structure explained with reference to Fig. 11 can be achieved by using both stereo tools (for example, both stereo tools
USAC) and re-distributing channels between them. Hence, no additional pre / post processing step is required, and the bit stream syntax for data transmission of useful tools remains unchanged (e.g., substantially unchanged compared to the USAC standard). This idea leads to the encoder structure shown in Fig.
12.
[0099] Fig. 12 shows a schematic block diagram of an audio encoder 1200 in accordance with an embodiment of the invention. The audio signal encoder 1200 is adapted to receive the first channel signal 1210, the second channel signal 1212, the third channel signal 1214 and the fourth channel signal 1216. The audio encoder 1200 is adapted to provide a 1220 bit stream for the first channel pair element and a 1222 bit stream for the second channel pair element.
[0100] The audio encoder 1200 includes a first multi-channel encoder 1230, which is an MPEG-surround 2-1-2 encoder or a unified stereo encoder, and which receives the first channel signal 1210 and the second channel signal 1212. In addition, the first multi-channel 1230 encoder provides the first downmix 1232 signal, 1236 MPEG surround useful data and, optionally, the first 1234 residual signal. The 1200 audio encoder also includes a second 1240 multi-channel encoder, which is a 2-1-2 MPEG encoder or a unified stereo encoder that receives the third channel 1214 signal and the fourth channel 1216 signal. The second multi-channel encoder 1240 provides the first downmix signal 1242, the 1246 MPEG surround signal and, optionally, the second 1244 residual signal.
[0101] The audio encoder 1200 also includes a first stereo coding 1250, which is a composite predictive stereo coding. The first stereo 1250 encoding receives the first downmix signal 1232 and the second downmix signal 1242. The first stereo coding 1250 provides a jointly coded representation of the first downmix signal 12322 and the second downmix signal 1242, wherein the jointly coded representation 1252 may include a representation of the (common) downmix signal (first downmix signal 1232 and the second downmix signal 1232). second downmix signal 1242) and common residual signal (first downmix signal 1232 and second downmix signal 1242). In addition, (first) complex predictive 1250 stereo coding provides complex predictive user data 1254, which typically includes one or more complex prediction coefficients. In addition, the 1200 audio encoder also includes a second 1260 stereo encoding which is a composite predictive stereo encoding. The second stereo coding 1260 receives the first residual signal 1234 and the second residual signal 1244 (or zero input values if there is no residual signal provided by multi-channel encoders 1230, 1240). The second stereo coding 1260 provides a jointly coded representation of 1262 the first residual signal 1234 and the second residual signal 1244, which may, for example, include a (common) downmix signal (first residual signal 1234 and second residual signal 1244) and a common residual signal (first signal 1234 residual and second residual signal 1244). In addition, 1260 stereo predictive coding provides complex predictive 1264 user data that typically includes one or more prediction coefficients.
[0102] In addition, the audio encoder 1200 includes a psychoacoustic model 1270 that provides information that controls the first composite stereo predictive 1250 encoding and the second composite predictive stereo 1260 encoding. For example, the information provided by the psychoacoustic model 1270 may describe which frequency bands or containers ( bins) frequencies are of great psychoacoustic importance and should be coded with high accuracy. However, it should be noted that the use of information provided by the 1270 psychoacoustic model is optional.
[0103] Additionally, the audio signal encoder 1200 includes a first encoder and a 1280 multiplexer that receives jointly coded representation 1252 from the first 1250 stereo composite encoding, 1254 predictive composite user data from the first 1250 stereo composite encoding of 1236 MPEG surround user data from the first 1230 multi-channel encoder audio 1230. In addition, the first 1280 encoding and multiplexing can receive information from the 1270 psychoacoustic model, which describes, for example, which coding precision should be applied to which frequency band or frequency subbands, taking into account the effects of psychoacoustic masking and the like. Accordingly, the first 1280 encoding and multiplexing provides the first 1220 bit stream of the first channel pair element.
[0104] Additionally, the audio signal encoder 1200 includes a second coding and multiplexing 1290 that is adapted to receive a co-coded representation 1262 provided by the second composite predictive 1260 stereo coding, the composite predictive user data 1264 provided by the second composite predictive 1260 stereo coding, and the data 1240 usable MPEG surround provided by a second 1240 multi-channel audio encoder. In addition, the second 1290 coding and multiplexing may receive information from the 1270 psychoacoustic model. Accordingly, the second coding and multiplexing 1290 provides a stream of 1222 bits of the second channel pair element.
[0105] Referring to the functionality of the audio encoder 1200, reference should be made to the above explanations as well as to the explanations regarding the audio encoders according to Figs. 2, 3, 5 and 6.
[0106] In addition, it should be noted that this concept can be extended to the use of multiple MPEG surround boxes for joint coding of horizontal, vertical or other geometrically related channels and combining downmix and debris signals with complex predictive stereo pairs, considering their geometric and perceptive properties. This leads to a generalized decoder structure.
[0107] The following describes the implementation of the four channel element. The three-dimensional audio coding system uses a hierarchical combination of four channels to form a four-channel element (QCE). QCE consists of two elements of the USAC channel pair (CPE) (or provides two elements of the USAC channel pair or receives elements of the USAC channel pair). Pairs of vertical channels are combined using MPS 2-1-2 or unified stereo. Downmix channels are coded together in the first element of the channel pair, CPE. If residual coding is used, the residual signals are jointly coded in the second element of the CPE channel pair, otherwise the signal in the second CPE is set to zero. Both elements of CPE channel pairs use complex prediction for common stereo coding, including left-right and center-side coding. To preserve the perceptive stereo properties of the high frequency signal part, stereo SBR (spectral band replication) is used between the upper left / right pair of channels and the lower left / right pair of channels, an additional step protecting against the use of SBR.
[0108] A possible decoder structure will be described with reference to Fig. 13, which shows a schematic block diagram of an audio decoder according to an embodiment of the invention. The audio decoder 1300 is adapted to receive a first stream 1310 bit representing the first element of the channel pair and a second stream 1312 bit representing the second element of the channel pair. However, the first 1310 bit stream and the second 1312 bit stream may be included in the common overall bit stream.
[0109] The audio decoder 1300 is adapted to provide the first channel signal 1320 with extended bandwidth, who can e.g, represent the bottom left position of the audio scene, a second channel signal 1322 with extended bandwidth, who can e.g, show the top left position of the audio scene, third channel signal 1324 with extended bandwidth, who can e.g, be associated with the lower right position of the audio scene and the fourth channel signal 1326 with extended bandwidth, who can e.g, be associated with the top right position of the audio scene.
[0110] The audio decoder 1300 includes decoding 1330 the first bit stream, which is adapted to receive a stream 1310 bit for the first channel pair element and to provide, based thereon, a jointly encoded representation of two downmix signals, composite predictive data 1334 useful, data 1336 useful MPEG surround and 1338 data useful for replicating spectrum bandwidth. The audio decoder 1300 also includes a first composite predictive 1340 stereo decoding that is adapted to receive a jointly coded representation 1332 and composite useful data 1334 and to provide, based on this first downmix signal 1342 and the second downmix signal 1344. Similarly, the 1300 audio decoder includes a second 1350-bit stream decoding that is adapted to receive a 1312-bit stream for the second channel element and to provide, based on them, a jointly encoded representation of 1352 two residual signals, composite predictive useful data 1354 useful data 1356 MPEG useful data surround and content 1358 bits replicate viewband width. The audio decoder also includes a second composite predictive 1360 stereo decoding that provides the first 1362 residual signal and the second 1364 residual signal based on the jointly coded representation of 1352 and the composite predictive user data 1354.
[0111] In addition, the 1300 audio decoder includes first MP70 multi-channel decoding of the surround type that is MPEG 2-1-2 surround decoding or unified stereo decoding. The first 1370 MPEG multichannel surround decoding receives the first downmix signal 1342, the first residual 1362 signal (optional) and the 1336 MPEG surround user data and provides, based on this, the first 1372 audio channel signal and the second 1374 audio channel signal. The 1300 audio decoder also includes second MPEG surround 1380 multi-channel decoding, which is 2-1-2 MPEG multi-channel decoding or unified stereo multi-channel decoding. The second 1380 MPEG surround type decode receives the second downmix 1344 signal and the second residual 1364 signal (optional) as well as the 1356 useful MPEG surround data and provides, based on them, the third 1382 audio channel signal and the fourth 1384 audio channel signal. The audio decoder 1300 also includes a first spectral replication 1390 of the stereo spectrum width that is adapted to receive a first audio channel signal 1372 and a third audio channel signal 1382 as well as data 1338 useful spectral bandwidth replication and to provide a first signal 1320 therefrom. channel with extended bandwidth and third signal 1324 with extended bandwidth. In addition, the audio decoder includes a second spectral replication 1394 stereo bandwidth that is adapted to receive a second audio channel 1374 signal and a fourth audio channel 1384 signal as well as data 1358 useful spectral bandwidth replication and to provide a second 1322 signal based thereon channel with extended bandwidth and fourth signal 1326 with extended bandwidth.
[0112] Regarding the functionality of the 1300 audio decoder, reference should be made to the above discussion as well as to the discussion of the audio decoder according to Figs. 2, 3, 5 and 6.
[0113] Hereinafter, an example of a bit stream that can be used to encode / decode audio described herein with reference to Fig. 14a and 14b. It should be noted that the bit stream may be, for example, the bit stream extension used in the unified speech and audio coding (USAC), which is described in the above-mentioned standard (ISO / IEC 23003-3: 2012). For example, data 1236, 1246, 1336, 1356 useful MPEG surround and composite predictive data 1254, 1264, 1334, 1354 useful can be transmitted as for older elements of the channel pair (i.e. for channel pair elements in accordance with the USAC standard). To signal the use of the QCE four channel element, the configuration of the USAC channel pair can be extended by two bits as shown in Fig. 14a. In other words, two bits marked "qcelndex" can be added to the USAC bit stream element "UsacChannelPairElementConfig ()". The meaning of the parameter represented by the "qcelndex" bits can be defined, for example, as shown in the table in Fig. 14b.
[0114] For example, two channel pair elements that make up the QCE may be transmitted as successive elements, first a CPE containing downmix channels and content
MPS for the first MPS field, secondly a CPE containing the residual signal (or zero audio signal for MPS 2 -1-2 coding) and the MPS content for the second MPS field.
[0115] In other words, there is only a small signaling overhead compared to the conventional USAC bit stream for transmitting the QCE four channel element.
[0116] However, of course, various bit stream formats can also be used.
12. Coding / decoding environment [0117] In the following, an audio coding / decoding environment in which the concepts of the present invention can be applied will be described.
[0118] The 3D audio codec system in which the concepts of the present invention can be used is based on the USEG MPEG-D codec for decoding channel and object signals. To increase the encoding efficiency of a large number of objects, MPEG SAOC technology has been adapted. Three types of rendering modules perform the tasks of rendering objects to channels, rendering channels to headphones, or rendering channels to a different speaker configuration. When object signals are explicitly transmitted or parametrically coded using SAOC, the corresponding object metadata information is compressed and multiplexed to the 3D audio bit stream.
[0119] Fig. 15 is a basic block diagram of such an audio encoder and Fig. 16 is a basic block diagram of such an audio decoder. In other words, Figs. 15 and 16 show different algorithm blocks of the 3D audio system.
[0120] Referring now to Fig. 15, which shows the basic block diagram of the 3D 1500 encoder, some details will be explained. Encoder 1500 includes an optional pre-rendering module / mixer 1510 that receives one or more 1512 channel signals and one or more 1514 object signals and provides, on the basis thereof, one or more 1516 channel signals as well as one or more 1518, 1520 object signals . The audio encoder also includes the USAC 1530 encoder and optionally the SAOC 1540 encoder. The SAOC 1540 encoder is adapted to provide one or more SAO 1542 transport channels and SAOC 1544 side information based on one or more 1520 objects provided to the SAOC encoder. In addition, the USAC 1530 encoder is adapted to receive channel signals 1516 comprising channels and pre-rendered objects from the pre-render module / mixer to receive one or more object signals 1518 from the pre-render module / mixer and receive one or more SAOC 1542 transport channels and side information SAOC 1544, and based on them provide a 1532 encoded representation. Furthermore, the audio encoder 1500 also includes an object metadata encoder 1550 that is adapted to receive metadata of the 1552 object (which can be evaluated by the pre-rendering module // mixer 1510) and to encode the object metadata to obtain the encoded object metadata 1554. The encoded metadata is also received by the USAC 1530 encoder and used to provide the 1532 encoded representation.
[0121] Some details of individual components of the audio encoder 1500 will be described below.
[0122] Referring now to Fig. 16, an audio decoder 1600 will be described. The audio decoder 1600 is adapted to receive a coded representation 1610 and to provide, based on this, multi-channel speaker signals 1612, headphone signals 1614 and / or speaker signals 1616 in an alternative format (for example, in 5.1 format).
[0123] The audio decoder 1600 includes a USAC decoder 1620 and provides one or more channel signals 1622, one or more pre-rendered object signals 1624, one or more object signals 1626, one or more transport channels 1628 SAOC, side information 1602 SAOC and compressed information 1632 object metadata based on the 1610 encoded representation. Audio decoder 1600 also includes an object rendering module 1640 that is adapted to provide one or more rendered signals of object 1642 based on object signal 1626 and object 1644 information metadata, wherein object information 1644 metadata is provided by object 1650 metadata decoder based on information about 1632 compressed metadata. The audio decoder 1600 also includes, optionally, a SAOC 1660 decoder, which is adapted to receive the SAOC 1628 transport channel and SAOC 1630 side information, and to provide, based on this one or more rendered signals of the object 1662. Audio decoder 1600 also includes a 1670 mixer that is adapted to receive channel 1622 signals, pre-rendered object signals 1624, rendered object signals 1642, and rendered object signals 1662, and to provide, based on this, a plurality of mixed channel 1672 signals that may an example is 1612 multi-channel speaker signals. Audio decoder 1600 may, for example, also include a binaural 1680 signal that is adapted to receive mixed channel 1672 signals and to provide headphone signals 1614 therefrom. In addition, the audio decoder 1600 may include a 1690 format conversion that is adapted to receive mixed channel signals 1672 and distribution information 1692 and to provide, based on this, a loudspeaker signal 1616 for an alternative loudspeaker position.
[0124] Some details of the audio encoder elements will be described below
1500 and 1600 audio decoder.
Pre-rendering module / mixer [0125] Pre-rendering module 1510 can optionally be used to convert the channel plus the scene of inserting the object into the channel scene before coding. It can functionally be, for example, identical to the rendering module / object mixer described below. Pre-rendering objects can, for example, provide a deterministic entropy of the signal at the encoder input that is substantially independent of the number of simultaneously active object signals. Pre-rendering objects does not require the transmission of object metadata. Discrete object signals are rendered to the channel layout whose encoder is adapted for use. The object weights for each channel are derived from the associated object metadata (OAM) 1552.
USAC Core Codec [0126] Core coders 1530, 1620 for loudspeaker signals, discrete object signals, object downmix signals, and pre-rendered signals are based on USEG MPEG-D technology. It supports multi-coding, creating channel and object mapping information based on geometric and semantic information about the input channel and object assignment. This mapping information describes how input channels and objects are mapped to USAC channel elements (CPE, SCE, LFE) and the corresponding information is sent to the decoder. All additional content, such as SAOC data or object metadata, was passed on through extension elements and included in the encoder speed control.
[0127] Object encoding is possible in various ways, depending on the speed / distortion requirements and interactivity requirements for the rendering module. The following variants of object encoding are possible:
1. Pre-rendered objects: Object signals are pre-rendered and mixed with 22.2 channel signals before coding. The next coding chain sees 22.2 channel signals.
2. Discrete object wave shapes: objects are delivered in the form of monophonic waves to the encoder. The encoder uses SCE elements of single channel modules to transfer objects in addition to channel signals. Decoded objects are rendered and mixed on the receiver side. The compressed metadata of the object is sent to the receiver / renderer along the side.
3. Parametric object wave shapes: object properties and their interrelationships are described using SAOC parameters. The object's downmix is encoded using USAC. Parametric information is sent along the side. The number of downmix channels is selected depending on the number of objects and the total data rate. Information about the metadata of compressed objects is sent to the SAOC renderer.
SAOC [0128] The SAOC 1540 encoder and SAOC 1660 decoder for object signals are based on MPEG SAOC technology. The system is able to play, modify and render many audio objects based on a smaller number of transmitted channels and additional parametric data (OLD differences at the object level, correlations between IOC objects, DMG gains from downmix). Additional parametric data show a much lower data rate than required for the transmission of individual objects, making coding very efficient. The SAOC encoder accepts object / channel signals as mono shapes as input and sends parametric information (which is packaged to 3Daudio 1532, 1610 bit stream) and SAOC transport channels (which are encoded using single channel elements and transmitted).
[0129] The SAOC decoder 1600 reconstructs object / channel signals from SAOC 1628 decoded transport channels and parametric information 1630 and generates an output audio scene based on the reproduction system, information about the decompressed object metadata and optional information about user interaction.
Metadata Object Codec [0130] For each object, the associated metadata that determine the geometric position and volume of the object in 3D space are effectively encoded by quantizing the object's properties in time and space. The compressed metadata of the cOAM 1554, 1632 is sent to the receiver as side information.
Object rendering module / Mixer [0131] The object rendering module uses metadata of the compressed object to generate waveforms of objects according to a given reproduction format. Each object is rendered to specific output channels according to its metadata. The result of this block results from the sum of partial results. If both channel-based content and discrete / parametric objects are decoded, the channel-based waveforms and rendered object runs are mixed before outputting the resulting waveforms (or before passing them to a postprocessor module such as a binaural rendering module or speaker rendering module).
Binaural Rendering Module [0132] Binaural Rendering Module 1680 produces a binaural downmix of multi-channel audio material such that each input channel is represented by a virtual sound source. Processing takes place in a frame in the field of QMF. Binaural is based on measured binaural impulse responses in the room.
Speaker rendering module / format conversion [0133] The speaker rendering module 1690 performs a conversion between the configuration of the transmitted channel and the desired playback format. Therefore, it is called "format converter". The format converter performs conversions to a smaller number of output channels, i.e. creates downmixes. The system automatically generates optimized downmix matrices for a given combination of input and output formats and uses these matrices in the downmix process. The format converter allows for standard speaker configurations as well as for random configurations with non-standard speaker positions.
[0134] Fig. 17 shows the block diagram of the format converter. As can be seen, the 1700 format converter receives mixer output signals 1710, e.g., mixed channel signals 1672, and provides loudspeaker signals 1712, e.g., 1616 loudspeaker signals. The format converter includes a downmix 1720 process in the QMF domain and a downmix configurator 1730, wherein the downmix configurator provides configuration information for the downmix 1720 process based on mixer output 1732 information and reproduction system information 1734.
[0135] In addition, it should be noted that the concepts described above, e.g., audio encoder 100, audio decoder 200 or 300, audio encoder 400, audio decoder 500 or 600, methods 700, 800, 900 or 1000, audio encoder 1100 or 1200, and decoder audio 1300 can be used in audio encoder 1500 and / or in audio decoder 1600. For example, the aforementioned audio encoders / decoders can be used to encode or decode channel signals that are associated with different spatial positions.
13. Alternative embodiments [0136] Some additional embodiments will be described below.
[0167] Referring now to Figs. 18 to 21, additional embodiments of the invention will be explained.
[0138] It should be noted that the so-called Quad Channel (four channel) element (QCE) can be considered as an audio decoder tool that can be used, for example, to decode three-dimensional audio content.
[0139] In other words, the Quad Channel (QCE) element is a method of co-coding four channels for more efficient coding of horizontally and vertically arranged channels. QCE consists of two consecutive CPEs and is created by a hierarchical combination of the Joint Stereo Tool with the Complex Stereo Prediction Tool in the horizontal direction and the MPEG Surround based stereo tool in the vertical direction. This is achieved by enabling both stereo tools and switching the output channels between tools. Stereo SBR is executed in the horizontal direction to maintain high frequency left-hand relations.
[0140] Fig. 18 shows the topological structure of QCE. It should be noted that the QCE of Fig. 18 is very similar to the QCE of Fig. 11, so that reference is made to the above explanations. However, it should be noted that in QCE in Fig. 18 it is not necessary to use a psychoacoustic model when performing complex stereo forecasts (while such use is of course possibly optional). In addition, it can be seen that the first spectral replication of the bandwidth (Stereo SBR) is based on the lower left channel and the lower right channel, and that the second spectral stereo frequency replication (Stereo SBR) is performed based on the upper left channel and upper right channel.
[0141] Some terms and definitions that may apply in some embodiments will be given below.
[0142] The qceIndex data element indicates the QCE CPE mode. For the meaning of the qcelndex bit stream variable, refer to Fig. 14b. Note that qceIndex describes whether two consecutive UsacChannelPairElement () elements are treated as a four-channel element (QCE). The different QCE modes are shown in Fig. 14b. The qceIndex value will be the same for two consecutive elements forming one QCE.
[0143] In the following, some help elements will be defined that can be used in some embodiments of the invention:
cplx_out_dmx_L [] the first channel of the first CPE after complex stereo decoding cplx_out_dmx_R [] the second channel of the first CPE after complex stereo decoding cplx_out_res_L [] the second CPE after complex stereo decoding (zero if qceIndex = 1) cplx_out_res_R [] the second channel of the second CPod decoding zero if qcelndex = 1) mps_out_L_1 [] first output channel of the first MPS field mps_out_L_2 [] second output channel of the first MPS field mps_out_R_1 [] first output channel of the second MPS field mps_out_R_2 [] second output channel of the second MPS field sbr_out_L_1 [] first output channel of the first field Stereo SBR sbr_out_R_1 [] second output channel of the first field Stereo SBR sbr_out_L_2 [] first output channel of the second field Stereo SBR sbr_out_R_2 [] second output channel of the second field Stereo SBR [0144] Below will beexplained decoding process, which is carried out in the embodiment of the present invention.
[0145] The syntax element (or bit stream element or data element) qceIndex in UsacChannelPairElementConfig () indicates whether the CPE belongs to QCE and whether residual encoding is used. In the case where qcelndex is not equal to 0, the current CPE creates QCE along with its next element, which will be a CPE having the same qceIndex. Stereo SBR is always used for QCE, so stereoConfigIndex is 3, and bsStereoSbr will be 1.
[0146] For qceIndex == 1, the second CPE only contains content for MPEG Surround and SBR and no corresponding audio signal data, and the bsResidualCoding syntax element is set to 0.
[0147] The presence of the residual signal in the second CPE is indicated by qceIndex == 2. In this case, the syntax element bsResidualCoding is 1.
[0148] However, various and possible simplified signaling schemes can also be used.
[0149] Common Stereo decoding with the possibility of complex stereo prediction is performed as described in ISO / IEC 23003-3, chapter 7.7. The output of the first CPE are the MPS downmix signals cplx_out_dmx_L [] and cplx_out_dmx_R []. If residual encoding (i.e. qceIndex == 2) is used, the output of the second CPE are MPS residual signals cplx_out_res_L [], cplx_out_res_R [], if no residual signal was passed (i.e. QceIndex == 1), the zero signals are inserted.
[0150] Before applying MPEG Surround decoding, the second channel of the first element (cplx_out_dmx_R []) and the first channel of the second element (cplx out_res_L []) are exchanged.
[0151] MPEG Surround decoding takes place as described in ISO / IEC 230033, section 7.11. If residual coding is used, however, the decoding may be modified compared to conventional MPEG surround decoding in some embodiments. MPEG Surround decode without residual using SBR, as defined in ISO / IEC 23003-3, subsection 7.11.2.7 (Figure 23), has been modified so that Stereo SBR is also used for bsResidualCode == 1, resulting in decoder schemes shown in Fig. 19. In Fig. 19 block diagram of the audio encoder is shown for bsResidualCoding == 0 and bsStereoSbr == 1.
[0152] As can be seen in Fig. 19, the USAC 2010 core decoder provides a downmix signal (DMX) 2012 to the MPS (MPEG Surround) 2020 decoder that provides the first decoded audio signal 2022 and the second decoded audio signal 2024. The SBR 2030 stereo decoder receives the first decoded audio signal 2022 and the second decoded audio signal 2024 and provides, based on this, the left 2032 extended bandwidth audio signal and the right 2034 extended bandwidth audio signal.
[0153] Before using Stereo SBR, the second channel of the first element (mps_out_L_2 []) and the first channel of the second element (mps_out_R_1 []) are exchanged to allow SBR creation on the left and right. After using Stereo SBR, the second output channel of the first item (sbr_out_R_1 []) and the first channel of the second item (sbr_out_L_2 []) are changed again to restore the order of the input channels.
[0154] The structure of the QCE decoder is shown in Fig. 20, which shows the schemes of the QCE decoder.
[0155] It should be noted that the basic block diagram of Fig. 20 is very similar to the schematic block diagram of Fig. 13, so that it also refers to the above explanations. Furthermore, it should be noted that some signal markers have been added in Fig. 20, referring to the definitions in this section. In addition, the final re-sorting of channels that is performed after Stereo SBR is shown.
[0156] Fig. 21 schematically shows a block diagram of a Quad Channel 2200 four channel encoder in accordance with an embodiment of the present invention. In other words, a four-channel encoder (four-channel element), which can be considered as a core encoder tool, is shown in Fig. 21.
[0157] The Quad Channel 2200 four channel encoder includes a first SBR 2210 Stereo which receives the first left channel input 2212 and a second left channel input 2214, and which provides, on this basis, the first SBR 2215 content, the left channel SBR 2216 output signal and the first signal right channel output SBR 2218. In addition, the Quad Channel 2200 encoder includes a second Stereo SBR which receives a second left channel input 2222 and a second right channel input 2224, and which provides, based on this first SBR 2225 content, the first left channel SBR 2226 output signal and the first output signal 2228 SBR of the right channel.
[0158] The Quad Channel 2200 encoder includes the first MPEGSurround 2230 multi-channel encoder (MPS 2-1-2 or Unified Stereo) that receives the first SBR 2226 left channel output signal and the second SBR 2226 left channel output signal, and which provides, based on this , first MPS 2232 content, MPEG Surround downmix signal 2234 of the left channel, and optional residual 22EG Left MPEG Surround signal. The Quad Channel 2200 encoder also includes a second MPEG-Surround 2240 multi-channel encoder (MPS 2-1-2 or Unified Stereo) that receives the 2236 SBR output of the first right channel and the 22R SBR output signal of the second right channel, and which provides, based on of this first MPS 2242 content, MPEG Surround downmix signal 2244 of the right channel and, optionally, residual 22EG MPEG Surround of the right channel.
[0159] The Quad Channel 2200 encoder includes the first complex coding 2250 of a predictive stereo signal that receives the left channel MPEG Surround downmix signal 2234 and the right channel MPEG Surround downmix signal 2244, and which provides, based on this, the complex prediction content 2252 and a jointly coded representation 2254 2234 MPEG Surround downmix signal of the left channel and 2244 MPEG Surround downmix signal of the right channel. The Quad Channel 2200 encoder includes a second complex coding 2260 of a predictive stereo signal that receives a 2236 MPEG Surround left channel signal and a 2246 MPEG Surround right channel signal, and provides complex 2262 predictive content and a jointly coded 2264 representation of the 2236 MPEG Surround downmix signal left channel and MPEG Surround 2246 right channel downmix signal.
[0160] The Quad Channel encoder also includes a first bit stream coding 2270 that receives the jointly coded representation 2254, the composite prediction content 2252m, the MPS 2232 content and the SBR content 2215 and provides, on its basis, a bit stream portion representing an element of the first channel pair. The Quad Channel encoder also includes a 2280 encoding of the second bit stream that receives the jointly coded representation 2264, complex prediction content 2262, MPS content 2242 and SBR content 2225, and provides, based on this, a bit stream portion representing an element of the first channel pair.
14. Implementation alternatives [0161] Although some aspects have been described in the context of the device, it is clear that these aspects also provide a description of the respective method in which the block or device corresponds to the method step or feature of the method step. Similarly, aspects described in the context of a method step also provide a description of the respective block or element or feature of the respective device. Some or all of the method steps may be performed using (or using) a hardware device, such as, for example, a microprocessor, programmable computer or electronic circuit. In some embodiments, one or more of the most important steps of the method can be performed by such a device.
[0162] The encoded audio signal of the invention may be stored on a digital storage medium or may be transmitted on a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.
[0163] Depending on certain implementation requirements, embodiments of the invention can be implemented in hardware or in software. The implementation can be carried out using a digital storage medium, e.g. floppy disks, DVDs, Blu-Ray discs, CDs, ROMs, PROMs, EPROMs, EEPROMs or FLASH memories, with electronically read control signals on it that cooperate (or are capable of to cooperate) with a programmable computer system so that the appropriate way is performed. Therefore, the digital storage medium can be computer readable.
[0164] Some embodiments of the invention include a data carrier having electronically readable control signals that are capable of interacting with a programmable computer system such that one of the methods described herein is performed.
[0165] In general, embodiments of the present invention may be implemented as a computer program product with a program code, wherein the program code operates to perform one of the ways when the computer program product runs on the computer. For example, the program code may be stored on a machine readable medium.
[0166] Other embodiments include a computer program for performing one of the methods described herein, stored on a machine readable medium.
[0167] In other words, an embodiment of the method of the invention is therefore a computer program having a program code for performing one of the methods described herein when the computer program runs on a computer.
[0168] A further embodiment of the methods of the invention is therefore a data carrier (or digital data carrier or computer readable medium) comprising a computer program recorded therein for performing one of the methods described herein. The recording medium, digital recording medium or recorded medium is usually material and / or non-transient.
[0169] Another embodiment of the method of the invention is thus a data stream or a sequence of signals representing a computer program for performing one of the methods described herein. The data stream or signal sequence may, for example, be configured for transmission over a data connection, e.g., via the Internet.
[0170] Another embodiment includes processing means, e.g., a computer or programmable logic device, adapted or configured to perform one of the methods described herein.
[0171] Another embodiment includes a computer on which a computer program is installed to perform one of the methods described herein.
[0172] Another embodiment of the invention includes a device or system adapted to transmit (e.g. electronically or optically) a computer program for carrying out one of the methods described herein to a receiver. The receiver may for example be a computer, mobile device, memory device or the like. The device or system may, for example, comprise a file server for transmitting the computer program to the receiver.
[0173] In some embodiments, the programmable logic device (e.g., programmable gate matrix) may be used to perform some or all of the functionalities of the methods described herein. In some embodiments, the programmable gate matrix may cooperate with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by any hardware device.
[0174] The above described embodiments are merely illustrative of the principles of the present invention. It is understood that the modifications and variants of the systems and details described herein will be apparent to those skilled in the art. Therefore, the intention is to limit only by the scope of the following claims, and not to the specific details provided to describe and explain the present embodiments.
15. Conclusions [0175] Some conclusions will be presented below.
[0176] Embodiments of the invention are based on the assumption that, in order to take into account signal dependencies between vertically and horizontally distributed channels, the four channels may be coded together by a hierarchical combination of common stereo coding tools. For example, vertical channel pairs are combined using MPS 2-1-2 and / or unified stereo with limited bandwidth or full residual coding band. To meet perceptional requirements for binaural unmaskings, the output downmixes are, for example, co-coded using complex predictions in the MDCT domain, which includes left-right and center-side coding. If residual signals are present, they are combined horizontally using the same method.
[0177] Furthermore, it should be noted that the embodiments of the invention overcome some or all of the disadvantages of the prior art. The embodiments according to the invention are adapted to the 3D audio context, the speaker channels being arranged in several height layers, resulting in a pair of horizontal and vertical channels. It was found that joint coding of only two channels, as defined in USAC, is not sufficient to consider spatial and perceptual channel relationships. However, this problem is overcome by the embodiments of the invention.
[0178] Furthermore, conventional MPEG surround is used in an additional pre / post processing step, so that residual signals are transmitted individually without the possibility of joint stereo coding, e.g. to examine the relationship between the left and right root residual signals. In contrast, embodiments of the invention allow efficient encoding / decoding by using such dependencies.
[0179] To further complete, the embodiments of the invention form an encoding and decoding computer device, method or program as described herein.
references:
[0180]
1. [1] ISO / IEC 23003-3: 2012 - information Technology - MPEG Audio Technologies, Part 3: Unified Speech and Audio Coding;
2. [2] ISO / IEC 23003-1: 2007 - Information Technology - MPEG Audio Technologies, Part 1: MPEG Surround
Fraunhofer-Gesellschaft zur Forderung der angewandten Forschung eV, Germany;
Proxy:
EP 3 022 735 B1 Z-16499/17
21 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21
80 members in 19 offices
Priority claims14
| Document | Office | Kind | Date |
|---|---|---|---|
| 13177376 | European Patent Office (EPO) | A | |
| 13177376 | European Patent Office (EPO) | A | |
| 13189305 | European Patent Office (EPO) | A | |
| 13189305 | European Patent Office (EPO) | A | |
| 14739141 | European Patent Office (EPO) | A | |
| 2014064915 | European Patent Office (EPO) | W | |
| 2014064915 | European Patent Office (EPO) | W | |
| 13177376 | – | – | – |
| 13189305 | – | – | – |
| 147391411 | – | – | – |
| EP20130177376 | – | – | – |
| EP20130189305 | – | – | – |
| EP20140739141 | – | – | – |
| WO2014EP64915 | – | – | – |
Members80
| Document | Office | Kind | |
|---|---|---|---|
| EP2830051A2 | European Patent Office (EPO) | A2 | |
| EP2830052A1 | European Patent Office (EPO) | A1 | |
| CA2917770A1 | Canada | A1 | |
| CA2918237A1 | Canada | A1 | |
| WO2015010926A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2015010934A1 | World Intellectual Property Organization (WIPO) | A1 | |
| EP2830051A3 | European Patent Office (EPO) | A3 | |
| TW201514972A | Taiwan Province of China | A | |
| TW201514973A | Taiwan Province of China | A | |
| AR097011A1 | Argentina | A1 | |
| AR097012A1 | Argentina | A1 | |
| SG11201600468SA | Singapore | A | |
| AU2014295282A1 | Australia | A1 | |
| AU2014295360A1 | Australia | A1 | |
| KR20160033777A | Republic of Korea | A | |
| KR20160033778A | Republic of Korea | A | |
| MX2016000939A | Mexico | A | |
| MX2016000858A | Mexico | A | |
| CN105580073A | China | A | |
| CN105593931A | China | A | |
| EP3022734A1 | European Patent Office (EPO) | A1 | |
| EP3022735A1 | European Patent Office (EPO) | A1 | |
| TWI544479B | Taiwan Province of China | B | |
| US2016247508A1 | United States of America | A1 | |
| US2016247509A1 | United States of America | A1 | |
| TWI550598B | Taiwan Province of China | B | |
| US2016275957A1 | United States of America | A1 | |
| JP2016529544A | Japan | A | |
| JP2016530788A | Japan | A | |
| JP6117997B2 | Japan | B2 | |
| ZA201601078B | South Africa | B | |
| BR112016001137A2 | Brazil | A2 | |
| BR112016001141A2 | Brazil | A2 | |
| AU2014295282B2 | Australia | B2 | |
| EP3022734B1 | European Patent Office (EPO) | B1 | |
| RU2016105702A | Russian Federation | A | |
| RU2016105703A | Russian Federation | A | |
| ZA201601080B | South Africa | B | |
| EP3022735B1 | European Patent Office (EPO) | B1 | |
| AU2014295360B2 | Australia | B2 | |
| PT3022734T | Portugal | T | |
| PT3022735T | Portugal | T | |
| ES2649194T3 | Spain | T3 | |
| ES2650544T3 | Spain | T3 | |
| KR101823278B1 | Republic of Korea | B1 | |
| PL3022734T3 | Poland | T3 | |
| PL3022735T3This record | Poland | T3 | |
| KR101823279B1 | Republic of Korea | B1 | |
| US9940938B2 | United States of America | B2 | |
| US9953656B2 | United States of America | B2 | |
| JP6346278B2 | Japan | B2 | |
| MX357667B | Mexico | B | |
| MX357826B | Mexico | B | |
| RU2666230C2 | Russian Federation | C2 | |
| US10147431B2 | United States of America | B2 | |
| RU2677580C2 | Russian Federation | C2 | |
| US2019108842A1 | United States of America | A1 | |
| US2019378522A1 | United States of America | A1 | |
| CN105580073B | China | B | |
| CN105593931B | China | B | |
| CN111105805A | China | A | |
| CN111128205A | China | A | |
| CN111128206A | China | A | |
| US10741188B2 | United States of America | B2 | |
| US10770080B2 | United States of America | B2 | |
| CA2917770C | Canada | C | |
| MY181944A | Malaysia | A | |
| US2021056979A1 | United States of America | A1 | |
| US2021233543A1 | United States of America | A1 | |
| CA2918237C | Canada | C | |
| BR112016001141B1 | Brazil | B1 | |
| US11488610B2 | United States of America | B2 | |
| BR112016001137B1 | Brazil | B1 | |
| US11657826B2 | United States of America | B2 | |
| US2024029744A1 | United States of America | A1 | |
| CN111128206B | China | B | |
| CN111105805B | China | B | |
| US12380899B2 | United States of America | B2 | |
| US20260024535A1 | United States of America | A1 | |
| CN111128205B | China | B |
Numbers
- Publication
- 3022735
- Publication, DOCDB
- 3022735
- Publication, EPODOC
- PL3022735T
- Application
- 14739141
- Application, DOCDB
- 14739141
- Application, EPODOC
- PL20140739141T
Titles2
- English
- AUDIO ENCODER, AUDIO DECODER, METHODS AND COMPUTER PROGRAM USING JOINTLY ENCODED RESIDUAL SIGNALS
- Polish
- KODER AUDIO, AUDIO DEKODER, SPOSOBY I PROGRAM KOMPUTEROWY UŻYWAJĄCY WSPÓLNIE KODOWANYCH SYGNAŁÓW RESZTKOWYCH
Classification
- CPC, 8
- G10L19/0017
- G10L19/008
- G10L21/038
- H04S3/008
- H04S7/30
- H04S2400/01
- H04S2400/03
- H04S2420/03
- IPC, 3
- G10L19 008
- G10L19 00
- G10L21 038
