Audio decoder, audio encoder, method for providing at least four audio channel signals on the basis of an encoded representation, method for providing an encoded representation on the basis of at least four audio channel signals and computer program using a bandwidth extension
40 claims: 15 independent, 25 dependent
- 1Zastrzeżenia patentowe 1. Dekoder audio (500;600;1300;1600;2000) do dostarczania co najmniej czterech sygnałów (520, 522, 524, 526) kanału audio o rozszerzonej szerokości pasma na podstawie kodowanej reprezentacji (510;610, 682;1310, 1312), przy czym dekoder audio jest przystosowany do dostarczania pierwszego sygnału downmixu (532;632;1342) i drugiego sygnału downmixu (534;634;1344) na podstawie wspólnie kodowanej reprezentacji (510;610;1310) pierwszego sygnału downmixu i drugiego sygnału downmixu z wykorzystaniem dekodowania wielokanałowego (530;630;1340);przy czym dekoder audio jest przystosowany do dostarczania co najmniej pierwszego sygnału (542;642;1372) kanału audio i drugiego sygnału (544;644;1374) kanału audio na podstawie pierwszego sygnału downmixu z wykorzystaniem dekodowania wielokanałowego (540;640;1370);przy czym dekoder audio jest przystosowany do dostarczania co najmniej trzeciego sygnału (556;656;1382) kanału audio i czwartego sygnału (558;658;1384) kanału audio na podstawie drugiego sygnału downmixu z wykorzystaniem dekodowania wielokanałowego ( 550;650;1380);przy czym dekoder audio jest przystosowany do wykonywania pierwszego wspólnego wielokanałowego rozszerzenia szerokości pasma (560;660;1390) na podstawie pierwszego sygnału kanału audio i trzeciego sygnału kanału audio, w celu uzyskania pierwszego sygnału kanału o rozszerzonej szerokości pasma (520;620;1320) i trzeciego sygnału o rozszerzonej szerokości pasma (524;624;1324), przy czym wielokanałowe rozszerzenie szerokości pasma wykorzystuje zależność między pierwszym sygnałem kanału audio i trzecim sygnałem kanału audio;i przy czym dekoder audio jest przystosowany do wykonywania drugiego wspólnego wielokanałowego rozszerzenia szerokości pasma (570;670;1394) na podstawie drugiego sygnału kanału audio i czwartego sygnału kanału audio, aby otrzymać drugi sygnał kanału o rozszerzonej szerokości pasma (522;622;1322) i czwarty sygnał kanału o rozszerzonej szerokości pasma (526;626;1326).
- 2Dekoder audio według zastrz. 1, przy czym pierwszy sygnał downmixu i drugi sygnał downmixu są powiązane z różnymi położeniami poziomymi lub położeniami azymutu sceny audio.
- 3Dekoder audio według zastrz. 1 albo 2, przy czym pierwszy sygnał downmixu jest powiązany z lewą stroną sceny audio, i przy czym drugi sygnał downmixu jest powiązany z prawą stroną sceny audio.
- 4Dekoder audio według jednego z zastrz. 1 do 3, przy czym pierwszy sygnał kanału audio i drugi sygnał kanału audio są powiązane z pionowo sąsiednimi położeniami sceny audio, oraz przy czym trzeci sygnał kanału audio i czwarty sygnał kanału audio są powiązane z pionowo sąsiednimi położeniami sceny audio.
- 5Dekoder audio według jednego z zastrz. 1 do 4, przy czym pierwszy sygnał kanału audio i trzeci sygnał kanału audio są powiązane z pierwszą wspólną płaszczyzną poziomą lub pierwszą wspólną wysokością sceny audio, ale z różnymi położeniami poziomymi lub położeniami azymutu sceny audio, przy czym drugi sygnał kanału audio i czwarty sygnał kanału audio są powiązane z drugą wspólną płaszczyzną poziomą lub drugą wspólną wysokością sceny audio, ale z różnymi położeniami poziomymi lub położeniami azymutu sceny audio, przy czym pierwsza wspólna płaszczyzna pozioma lub pierwsza wspólna wysokość jest różna od drugiej wspólnej płaszczyzny poziomej lub drugiej wspólnej wysokości.
- 6Dekoder audio według zastrz. 5, przy czym pierwszy sygnał kanału audio i drugi sygnał kanału audio są powiązane z pierwszą wspólną płaszczyzną pionową lub pierwszym wspólnym położeniem azymutu sceny audio, ale z różnymi położeniami pionowymi lub wysokościami sceny audio, oraz przy czym trzeci sygnał kanału audio i czwarty sygnał kanału audio są powiązane z drugą wspólną płaszczyzną pionową lub drugim wspólnym położeniem azymutu sceny audio, ale z różnymi położeniami pionowymi lub wysokościami sceny audio, przy czym pierwsza wspólna pionowa płaszczyzna lub pierwsze położenie azymutu jest różne od drugiej wspólnej pionowej płaszczyzny lub drugiego położenia azymutu.
- 7Dekoder audio według jednego z zastrz. 1 do 6, przy czym pierwszy sygnał kanału audio i drugi sygnał kanału audio są powiązane z lewą stroną sceny audio, i przy czym trzeci sygnał kanału audio i czwarty sygnał kanału audio są powiązane z prawą stroną sceny audio.
- 8Dekoder audio według jednego z zastrz. 1 do 7, przy czym pierwszy sygnał kanału audio i trzeci sygnał kanału audio są powiązane z częścią dolną sceny audio, oraz przy czym drugi sygnał kanału audio i czwarty sygnał kanału audio są powiązane z częścią górną sceny audio.
- 9Dekoder audio według jednego z zastrz. 1 do 8, przy czym dekoder audio jest przystosowany do wykonywania podziału poziomego, gdy dostarcza pierwszy sygnał downmixu i drugi sygnał downmixu na podstawie wspólnie kodowanej reprezentacji pierwszego sygnału downmixu i drugiego sygnału downmixu przy użyciu dekodowania wielokanałowego.
- 10Dekoder audio według jednego z zastrz. 1 do 9, przy czym dekoder audio jest przystosowany do wykonywania pionowego podziału, dostarczając co najmniej pierwszy sygnał kanału audio i drugi sygnał kanału audio na podstawie pierwszego sygnału downmixu wykorzystującego dekodowanie wielokanałowe;i przy czym dekoder audio jest przystosowany do wykonywania podziału pionowego, gdy zapewnia co najmniej trzeci sygnał kanału audio i sygnał czwartego kanału audio na podstawie drugiego sygnału downmixu z wykorzystaniem dekodowania wielokanałowego.
- 11Dekoder audio według jednego z zastrz. 1 do 10, przy czym dekoder audio jest przystosowany do wykonywania rozszerzania szerokości pasma stereo na podstawie pierwszego sygnału kanału audio i trzeciego sygnału kanału audio, aby uzyskać pierwszy sygnał kanału o rozszerzonej szerokości pasma i trzeci sygnał kanału o rozszerzonej szerokości pasma, przy czym pierwszy sygnał kanału audio i trzeci sygnał kanału audio przedstawiają pierwszą parę kanałów lewy/prawy;i przy czym dekoder audio jest przystosowany do wykonywania rozszerzania szerokości pasma stereo na podstawie drugiego sygnału kanału audio i czwartego sygnału kanału audio w celu uzyskania drugiego sygnału rozszerzonego o szerokości pasma i czwartego sygnału o rozszerzonej szerokości pasma szerokości pasma, przy czym drugi sygnał kanału audio i czwarty sygnał kanału audio przedstawiają drugą lewą/prawą parę kanałów.
- 12Dekoder audio według jednego z zastrz. 1 do 11, przy czym dekoder audio jest przystosowany do dostarczania pierwszego sygnału downmixu i drugiego sygnału downmixu na podstawie wspólnie kodowanej reprezentacji pierwszego sygnału downmixu i drugiego sygnału downmixu z wykorzystaniem opartego na predykcji dekodowania wielokanałowego.
- 13Dekoder audio według jednego z zastrz. 1 do 12, przy czym dekoder audio jest przystosowany do dostarczania pierwszego sygnału downmixu i drugiego sygnału downmixu na podstawie wspólnie kodowanej reprezentacji pierwszego sygnału downmixu i drugiego sygnału downmixu z wykorzystaniem dekodowania wielokanałowego wspomaganego sygnałem resztkowym.
- 14Dekoder audio według jednego z zastrz. 1 do 13, przy czym dekoder audio jest przystosowany do dostarczania co najmniej pierwszego sygnału kanału audio i drugiego sygnału kanału audio na podstawie pierwszego sygnału downmixu z wykorzystaniem opartego na parametrach dekodowania wielokanałowego;przy czym dekoder audio jest przystosowany do dostarczania co najmniej trzeciego sygnału kanału audio i czwartego sygnału kanału audio na podstawie drugiego sygnału downmixu z wykorzystaniem dekodowania wielokanałowego opartego na parametrach.
- 15Dekoder audio według zastrz. 14, przy czym oparte na parametrach dekodowanie wielokanałowe jest przystosowane do oceny jednego lub większej liczby parametrów opisujących pożądaną korelację między dwoma kanałami i/lub różnic poziomów między dwoma kanałami w celu zapewnienia dwóch lub większej liczby kanałów audio sygnały na podstawie odpowiedniego sygnału downmixu.
- 16Dekoder audio według jednego z zastrz. 1 do 15, przy czym dekoder audio jest przystosowany do dostarczania co najmniej pierwszego sygnału kanału audio i drugiego sygnału kanału audio na podstawie pierwszego sygnału downmixu z wykorzystaniem dekodowania wielokanałowego wspomaganego sygnałem resztkowym;i przy czym dekoder audio jest przystosowany do dostarczania co najmniej trzeciego sygnału kanału audio i czwartego sygnału kanału audio na podstawie drugiego sygnału downmixu z wykorzystaniem dekodowania wielokanałowego wspomaganego sygnałem resztkowym.
- 17Dekoder audio według jednego z zastrz. 1 do 16, przy czym dekoder audio jest przystosowany do dostarczania pierwszego sygnału resztkowego, który jest wykorzystywany do dostarczania co najmniej pierwszego sygnału kanału audio i drugiego sygnału kanału audio, i drugiego sygnału resztkowego, który jest wykorzystywany do dostarczania co najmniej trzeciego sygnału kanału audio i czwarty sygnał kanału audio, na podstawie wspólnie zakodowanej reprezentacji pierwszego sygnału resztkowego i drugiego sygnału resztkowego z wykorzystaniem dekodowania wielokanałowego.
- 18Dekoder audio według zastrz. 17, przy czym pierwszy sygnał resztkowy i drugi sygnał resztkowy są powiązane z różnymi położeniami poziomymi lub położeniami azymutu sceny audio.
- 19Dekoder audio według zastrz. 17 albo 18, przy czym pierwszy sygnał resztkowy jest powiązany z lewą stroną sceny audio, i przy czym drugi sygnał resztkowy jest powiązany z prawą stroną sceny audio.
- 20Koder audio (400;1500;2200) do dostarczania zakodowanej reprezentacji (420;1532;2272,2282) na podstawie co najmniej czterech sygnałów kanału audio (410, 412;1512, 1514;2212, 2222, 2214, 2224 ), przy czym koder audio jest przystosowany do otrzymywania pierwszego zestawu (2215) wspólnych parametrów rozszerzania szerokości pasma na podstawie pierwszego sygnału kanału audio (410;2212) i trzeciego sygnału kanału audio (414, 2214);przy czym koder audio jest przystosowany do otrzymywania drugiego zestawu (2225) wspólnych parametrów rozszerzania szerokości pasma na podstawie drugiego sygnału kanału audio (412;2222) i czwartego sygnału kanału audio (416;2224);przy czym koder audio jest przystosowany do wspólnego kodowania co najmniej pierwszego sygnału kanału audio i drugiego sygnału kanału audio z wykorzystaniem kodowania wielokanałowego (450;2230), aby uzyskać pierwszy sygnał downmixu (452;2234);przy czym koder audio jest przystosowany do wspólnego kodowania co najmniej trzeciego sygnału kanału audio i czwartego sygnału kanału audio z wykorzystaniem kodowania wielokanałowego (460;2240), aby uzyskać drugi sygnał downmixu (462;2244);i przy czym koder audio jest przystosowany do wspólnego kodowania pierwszego sygnału downmixu i drugiego sygnału downmixu z wykorzystaniem kodowania wielokanałowego (470;2250), aby uzyskać zakodowaną reprezentację sygnałów downmixu.
- 21Koder audio według zastrz. 20, przy czym pierwszy sygnał downmixu i drugi sygnał downmixu są powiązane z różnymi położeniami poziomymi lub położeniami azymutu sceny audio.
- 22Koder audio według jednego z zastrz. 20 albo 21, przy czym pierwszy sygnał downmixu jest powiązany z lewą stroną sceny audio, i przy czym drugi sygnał downmixu jest powiązany z prawą stroną sceny audio.
- 23Koder audio według jednego z zastrz. 20 do 22, przy czym pierwszy sygnał kanału audio i drugi sygnał kanału audio są powiązane z pionowo sąsiednimi położeniami sceny audio, i przy czym trzeci sygnał kanału audio i czwarty sygnał kanału audio są powiązane z pionowo sąsiednimi położeniami sceny audio.
- 24Koder audio według jednego z zastrz. 20 do 23, przy czym pierwszy sygnał kanału audio i trzeci sygnał kanału audio są powiązane z pierwszą wspólną płaszczyzną poziomą lub pierwszą wysokością sceny audio, ale różnymi położeniami poziomymi lub położeniami azymutu sceny audio, przy czym drugi sygnał kanału audio i czwarty sygnał kanału audio są powiązane z drugą wspólną płaszczyzną poziomą lub drugą wysokością sceny audio, ale z różnymi położeniami poziomymi lub położeniami azymutu sceny audio, przy czym pierwsza wspólna płaszczyzna pozioma lub pierwsza wysokość jest różna od drugiej wspólnej płaszczyzny poziomej lub drugiej wysokości.
- 25Koder audio według zastrz. 24, przy czym pierwszy sygnał kanału audio i drugi sygnał kanału audio są powiązane z pierwszą wspólną płaszczyzną pionową lub pierwszym położeniem azymutu sceny audio, ale z różnymi położeniami w pionie lub wysokościami sceny audio, oraz przy czym trzeci sygnał kanału audio i czwarty sygnał kanału audio są powiązane z drugą wspólną płaszczyzną pionową lub z drugimi położeniami azymutu sceny audio, ale z różnymi położeniami pionowymi lub wysokościami sceny audio, przy czym pierwsza wspólna pionowa płaszczyzna lub pierwsze położenie azymutu jest różne od drugiej wspólnej pionowej płaszczyzny lub drugiego położenia azymutu.
- 26Koder audio według jednego z zastrz. 20 do 25, przy czym pierwszy sygnał kanału audio i drugi sygnał kanału audio są powiązane z lewą stroną sceny audio, oraz przy czym trzeci sygnał kanału audio i czwarty sygnał kanału audio są powiązane z prawą stroną sceny audio.
- 27Koder audio według jednego z zastrz. 20 do 26, przy czym pierwszy sygnał kanału audio i trzeci sygnał kanału audio są powiązane z częścią dolną sceny audio, i przy czym drugi sygnał kanału audio i czwarty sygnał kanału audio są powiązane z częścią górną sceny audio.
- 28Koder audio według jednego z zastrz. 20 do 27, przy czym koder audio jest przystosowany do wykonywania łączenia poziomego, gdy zapewnia zakodowaną reprezentację sygnałów downmixu na podstawie pierwszego sygnału downmixu i drugiego sygnału downmixu z wykorzystaniem kodowania wielokanałowego.
- 29Koder audio według jednego z zastrz. 20 do 28, przy czym koder audio jest przystosowany do wykonywania łączenia w pionie podczas dostarczania pierwszego sygnału downmixu na podstawie pierwszego sygnału kanału audio i drugiego sygnału kanału audio z wykorzystaniem kodowania wielokanałowego;i przy czym koder audio jest przystosowany do wykonywania łączenia w pionie podczas dostarczania drugiego sygnału downmixu na podstawie trzeciego sygnału kanału audio i czwartego sygnału kanału audio z wykorzystaniem kodowania wielokanałowego.
- 30Koder audio według jednego z zastrz. 20 do 29, przy czym koder audio jest przystosowany do dostarczania wspólnie zakodowanej reprezentacji pierwszego sygnału downmixu i drugiego sygnału downmixu na podstawie pierwszego sygnału downmixu i drugiego sygnału downmixu z wykorzystaniem kodowania wielokanałowego opartego na predykcji.
- 31Koder audio według jednego z zastrz. 20 do 30, przy czym koder audio jest przystosowany do dostarczania wspólnie zakodowanej reprezentacji pierwszego sygnału downmixu i drugiego sygnału downmixu na podstawie pierwszego sygnału downmixu i drugiego sygnału downmixu z wykorzystaniem kodowania wielokanałowego wspomaganego sygnałem resztkowym.
- 32Koder audio według jednego z zastrz. 20 do 31, przy czym koder audio jest przystosowany do dostarczania pierwszego sygnału downmixu na podstawie pierwszego sygnału kanału audio i drugiego sygnału kanału audio z wykorzystaniem kodowania wielokanałowego opartego na parametrze;i przy czym koder audio jest przystosowany do dostarczania drugiego sygnału downmixu na podstawie trzeciego sygnału kanału audio i czwartego sygnału kanału audio z wykorzystaniem kodowania wielokanałowego opartego na parametrze.
- 33Koder audio według zastrz. 32, przy czym oparte na parametrach wielokanałowe kodowanie jest przystosowane do dostarczania jednego lub więcej parametrów opisujących pożądaną korelację między dwoma kanałami i/lub różnic poziomu między dwoma kanałami.
- 34Koder audio według jednego z zastrz. 20 do 33, przy czym koder audio jest przystosowany do dostarczania pierwszego sygnału downmixu na podstawie pierwszego sygnału kanału audio i drugiego sygnału kanału audio z wykorzystaniem kodowania wielokanałowego wspomaganego sygnałem resztkowym;i przy czym koder audio jest przystosowany do dostarczania drugiego sygnału downmixu na podstawie trzeciego sygnału kanału audio i czwartego sygnału kanału audio z wykorzystaniem kodowania wielokanałowego wspomaganego sygnałem resztkowym.
- 35Koder audio według jednego z zastrz. 20 do 34, przy czym koder audio jest przystosowany do zapewniania wspólnie zakodowanej reprezentacji pierwszego sygnału resztkowego, który jest uzyskiwany, gdy wspólnie kodowany jest co najmniej pierwszy sygnał kanału audio i drugi sygnał kanału audio, i drugiego resztkowego, który jest, gdy wspólnie kodowany jest co najmniej trzeci sygnał kanału audio i czwarty sygnał kanału audio, z wykorzystaniem kodowania wielokanałowego.
- 36Koder audio według zastrz. 35, przy czym pierwszy sygnał resztkowy i drugi sygnał resztkowy są powiązane z różnymi położeniami poziomymi lub położeniami azymutu sceny audio.
- 37Koder audio według zastrz. 35 albo zastrz. 36, przy czym pierwszy sygnał resztkowy jest związany z lewą stroną sceny audio, i przy czym drugi sygnał resztkowy jest związany z prawą stroną sceny audio.
- 38Sposób (1000) do dostarczania co najmniej czterech sygnałów kanału audio o rozszerzonej szerokości pasma szerokości w oparciu o zakodowaną reprezentację, przy czym sposób obejmuje:dostarczanie (1010) pierwszego sygnału downmixu i drugiego sygnału downmixu na podstawie wspólnie kodowanej reprezentacji pierwszego sygnału downmixu i drugiego sygnału downmixu z wykorzystaniem dekodowania wielokanałowego;dostarczanie (1020) co najmniej pierwszego sygnału kanału audio i drugiego sygnału kanału audio na podstawie pierwszego sygnału downmixu z wykorzystaniem dekodowania wielokanałowego;dostarczanie (1030) co najmniej trzeciego sygnału kanału audio i czwartego sygnału kanału audio na podstawie drugiego sygnału downmixu z wykorzystaniem dekodowania wielokanałowego;wykonanie (1040) pierwszego wspólnego wielokanałowego rozszerzenia szerokości pasma na podstawie pierwszego sygnału kanału audio i trzeciego sygnału kanału audio, aby uzyskać pierwszy sygnał kanału o rozszerzonej szerokości pasma i trzeci sygnał kanału o rozszerzonej szerokości pasma, przy czym rozszerzenie szerokości pasma kanału wykorzystuje zależność między pierwszym sygnałem kanału audio i trzecim sygnałem kanału audio;i wykonanie (1050) drugiego wspólnego wielokanałowego rozszerzenia szerokości pasma na podstawie drugiego sygnału kanału audio i czwartego sygnału kanału audio, aby uzyskać drugi sygnał kanału rozszerzonego o szerokość pasma i czwarty sygnał kanału rozszerzonego o szerokość pasma.
- 39Sposób (900) do zapewniania kodowanej reprezentacji na podstawie co najmniej czterech sygnałów kanału audio, który to sposób obejmuje:uzyskanie (920) pierwszego zestawu wspólnych parametrów rozszerzania szerokości pasma na podstawie pierwszego sygnału kanału audio i trzeciego sygnału kanału audio;uzyskanie (930) drugiego zestawu wspólnych parametrów rozszerzania szerokości pasma na podstawie drugiego sygnału kanału audio i czwartego sygnału kanału audio;wspólne kodowanie (930) co najmniej pierwszego sygnału kanału audio i drugiego sygnału kanału audio z wykorzystaniem kodowania wielokanałowego, w celu uzyskania pierwszego sygnału downmixu;wspólne kodowanie (940) co najmniej trzeciego sygnału kanału audio i czwartego sygnału kanału audio z wykorzystaniem kodowania wielokanałowego, w celu uzyskania drugiego sygnału downmixu;i wspólne kodowanie (950) pierwszego sygnału downmixu i drugiego sygnału downmixu z wykorzystaniem kodowania wielokanałowego, w celu uzyskania zakodowanej reprezentacji sygnałów downmixu.
- 40Program komputerowy przystosowany do wykonywania sposobu według zastrz. 38 albo 39, gdy program komputerowy działa na komputerze. Fraunhofer-Gesellschaft zur Fórderung der angewandten Forschung e.V., Niemcy; Pełnomocnik:ΕΡ3022734 Ζ-16506 1/21 FIG1 ΕΡ3022734 Ζ-16506 2/21 OJ FIG 2 ΕΡ3022734 Ζ-16506 FIG 3 ΕΡ3022734 Ζ-16506 4/21 ΕΡ3022734 Ζ-16506 FIG 5 ΕΡ3022734 Ζ-16506 6/21 predykcji) ΕΡ3022734 Ζ-16506 7/21 FIG 6Β ΕΡ3022734 Ζ-16506 8/21 FIG 7 ΕΡ3022734 Ζ-16506 800 9/21 FIG 8 ΕΡ3022734 Ζ-16506 10/21 FIG 9 ΕΡ3022734 Ζ-16506 11/21 1000 FIG 10 ΕΡ3022734 Ζ-16506 12/21 co 1100 GD CXI ΕΡ3022734 Ζ-16506 13/21 1200 ΕΡ3022734 Ζ-16506 14/21 EP3022734 Z-16506 15/21 UsacChannelPairElementConfig (sbrRatiolndex) { UsacCoreConfig ();if (sbrRatiolndex 0) { SbrConfig ();stereoConfiglndex;2 } else { stereoConfiglndex = 0;l· if (stereoConfiglndex 0) { Mps212Config(stereoConfiglndex);l· + qcelndex 2 l· uimsbf uimsbf FIG 14A FIG14B ΕΡ3022734 Ζ-16506 16/21 1500 Widok z góry kodera audio 3D FIG 15 ΕΡ3022734 Ζ-16506 ΕΡ3022734 Ζ-16506 18/21 Sygnały głośnika Odtworzenie układu Budowa modułu konwersji formatu FIG 17 ΕΡ3022734 Ζ-16506 19/21 FIG 18 2010 2020 2022 2030 2032 2024 2034 FIG 19 ΕΡ3022734 Ζ-16506 20/21 Schemat dekodera QCE ΕΡ3022734 Ζ-16506 21/21
Independent claims40
283 paragraphs in 6 sections, as filed
THE REPUBLIC OF POLAND (12) TRANSLATION OF THE EUROPEAN PATENT (19) PL (11) PL / EP 3022734
<img file="PL3022734T3_D0001.tif" />
The Patent Office of the Republic of Poland (96) Date and number of the European patent application: 14.07.2014 14738535.5 (97) The grant of the European patent was announced:
23.08.2017 European Patent Bulletin 2017/34 EP 3022734 B1 (13) T3 (51) Int.CI.
G10L 19/008 (2013.01)
G10L 21/038 (2013.01)
G10L 19/00 (2013.01) (54) Title of the invention:
AUDIO DECODER, AUDIO ENCODER, WAY OF DELIVERY OF AT LEAST FOUR AUDIO CHANNEL SIGNALS ON THE BASIS OF CODED REPRESENTATIONS, WAY OF DELIVERING STORED REPRESENTATIONS ON THE BASIS OF AT LEAST FOUR AUDIO PROGRAMS WITH AT LEAST FOUR SIGNAL PROGRAMS
Priority:
2013-07-22 EP 13177376
October 18, 2013 EP 13189306 (43) Application announced:
25.05.2016 in the European Patent Bulletin No. 2016/21 (45) The following was announced about the submission of the translation of the patent:
31.01.2018 Patent Office News 2018/01 (73) Patent holder:
Fraunhofer-Gesellschaft zur Fórderung der angewandten Forschung eV, MOnchen, DE (72) Inventor (s):
What what
CM CM
What about (74)
SASCHA DICK, NOrnberg, DE CHRISTIAN ERTEL, Eckental, DE CHRISTIAN HELMRICH, Erlangen, DE JOHANNES HILPERT, Nurnberg, DE ANDREAS HÓLZER, Erlangen, DE ACHIM KUNTZ, Hemhofen, DE
Proxy:
thing, pat. Eliza Stypińska
LDS ŁAZEWSKI DEPOSIT AND PARTNERS ul. Okopowa 58/72 01-042 Warsaw
Attention:
Within nine months from the publication of the information on the grant of a European patent, any person may file an objection to the European Patent Office against the granted European patent. The opposition is filed in the form of a written statement of reasons.It is considered to have been brought only upon payment of the opposition fee (Art 99 ( 1) European Patent Convention)
ΕΡ 3 022 734 BI
Z-16506/17
AUDIO DECODER, AUDIO ENCODER, WAY OF SUPPLYING AT LEAST FOUR 5 AUDIO CHANNEL SIGNALS BASED ON CODED REPRESENTATIONS,
WAY OF DELIVERING STORED REPRESENTATIONS BASED ON AT LEAST FOUR SIGNALS OF AUDIO CHANNELS AND COMPUTER PROGRAM USING BAND EXTENSION
Description Technical field
[0001] An embodiment of the invention creates an audio decoder for providing at least four extended bandwidth channel signals based on an encoded representation.
[0002] Another embodiment of the invention creates an audio encoder for providing an encoded representation based on at least four audio channel signals.
[0003] Another embodiment of the invention creates a method for providing at least four audio channel signals based on an encoded representation.
[0004] Another embodiment of the invention creates a method for providing a coded representation based on at least four audio channel signals.
[0005] Another embodiment of the invention creates a computer program for performing one of the methods.
[0006] In general, embodiments of the invention relate to a common coding of n channels.
Background of the invention
[0007] The demand for storing and transmitting audio content has steadily increased in recent years. Moreover, the quality requirements for the storage and transmission of audio content are also constantly increasing. Accordingly, the concepts of encoding and decoding audio content have been improved. For example, so-called "advanced audio coding (AAC)" has been developed, which is described, for example, in the International Standard ISO / IEC 13818-7: 2003. In addition, several spatial extensions have been created, such as for example the term "MPEG Surround, which is described for example in the international standard ISO / IEC 23003-1: 2007. Moreover, additional improvements to the encoding and decoding of spatial information of audio signals are described in the international standard ISO / IEC 23003-2: 2010, which deals with so-called spatial audio object coding (SAOC).
[0008] Moreover, a flexible audio encoding / decoding concept, which provides the ability to encode both general audio signals and speech signals with good encoding efficiency and handling multi-channel audio signals, is specified in the international standard ISO / IEC 23003-3: 2012, which describes the so-called concept of "unified speech and audio coding (USAC)".
[0009]
In MPEG USAC [1], the common stereo coding of the two channels is performed using composite prediction, MPS 2-1-1, or a unified stereo signal with band-limited or full-band residual signals.
MPEG surround [2] hierarchically combines OTT and TTT frames to co-encode multi-channel audio with or without transmission of residual signals. Moreover, US 2012/0070007 A1 discloses multi-channel encoding / decoding using bandwidth extension.
[0010] There is, however, a desire to provide an even more advanced concept for efficiently encoding and decoding three-dimensional audio scenes.
Summary of the invention
[0011] An embodiment according to the invention creates an audio decoder for providing at least four extended bandwidth audio channel signals based on an encoded representation. The audio decoder is adapted to provide the first downmix signal and the second downmix signal on the basis of a jointly coded representation of the first downmix signal and the second downmix signal using the (first) multi-channel decoding. The audio decoder is adapted to provide at least the first audio channel signal and the second audio channel signal from the first downmix signal using the (second) multi-channel decoding and provide at least the third audio channel signal and the fourth audio channel signal based on the second downmix signal using ( third) of multi-channel decoding. The audio decoder is adapted to perform the first common multi-channel bandwidth extension based on the first audio channel signal and the third audio channel signal to obtain the first bandwidth extended channel signal and the third bandwidth extended channel signal, why the multi-channel bandwidth extension uses the relationship between the first signal of the audio channel and the third signal of the audio channel. Moreover, the audio decoder is adapted to perform a second common multi-channel bandwidth extension based on the second audio channel signal and the fourth audio channel signal to obtain the second signal with extended bandwidth and the fourth signal with extended bandwidth.
[0012] This embodiment of the invention is based on the finding that particularly good bandwidth extension results can be obtained in a hierarchical audio decoder if the audio channel signals which are obtained from different downmix signals in the second step of the audio decoder used in a multi-channel bandwidth extension wherein the different downmix signals are derived from the jointly coded representation in the first step of the audio decoder. It has been found that a particularly good sound quality can be obtained if the downmix signals which are associated with perceptually particularly important positions of the audio scene are separated in the first stage of the hierarchical audio decoder, while spatial positions which are not so important for the auditory experience are separated in a second stage of the hierarchical audio decoder. Furthermore, it has been found that audio channel signals that are associated with different perceptually important positions of the audio scene (e.g. audio scene positions where the relationship between the signals from these positions is perceptually important) should be processed together as a multi-channel bandwidth extension, since the multi-channel bandwidth extension the bandwidth extension can consistently take into account the relationships and differences between signals from these acoustically important positions. This is achieved by performing a multi-channel bandwidth extension based on a first audio channel signal (which is derived from the first downmix signal in the second stage of the hierarchical audio decoder) and based on the third audio channel signal which comes from the second downmix signal in the second stage of the hierarchical audio decoder. , to obtain two extended bandwidth channel signals (namely, a first extended bandwidth channel signal and a third extended bandwidth channel signal). Correspondingly, the (common) multi-channel bandwidth extension is performed based on the audio channel signals that come from the different downmix signals in the second stage of the hierarchical multi-channel decoder, such that the relationship between the first audio channel signal and the third audio channel signal is similar to (or determined by) the relationship between the first downmix signal and the second downmix signal. Thus, the multi-channel bandwidth extension may take advantage of this relationship (e.g., between the first audio channel signal and the third audio channel signal), which is substantially determined by deriving the first downmix signal and the second downmix signal from a common coded representation of the first downmix signal and the second downmix signal. using multi-channel decoding which is performed in the first step of the audio decoder. Accordingly, the multi-channel bandwidth extension can use this relationship, which can be reproduced with good accuracy in the first step of the hierarchical audio decoder, so that a particularly good hearing sensation is achieved.
[0013] In a preferred embodiment, the first downmix signal and the second downmix signal are associated with different horizontal positions (or azimuth positions) of the audio scene. It has been found that distinguishing between different horizontal audio positions (or azimuth positions) is particularly important as the human auditory system is particularly sensitive with respect to different horizontal positions. Accordingly, it is preferable to separate the downmix signals associated with the different horizontal positions of the audio scene in the first step of the hierarchical audio decoder, since the processing in the first step of the hierarchical audio decoder is typically more precise than the processing in the following steps. Moreover, consequently, the first audio channel signal and the third audio channel signal that are used together in the (first) multi-channel bandwidth extension are associated with different horizontal positions of the audio scene (since the first audio channel signal is derived from the first downmix signal and the third the audio channel signal is derived from the second downmix signal in a second step of the hierarchical audio decoder), which allows the (first) multi-channel bandwidth extension to be well suited to the human ability to distinguish between different horizontal positions. Similarly, the (second) multi-channel bandwidth extension that is performed on the basis of the second audio channel signal and the fourth audio channel signal acts on the audio channel signals that are associated with different horizontal positions of the audio scene, such that the (second) multi-channel bandwidth extension bandwidths can also be well suited to the psychoacoustically important relationship between the audio channel signals associated with different horizontal positions of the audio scene. Accordingly, a particularly good sound impression can be achieved.
[0014] In a preferred embodiment, the first downmix signals are associated with the left side of the audio scene and the second downmix signal is associated with the right side of the audio scene. Accordingly, the first audio channel signal is typically also associated with the left side of the audio scene and the third audio channel signal is associated with the right side of the audio scene such that the (first) multi-channel bandwidth extension acts (preferably jointly) on the audio channel signals. from different sides of the audio scene and can therefore be well adapted to human left / right perception. The same is true for the (second) multi-channel bandwidth extension which operates on the basis of the second audio channel signal and the fourth audio channel signal.
[0015] In a preferred embodiment, the first audio channel signal and the second audio channel signal are associated with vertically adjacent positions of the audio scene. Similarly, the third audio channel signal and the fourth audio channel signal are associated with vertically adjacent positions of the audio scene. It has been found advantageous to separate the audio channel signals related to vertically adjacent positions of the audio scene in the second stage of the hierarchical audio decoder. Moreover, it has been found that audio channel signals are typically not severely degraded by separating audio channel signals associated with vertically adjacent positions, so that input signals to multi-channel bandwidth extensions are still well suited to multi-channel bandwidth extension (e.g. stereo bandwidth extension).
[0016] In a preferred embodiment, the first audio channel signal and the third audio channel signal are associated with a first common horizontal plane (or first common height) of the audio scene but with different horizontal positions (or azimuth positions) of the audio scene, and the second channel signal the audio and the fourth audio channel signal are associated with a second common horizontal plane (or second common height) of the audio scene, but with different horizontal positions (or azimuth positions) of the audio scene. In this case, the first common horizontal plane (or height) differs from the second common horizontal plane (or height). It has been found that the multi-channel bandwidth extension can be performed with extremely good result quality based on two audio channel signals that are associated with the same horizontal plane (or height).
[0017] In a preferred embodiment, the first audio channel signal and the second audio channel signal are associated with a first common vertical plane (or common azimuth position) of the audio scene, but with different vertical positions (or heights) of the audio scene. Similarly, the third audio channel signal and the fourth audio channel signal are associated with a second common vertical plane (or common azimuth position) of the audio scene, but with different vertical positions (or heights) of the audio scene. In this case, the first common vertical plane (or azimuth position) is preferably different from the second common vertical plane (or azimuth positions). It has been found that splitting (or splitting) the audio channel signals related to a common vertical plane (or azimuth position) can be performed with good results using the second stage of the hierarchical audio decoder, while the separation (or splitting) between the audio channel signals related to different verticals the planes (or azimuth positions) can be performed with good quality results using the first step of the hierarchical audio decoder.
[0018] In a preferred embodiment, the first audio channel signal and the second audio channel signal are associated with the left side of the audio scene, and the third audio channel signal and the fourth audio channel signal are associated with the right side of the audio scene. This configuration allows for a particularly good multi-channel bandwidth extension which takes advantage of the relationship between an audio channel signal associated with the left side and an audio channel signal associated with the right side, and is therefore well suited to human ability to distinguish between audio coming from the left and audio. on the right.
[0019] In a preferred embodiment, the first audio channel signal and the third audio channel signal are associated with the bottom of the audio scene and the second audio channel signal and the fourth audio channel signal are associated with the top of the audio scene. Such a spatial distribution of the audio channel signals has been found to produce particularly good hearing results.
[0020] In a preferred embodiment, the audio decoder is adapted to perform the horizontal division by providing the first downmix signal and the second downmix signal based on a jointly coded representation of the first downmix signal and the second downmix signal using multi-channel decoding. It has been found that performing the horizontal partitioning of the first step of the hierarchical audio decoder gives a particularly good listening experience because the processing performed in the first step of the hierarchical audio decoder can typically be performed with higher efficiency than the processing performed in the second step of the hierarchical audio decoder. Moreover, performing the horizontal split in the first step of the audio decoder gives a good listening experience as the human auditory system is more sensitive with respect to the horizontal position of the audio object compared to the vertical position of the audio object.
[0021] In a preferred embodiment, the audio decoder is arranged to perform vertical splitting when it provides at least the first audio channel signal and the second audio channel signal based on the first downmix signal using multi-channel decoding. Likewise, the audio decoder is preferably adapted to perform a vertical division when it provides at least a third audio channel signal and a fourth audio channel signal based on the second downmix signal using multi-channel decoding. Performing vertical split in the second step of the hierarchical decoder has been found to produce a good auditory impression as the human auditory system is not particularly sensitive to the vertical position of the sound source (or audio object).
[0022] In a preferred embodiment, the audio decoder is arranged to perform a stereo bandwidth extension based on the first audio channel signal and the third audio channel signal to obtain a first bandwidth extended channel signal and a third extended bandwidth channel signal, wherein the first audio channel signal and the third audio channel signal represent a first left / right channel pair. Similarly, the audio decoder is arranged to perform a stereo bandwidth extension based on the second audio channel signal and the fourth audio channel signal to obtain the second bandwidth extended channel signal and the fourth extended bandwidth channel signal, the second audio channel signal and the fourth audio channel signal. the audio channel signal represents the second left / right channel pair. It has been found that the stereo bandwidth extension produces a particularly good listening experience since the stereo bandwidth extension can take into account the relationship between the left stereo channel and the right stereo channel and perform a bandwidth extension depending on this relationship.
[0023] In a preferred embodiment, the audio decoder is adapted to provide the first downmix signal and the second downmix signal based on a jointly coded representation of the first downmix signal and the second downmix signal using prediction based multi-channel decoding. It has been found that the use of prediction based multi-channel decoding in the first stage of the hierarchical audio decoder provides a good trade-off between bit rate and quality. It has been found that the use of the prediction results in a good reconstruction of the differences between the first downmix signal and the second downmix signal, which is important for left / right audio object discrimination.
[0024] For example, the audio decoder may be adapted to evaluate a prediction parameter describing a contribution of a signal element that comes from the signal element of the previous frame to downmix the current frame. Accordingly, the intensity of the contribution of the signal element that comes from the signal element of the previous frame can be adjusted based on a parameter that is included in the coded representation.
For example, the prediction based multi-channel decoding may operate in the MDCT domain, such that the prediction based multi-channel decoding may be well adapted and easily combined with the audio decoding step which provides an input for multi-channel decoding which outputs the first downmix signal and second downmix signal. Preferably, but not necessarily, the prediction based multi-channel decoding may be a stereo prediction of USAC, which facilitates the implementation of the audio decoder.
[0026] In a preferred embodiment, the audio decoder is adapted to provide the first downmix signal and the second downmix signal on the basis of a jointly coded representation of the first downmix signal and the second downmix signal using a residual-assisted multi-channel decoding. The use of residual-assisted multi-channel decoding allows a particularly precise reconstruction of the first downmix signal and the second downmix signal, which in turn improves the perception of left-right positions from the audio channel signals and consequently from the extended bandwidth channel signals.
[0027] In a preferred embodiment, the audio decoder is adapted to provide at least the first audio channel signal and the second audio channel signal based on the first downmix signal using parameter based multi-channel decoding. Moreover, the audio decoder is adapted to provide at least the third audio channel signal and the fourth audio channel signal based on the second downmix signal using parameter based multi-channel decoding. It has been found that the use of parameter based multi-channel decoding fits well with the second stage of the hierarchical audio decoder. It has been found that parameter-based multi-channel decoding provides a good compromise between audio quality and bit rate. Although the reproduction quality of parameter-based multi-channel decoding is usually not as good as the reproduction quality of prediction-based (and possibly residual) multi-channel decoding, it has been found that the use of parameter-based multi-channel decoding is usually sufficient because the human auditory system is not particularly sensitive to the vertical position (or height) of the audio object, which is preferably determined by the spread (or separation) between the first signal of the audio channel and the second signal of the audio channel, or between the third signal of the audio channel and the fourth signal of the audio channel.
[0028] In a preferred embodiment, parameter based multi-channel decoding is adapted to evaluate one or more parameters describing a desired correlation (or covariance) between two channels and / or level differences between two channels to provide two or more audio channel signals based on the relevant downmix signal. It has been found that the use of parameters that describe, for example, a desired correlation between two channels and / or level differences between the two channels, is well suited for splitting (or splitting) the signals of the first audio channel and the second audio channel (which are typically associated with different vertical positions of the audio scene) and for splitting (or separating) between the third audio channel signal and the fourth audio channel signal (which are also usually associated with different vertical positions).
[0029] For example, parameter based multi-channel decoding may operate in the QMF domain. Accordingly, parameter-based multi-channel decoding may be well suited - and easy to combine with multi-channel bandwidth extension which may also - but not necessarily - operate in the QMF domain.
[0030] For example, parameter based multi-channel decoding may be MPEG surround 2-1-2 decoding or unified stereo decoding. The use of such coding concepts may facilitate implementation as these decoding concepts may already be present in legacy audio decoders.
[0031] In a preferred embodiment, the audio decoder is adapted to provide at least the first audio channel signal and the second audio channel signal based on the first downmix signal using residual multi-channel decoding. Furthermore, the audio decoder may be adapted to provide at least the third audio channel signal and the fourth audio channel signal based on the second downmix signal using residual-assisted multi-channel decoding. By using residual-assisted multi-channel decoding, the sound quality can even be improved as the separation between the first audio channel signal and the second audio signal and / or the separation between the third audio channel signal and the fourth audio channel can be performed with a particularly high quality.
In a preferred embodiment, the audio decoder may be adapted to provide a first residual signal that is used to provide at least the first audio channel signal and the second audio channel signal and a second residual signal that is used to provide at least a third channel signal. audio and fourth channel audio signal, based on the jointly coded representation of the first residual signal and the second residual signal using multi-channel decoding. Accordingly, the hierarchical decoding concept can be extended to provide two residual signals, one of which is used to provide a first audio channel signal and a second audio channel signal (but which is typically not used to provide a third audio channel signal and a fourth audio channel signal). one of them is used for providing the third audio channel signal and the fourth audio channel signal (but preferably is not used for providing the first audio channel signal and the second audio channel signal).
[0033] In a preferred embodiment, the first residual signal and the second residual signal may be associated with different horizontal positions (or azimuth positions) of the audio scene. Accordingly, the provision of the first residual signal and the second residual that is performed in the first step of the hierarchical audio decoder may perform horizontal split (or separation), and it has been found that particularly good horizontal splitting (or separation) may be performed in the first step. the hierarchical audio decoder (compared to the processing performed in the second stage of the hierarchical audio decoder). Correspondingly, the horizontal separation, which is especially important for the human listener, is performed in the first step of hierarchical audio decoding, which provides a particularly good reproduction, so that a good auditory impression can be obtained.
[0034] In a preferred embodiment, the first residual signal is associated with the left side of the audio scene and the second residual signal is associated with the right side of the audio scene which matches the human positional sensitivity.
[0035] An embodiment of the invention creates an audio encoder providing an encoded representation based on at least four audio channel signals. The audio encoder is adapted to derive the first set of common bandwidth extension parameters from the first audio channel signal and the third audio channel signal. The audio encoder is also adapted to derive the second set of common bandwidth extension parameters from the second audio channel signal and the fourth audio channel signal. The audio encoder is adapted to co-code at least the first audio channel signal and the second audio channel signal using multi-channel coding to obtain the first downmix signal and co-code the at least the third audio channel signal and the fourth audio channel signal by multi-channel coding to obtain a second downmix signal. downmix signal. Moreover, the audio encoder is adapted to co-code the first downmix signal and the second downmix signal by multi-channel coding to obtain an encoded representation of the downmix signals.
[0036] This embodiment is based on the assumption that the first set of common bandwidth extension parameters should be obtained from the audio channel signals that are represented by different downmix signals that are only coded together in the second step of the hierarchical audio encoder. In parallel with the above-discussed audio decoder, the relationship between the audio channel signals, which are connected only in the second step of hierarchical audio encoding, can be reproduced with particularly high accuracy on the audio decoder side. Therefore, it has been found that two audio signals that are efficiently combined only in the second step of the hierarchical encoder are well suited to obtain a common bandwidth extension parameter set, since multi-channel bandwidth extension can best be applied to. audio channel signals, the relationship between which is well reconstructed on the audio decoder side. Accordingly, it has been found that it is better, in terms of achievable audio quality, to derive a common band extension parameter set from such audio channel signals that are connected only in the second stage of the hierarchical audio encoder compared to obtaining a common band extension parameter set from such signals. audio channels that are combined in the first stage of the hierarchical audio encoder. However, it has also been found that the best audio quality can be obtained by deriving common bandwidth extension parameter sets from the audio channel signals before they are jointly encoded in the first step of the hierarchical audio encoder.
[0037] In a preferred embodiment, the first downmix signal and the second downmix signal are associated with different horizontal positions (or azimuth positions) of the audio scene. This concept is based on the assumption that the best audibility sensation can be obtained if the signals associated with different horizontal positions are only coded together in the second stage of the hierarchical audio encoder.
[0038] In a preferred embodiment, the first downmix signal is associated with the left side of the audio scene and the second downmix signal is associated with the right side of the audio scene. Thus, such multi-channel signals which are associated with different sides of the audio scene are used to provide common bandwidth extension parameter sets. Consequently, the parameter sets of common bandwidth extension are well suited to human ability to distinguish audio sources on different sides.
[0039] In a preferred embodiment, the first audio channel signal and the second audio channel signal are associated with vertically adjacent positions of the audio scene. Moreover, the third audio channel signal and the fourth audio channel signal are also associated with vertically adjacent positions of the audio scene. It has been found that a good audio impression can be obtained if the audio channel signals that are associated with vertically adjacent positions of the audio scene are collectively encoded in a first step of the hierarchical encoder, while it is better to derive common bandwidth extension parameter sets from the audio channel signals that are they are not associated with vertically adjacent locations (but which are associated with different horizontal or different azimuth locations).
[0040] In a preferred embodiment, the first audio channel signal and the third audio channel signal are associated with a first common horizontal plane (or first common height) of the audio scene, but with different horizontal positions (or azimuth positions) of the audio scene, and the second channel signal the audio and the fourth audio channel signal are associated with a second common horizontal plane (or second common height) of the audio scene, but with different horizontal positions (or azimuth positions) of the audio scene, the first horizontal plane being different from the second horizontal plane. It has been found that particularly good audio coding results (and, consequently, audio decoding results) can be achieved by using such spatial association of the audio channel signals.
[0041] In a preferred embodiment, the first audio channel signal and the second audio channel signal are associated with the first vertical plane (or first azimuth position) of the audio scene but with different vertical positions (or heights) of the audio scene. Moreover, the third audio channel signal and the fourth audio channel signal are associated preferably with the second vertical plane (or second azimuth position) of the audio scene, but with different vertical positions (or heights) of the audio scene, the first common vertical plane being different from the second common vertical plane. It has been found that such spatial association of the audio channel signals provides good audio coding quality.
[0042] In a preferred embodiment, the first audio channel signal and the second audio channel signal are associated with the left side of the audio scene, and the third audio channel signal and the fourth audio channel signal are associated with the right side of the audio scene. Accordingly, a good listening experience can be obtained while the decoding is usually bit rate efficient.
[0043] In a preferred embodiment, the first audio channel signal and the third audio channel signal are associated with the bottom of the audio scene and the second audio channel signal and the fourth audio channel signal are associated with the top of the audio scene. This setting also helps you achieve efficient audio encoding with a good hearing experience.
[0044] In a preferred embodiment, the audio encoder is adapted to make a horizontal call when it provides an encoded representation of the downmix signals on the basis of the first downmix signal and the second downmix signal using multi-channel coding. In parallel with the above explanations for the audio decoder, it has been found that a particularly good listening experience can be obtained if the horizontal joining is performed in the second stage of the audio encoder (compared to the first stage of the audio encoder), since the horizontal position of the audio object is of particular importance. for the listener, the second step of the hierarchical audio encoder typically corresponds to the first step of the hierarchical audio decoder described above.
[0045] In a preferred embodiment, the audio encoder is arranged to make a vertical connection when providing the first downmix signal from the first audio channel signal and the second audio channel signal using multi-channel decoding. Moreover, the audio decoder is preferably adapted to perform vertical combining when it provides the second downmix signal based on the third audio channel signal and the fourth audio channel signal. Correspondingly, a vertical connection is performed in the first step of the audio encoder. This is advantageous because the vertical position of the audio object is usually not as important to the human listener as the horizontal position of the audio object, so that the degradation of reproduction that is caused by hierarchical coding (and consequently, hierarchical decoding) can be kept to a reasonably small extent.
[0046] In a preferred embodiment, the audio encoder is adapted to provide a jointly coded representation of the first downmix signal and the second downmix signal on the basis of the first downmix signal and the second downmix signal using a prediction based multi-channel coding. It has been found that such predictive multi-channel coding is well suited to the common coding which is pre-formed in the second step of the hierarchical coder. Refer to the explanations for the audio decoder above which also apply in parallel here.
[0047] In a preferred embodiment, a prediction parameter describing the contribution of the signal element that was obtained using the signal element of the previous frame to provide the downmix signal of the current frame is provided using a prediction based multi-channel coding. Accordingly, a good signal reconstruction can be obtained at that side of the audio encoder that uses this prediction parameter describing the contribution of the signal element that originated from the signal element of the previous frame to downmix the current frame.
[0048] In a preferred embodiment, the prediction based multi-channel coding operates in the MDCT domain. Accordingly, the prediction based multi-channel coding is well suited for post-coding the prediction based multi-channel coding output (e.g., a common downmix signal), this post-coding typically being performed in the MDCT domain to keep blocking artifacts reasonably low.
[0049] In a preferred embodiment the prediction based multi-channel coding is USAC compound stereo prediction coding. The use of complex USAC predictive stereo coding facilitates implementation as existing hardware and / or program code can be easily reused to implement a hierarchical audio encoder.
[0050] In a preferred embodiment, the audio encoder is adapted to provide a jointly coded representation of the first downmix signal and the second downmix signal on the basis of the first downmix signal and the second downmix signal using the residual signal assisted multi-channel coding. Accordingly, a particularly good reproduction quality can be achieved on the audio decoder side.
[0051] In a preferred embodiment, the audio encoder is adapted to provide the first downmix signal on the basis of the first audio channel signal and the second audio channel signal using parameter based multi-channel coding. Moreover, the audio encoder is adapted to drive the second downmix signal on the basis of the third audio channel signal and the fourth audio channel signal using parameter based multi-channel coding. It has been found that the use of parameter-based multi-channel coding provides a good compromise between the reproduction quality and the bit rate when using a hierarchical audio encoder in the first step.
[0052] In a preferred embodiment, the parameter based multi-channel coding is adapted to provide one or more parameters describing a desired correlation between the two channels and / or level differences between the two channels. Accordingly, it is possible to encode efficiently at a moderate bit rate without significantly degrading the audio quality.
[0053] In a preferred embodiment, the parameter based multi-channel coding operates in a QMF domain which is well suited for pre-processing which may be performed on the audio channel signals.
[0054] In a preferred embodiment, the parameter based multi-channel encoding is MPEG surround 2-1-2 encoding or unified stereo encoding. The use of such coding concepts can significantly reduce the implementation effort.
[0055] In a preferred embodiment, the audio encoder is adapted to provide the first downmix signal on the basis of the first audio channel signal and the second audio channel signal using a residual multi-channel coding. Moreover, the audio encoder may be adapted to provide the second downmix signal based on the third audio channel signal and the fourth audio channel signal using a residual-assisted multi-channel coding. Accordingly, it is possible to achieve even better sound quality.
[0056] In a preferred embodiment, the audio encoder is adapted to provide a jointly coded representation of the first residual signal that is obtained while at least the first audio channel signal and the second audio channel signal are jointly coded, and the second residual signal that is obtained. when at least the third audio channel signal and the fourth audio channel signal are jointly encoded using multi-channel coding. It has been found that the hierarchical coding concept can even be applied to the residual signals that are provided in the first step of hierarchical audio coding. By using the common coding of the residual signals, relationships (or correlations) between the audio channel signals can be exploited as these relationships (or correlations) are typically also reflected in the residual signals.
[0057] In a preferred embodiment, the first residual signal and the second residual signal are associated with different horizontal positions (or azimuth positions) of the audio scene. Accordingly, the relationships between the residual signals may be coded with good precision in the second step of hierarchical coding. This makes it possible to recreate the relationship (or correlation) between different horizontal positions (or azimuth positions) with a good hearing impression on the audio decoder side.
[0058] In a preferred embodiment, the first residual signal is associated with the left side of the audio scene and the second residual signal is associated with the right side of the audio scene. Correspondingly, the joint coding of the first residual signal and the second residual signal, which are associated with different horizontal positions (or azimuth positions) of the audio scene, is performed in a second audio encoder step which allows high quality reproduction to the side of the audio decoder.
[0059] A preferred embodiment of the invention creates a method for providing at least four extended bandwidth audio channel signals based on an encoded representation. The method includes providing a first downmix signal and a second downmix signal based on a jointly coded representation of the first downmix signal and the second downmix signal by (first) multi-channel decoding. The method also includes providing at least a first audio channel signal and a second audio channel signal based on the first downmix signal from using (second) multi-channel decoding and providing at least a third signal the audio channel and the fourth audio channel signal from the second downmix signal using (third) multi-channel decoding The method also includes performing a (first) common multi-channel bandwidth extension based on the first audio channel signal and the third audio channel signal to obtain a first signal with extended bandwidth and a third signal with extended bandwidth, wherein the channel bandwidth extension uses a relationship between the first audio channel signal and the third audio channel signal. The method also includes performing a (second) multi-channel bandwidth extension based on the second audio channel signal and the fourth audio channel signal to obtain a second channel signal having a bandwidth extended bandwidth. This method is based on the same considerations as the audio decoder described above.
[0060] A preferred embodiment of the invention creates a method for providing an encoded representation based on at least four audio channel signals. The method includes deriving a first set of common bandwidth extension parameters based on the first audio channel signal and the third audio channel signal. The method also includes deriving the second set of common bandwidth extension parameters based on the second audio channel signal and the fourth audio channel signal. The method further comprises coding the at least the first audio channel signal and the second audio channel signal jointly using multi-channel coding to obtain a first downmix signal and co-coding the at least the third audio channel signal and the fourth audio channel signal by multi-channel coding to obtain a second signal. downmix. The method further comprises coding the first downmix signal and the second downmix signal jointly using multi-channel coding to obtain an encoded representation of the downmix signals. This method is based on the same considerations as the audio encoder described above.
[0061] Further embodiments of the invention create computer programs for performing the methods listed herein.
Brief description of the figures
[0062] Embodiments of the present invention will then be described with reference to the accompanying Figures in which:
Fig. 1 shows a basic block diagram of an audio encoder according to an embodiment of the present invention;
Fig. 2 shows a basic block diagram of an audio decoder according to an embodiment of the present invention;
Fig. 3 shows a basic block diagram of an audio decoder, according to another embodiment of the present invention;
Fig. 4 shows a basic block diagram of an audio encoder according to an embodiment of the present invention;
Fig. 5 shows a basic block diagram of an audio decoder according to an embodiment of the present invention;
Fig. 6 shows a basic block diagram of an audio decoder, according to another embodiment of the present invention;
Fig. 7 is a flowchart of a method for providing an encoded representation based on at least four audio channel signals, according to an embodiment of the present invention;
Fig. 8 is a flowchart of a method for providing at least four audio channel signals based on an encoded representation, according to an embodiment of the invention;
Fig. 9 is a flowchart of a method for providing an encoded representation based on at least four audio channel signals, according to an embodiment of the invention; and
Fig. 10 shows a block diagram of a method for providing at least four audio channel signals based on an encoded representation, according to an embodiment of the invention;
Fig. 11 shows a basic block diagram of an audio encoder according to an embodiment of the invention;
Fig. 12 shows a basic block diagram of an audio encoder, according to another embodiment of the invention;
Fig. 13 shows a basic block diagram of an audio decoder according to an embodiment of the invention;
Fig. 14a shows a bitstream syntax representation that can be used with the audio encoder according to Fig. 13;
Fig. 14b shows a tabular representation of various qcelndex values;
Fig. 15 shows a basic block diagram of a 3D audio encoder in which concepts of the present invention can be applied;
Fig. 16 shows a basic block diagram of a 3D audio decoder to which concepts of the present invention can be applied; and
Fig. 17 shows the principal block diagram of a format converter.
Fig. 18 shows a graphical representation of a topological structure of a Quad Channel Element (QCE), according to an embodiment of the present invention;
Fig. 19 shows a basic block diagram of an audio decoder according to an embodiment of the present invention;
Fig. 20 shows a detailed block diagram of a QCE decoder according to an embodiment of the present invention; and
Fig. 21 shows a detailed block diagram of a four channel component encoder according to an embodiment of the present invention.
Detailed description of the embodiments 1. The audio encoder according to Fig. 1
[0063] Fig. 1 shows a basic block diagram of an audio encoder, indicated in its entirety by 100. The audio encoder 100 is adapted to provide an encoded representation from at least four audio channel signals. The audio encoder 100 is adapted to receive the first audio channel signal 110, the second audio channel signal 112, the third audio channel signal 114, and the fourth audio channel signal 116. Moreover, the audio encoder 100 is adapted to provide an encoded representation of the first downmix signal 120 and the second downmix signal 122, as well as a jointly encoded representation 130 of the residual signals. The audio encoder 100 includes a residual-assisted multi-channel encoder 140, which is adapted to co-code the first audio channel signal 110 and the second audio channel signal 112 using the residual-assisted multi-channel coding to obtain the first downmix signal 120 and the first residual signal 142. The audio encoder 100 also includes a residual multi-channel encoder 150, which is adapted to co-encode at least the third audio channel signal 114 and the fourth audio channel signal 116 using residual-assisted multi-channel encoding to obtain a second downmix 122 signal and a second residual signal. 152. The audio decoder 100 also includes a multi-channel encoder 160 that is adapted to jointly encode the first residual signal 142 and the second residual signal 152 using multi-channel encoding to obtain a jointly encoded representation 130 of the residual signals 142, 152.
[0064] Regarding the functionality of the audio encoder 100, it should be noted that the audio encoder 100 performs hierarchical encoding in which the first audio channel signal 110 and the second audio channel signal 112 are collectively encoded using a residual multi-channel encoding 140 in which there is both the first downmix signal 120 and the first residual signal 142. The first residual signal 142 may, for example, describe the differences between the first audio channel signal 110 and the second audio channel signal 112, and / or may describe some or any signal characteristics that cannot be represented by the first downmix signal 120 and optional parameters that are it may be provided by the multi-channel encoder 140 assisted by the residual signal. In other words, the first residual signal 142 may be a residual signal which allows a refinement of the decoding result that can be obtained from the first downmix signal 120 and any possible parameters that can be provided by the residual multi-channel encoder 140. For example, the first residual signal 142 may permit at least a partial reconstruction of the waveform of the first audio channel signal 110 and the second audio channel signal 112 at the audio decoder side as compared to the usual reconstruction of the high level signal characteristic (e.g., correlation characteristic, covariance characteristic, features level differences and the like). Similarly, the residual-assisted multi-channel encoder 150 provides both the second downmix 122 and the second residual signal 152 based on the third audio channel signal 114 and the fourth audio channel signal 116, such that the second residual signal improves reconstruction of the third audio channel signal 114 and the fourth audio channel signal 116 at the audio decoder side. The second residual signal 152 may consequently serve the same functionality as the first residual signal 142. However, if the audio channel signals 110, 112, 114, 116 include some correlation, the first residual signal 142 and the second residual signal 152 are typically also correlated to some extent. . Accordingly, co-encoding the first residual 142 and the second residual 152 with the multi-channel encoder 160 typically involves high performance since multi-channel encoding of correlated signals typically reduces bit rate by using a relationship. Consequently, the first residual signal 142 and the second residual signal 152 can be encoded with good precision while maintaining a reasonable low bit rate of the jointly coded representation 130 of the residual signals.
In summary, the embodiment of Fig. 1 provides hierarchical multi-channel coding where good reproduction quality can be achieved by using residual multi-channel encoders 140,150, while the required bit rate can be moderately maintained by jointly coding the first residual 142 and the second. residual signal 152.
[0066] It is possible to further optionally improve the audio encoder 100. Some of these improvements will be described with reference to the Figures. 4, 11, and 12. However, it should be noted that the audio encoder 100 may also be adapted in parallel with the audio decoder described herein, with the functionality of the audio encoder typically being opposite to that of the audio decoder.
2. An audio decoder according to Fig. 2
[0067] Fig. 2 shows a basic block diagram of an audio decoder, which is indicated in its entirety by 200.
[0068] The audio decoder 200 is adapted to receive an encoded representation that includes an encoded representation 210 of the first residual signal and the second residual signal collectively. The audio decoder 200 also receives a representation of the first downmix signal 212 and the second downmix signal 214. The audio decoder 200 is adapted to provide a first audio channel signal 220, a second audio channel signal 222, a third audio channel signal 224, and a fourth audio channel signal 226.
[0069] The audio decoder 200 includes a multi-channel decoder 230, which is adapted to provide the first residual signal 232 and the second residual signal 234 based on a jointly coded representation 210 of the first residual signal 232 and the second residual signal 234. The audio decoder 200 also includes a (first) residual multi-channel decoder 240 that is adapted to provide the first audio channel signal 220 and the second audio channel signal 222 based on the first downmix signal 212 and the first residual signal 232 using multi-channel decoding. The audio decoder 200 also includes a (second) residual multi-channel decoder 250 that is adapted to provide the third audio channel signal 224 and the fourth audio channel signal 226 based on the second downmix signal 214 and the second residual signal 234.
[0070] Regarding the functionality of the audio decoder 200, it should be noted that the audio signal decoder 200 provides the first audio channel signal 220 and the second audio channel signal 222 based on the (first) residual common multi-channel decoding 240, wherein the decoding quality for the decoding is of the multi-channel signal is incremented by the first residual signal 232 (compared to decoding of the untreated residual). In other words, the first downmix signal 212 provides "rough information about the first audio channel signal 220 and the second audio channel signal 222, where, for example, the differences between the first audio channel signal 220 and the second audio channel signal 222 can be described by (optional) ) parameters that can be received by the multi-channel decoder 240 supported by the residual signal and by the first residual signal 232. Consequently, the first residual signal 232 may, for example, permit a partial reconstruction of the waveform of the first audio channel signal 220 and the second audio channel signal 222.
Similarly, the (second) residual multi-channel decoder 250 provides the third audio channel signal 224 in the fourth audio channel signal 226 based on the second downmix signal 214, the second downmix signal 214 may, for example, "roughly describe the third channel signal 224. audio and the fourth audio channel 226 signal. In addition, the differences between the third audio channel signal 224 and the fourth audio channel signal 226 may, for example, be described by (optional) parameters that may be received by the (second) multi-channel decoder 250 assisted by the residual signal and by the second residual signal 234. Accordingly, , evaluation of the second residual signal 234 may, for example, enable a partial reconstruction of the waveforms of the third audio channel signal 224 and the fourth audio channel signal 226. Accordingly, the second residual signal 234 may improve the reconstruction quality of the third audio channel signal 224 and the fourth audio channel signal 226.
[0072] However, the first residual signal 232 and the second residual signal 234 are derived from the jointly coded representation 210 of the first residual signal and the second residual signal. Such multi-channel decoding, which is performed by multi-channel decoder 230, allows for high decoding efficiency since the first audio channel signal 220, the second audio channel signal 222, the third audio channel signal 224, and the fourth audio channel signal 226 are typically similar or "correlated." Accordingly, the first residual signal 232 and the second residual 234 are typically also similar or "correlated, which may be used by deriving the first residual 232 and the second residual 234 from the jointly encoded representation 210 using multi-channel decoding.
[0073] Consequently, it is possible to obtain high quality decoding at a moderate bit rate by decoding the residual signals 232, 234 based on the jointly coded representation 210, and by using each of the residual signals to decode two or more audio channel signals.
[0074] In summary, the audio decoder 200 allows high coding efficiency by providing high quality audio channel signals 220, 222, 224, 226.
[0075] It should be noted that additional features and functions that may be implemented optionally in the audio decoder 200 will be described later with reference to Figs. 3, 5, 6 and 13. However, it should be noted that the audio encoder 200 may include the above-mentioned advantages without any additional modifications.
3. An audio decoder according to Fig. 3
[0076] Fig. 3 shows a basic block diagram of an audio decoder according to another embodiment of the present invention. The audio decoder of Fig. 3 indicated in its entirety by 300. The audio decoder 300 is similar to the audio decoder 200 according to Fig. 2, so the above explanations also apply. However, the audio decoder 300 is supplemented with additional functions and functionalities as compared to the audio decoder 200 as will be explained below.
[0077] The audio decoder 300 is adapted to receive a jointly coded representation 310 of the first residual signal and the second residual signal. Furthermore, the audio decoder 300 is adapted to take together the coded representation 360 of the first downmix signal and the second downmix signal. Moreover, the audio decoder 300 is adapted to provide a first audio channel signal 320, a second audio channel signal 322, a third audio channel signal 324, and a fourth audio channel signal 326. The audio decoder 300 includes a multi-channel decoder 330, which is adapted to receive a jointly coded representation 310 of the first residual signal and the second residual signal, and to provide a residual signal 332 and a second residual 334 based on this first residual signal. The audio decoder 300 also includes a (first) residual-assisted multi-channel decoding 340 that receives the first residual signal 332 and the first downmix signal 312, and provides the first audio channel signal 320 and the second audio channel signal 322. Audio decoder 300 also includes (second) multi-channel decoding 350 assisted by the residual signal, which is adapted to receive the second residual signal 334 and the second downmix signal 314, and for providing a third audio channel signal 324 and a fourth audio channel signal 326.
[0078] The audio decoder 300 also comprises another multi-channel decoder 370, which is adapted to receive a jointly coded representation 360 of the first downmix signal and the second downmix signal, and to provide, based on the first downmix signal 312 and the second downmix signal 314.
[0079] Further specific details of the audio decoder 300 will be described below. However, it should be noted that a proper audio decoder need not implement a combination of all of these additional features and functionalities. Instead, the functions and functions described below may be individually added to the audio decoder 200 (or any other audio decoder) to gradually improve the audio decoder 200 (or any other audio decoder).
[0080] In a preferred embodiment, the audio decoder 300 receives a jointly coded representation 310 of the first residual signal and the second residual signal, the jointly coded representation 310 may include the downmix signal of the first residual signal 332 and the second residual signal 334 and a common residual signal of the first residual signal 332 and the second residual 334. In addition, the jointly coded representation 310 may, for example, include one or more prediction parameters. Accordingly, the multi-channel decoder 330 may be a residual-assisted multi-channel decoder based on a prediction. For example, the multi-channel decoder 330 may be a USAC composite stereo prediction as described, for example, in the section "Composite stereo prediction of the international standard ISO / IEC 23003-3: 2012. For example, multi-channel decoder 330 may be adapted to evaluate a prediction parameter describing a contribution of a signal element that comes from a signal element of the previous frame to provide the first residual signal 332 and the second residual 334 for the current frame. Moreover, the multi-channel decoder 330 may be adapted to use a common residual (which is included in the jointly coded representation 310) with the first character to obtain the first residual signal 332 and to use a common residual (which is included in the jointly coded representation 310) with the first character. a second character that is opposite to the first character to obtain the second residual signal 334. Thus, the common residual signal may, at least partially, describe the differences between the first residual 332 and the second residual 334. However, the multi-channel decoder 330 may evaluate the downmix signal, the common residual signal, and one or more prediction parameters that are included in the total coded representation 310 to obtain a first residual signal 332 and a second residual signal 334 as described in the above-mentioned international standard ISO / IEC 23003-3: 2012. Further, it should be noted that the first residual signal 332 may be associated with the first horizontal position (or azimuth position), e.g., the left horizontal position, and that the second residual signal 334 may be associated with the second horizontal position (or azimuth position). for example, in the right vertical position of the audio stage.
[0081] The jointly encoded representation 360 of the first downmix signal and the second downmix signal preferably comprises the downmix signal of the first downmix signal and the second downmix signal, a common residual signal of the first downmix signal and the second downmix signal, and one or more prediction parameters. In other words, there is a "common downmix signal to which the first downmix signal 312 and the second downmix signal 314 are downmixed, and there is a" common residual signal that can describe, at least partially, the differences between the first downmix signal 312 and the second downmix signal 314. Decoder. multi-channel 370 is preferably a prediction-based residual multi-channel decoder, e.g. a USAC compound stereo prediction decoder. In other words, the multi-channel decoder 370 that provides the first downmix signal 312 and the second downmix signal 314 may be substantially identical to the multi-channel decoder 330 that provides the first residual 332 and the second residual 334, such that the above explanations and references also apply. Furthermore, it should be noted that the first downmix signal 312 is associated preferably with the first horizontal position or azimuth position (e.g., the left horizontal or azimuth position) of the audio scene, and that the second downmix signal 314 is associated preferably with the second horizontal or azimuth position. (for example, the right horizontal or azimuth position) of the audio scene. Accordingly, the first downmix signal 312 and the first residual signal 332 may be associated with the same first horizontal position or azimuth position (e.g., the left horizontal position), and the second downmix signal 314 and the second residual signal 334 may be associated with the same. the second horizontal or azimuth position (for example, the right horizontal position). Accordingly, both multi-channel decoder 370 and multi-channel decoder 330 may perform horizontal split (or horizontal splitting or horizontal distribution).
[0082] The multi-channel decoder 340 assisted by the residual signal may advantageously be parameter based and may consequently obtain one or more parameters 342 describing a desired correlation between the two channels (e.g., between the first audio channel signal 320 and the second audio channel signal 322) and / or level differences between the two channels. For example, the residual-assisted multi-channel decoding 340 may be based on MPEG-Surround encoding (as described, for example, in ISO / IEC 23003-1: 2007) with a residual signal extension or a "unified stereo decoding" decoder (as described, for example, in ISO / IEC 23003-3, Chapter 7.11 (Decoder) and Annex B.21 (description Encoder and definition of the term "Unified Stereo"). Accordingly, the residual multi-channel decoder 340 may provide a first audio channel signal 320 and a second audio channel signal 322, with the first audio channel signal 320 and the second audio channel signal 322 being associated with vertically adjacent positions of the audio scene. For example, the first audio channel signal may be associated with the lower left position of the audio scene, and the second signal of the audio channel may be associated with the upper left position of the audio scene (such that the first audio channel signal 320 and the second audio channel signal 322 are related e.g. with identical horizontal or azimuth positions of the audio scene, or with azimuth positions separated by no more than 30 degrees). In other words, the residual multi-channel decoder 340 may perform vertical split (or distribution or split).
[0083] The functionality of the residual multi-channel decoder 350 may be identical to that of the residual multi-channel decoder 340, where the third audio channel signal may for example be associated with the lower right position of the audio scene and wherein the fourth audio channel signal may e.g. with the top right position of the audio scene. In other words, the third audio channel signal and the fourth audio channel signal may be associated with vertically adjacent positions of the audio scene and may be associated with the same horizontal position or azimuth position of the audio scene, with the residual multi-channel decoder 350 performing a vertical split (or splitting). or distribution).
[0084] In summary, the audio decoder 300 according to Fig. 3 performs hierarchical audio decoding in which left-right splitting is performed in the first steps (multi-channel decoder 330, multi-channel decoder 370), and in which high-low splitting is performed in a second. stage (multi-channel decoders 340,350 residual assisted). In addition, the residual signals 332, 334 are also encoded using the jointly coded representation 310 as well as the downmix signals 312, 314 (jointly coded representation 360). Thus, correlations between different channels are used both for encoding (and decoding) the downmix signals 312, 314 and for encoding (and decoding) residual signals 332, 334. Accordingly, high coding efficiency is achieved and correlations between the signals are well used. .
4. The audio encoder according to Fig. 4
[0085] Fig. 4 shows a basic block diagram of an audio encoder, according to another embodiment of the present invention. The audio signal encoder of Fig. 4 is indicated in its entirety by 400. The audio encoder 400 is adapted to receive four audio channel signals, namely a first audio channel signal 410, a second audio channel signal 412, a third audio channel signal 414 and a fourth audio channel signal 416. audio. Moreover, the audio signal encoder 400 is adapted to provide an encoded representation from the signals 410, 412, 414 and 416 of the audio channel, said encoded representation including an encoded representation 420 of the two downmix signals as well as an encoded representation of the first set of common width extension parameters 422. bandwidth and a second set of 424 common bandwidth extension parameters. Audio encoder 400 includes a first bandwidth extension parameter extractor 430, which is adapted to obtain a first set 422 of common bandwidth extraction parameters from the first audio channel signal 410 and the third audio channel signal 414. The audio encoder 400 also includes a second bandwidth extension parameter extractor 440, which is adapted to obtain a second set of common bandwidth extension parameters 424 from the second audio channel signal 412 and the fourth audio channel signal 416.
[0086] Furthermore, the audio signal encoder 400 includes a (first) multi-channel encoder 450 which is adapted to jointly encode at least the first audio channel signal 410 and the second audio channel signal 412 using multi-channel encoding to obtain a first downmix signal 452. Moreover, the audio encoder 400 also includes a (second) multi-channel encoder 460 that is adapted to jointly encode at least the third audio channel signal 414 and the fourth audio channel signal 416 using multi-channel encoding to obtain a second downmix signal 462. Furthermore, the audio encoder 400 also includes a (third) multi-channel encoder 470, which is adapted to jointly encode the first downmix signal 452 and the second downmix signal 462 by multi-channel encoding to obtain a jointly encoded representation 420 of the downmix signals.
[0087] As to the functionality of the audio encoder 400, it should be noted that the audio encoder 400 performs hierarchical multi-channel coding in which the first audio channel signal 410 and the second audio channel signal 412 are combined in a first step, and wherein the third audio channel signal 414 and the fourth audio channel signal are combined. The audio channels 416 are also combined in a first step to thereby obtain the first downmix signal 452 and the second downmix signal 462. The first downmix signal 452 and the second downmix signal 462 are then encoded together in a second step. However, it should be noted that the first bandwidth extension parameter extractor 430 provides the first common bandwidth extension parameter set 422 based on audio channel signals 410, 414 that are served by different multi-channel encoders 450, 460 in the first step of hierarchical multi-channel coding. Similarly, the second width extension parameter extractor 440 provides a second set 424 of common bandwidth extraction parameters based on different audio channel signals 412, 416 that are served by different multi-channel encoders 450, 460 in the first processing step. This specific processing order has the advantage that the bandwidth extension parameter sets 422,424 are based on channels that are only connected in the second step of hierarchical coding (i.e. in the multichannel 470 encoder). This is advantageous because it is desirable to combine such audio channels in a first step of hierarchical coding, the relationship of which is not very important with regard to the perception of the positions of the sound source. Rather, it is recommended that the relationship between the first downmix signal and the second downmix signal mainly determines the perception of the audio source location, since the relationship between the first downmix signal 452 and the second downmix signal 462 may be better preserved than the relationship between the individual audio channel signals 410,412,414,416. In other words, it has turned out that it is desirable that the first set of common bandwidth extension parameters 422 is based on two audio channels (audio channel signals) that contribute to different downmix signals 452, 462, and that the second set 424 of common bandwidth extension parameters is bandwidth is provided based on audio channel signals 412,416, which also contribute to different downmix signals 452,462, which is achieved by the above-described processing of the audio channel signals in hierarchical multi-channel coding. Consequently, the first common bandwidth extension parameter set 422 is based on a similar channel relationship compared to the channel relationship between the first downmix signal 452 and the second downmix signal 462, the latter typically dominating the spatial impression generated at the audio decoder side. Correspondingly, the provision of the first set of bandwidth extension parameters 422 as well as the provision of the second set of bandwidth extension parameters 424 is well suited to the spatial auditory impression that is generated at the side of the audio decoder.
5. An audio decoder according to Fig. 5
[0088] Fig. 5 shows a basic block diagram of an audio decoder, according to another embodiment of the present invention. The audio decoder according to Fig. 5 is indicated in its entirety by 500.
[0089] The audio decoder 500 is adapted to receive a jointly coded representation 510 of the first downmix signal and the second downmix signal. In addition, the audio decoder 500 is adapted to provide the first extended bandwidth channel signal 520, the second extended bandwidth channel signal 522, the third extended bandwidth channel signal 524, and the fourth extended bandwidth channel signal 526.
[0090] Audio decoder 500 includes a (first) multi-channel decoder 530, which is adapted to provide the first downmix signal 532 and the second downmix signal 534 based on a jointly coded representation 510 of the first downmix signal and the second signal downmix by multi-channel decoding. The audio decoder 500 also includes a (second) multi-channel decoder 540, which is adapted to provide at least a first audio channel signal 542 and a second audio channel signal 544 based on the first downmix signal 532 using multi-channel decoding. The audio decoder 500 also includes a (third) multi-channel decoder 550, which is adapted to provide at least a third audio channel signal 556 and a fourth audio channel signal 558 from the second downmix signal 544 using multi-channel decoding. Further, the audio decoder 500 includes a (first) multi-channel bandwidth extension 560 that is adapted to perform multi-channel bandwidth extension based on the first audio channel signal 542 and the third audio channel signal 556 to obtain a first extended bandwidth channel signal 520 and a third signal. Extended Bandwidth Channel 524 Further, the audio decoder includes a (second) multi-channel bandwidth extension 570, which is adapted to perform multi-channel bandwidth extension based on the second audio channel signal 544 and the fourth audio channel signal 558 to obtain a second extended bandwidth channel signal 522 and a fourth extended bandwidth channel signal 526.
[0091] Regarding the functionality of audio decoder 500, it should be noted that audio decoder 500 performs hierarchical multi-channel decoding in which the separation between the first downmix signal 532 and the second downmix signal 534 is performed in a first step of the hierarchical decoding process and wherein the first audio channel signal 542 and the second audio channel signal 544 is output from the first downmix signal 532 in a second hierarchical decoding step, and wherein the third audio channel signal 556 and the fourth audio channel signal 558 are derived from the second downmix signal 550 in a second hierarchical decoding step. However, both the first multi-channel bandwidth extension 560 and the second multi-channel bandwidth extension 570 receive each one audio channel signal that is derived from the first downmix signal 532 and one audio channel signal that is derived from the second downmix signal 534. Since better channel separation is typically achieved by (first) multi-channel decoding 530 which is performed as the first step of hierarchical multi-channel decoding as compared to the second step of hierarchical decoding, it can be seen that each multi-channel bandwidth extension 560, 570 receives input signals that are are well separated (as they come from the first downmix signal 532 and the second downmix signal 534, which are well separated in the channels). Thus, the multi-channel bandwidth extension 560, 570 may take into account stereo characteristics which are important for the auditory experience and which are well represented by the relationship between the first downmix signal 532 and the second downmix signal 534, and therefore can provide a good hearing impression.
In other words, the cross structure of the audio decoder in which each of the steps 560, 570 of the multi-channel bandwidth extension receives the input signals from both (second step) of the multi-channel decoders 540, 550 allows for a good multi-channel bandwidth extension that takes into account the stereo compound between channels.
[0093] However, it should be noted that the audio decoder 500 may be supplemented with any of the functions and functions described herein in relation to the audio decoders of the Figures. 2, 3, 6 and 13, wherein it is possible to input individual features into the audio decoder 500 to gradually improve the performance of the audio decoder.
6. An audio decoder according to Fig. 6
[0094] Fig. 6 shows a basic block diagram of an audio decoder according to another embodiment of the present invention. The audio decoder according to Fig. 6 is indicated in its entirety by 600. The audio decoder 600 according to Fig. 6 is similar to the audio decoder 500 according to Fig. 5, so that the above explanations also apply. However, the audio decoder 600 has been supplemented with some features and functionalities that may also be incorporated, individually or in combination, into the audio decoder 500 for enhancement.
[0095] The audio decoder 600 is adapted to receive the coded representation 610 of the first downmix signal and the second downmix signal together and to provide the first extended bandwidth signal 620, the second extended bandwidth signal 622, the third extended bandwidth signal 624 and the fourth signal. 626 with extended bandwidth. The audio decoder 600 includes a multi-channel decoder 630, which is adapted to receive a jointly coded representation 610 of the first downmix signal and the second downmix signal, and to provide, based on the first downmix signal 632 and the second downmix signal 634. The audio decoder 600 further comprises a multi-channel decoder 640 which is adapted to receive the first downmix signal 632 and provide, based on the first audio channel signal 542 and the second audio channel signal 544. The audio decoder 600 also includes a multi-channel decoder 650, which is adapted to receive the second downmix signal 634 and provide the third audio channel signal 656 and the fourth audio channel signal 658. Audio decoder 600 also includes a (first) multi-channel bandwidth extension 660, which is adapted to receive the first audio channel signal 642 and the third audio channel signal 656 and provide, based on the first extended bandwidth channel signal 620 and the third channel signal 624. with extended bandwidth. In addition, the (second) multi-channel bandwidth extension 670 receives the second audio channel signal 644 and the fourth audio channel signal 658, and provides therefrom a second extended bandwidth channel signal 622 and a fourth extended bandwidth channel signal 626.
[0096] The audio decoder 600 also includes a further multi-channel decoder 680, which is adapted to receive a jointly coded representation 682 of the first residual signal and the second residual signal, and which provides a first residual signal 684 for use by multi-channel decoder 640 and the second residual signal based thereon. 686 for use by the 650 multi-channel decoder.
[0097] The multi-channel decoder 630 is preferably a residual-assisted multi-channel decoder based on a prediction. For example, multi-channel decoder 630 may be substantially identical to multi-channel decoder 370 described above. For example, multi-channel decoder 630 may be a composite USAC stereo predictive decoder as mentioned above and as described in the above USAC standard. Accordingly, the jointly encoded representation 610 of the first downmix signal and the second downmix signal may, for example, include a (common) downmix signal of the first downmix signal and the second downmix signal, a (common) residual signal of the first downmix signal and the second downmix signal and one or more prediction parameters that are are evaluated by the multi-channel decoder 630.
[0098] Furthermore, it should be noted that the first downmix signal 632 may e.g. be associated with the first horizontal position or the azimuth position (e.g., left horizontal position) of the audio scene and that the second downmix signal 634 may e.g. be associated with the second horizontal position. or the azimuth position (for example, the right-hand horizontal position) of the audio scene.
[0099] Moreover, the multi-channel decoder 680 may, for example, be a prediction based multi-channel decoder associated with a residual signal. The multi-channel decoder 680 may be substantially identical to the multi-channel decoder 330 described above. For example, the multi-channel decoder 680 may be a composite USAC stereo predictive decoder as mentioned above. Consequently, the jointly coded representation 682 of the first residual signal and the second residual signal may include a (common) downmix signal of the first residual signal and the second residual signal, a (common) residual signal of the first residual signal and the second residual signal and one or more prediction parameters that are rated by a 680 multi-channel decoder. Further, it should be noted that the first residual signal 684 may be associated with the first horizontal position or azimuth position (e.g., left horizontal position) of the audio scene, and that the second residual signal 686 may be associated with the second residual horizontal position or azimuth position (e.g. right (H) position) of the audio scene.
[0100] The multi-channel decoder 640 may, for example, be a parameter-based multi-channel decoding, such as MPEG surround multi-channel decoding as described above and in the referenced standard. However, in the presence of the (optional) multi-channel decoder 680 and the (optional) first residual 684, the multi-channel decoder 640 may be a residual multi-channel decoder based on a parameter, such as, for example, a unified stereo decoder. Thus, the multi-channel decoder 640 may be substantially identical to the multi-channel decoder 340 described above, and the multi-channel decoder 640 may, for example, obtain the parameters 342 described above.
[0101] Similarly, multi-channel decoder 650 may be substantially identical to multi-channel decoder 640. Accordingly, multi-channel decoder 650 may, for example, be parameter-driven and may optionally be residual (in presence of optional multi-channel decoder 680).
[0102] Moreover, it should be noted that the first audio channel signal 642 and the second audio channel signal 644 are associated preferably with the vertically adjacent spatial positions of the audio scene. For example, the first signal of audio channel 642 is associated with the lower-left position of the audio scene, and the second signal of audio channel 644 is associated with the upper-left position of the audio scene. Correspondingly, the multi-channel decoder 640 performs a vertical division (or separation or distribution) of the audio content described by the first downmix signal 632 (and, optionally, by the first residual signal 684). Likewise, the third audio channel signal 656 and the fourth audio channel signal 658 are associated with vertically adjacent positions of the audio scene and are associated preferably with the same horizontal position or azimuth of the audio scene. For example, the third audio channel signal 656 is preferably associated with the lower right position of the audio scene, and the fourth audio channel signal 658 is preferably associated with the upper right position of the audio scene. Thus, multi-channel decoder 650 performs vertical splitting (or separating or distributing) of the audio content described by the second downmix signal 634 (and, optionally, the second residual signal 686).
[0103] However, the first multi-channel bandwidth extension 660 receives the first audio channel signal 642 and the third audio channel 656 which are related to the lower left positions and lower right positions of the audio scene. Correspondingly, the first multi-channel bandwidth extension 660 performs the multi-channel bandwidth extension based on two audio channel signals that are associated with the same horizontal plane (e.g., a lower horizontal plane) or height of the audio scene and different sides (left / right) of the audio scene. Accordingly, the multi-channel bandwidth extension may account for stereo characteristics (e.g., human stereo perception) when performing bandwidth extension. Likewise, the second multi-channel bandwidth extension 670 may also take into account the stereo characteristics because the second multi-channel bandwidth extension acts on the audio channel signals of the same horizontal plane (e.g., horizontal top plane) or height, but at different horizontal positions (sides) (left). / right) of the audio scene.
[0104] To further illustrate, hierarchical audio decoder 600 includes a structure in which left / right partitioning (or splitting or distribution) is performed in a first step (multi-channel decoding 630, 680), and the vertical partitioning (splitting or distribution) is performed in a second step (multi-channel 640, 650 decoding), and wherein the multi-channel bandwidth extension operates on a left / right signal pair (multi-channel bandwidth extension 660, 670). Such "crossing of the decoding paths allows left / right separation, which is particularly important for the auditory impression (e.g., more important than the upper / lower partition), can be performed in the first processing step of the hierarchical audio decoder and that the multi-channel bandwidth extension it can also be made on a pair of left-right audio signals, again giving a particularly good auditory impression. The up / down split is performed as an intermediate state between the left-right separation and the multi-channel bandwidth extension, which allows the output of four audio channel signals (or bandwidth extended channel signals) without significantly degrading the auditory experience.
7. The method according to Fig. 7
[0105] Fig. 7 shows a flowchart of a method 700 for providing an encoded representation based on at least four audio channel signals.
[0106] The method 700 includes co-coding 710 of at least the first audio channel signal and the second audio channel signal using residual multi-channel coding to obtain a first downmix signal and a first residual signal. The method also includes co-coding 720 at least the third audio channel signal and the fourth audio channel signal using residual multi-channel coding to obtain a second downmix signal and a second residual signal. The method further comprises co-coding 730 the first residual signal and the second residual signal using multi-channel coding to obtain an encoded representation of the residual signals. However, it should be noted that the method 700 may be supplemented by any of the features and functionalities described herein with respect to audio encoders and audio decoders.
8. The method according to Fig. 8
[0107] Fig. 8 is a flowchart of a method 800 for providing at least four audio channel signals based on an encoded representation.
[0108] The method 800 includes providing 810 a first residual signal and a second residual signal based on the co-coded representation of the first residual signal and the second residual signal using multi-channel decoding. The method 800 also includes providing 820 a first audio channel signal and a second audio channel signal based on the first downmix signal and the first residual signal using residual-assisted multi-channel decoding. The method also includes providing 830 a third audio channel signal and a fourth audio channel signal based on the second downmix signal and the second residual signal using residual-assisted multi-channel decoding.
[0109] Additionally, it should be noted that the method 800 may be supplemented with any of the features and functionalities described herein with respect to audio decoders and audio encoders.
9. The method according to Fig. 9
[0110] Fig. 9 is a flowchart of a method 900 for providing an encoded representation based on at least four audio channel signals.
[0111] The method 900 includes obtaining 910 a first set of common bandwidth extension parameters based on the first audio channel signal and the third audio channel signal. The method 900 also includes deriving 920 a second set of common bandwidth extension parameters based on the second audio channel signal and the fourth audio channel signal. The method also includes coding at least the first audio channel signal and the second audio channel signal together using multi-channel coding to obtain a first downmix signal and jointly coding 940 the at least a third audio channel signal and a fourth audio channel signal using multi-channel coding to obtain a second downmix signal. The method also includes co-coding 950 the first downmix signal and the second downmix signal using multi-channel coding to obtain an encoded representation of the downmix signals.
[0112] It should be noted that some steps of method 900 that do not include k dependencies may be performed in any order or in parallel. Additionally, it should be noted that the method 900 may be supplemented by any of the features and functionalities described herein with respect to audio encoders and audio decoders.
10. The method according to Fig. 10
[0113] Fig. 10 shows a flowchart of a method 1000 for providing at least four audio channel signals based on an encoded representation.
[0114] The method 1000 includes providing 1010 a first downmix signal and a second downmix signal based on the co-coded representation of the first downmix signal and the second downmix signal using multi-channel decoding, providing 1020 the at least a first audio channel signal and a second audio channel signal based on the first downmix signal using multi-channel decoding, providing 1030 the at least a third audio channel signal and a fourth audio channel signal from the second downmix signal using multi-channel decoding, performing 1040 a multi-channel bandwidth extension based on the first audio channel signal and the third audio channel signal to obtain a first channel signal with extended bandwidth and a third signal with extended bandwidth, and performing 1050 a multi-channel bandwidth extension based on the second audio channel signal and the fourth audio channel signal to obtain a second channel signal with extended bandwidth and a fourth channel signal with extended bandwidth.
[0115] It should be noted that some steps of the method 1000 may be performed concurrently or in a different order. Additionally, it should be noted that the method 1000 may be supplemented with any of the features and functionalities described herein with respect to an audio encoder and an audio decoder.
11. Embodiments according to Figs. 11, 12 and 13
[0116] Certain additional embodiments of the present invention and underlying considerations will be described below.
[0117] Fig. 11 shows a schematic block diagram of an audio encoder 1100 according to an embodiment of the invention. The audio encoder 1100 is adapted to receive the lower left channel signal 1110, the upper left channel signal 1112, the lower right channel signal 1114 and the upper right channel signal 1116.
[0118] The audio encoder 1100 includes a first multi-channel audio encoder (or encoding) 1120 that is an MPEG surround 2-1-2 audio encoder (or encoding) or unified stereo audio encoder (or encoding) and that receives the lower left channel signal 1110 and upper left channel signal 1112. The first multi-channel audio encoder 1120 provides a left downmix signal 1122 and optionally a left residual signal 1124. Additionally, the audio encoder 1100 includes a second multi-channel encoder (or encoding) 1130 that is an MPEG-surround 2-1-2 encoder (or encoding) or a unified stereo encoder (or encoding) that receives the lower right channel signal 1114 and the right signal 1116. upper channel. A second multi-channel audio encoder 1130 provides a right downmix signal 1132 and, optionally, a right residual signal 1134. The audio encoder 1100 also includes a stereo encoder (or encoding) 1140 that receives a left downmix signal 1122 and a right downmix signal 1132. Moreover, the first stereo encoding 1140, which is complex prediction stereo coding, receives information 1142 about the psychoacoustic model from the psychoacoustic model. For example, psycho model information 1142 may describe the psychoacoustic significance of different frequency bands or frequency subbands, psychoacoustic masking effects, and the like. The stereo coding 1140 provides a "channel pair element (CPE) downmix" which is denoted by 1144 and which describes a left downmix signal 1122 and a right downmix signal 1132 1132 in co-coded form. Additionally, the audio encoder 1100 optionally includes a second stereo encoder (or encoding) 1150 that is adapted to receive an optional left residual signal 1124 and an optional right residual signal 1134, as well as psychoacoustic model information 1142. The second stereo encoding 1150, which is composite predictive stereo coding, is adapted to provide a "residual CPE" channel pair element that represents a left residual signal 1124 and a right residual signal 1134 in a co-coded form.
[0119] The encoder 1100 (as well as other audio encoders described herein) is based on the assumption that horizontal and vertical signal relationships are used by a hierarchical combination of available USAC stereo tools (i.e., coding concepts that are available in USAC coding) . Vertically adjacent channel pairs are combined using MPEG surround 2-1-2 or unified stereo (labeled 1120 and 1130) with band limited or fullband residual (labeled 1124 and 1134). The output of each vertical channel pair of the channel pairs is a downmix signal 1122, 1132, and for the unified stereo a residual signal 1124, 1134. In order to meet the perceptual requirements for binaural mask removal. binaural unmasking), both downmix signals 1122, 1132 are horizontally combined and co-coded using composite prediction (coder 1140) in the MDCT domain, which includes left-right and center-side coding capabilities. The same method can be applied to the horizontally linked residual signals 1124,1134. This concept is illustrated in Fig. 11.
[0120] The hierarchical structure explained with reference to Fig. 11 may be achieved by using both stereo tools (for example, both USAC stereo tools) and again separating the channels between them. Hence, no additional pre / post-processing step is required, and the bitstream syntax for utility data transmission remains unchanged (e.g., substantially unchanged from the USAC standard). This idea leads to the encoder structure shown in Fig. 12.
[0121] Fig. 12 shows a schematic block diagram of an audio encoder 1200, according to an embodiment of the invention. The audio signal encoder 1200 is adapted to receive a first channel signal 1210, a second channel signal 1212, a third channel signal 1214, and a fourth channel signal 1216. The audio encoder 1200 is adapted to provide a bit stream 1220 for a first channel pair member and a bit stream 1222 for the second channel pair member.
[0122] The audio encoder 1200 includes a first multi-channel encoder 1230 that is an MPEG-surround 2-1-2 encoder or a unified stereo encoder and that receives the first channel signal 1210 and the second channel signal 1212. In addition, the first multi-channel encoder 1230 provides a first downmix signal 1232, MPEG surround payload 1236 and, optionally, a first residual signal 1234. The audio encoder 1200 also includes a second multi-channel encoder 1240 which is an MPEG surround 2-1-2 encoder or a unified stereo encoder and which receives the third channel signal 1214 and the fourth channel signal 1216. The second multi-channel encoder 1240 provides a first downmix signal 1242, an MPEG surround signal 1246, and, optionally, a second residual signal 1244.
[0123] The audio encoder 1200 also includes a first stereo encoding 1250 which is complex predictive stereo encoding. The first stereo encoding 1250 receives the first downmix signal 1232 and the second downmix signal 1242. The first stereo encoding 1250 provides a co-coded representation 1252 of the first downmix signal 1232 and the second downmix signal 1242, wherein the co-coded representation 1252 may include a representation of the (common) downmix signal (first downmix signal 1232 and second downmix signal 1232). second downmix signal 1242) and a common residual signal (first downmix signal 1232 and second downmix signal 1242). Additionally, the (first) composite predictive stereo coding 1250 provides complex predictive payload data 1254 that typically includes one or more composite prediction coefficients. In addition, the audio encoder 1200 also includes a second stereo encoding 1260 that is complex predictive stereo encoding. The second stereo encoding 1260 receives the first residual signal 1234 and the second residual signal 1244 (or zero inputs if there is no residual signal provided by multi-channel encoders 1230, 1240). The second stereo encoding 1260 provides a co-coded representation 1262 of the first residual 1234 and the second residual 1244, which may, for example, include a (common) downmix signal (first residual 1234 and second residual 1244) and a common residual (first residual signal 1234). residual and second residual 1244). Additionally, the predictive stereo encoding 1260 provides complex predictive payload data 1264 that typically includes one or more prediction coefficients.
[0124] In addition, the audio encoder 1200 includes a psychoacoustic model 1270 that provides information that controls the first composite predictive stereo encoding 1250 and the second composite predictive stereo encoding 1260. For example, the information provided by the psychoacoustic model 1270 may describe which frequency bands or frequency bins are of high psychoacoustic importance and should be encoded with high accuracy. However, it should be noted that the use of the information provided by the psychoacoustic model 1270 is optional.
[0125] In addition, the audio signal encoder 1200 includes a first encoder and a mux 1280 that receives a co-coded representation 1252 from the first stereo composite encoding 1250, predictive composite payload 1254 from the first stereo composite encoding 1250 of MPEG surround payload 1236 from the first multi-channel encoder 1230. audio 1230. Additionally, the first encoding and multiplexing 1280 can receive information from a psychoacoustic model 1270 that describes, for example, which coding precision should be applied to which frequency band or frequency subbands, taking into account psychoacoustic masking effects and the like. Correspondingly, the first encoding and multiplexing 1280 provides a first bit stream 1220 of a first element of a channel pair.
[0126] In addition, the audio signal encoder 1200 includes a second encoding and multiplexing 1290, which is adapted to receive the co-coded representation 1262 provided by the second predictive stereo encoding 1260, the composite predictive payload data 1264 provided by the second predictive stereo encoding 1260, and data. 1240 utility MPEG surround provided by the second multi-channel audio encoder 1240. In addition, the second encoding and multiplexing 1290 may receive information from the psychoacoustic model 1270. Correspondingly, the second coding and multiplexing 1290 provides a bitstream 1222 of a second element of a channel pair.
[0127] Referring to the functionality of the audio encoder 1200, reference is made to the above explanations as well as to the explanations with regard to the audio encoders according to Figs. 2, 3, 5 and 6.
[0128] In addition, it should be noted that this concept can be extended to use multiple MPEG surround boxes to co-encode horizontal, vertical, or other geometrically related channels and combine downmix signals and residuals with complex prediction stereo pairs. taking into account their geometric and perceptual properties. This leads to a generalized decoder structure.
[0129] An embodiment of a four channel element is described below. A three dimensional audio coding system uses a hierarchical combination of four channels to form a four channel element (QCE). The QCE consists of two USAC channel pair (CPE) elements (or provides two USAC channel pair elements or receives the USAC channel pair elements). Vertical channel pairs are combined using MPS 2-1-2 or Unified Stereo. The downmix channels are co-coded in a first member of a channel pair, CPE. If residual coding is used, the residual signals are co-coded in the second member of the CPE channel pair, otherwise the signal in the second CPE is set to zero. Both elements of the CPE channel pairs use composite prediction for joint stereo coding, including left-right and center-side coding capabilities. To maintain the perceptual stereo properties of a high frequency signal part, stereo SBR (Spectral Band Replication) is used between the upper left / right channel pair and the lower left / right channel pair, an additional step to prevent the use of SBR.
[0130] A possible decoder structure will be described with reference to Fig. 13, which shows a schematic block diagram of an audio decoder according to an embodiment of the invention. The audio decoder 1300 is adapted to receive a first bit stream 1310 representing the first element of a channel pair and a second bit stream 1312 representing the second element of a channel pair. However, the first bit stream 1310 and the second bit stream 1312 may be included in a common overall bit stream.
[0131] The audio decoder 1300 is adapted to provide the first extended bandwidth channel signal 1320, which may, for example, represent the low-left position of the audio scene, the second extended bandwidth channel signal 1322, which may, for example, represent the upper-left channel. the position of the audio scene, third channel signal 1324 with extended bandwidth, which may, for example, be associated with the lower right position of the audio scene and the fourth extended bandwidth channel signal 1326 which may, for example, be associated with the upper right position of the audio scene.
[0132] The audio decoder 1300 includes decoding 1330 of the first bit stream, which is adapted to receive the bit stream 1310 for the first channel pair element and to provide, based thereon, a jointly coded representation of the two downmix signals, composite prediction data 1334 of useful data 1336. MPEG surround and data 1338 are useful for replicating spectral bandwidth. The audio decoder 1300 also includes a first composite predictive stereo decoding 1340, which is adapted to receive the jointly coded representation 1332 and the composite data 1334 and provide, based on the first downmix signal 1342 and the second downmix signal 1344. Similarly, the audio decoder 1300 includes a second decoding of a bit stream 1350 which is adapted to receive a bit stream 1312 for the second channel element and to provide, therefrom, a jointly coded representation 1352 of two residual signals, composite predictive payload data 1354, MPEG payload 1356 surround and the content of 1358 bits replication of the view bandwidth. The audio decoder also includes a second compound predictive stereo decoding 1360 that provides a first residual 1362 and a second residual 1364 based on the co-coded representation 1352 and the composite predictive payload 1354.
[0133] In addition, the audio decoder 1300 includes a first multi-channel MPEG surround decoding 1370, which is MPEG 2-1-2 surround decoding or unified stereo decoding. The first multi-channel MPEG surround decoding 1370 receives the first downmix signal 1342, the first residual signal 1362 (optional), and the MPEG surround payload data 1336 and provides a first audio channel signal 1372 and a second audio channel signal 1374 based thereon. The audio decoder 1300 also includes a second MPEG surround 1380 multi-channel decoding, which is MPEG 2-1-2 multi-channel surround decoding or unified multi-channel stereo decoding. The second MPEG surround multi-channel decoder 1380 receives the second downmix signal 1344 and the second residual signal 1364 (optional) as well as the useful MPEG surround data 1356 and provides a third audio channel signal 1382 and a fourth audio channel signal 1384 therefrom. The audio decoder 1300 also includes a first stereo spectral replication 1390 that is adapted to receive the first audio channel signal 1372 and the third audio channel signal 1382 as well as data 1338 of useful spectral bandwidth replication and to provide a first signal 1320 therefrom. a channel with extended bandwidth; and a third signal 1324 with extended bandwidth. Additionally, the audio decoder includes a second stereo bandwidth replication 1394 which is adapted to receive the second audio channel signal 1374 and the fourth audio channel signal 1384 as well as data 1358 of useful spectral replication bandwidth and to provide a second signal 1322 based thereon. channel with extended bandwidth; and a fourth signal 1326 with extended bandwidth.
[0134] Referring to the functionality of the audio decoder 1300, reference is made to the above discussion as well as the discussion of the audio decoder according to Figs. 2, 3, 5, and 6.
[0135] An example of a bitstream that can be used for the audio encoding / decoding described herein with reference to Figs. 14a and 14b will be described below. It should be noted that the bitstream may be, for example, a bitstream extension used in Unified Speech and Audio Coding (USAC) which is described in the above-mentioned standard (ISO / IEC 23003-3: 2012). For example, useful MPEG surround data 1236, 1246, 1336, 1356 and composite predictive data 1254, 1264, 1334, 1354 may be transmitted as for legacy channel pair members (i.e., for channel pair members according to USAC standard). For signaling the use of a four-channel QCE, the formation of a USAC pair may be extended by two bits as shown in Fig. 14a. In other words, two bits labeled "qcelndex" may be added to the USAC bitstream element "UsacChannelPairElementConfigO." The meaning of the parameter represented by the "qcelndex bits" can be defined, for example, as shown in the table in Fig. 14b.
[0136] For example, the two elements of a channel pair that make up the QCE may be transmitted as consecutive elements, first, the CPE containing the downmix channels and the MPS content for the first MPS field, and second, the CPE containing the residual signal (or null audio signal for MPS 2 -1 -2 encodings) and MPS content for the second MPS field.
[0137] In other words, there is only a small signaling overhead compared to a conventional USAC bitstream for transmitting a QCE Quad Channel.
[0138] However, of course also different bitstream formats can be used.
12. Encoding / decoding environment
[0139] In the following, an audio encoding / decoding environment will be described in which the concepts of the present invention can be applied.
[0140] The 3D audio codec system to which the concepts of the present invention can be applied is based on an MPEG-D USAC codec for decoding channel and object signals. In order to increase the efficiency of encoding a large number of objects, the MPEG SAOC technology was adapted. The three types of renderers perform the tasks of rendering objects to channels, rendering channels to headphones, or rendering channels to a different speaker configuration. When object signals are explicitly transmitted or parametrically encoded with SAOC, the respective object metadata information is compressed and multiplexed into a 3D audio bitstream.
[0141] Fig. 15 shows a basic block diagram of such an audio encoder and Fig. 16 shows a basic block diagram of such an audio decoder. In other words, Figs. 15 and 16 show different algorithmic blocks of a 3D audio system.
[0142] Now referring to Fig. 15, which shows a basic block diagram of a 3D encoder 1500, some details will be explained. The encoder 1500 includes an optional preconditioner / mixer 1510 that receives one or more channel signals 1512 and one or more object signals 1514 and provides, therefrom, one or more channel signals 1516, as well as one or more object signals 1518,1520. . The audio encoder also includes a USAC 1530 encoder and an optional SAOC 1540 encoder. The SAOC encoder 1540 is adapted to provide one or more SAO transport channels 1542 and SAOC side information 1544 based on one or more entities 1520 provided to the SAOC encoder. Further, the USAC encoder 1530 is adapted to receive channel signals 1516 including channels and pre-rendered objects from the pre-renderer / mixer, to receive one or more object signals 1518 from the pre-renderer // mixer and receive one or more transport channels.
SAOC 1542 and SAOC 1544 side information, and based thereon provide an encoded representation of 1532. Moreover, the audio encoder 1500 also includes an object-oriented metadata encoder 1550 which is adapted to receive metadata of an object 1552 (which can be evaluated by a preconditioner // mixer 1510) and to encode the object metadata to obtain the encoded 1554 object metadata. The encoded metadata is also received by the USAC encoder 1530 and used to provide the encoded representation 1532.
[0143] Some details about individual elements of the audio encoder 1500 will be described below.
[0144] Referring now to Fig. 16, the audio decoder 1600 will be described. The audio decoder 1600 is adapted to receive the encoded representation 1610 and to provide, based thereon, multi-channel speaker signals 1612, headphone signals 1614, and / or speaker signals 1616 in an alternate alternative. format (for example, 5.1 format).
[0145] The audio decoder 1600 includes a USAC decoder 1620, and provides one or more channel signals 1622, one or more pre-rendered object signals 1624, one or more object signals 1626, one or more transport channels 1628 SAOC, SAOC side information 1602, and compressed information. Object 1632 metadata based on the encoded representation of 1610. The audio decoder 1600 also includes an object renderer 1640 that is adapted to provide one or more rendered object signals 1642 based on the object signal 1626 and the object metadata information 1644, wherein the object metadata information 1644 is provided by the object metadata decoder 1650 based on the information about the object. Compressed Metadata for Object 1632. The audio decoder 1600 also includes, optionally, a SAOC decoder 1660, which is adapted to receive the SAOC transport channel 1628 and SAOC side information 1630, and to provide one or more rendered object signals 1662 based thereon. The audio decoder 1600 also includes a mixer 1670 that is adapted to receive channel signals 1622, pre-rendered object signals 1624, rendered object signals 1642 and rendered object signals 1662, and to provide, based thereon, a plurality of mixed-channel signals 1672 that may be mixed-channel. an example would be multi-channel loudspeaker signals 1612. The audio decoder 1600 may, for example, also include a binaural signal 1680 that is adapted to receive the mixed channel signals 1672 and provide the handset signals 1614 therefrom. In addition, the audio decoder 1600 may include a format conversion 1690 that is configured to receive mixed channel signals 1672 and layout information 1692 and to provide a loudspeaker signal 1616 based thereon for the alternate loudspeaker positioning.
[0146] In the following, some details about components of the audio encoder 1500 and the audio decoder 1600 will be described.
Pre-renderer / mixer
[0147] The pre-renderer 1510 may be optionally used to convert the channel plus the object insertion scene into the channel scene prior to encoding. For example, it may be functionally identical to the renderer / object mixer described below. Pre-rendering the objects can, for example, provide a deterministic signal entropy at the encoder input that is substantially independent of the number of simultaneously active object signals. Object metadata transmission is not required in pre-rendering objects. The discrete object signals are rendered to the channel layout whose encoder is adapted to be used. The object weights for each channel are obtained from the associated Object Metadata (OAM) 1552.
USAC core codec
[0148] Core coders 1530, 1620 for loudspeaker signals, discrete object signals, object downmix signals and pre-rendered signals are based on MPEG-D USAC technology. It supports multi-signal coding to create channel and object mapping information from the geometric and semantic information about the input channel and object assignment. This mapping information describes how input channels and objects are mapped to USAC channel elements (CPE, SCE, LFE), and the corresponding information is sent to the decoder. All additional content such as SAOC data or object metadata was passed through extension elements and was included in the encoder speed control.
[0149] The coding of objects is possible in various ways, depending on the speed / distortion requirements and the interactivity requirements of the renderer. The following coding variants for objects are possible:
1. 1. Pre-rendered objects: Object signals are pre-rendered and mixed with 22.2-channel signals before encoding. The next coding chain sees 22.2 channel signals.
2. 2. Discrete Object Waveforms: Objects are delivered as monophonic waves to the encoder. The encoder uses SCEs of single channel modules to transmit objects in addition to the channel signals. The decoded objects are rendered and blended on the receiver side. The compressed metadata of the object is sent to the listener / renderer along the side.
3. 3. Parametric waveforms of an object: Object properties and their interrelationships are described using SAOC parameters. The downmix of the object signals is USAC encoded. Parametric information is sent along the side. The number of downmix channels is chosen depending on the number of objects and the total data rate. The metadata information of the compressed objects is passed to the SAOC renderer.
SAOC
[0150] The SAOC encoder 1540 and the SAOC decoder 1660 for the object signals are based on MPEG SAOC technology. The system is capable of reproducing, modifying and rendering multiple audio objects based on fewer transmitted channels and additional parametric data (object level OLD differences, correlations between IOCs, DMG downmix gains). The additional parametric data has a much lower data rate than that required for the transmission of individual objects, which makes the encoding very efficient. The SAOC encoder takes object / channel signals as input as monophonic shapes and outputs the parametric information (which is packaged into 3Daudio bitstream 1532, 1610) and SAOC transport channels (which are encoded with single channel elements and transmitted).
[0151] The SAOC decoder 1600 reconstructs the object / channel signals from the decoded SAOC transport channels 1628 and parametric information 1630 and generates an output audio scene based on the reproduction circuitry, decompressed object metadata information, and optionally user interaction information.
Codec of Metadata Objects
[0152] For each object, the associated metadata that specifies the geometric position and volume of the object in 3D space is efficiently encoded by quantizing the object's properties over time and space. Compressed metadata of COAM 1554,1632 is sent to the receiver as side information.
Object Renderer / Mixer
[0153] The object renderer uses the metadata of the compressed object to generate a waveform of the object according to a given reproduction format. Each object is rendered to the specified output channels according to its metadata. The result of this block is the sum of the partial scores. If both channel-based content and discrete / parametric objects are decoded, the channel-based waveforms and the rendered object waveforms are mixed before outputting the resulting waveforms (or before feeding them to a post-processor module such as a binaural renderer or speaker renderer).
Binaural Renderer
[0154] The binaural renderer 1680 produces a binaural downmix of multi-channel audio material such that each input channel is represented by a virtual audio source. The processing takes place in the frame in the QMF domain. Binaural hearing is based on measured room binaural impulse responses.
Speaker renderer / format conversion
[0155] The loudspeaker renderer 1690 performs a conversion between the configuration of the transmitted channel and the desired reproduction format. Hence, it is called a "format converter. The format converter converts to a smaller number of output channels, ie creates downmixes. The system automatically generates optimized downmix matrices for a given combination of input and output formats and applies these matrices to the downmix process. The format converter allows standard speaker setups as well as random setups with non-standard speaker positions.
[0156] Fig. 17 shows a block diagram of the format converter. As can be seen, format converter 1700 receives the outputs of the mixer 1710, e.g., mixed channel signals 1672, and provides loudspeaker signals 1712, e.g., loudspeaker signals 1616. The format converter includes a QMF downmix process 1720 and a downmix configurator 1730, the downmix configurator 1730 providing configuration information for the downmix process 1720 based on the mixer output circuit information 1732 and the reproduction circuit information 1734.
[0157] Moreover, it should be noted that the terms described above, e.g., audio encoder 100, audio decoder 200 or 300, audio encoder 400, audio decoder 500 or 600, methods 700, 800, 900, or 1000, audio encoder 1100 or 1200, and decoder the audio 1300 may be used at the audio encoder 1500 and / or at the audio decoder 1600. For example, the aforementioned audio encoders / decoders may be used to encode or decode channel signals that are associated with different spatial positions.
13. Alternative Embodiments
[0158] Some additional embodiments will be described below.
[0159] Referring now to Figs. 18 to 21, additional embodiments of the invention will be explained.
[0160] It should be noted that a so-called "Quad Channel Element (QCE)" can be considered as an audio decoder tool which can be used for example for decoding three dimensional audio content.
[0161] In other words, a Quad Channel (QCE) is a four channel coding method for more efficiently coding horizontally and vertically spaced channels. The QCE consists of two consecutive CPEs and is created by hierarchically combining the Joint Stereo Tool with the Complex Stereo Prediction Tool capability in the horizontal direction and the MPEG Surround based stereo tool in the vertical direction. This is achieved by incorporating both stereo tools and by swapping the output channels between tool uses. Stereo SBR is performed in the horizontal direction to preserve the left-hand high frequency relationships.
[0162] Fig. 18 shows the topological structure of the QCE. Note that the QCE of Fig. 18 is very similar to the QCE of Fig. 11 so reference is made to the above explanations. It should be noted, however, that in the QCE of Fig. 18 it is not necessary to use a psychoacoustic model when making complex stereo predictions (while such use is of course possibly optional). Moreover, it can be seen that the first spectral bandwidth replication (Stereo SBR) is performed based on the lower left channel and the lower right channel, and that this second stereo spectral frequency replication (Stereo SBR) is performed based on the upper left and upper right channels.
[0163] Some terms and definitions that may be used in some embodiments will be given below.
[0164] Data element qcelndex indicates QCE CPE mode. With regard to the significance of the qcelndex bitstream variable, reference is made to Fig. 14b. Note that qcelndex describes whether two consecutive UsacChannelPairElement () elements are treated as a four-channel element (QCE). The different QCE modes are shown in Fig. 14b. The qcelndex value will be the same for two consecutive QCE elements.
[0165] Some of the aids that may be used in some embodiments of the invention will be defined below:
cplx_out_dmx_L [] first channel of the first CPE after composite stereo decoding cplx_out_dmx_R [] second channel of the first CPE after composite stereo decoding cplx_out_res_L [] second CPE after composite stereo decoding (zero if qcelndex = 1) cplx_out_res_R [] second channel of the second stereo CPE (second channel of the second CPE after composite stereo decoding zero if qcelndex = 1) mps_out_L_l [] first output channel of the first MPS field mps_out_L_2 [] second output channel of the first MPS field mps_out_R_l [] first output channel second MPS field mps_out_R_2 [] second output channel of the second MPS field sbr_out_L_l [] first output channel of the first SBR Stereo field sbr_out_R_l [] second output channel of the first SBR Stereo field sbr_out_L_2 [] first output channel of the second Stereo SBR field sbr_out_R_2 [] second output channel of the second field Stereo SBR
[0166] The decoding process which is performed in an embodiment of the invention will be explained below.
[0167] The syntax element (or bitstream element or data element) qcelndex in the UsacChannelPairElementConfig () indicates whether the CPE belongs to QCE and whether residual encoding is used. In case qcelndex is not equal to 0, the current CPE creates a QCE with its next element which will be the CPE having the same qcelndex. Stereo SBR is always used for QCE, therefore stereoConfigIndex is 3 and bsStereoSbr will be 1.
[0168] In the case of qcelndex == 1, only the contents for MPEG Surround and SBR and no corresponding audio signal data are included in the second CPE, and the syntax element bsResidualCoding is set to 0.
[0169] The presence of a residual signal in the second CPE is indicated by qcelndex == 2. In this case, the bsResidualCoding syntax element has the value 1.
[0170] However, also different and possible simplified signaling schemes may be used.
[0171] Decoding of Common Stereo with compound stereo prediction capability is performed as described in ISO / IEC 23003-3, Section 7.7. The resulting output of the first CPE is the MPS downmix signals cplx_out_dmx_L [] and cplx_out_dmx_R []. If a residual encoding (i.e. qcelndex == 2) is used, the output of the second CPE are residual signals MPS cplx_out_res_L [], cplx_out_res_R [], if no residual signal is passed (i.e. Qcelndex == 1), zero signals are inserted.
[0172] Before MPEG Surround decoding is applied, the second channel of the first item (cplx_out_dmx_R []) and the first channel of the second item (cplx out_res_L []) are swapped.
[0173] MPEG Surround decoding is performed as described in ISO / IEC 230033, subchapter 7.11. However, if residual encoding is used, the decoding may be modified compared to conventional MPEG surround decoding in some embodiments. Residual MPEG Surround decoding using SBR, as defined in ISO / IEC 23003-3, subsection 7.11.2.7 (Figure 23), has been modified such that Stereo SBR is also used for bsResidualCode == 1, resulting in decoder schemes shown in Fig. 19. Fig. 19 shows a block diagram of an audio encoder for bsResidualCoding == 0 and bsStereoSbr == 1.
[0174] As can be seen from Fig. 19, USAC core decoder 2010 provides a downmix (DMX) signal 2012 to a MPS (MPEG Surround) decoder 2020, which provides a first decoded audio signal 2022 and a second decoded audio signal 2024. The SBR stereo decoder 2030 receives the first decoded audio signal 2022 and the second decoded audio signal 2024 and provides, based thereon, a left bandwidth extended audio signal 2032 and a right bandwidth extended audio signal 2034.
[0175] Before using Stereo SBR, the second channel of the first item (mps_out_L_2 []) and the first channel of the second item (mps_out_R_l []) are swapped to allow left and right SBR formation. After using Stereo SBR, the second output channel of the first item (sbr_out_R_l []) and the first channel of the second item (sbr_out_L_2 []) are swapped again to restore the order of the input channels.
[0176] A QCE decoder structure is shown in Fig. 20, which shows diagrams of a QCE decoder.
[0177] It should be noted that the essential block diagram of Fig. 20 is very similar to the schematic block diagram of Fig. 13, so reference is also made to the above explanations. Moreover, it should be noted that certain signal tags have been added in Fig. 20 where reference is made to the definitions in this section. In addition, it shows the final channel re-sort that is performed after Stereo SBR.
[0178] Fig. 21 shows a schematic block diagram of a Quad Channel encoder 2200 according to an embodiment of the present invention. In other words, a four-channel encoder (four-channel component), which can be considered a core encoder tool, is shown in Fig. 21.
[0179] Quad Channel encoder 2200 includes a first Stereo SBR 2210 which receives the first left channel input 2212 and the second left channel input 2214, and which provides the first SBR content 2215, the left channel SBR output 2216 and the first signal based thereon. the output right channel SBR 2218. In addition, the Quad Channel encoder 2200 includes a second Stereo SBR that receives the second left channel input 2222 and the second right channel input 2224, and provides, based thereon, the first SBR content 2225, the first left channel SBR output 2226, and the first left channel output signal. 2228 SBR of the right channel.
[0180] The Quad Channel encoder 2200 includes a first MPEGSurround multi-channel encoder 2230 (MPS 2-1-2 or Unified Stereo) that receives the first left channel SBR output 2216 and the left channel SBR output 2226, and provides based thereon. , the first content of MPS 2232, left channel MPEG Surround downmix signal 2234 and optionally left channel MPEG Surround downmix signal 2236. The Quad Channel encoder 2200 also includes a second MPEG-Surround (MPS 2-1-2 or Unified Stereo) multi-channel encoder 2240 which receives SBR output 2236 of the first right channel and SBR output 2228 of the second right channel, and which provides, based on of said first MPS content 2242, a right channel MPEG surround downmix signal 2244 and, optionally, a right channel MPEG surround downmix signal 2246.
[0181] The Quad Channel encoder 2200 includes a first composite prediction stereo coding 2250, which receives the left channel MPEG Surround downmix signal 2234 and the right channel MPEG Surround downmix signal 2244, and provides, based thereon, the composite prediction content 2252 and the co-coded representation. 2254 Left channel MPEG Surround downmix signal 2244 and Right channel MPEG Surround downmix signal 2244. The Quad Channel encoder 2200 includes a second predictive stereo composite encoder 2260, which receives left channel MPEG Surround residual 2236 and right channel MPEG Surround residual 2246, and provides the composite predictive content 2262 and a co-coded representation 2264 of the MPEG Surround downmix signal 2236. the left channel and the MPEG Surround downmix signal 2246 of the right channel.
[0182] The Quad Channel encoder also includes a first bit stream coding 2270 that receives the co-coded representation 2254, composite prediction content 2252m, MPS content 2232 and SBR content 2215 and provides based thereon a portion of a bit stream representing an element of the first channel pair. The Quad Channel encoder also includes second bit stream encoding 2280 that receives the co-coded representation 2264, composite prediction content 2262, MPS content 2242, and SBR content 2225, and provides, based thereon, a portion of the bitstream representing an element of the first channel pair.
14. Implementation alternatives
[0183] While some aspects have been described in the context of an apparatus, it is clear that these aspects also provide a description of a corresponding method wherein the block or device corresponds to a method step or a feature of a method step. Likewise, aspects described in the context of a method step also provide a description of the respective block or element or feature of the respective device. Some or all of the steps of the method may be performed with (or by using) a hardware device such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important steps of the method may be performed by such a device.
[0184] The encoded audio signal according to the invention may be stored on a digital storage medium or may be transmitted on a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.
[0185] Depending on certain implementation requirements, embodiments of the invention can be implemented in hardware or in software. The implementation may be via a digital storage medium, for example floppy disks, DVDs, Blu-Ray, CD, ROM, PROM, EPROM, EEPROM or FLASH, with electronically readable control signals on it that cooperate (or are capable of to cooperate) with a programmable computer system so that it is performed in an appropriate manner. Accordingly, the digital storage medium can be computer readable.
[0186] Some embodiments according to the invention include a data carrier having electronically readable control signals which are operable with a programmable computer system such that one of the methods described herein is performed.
[0187] In general, embodiments of the present invention can be implemented as a computer program product with a program code, the program code operating to perform one of the methods when the computer program product runs on a computer. For example, the program code may be stored on a machine readable medium.
[0188] Other embodiments include a computer program for performing one of the methods described herein, stored on a machine-readable medium.
[0189] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.
[0190] Another embodiment of the inventive methods is, therefore, a data medium (or a digital data medium or a computer readable medium) having a computer program stored thereon for performing one of the methods described herein. The data medium, digital data medium or recorded medium are usually tangible and / or non-transitory.
[0191] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals may e.g. be configured to be transmitted over a data connection, e.g. over the Internet.
[0192] Another embodiment includes processing means, for example, a computer or programmable logic device adapted or configured to perform one of the methods described herein.
[0193] Another embodiment includes a computer on which the computer program for performing one of the methods described herein is installed.
[0194] Another embodiment of the invention includes an apparatus or system adapted to transmit (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver can be, for example, a computer, mobile device, memory device or the like. The device or system may, for example, include a file server for transmitting a computer program to the receiver.
[0195] In some embodiments, a programmable logic device (e.g., a programmable gate array) may be used to perform some or all of the functionality of the methods described herein. In some embodiments, the programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware device.
[0196] The above described embodiments are merely illustrative of the principles of the present invention. It is understood that the modifications and variations of the systems and details described herein will be apparent to those skilled in the art. Accordingly, it is intended to be limited only by the scope of the following patent claims, and not by the specific details presented to describe and explain the present embodiments.
15. Conclusions
[0197] Some conclusions will be made below.
[0198] The embodiments of the invention are based on the assumption that, in order to account for the signal relationship between the vertically and horizontally distributed channels, the four channels can be coded together by hierarchically combining common stereo coding tools. For example, pairs of vertical channels are combined using MPS 2-1-2 and / or unified stereo with limited bandwidth or full bandwidth residual coding. In order to meet the perceptual requirements for binaural unmasking, the output downmixes are, for example, co-coded using complex predictions in the MDCT domain, which includes left-right and middle-side coding capabilities. If residual signals are present, they are combined horizontally using the same method.
[0199] Furthermore, it will be appreciated that the embodiments of the invention overcome some or all of the disadvantages of the prior art. Embodiments according to the invention are adapted to a 3D audio context, with the speaker channels being arranged in several height layers resulting in pairs of horizontal and vertical channels. It has been found that the common coding of only two channels as defined in USAC is not sufficient to consider spatial and perceptual relationships between the channels. However, this problem has been overcome by the embodiments according to the invention.
[0200] Moreover, conventional MPEG surround is used in an additional pre / post processing step such that the residual signals are transmitted individually without being able to be co-coded in stereo, e.g. to investigate the relationship between a left and a right root residual signal. In contrast, the embodiments of the invention allow for efficient encoding / decoding by using such dependencies.
[0201] To complete further, embodiments of the invention create an encoding and decoding apparatus, method, or computer program as described herein.
Reference: [0202]
1. [1] ISO / IEC 23003-3: 2012 - Information Technology - MPEG Audio Technologies, Part 3: Unified Speech and Audio Coding;
2. [2] ISO / IEC 23003-1: 2007 - Information Technology - MPEG Audio Technologies, Part 1: MPEG Surround
Fraunhofer-Gesellschaft zur Fórderung der angewandten Forschung eV, Germany; Proxy:
ΕΡ 3 022 734 BI
Ζ-16506/17
Contents6
48 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41 Sheet 42 Sheet 43 Sheet 44 Sheet 45 Sheet 46 Sheet 47 Sheet 48
80 members in 19 offices
Priority claims11
| Document | Office | Kind | Date |
|---|---|---|---|
| 13177376 | European Patent Office (EPO) | A | |
| 13177376 | European Patent Office (EPO) | A | |
| 13189306 | European Patent Office (EPO) | A | |
| 13189306 | European Patent Office (EPO) | A | |
| 14738535 | European Patent Office (EPO) | A | |
| 2014065021 | European Patent Office (EPO) | W | |
| 2014065021 | European Patent Office (EPO) | W | |
| EP20130177376 | – | – | – |
| EP20130189306 | – | – | – |
| EP20140738535 | – | – | – |
| WO2014EP65021 | – | – | – |
Members80
| Document | Office | Kind | |
|---|---|---|---|
| EP2830051A2 | European Patent Office (EPO) | A2 | |
| EP2830052A1 | European Patent Office (EPO) | A1 | |
| CA2917770A1 | Canada | A1 | |
| CA2918237A1 | Canada | A1 | |
| WO2015010926A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2015010934A1 | World Intellectual Property Organization (WIPO) | A1 | |
| EP2830051A3 | European Patent Office (EPO) | A3 | |
| TW201514972A | Taiwan Province of China | A | |
| TW201514973A | Taiwan Province of China | A | |
| AR097011A1 | Argentina | A1 | |
| AR097012A1 | Argentina | A1 | |
| SG11201600468SA | Singapore | A | |
| AU2014295282A1 | Australia | A1 | |
| AU2014295360A1 | Australia | A1 | |
| KR20160033777A | Republic of Korea | A | |
| KR20160033778A | Republic of Korea | A | |
| MX2016000939A | Mexico | A | |
| MX2016000858A | Mexico | A | |
| CN105580073A | China | A | |
| CN105593931A | China | A | |
| EP3022734A1 | European Patent Office (EPO) | A1 | |
| EP3022735A1 | European Patent Office (EPO) | A1 | |
| TWI544479B | Taiwan Province of China | B | |
| US2016247508A1 | United States of America | A1 | |
| US2016247509A1 | United States of America | A1 | |
| TWI550598B | Taiwan Province of China | B | |
| US2016275957A1 | United States of America | A1 | |
| JP2016529544A | Japan | A | |
| JP2016530788A | Japan | A | |
| JP6117997B2 | Japan | B2 | |
| ZA201601078B | South Africa | B | |
| BR112016001137A2 | Brazil | A2 | |
| BR112016001141A2 | Brazil | A2 | |
| AU2014295282B2 | Australia | B2 | |
| EP3022734B1 | European Patent Office (EPO) | B1 | |
| RU2016105702A | Russian Federation | A | |
| RU2016105703A | Russian Federation | A | |
| ZA201601080B | South Africa | B | |
| EP3022735B1 | European Patent Office (EPO) | B1 | |
| AU2014295360B2 | Australia | B2 | |
| PT3022734T | Portugal | T | |
| PT3022735T | Portugal | T | |
| ES2649194T3 | Spain | T3 | |
| ES2650544T3 | Spain | T3 | |
| KR101823278B1 | Republic of Korea | B1 | |
| PL3022734T3This record | Poland | T3 | |
| PL3022735T3 | Poland | T3 | |
| KR101823279B1 | Republic of Korea | B1 | |
| US9940938B2 | United States of America | B2 | |
| US9953656B2 | United States of America | B2 | |
| JP6346278B2 | Japan | B2 | |
| MX357667B | Mexico | B | |
| MX357826B | Mexico | B | |
| RU2666230C2 | Russian Federation | C2 | |
| US10147431B2 | United States of America | B2 | |
| RU2677580C2 | Russian Federation | C2 | |
| US2019108842A1 | United States of America | A1 | |
| US2019378522A1 | United States of America | A1 | |
| CN105580073B | China | B | |
| CN105593931B | China | B | |
| CN111105805A | China | A | |
| CN111128205A | China | A | |
| CN111128206A | China | A | |
| US10741188B2 | United States of America | B2 | |
| US10770080B2 | United States of America | B2 | |
| CA2917770C | Canada | C | |
| MY181944A | Malaysia | A | |
| US2021056979A1 | United States of America | A1 | |
| US2021233543A1 | United States of America | A1 | |
| CA2918237C | Canada | C | |
| BR112016001141B1 | Brazil | B1 | |
| US11488610B2 | United States of America | B2 | |
| BR112016001137B1 | Brazil | B1 | |
| US11657826B2 | United States of America | B2 | |
| US2024029744A1 | United States of America | A1 | |
| CN111128206B | China | B | |
| CN111105805B | China | B | |
| US12380899B2 | United States of America | B2 | |
| US20260024535A1 | United States of America | A1 | |
| CN111128205B | China | B |
Numbers
- Publication, DOCDB
- 3022734
- Publication, EPODOC
- PL3022734T
- Application
- 738535
- Application, DOCDB
- 14738535
- Application, EPODOC
- PL20140738535T
Titles2
- English
- AUDIO DECODER, AUDIO ENCODER, METHOD FOR PROVIDING AT LEAST FOUR AUDIO CHANNEL SIGNALS ON THE BASIS OF AN ENCODED REPRESENTATION, METHOD FOR PROVIDING AN ENCODED REPRESENTATION ON THE BASIS OF AT LEAST FOUR AUDIO CHANNEL SIGNALS AND COMPUTER PROGRAM USING A BANDWIDTH EXTENSION
- Polish
- DEKODER AUDIO, AUDIO KODER, SPOSÓB DOSTARCZANIA NAJMNIEJ CZTERECH SYGNAŁÓW KANAŁÓW AUDIO NA PODSTAWIE KODEKSOWANYCH REPREZENTACJI, SPOSÓB DOSTARCZANIA ZAPISANYCH REPREZENTACJI NA PODSTAWIE CO NAJMNIEJ CZTERECH SYGNAŁÓW KANAŁÓW AUDIO I PROGRAMU KOMPUTEROWEGO ZA POMOCĄ ROZSZERZENIA PASMOWEGO
Classification
- CPC, 8
- G10L19/0017
- G10L19/008
- G10L21/038
- H04S3/008
- H04S7/30
- H04S2400/01
- H04S2400/03
- H04S2420/03
- IPC, 3
- G10L19 008
- G10L19 00
- G10L21 038
