Audio signal decoder, method for decoding an audio signal and computer program using cascaded audio object processing stages
Abstract
This record has no abstract on file.
Term
3.7 yearsto projected expiry
Projected expiry 23 June 2030, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
1 claim: 1 independent, 0 dependent
- 1Patent claims Zastrzeżenia patentowe 1. Audio signal decoder (100; 200; 500; 590) for providing an upmix signal representation based on a downmix signal representation (112; 210; 510; 510) and parametric object information (110; 212; 512; 512a), the signal decoder audio includes:1. Dekoder sygnałów audio (100;200;500;590) do dostarczania reprezentacji sygnału upmixu w oparciu o reprezentację sygnału downmixu (112;210;510;510) i obiektową informację parametryczną (110;212;512;512a), przy czym dekoder sygnału audio zawiera: an object separator (130;260;520;520a) configured to break down the downmix signal representation for providing the first audio information (132;262;562;562a) describing the first set of one or more audio objects from the first type of audio objects and the second audio information ( 134;264;564;564a) describing a second set of one or more audio objects from the second type of audio objects based on the downmix signal representation and using at least a portion of the parametric object information, wherein the second audio information is audio information describing the audio objects of the second type of audio objects in a combined manner;separator obiektów (130;260;520;520a) skonfigurowany do rozkładania reprezentacji sygnału downmixu dla dostarczania pierwszej informacji audio (132;262;562;562a) opisującej pierwszy zestaw jednego lub większej liczby obiektów audio z pierwszego typu obiektów audio i drugiej informacji audio (134;264;564;564a) opisującej drugi zestaw jednego lub większej liczby obiektów audio z drugiego typu obiektów audio w oparciu o reprezentację sygnału downmixu i z użyciem przynajmniej części obiektowej informacji parametrycznej, przy czym druga informacja audio jest informacją audio opisującą obiekty audio drugiego typu obiektów audio w połączony sposób;an audio signal processor configured to receive the second audio information (134;264;564;564a) and to process the second audio information depending on the parametric object information to obtain a processed version (142;272;572;572a) of the second audio information;and a combining module (150;280;580;580a) of the audio signal configured to combine the first audio information with the processed version of the second audio information to obtain an upmix signal representation;procesor sygnału audio skonfigurowany do odbioru drugiej informacji audio (134;264;564;564a) i do przetwarzania drugiej informacji audio w zależności od obiektowej informacji parametrycznej, dla uzyskania przetworzonej wersji (142;272;572;572a) drugiej informacji audio;oraz moduł łączenia (150;280;580;580a) sygnału audio skonfigurowany do łączenia pierwszej informacji audio z przetworzoną wersją drugiej informacji audio dla uzyskania reprezentacji sygnału upmixu;przy czym dekoder sygnału audio jest skonfigurowany do dostarczania reprezentacji sygnału upmixu w oparciu o informację resztkową powiązaną z podzestawem obiektów audio reprezentowanym przez reprezentację sygnału downmixu, przy czym separator obiektów jest skonfigurowany do rozkładania reprezentacji sygnału downmixu dla dostarczania pierwszej informacji audio opisującej pierwszy zestaw jednego lub większej liczby obiektów audio z pierwszego typu obiektów audio, z którymi powiązana jest informacja resztkowa, i drugiej informacji audio opisującej drugi zestaw jednego lub większej liczby obiektów audio z drugiego typu obiektów audio, z którymi nie jest powiązana żadna informacja resztkowa, w oparciu o reprezentację sygnału downmixu i z użyciem informacji resztkowej;oraz przy czym procesor sygnału audio jest skonfigurowany do przetwarzania drugiej informacji audio dla realizacji indywidualnego obiektowego przetwarzania obiektów audio drugiego typu obiektów audio z użyciem obiektowej informacji parametrycznej powiązanej z więcej niż dwoma obiektami audio z drugiego typu obiektów audio;oraz przy czym informacja resztkowa opisuje zniekształcenie resztkowe, które zgodnie z oczekiwaniem pozostanie, jeśli obiekt audio z pierwszego typu obiektów audio jest wyizolowany z użyciem tylko obiektowej informacji parametrycznej. wherein the audio decoder is configured to provide an upmix signal representation based on the residual information associated with the subset of audio objects represented by the downmix signal representation, wherein the object separator is configured to break down the downmix signal representation to provide the first audio information describing the first set of one or more number of audio objects from the first type of audio objects, with which residual information is associated, and second audio information describing a second set of one or more audio objects of the second type of audio objects with which no residual information is associated, based on downmix signal representation and using residual information;and wherein the audio signal processor is configured to process the second audio information to perform individual object-oriented processing of audio objects of the second type of audio objects using object-oriented parametric information associated with more than two audio objects from the second type of audio objects;and wherein the residual information describes a residual distortion that is expected to remain if the audio object of the first type of audio object is isolated using only object-oriented parametric information. 2. Audio signal decoder (100;200;500;590) according to claim 1, wherein the object separator is configured to provide the first audio information using residual information such that one or more audio objects from the first type of audio objects are highlighted relative to the audio objects of the second type of audio objects in the first information audio, and wherein the object separator is configured to provide second audio information by using residual information in such a way that the audio objects of the second type of audio objects are highlighted relative to the audio objects of the first type of audio objects in the second audio information. 2. Dekoder sygnału audio (100;200;500;590) według zastrzeżenia 1, w którym separator obiektów jest skonfigurowany do dostarczania pierwszej informacji audio z użyciem informacji resztkowej w taki sposób, że jeden lub większa liczba obiektów audio z pierwszego typu obiektów audio jest uwypuklona względem obiektów audio z drugiego typu obiektów audio w pierwszej informacji audio, i w którym separator obiektów jest skonfigurowany do dostarczania drugiej informacji audio za pomocą użycia informacji resztkowej w taki sposób, że obiekty audio z drugiego typu obiektów audio są uwypuklone względem obiektów audio z pierwszego typu obiektów audio w drugiej informacji audio. 3. Audio signal decoder (100;200;500;570) according to one of the claims 1 to 3. Dekoder sygnału audio (100;200;500;570) według jednego z zastrzeżeń od 1 do 2, w którym procesor sygnału audio jest skonfigurowany do przetwarzania drugiej informacji audio (134;264;564;564a) w oparciu o obiektową informację parametryczną (110;212;512;512a) powiązaną z obiektami audio z drugiego typu obiektów audio w oparciu o obiektową informację parametryczną (110;212;512;512a) powiązaną z obiektami audio z pierwszego typu obiektów audio. The process of claim 2, wherein the audio signal processor is configured to process the second audio information (134;264;564;564a) based on the object parametric information (110;212;512;512a) associated with the audio objects of the second type of audio objects based on object parametric information (110;212;512;512a) associated with audio objects of the first type of audio objects. 4. Audio signal decoder (100;200;500;590) according to one of the claims 1 to 4. Dekoder sygnału audio (100;200;500;590) według jednego z zastrzeżeń od 1 do 3, w którym separator obiektów jest skonfigurowany do uzyskania pierwszej informacji audio (132;262;562;562a, XEAo) i drugiej informacji audio (134;264;564;564a, XoBJ) z użyciem łączenia liniowego jednego lub większej liczby kanałów sygnału downmixu reprezentacji sygnału downmixu i jednego lub większej liczby kanałów resztkowych, przy czym separator obiektów jest skonfigurowany do uzyskania parametrów łączenia do realizacji łączenia liniowego w oparciu o parametry downmixu powiązane z obiektami audio z pierwszego typu obiektów audio (m0... mNEAo-1;no... nNEAo-1) i w oparciu o współczynniki predykcji kanału (cj,0, cj,1) obiektów audio z pierwszego typu obiektów audio. The object of claim 3, wherein the object separator is configured to obtain first audio information (132;262;562;562a, XEAo) and second audio information (134;264;564;564a, XoBJ) using a linear combination of one or more downmix signal channels of the downmix signal representation and one or more residual channels, where the object separator is configured to obtain the connection parameters for performing the line connection based on the downmix parameters associated with the audio objects associated with the first type of audio objects (m0 ... mNEAo-1;no ... nNEAo-1) and based on channel prediction coefficients (cj, 0, cj, 1) of audio objects from the first type of audio objects. 5. Audio signal decoder (100;200;500;590) according to one of the claims 1 to 5. Dekoder sygnału audio (100;200;500;590) według jednego z zastrzeżeń od 1 do 4, w którym separator obiektów jest skonfigurowany do uzyskania pierwszej informacji audio i drugiej informacji audio zgodnie z gdzie gdzie gdzie XOBJ reprezentuje kanały drugiej informacji audio;4. The object of claim 4, wherein the object separator is configured to obtain the first audio information and the second audio information according to where where XOBJ represents channels of the second audio information;gdzie XEAO reprezentuje sygnały obiektów pierwszej informacji audio;where XEAO represents object signals of the first audio information;gdzie D'1 reprezentuje macierz, która jest odwróceniem rozszerzonej macierzy downmixu;where D '1 represents the matrix, which is the inversion of the extended downmix matrix;gdzie C opisuje macierz reprezentującą wiele współczynników X / predykcji kanału;gdzie l0 i r0 reprezentują kanały reprezentacji sygnału downmixu;gdzie res0 do resNEAO-1 reprezentują kanały resztkowe;i gdzie AEAO jest macierzą EAO wstępnego renderowania, której wyrazy opisują mapowanie wzbogaconych obiektów audio na kanały sygnału XEAO wzbogaconych obiektów audio;where C describes a matrix representing multiple X coefficients / channel prediction;where l0 and r0 represent downmix signal representation channels;where res0 to resNEAO-1 represent residual channels;and where A.EDA is an EAO pre-rendering matrix whose words describe the mapping of enriched audio objects to XEAO enriched audio object channels;przy czym separator obiektów jest skonfigurowany do uzyskania odwróconej macierzy wherein the object separator is configured to obtain an inverted matrix D '1downmix, as the inversion of the extended Ddownmix matrix, which is defined as D'1downmixu, jako odwrócenia rozszerzonej macierzy Ddownmixu, która jest zdefiniowana jako przy czym separator obiektów jest skonfigurowany do uzyskania macierzy C jako gdzie m0 do mNEAO-1 są wartościami downmixu powiązanymi z obiektami audio pierwszego typu obiektów audio;wherein the object separator is configured to obtain a C matrix as where m0 to mNEAO-1 are downmix values associated with the audio objects of the first type of audio objects;gdzie no to nNEAO-1 są wartościami downmixu powiązanymi z obiektami audio pierwszego typu obiektów audio;where nNEAO-1 are downmix values associated with the audio objects of the first type of audio objects;przy czym separator obiektów jest skonfigurowany do obliczania współczynników predykcji Cci c jako i i wherein the object separator is configured to calculate prediction coefficients Cthese c as and and przy czym separator obiektów jest skonfigurowany do uzyskania ograniczonych współczynników predykcji c7;0 i j ze współczynników predykcji ¥i c z użyciem algorytmu ograniczenia, lub do użycia współczynników predykcji c - i G jako współczynników predykcji c.--!i X;where the object separator is configured to obtain limited prediction coefficients c7;0 ij from the prediction coefficients ¥ i c using a constraint algorithm, or to use prediction coefficients c - and G as prediction coefficients c.--!and X;przy czym wielkości energii PLo, PRo, PLoRo, PLoCoj i PRoCoj są zdefiniowane jako i=0 P/toCj ~ nfiLDg + - nfiLDj >=o i·) gdzie parametry OLDL, OLDR i IOCL,R odpowiadają obiektom audio drugiego typu obiektów audio i są zdefiniowane zgodnie z gdzie d0,i i d1,i są wartościami downmixu powiązanymi z obiektami audio z drugiego typu obiektów audio;where the energy quantities PLo, PRo, PLoRo, PLoCoj and PRoCoj are defined as i = 0 P/ toCj ~ nfiLDg + - nfiLDj> = oi ·) where the parameters OLDL, OLDR and IOCL, R correspond to the audio objects of the second type of audio objects and are defined according to where d0, ii d1, and are the downmix values associated with the audio objects of the second type audio objects;gdzie OLDi są wartościami różnicy poziomów obiektów powiązanymi z obiektami audio z drugiego typu obiektów audio;where OLDi are the object level difference values associated with the audio objects of the second type of audio objects;gdzie N jest całkowitą liczbą obiektów audio;where N is the total number of audio objects;gdzie NEAO jest liczbą obiektów audio pierwszego typu obiektów audio;where NEAO is the number of audio objects of the first type of audio objects;gdzie IOC0,1 jest wartością korelacji międzyobiektowej powiązaną z parą obiektów audio drugiego typu obiektów audio;where IOC0,1 is the cross-object correlation value associated with the pair of audio objects of the second type of audio objects;gdzie eij i eL,R są wartościami kowariancji uzyskanymi z parametrów różnicy poziomów obiektów i parametrów korelacji międzyobiektowej;i gdzie eij są powiązane z parą obiektów audio z pierwszego typu obiektów audio, a eL,R jest powiązany z parą obiektów audio z drugiego typu obiektów audio. where eij and eL, R are covariance values obtained from object level difference parameters and inter-object correlation parameters;and where e and j are associated with a pair of audio objects from the first type of audio objects and eL, R is associated with a pair of audio objects from the second type of audio objects. 6. An audio signal decoder (100;200;500;590) according to one of claims 1 to 4, wherein the object separator is configured to obtain the first audio information and the second audio information according to wherein wherein XOBJ represents the channel of the second audio information;6. Dekoder sygnału audio (100;200;500;590) według jednego z zastrzeżeń od 1 do 4, w którym separator obiektów jest skonfigurowany do uzyskania pierwszej informacji audio i drugiej informacji audio zgodnie z gdzie gdzie XOBJ reprezentuje kanał drugiej informacji audio;gdzie XEAO reprezentuje sygnały obiektów pierwszej informacji audio;gdzie D'1 reprezentuje macierz, która jest odwróceniem rozszerzonej macierzy downmixu;where XEAO represents object signals of the first audio information;where D '1 represents the matrix, which is the inversion of the extended downmix matrix;gdzie C opisuje macierz reprezentującą wiele współczynników c. - c predykcji kanałów;gdzie do reprezentuje kanał reprezentacji sygnału downmixu;i gdzie reso do resNEAo-1 reprezentuje kanały resztkowe;i gdzie AEAo jest macierzą EAo wstępnego renderowania. where C describes a matrix representing many coefficients c. - c channel prediction;where do represents the downmix signal representation channel;and where reso to resNEAo-1 represents residual channels;and where A.EAO is the pre-rendering EAo matrix. 7. An audio decoder according to claim 6, wherein the object separator is configured to obtain an inverted D 'matrix1downmix as the inversion of the extended downmix matrix D, which is defined as 7. Dekoder sygnału audio według zastrzeżenia 6, w którym separator obiektów jest skonfigurowany do uzyskania odwróconej macierzy D'1downmixu jako odwrócenia rozszerzonej macierzy D downmixu, która jest zdefiniowana jako przy czym separator obiektów jest skonfigurowany do uzyskania macierzy C jako wherein the object separator is configured to obtain the matrix C as gdzie m0 to mNEAo-1 są wartościami downmixu powiązanymi z obiektami audio z pierwszego typu obiektów audio. where m0 is mNEAo-1 are the downmix values associated with the audio objects of the first type of audio objects. 8. An audio signal decoder (100;200;500;590) according to one of claims 1 to 4, wherein the object separator is configured to obtain the first audio information and the second audio information according to where XOBJ represents the channel of the second audio information;8. Dekoder sygnału audio (100;200;500;590) według jednego z zastrzeżeń od 1 do 4, w którym separator obiektów jest skonfigurowany do uzyskania pierwszej informacji audio i drugiej informacji audio zgodnie z gdzie XOBJ reprezentuje kanał drugiej informacji audio;gdzie XEAO reprezentuje sygnały obiektów pierwszej informacji audio;where XEAO represents object signals of the first audio information;gdzie gdzie m0 to mNEAO-1 są wartościami downmixu powiązanymi z obiektami audio z pierwszego typu obiektów audio;where m0 is mNEAO-1 are downmix values associated with the audio objects of the first type of audio objects;gdzie n0 do nNEAO-1 są wartościami downmixu powiązanymi z obiektami audio z pierwszego typu obiektów audio;where n0 to nNEAO-1 are downmix values associated with the audio objects of the first type of audio objects;gdzie OLDi są wartościami różnicy poziomów obiektów powiązanymi z obiektami audio z pierwszego typu obiektów audio;where OLDi are the object level difference values associated with the audio objects of the first type of audio objects;gdzie OLDL i OLDR są wspólnymi wartościami różnicy poziomów obiektów powiązanymi z obiektami audio drugiego typu obiektów audio;i gdzie AEAO jest macierzą EAO wstępnego renderowania. where OLDL and OLDR are common object level difference values associated with audio objects of the second type of audio objects;and where A.EDA is the pre-rendering EAO matrix. 9. An audio signal decoder according to one of claims 1 to 3, wherein the object separator is configured to obtain the first audio information and the second audio information according to 9. Dekoder sygnału audio według jednego z zastrzeżeń od 1 do 3, w którym separator obiektów jest skonfigurowany do uzyskania pierwszej informacji audio i drugiej informacji audio zgodnie z Χ "," =Μ¥“< Χ„„,=Μ¥“< ^ Ł10 where XOBJ represents the channel of the second audio information;^Ł10 gdzie XOBJ reprezentuje kanał drugiej informacji audio;gdzie XEAO reprezentuje sygnały obiektów pierwszej informacji audio;where XEAO represents object signals of the first audio information;gdzie gdzie m0 do mNEAO-1 są wartościami downmixu powiązanymi z obiektami audio z pierwszego typu obiektów audio;where m0 to mNEAO-1 are downmix values associated with the audio objects of the first type of audio objects;gdzie OLDi są wartościami różnicy poziomów obiektów powiązanymi z obiektami audio z pierwszego typu obiektów audio;where OLDi are the object level difference values associated with the audio objects of the first type of audio objects;gdzie OLDL jest wspólną wartością różnicy poziomów obiektów powiązanych z obiektami audio drugiego typu obiektów audio;i gdzie AEAO jest macierzą EAO wstępnego renderowania;where OLDL is the common value of the difference in the level of objects associated with the audio objects of the second type of audio objects;and where A.EDA is an EAO pre-rendering matrix;Energy Energy T ^ t Energy are used for the representation d0 where the single signal downmix SAOC matrix. T^t Energy są zastosowane dla reprezentacji d0 przy czym macierze pojedynczego sygnału downmixu SAOC. 10. Audio signal decoder (100;200;500;590) according to one of the claims 1 to 10. Dekoder sygnału audio (100;200;500;590) według jednego z zastrzeżeń od 1 do 9, w którym separator obiektów jest skonfigurowany do zastosowania macierzy renderowania dla pierwszej informacji audio (132;262;562;562a) do mapowania sygnałów obiektów pierwszej informacji audio na kanały audio reprezentacji (120;220, 222;562;562a) sygnału audio upmixu. 9. The object of claim 9, wherein the object separator is configured to use a rendering matrix for the first audio information (132;262;562;562a) to map the signals of the first audio information objects to the audio channels of the representation (120;220, 222;562;562a) of the upmix audio signal . 11. Audio signal decoder (100;200;500;590) according to one of the claims 1 to 11. Dekoder sygnału audio (100;200;500;590) według jednego z zastrzeżeń od 1 do 10, w którym procesor (140;270;570;570a) sygnału audio jest skonfigurowany do realizacji wstępnego przetwarzania stereo drugiej informacji audio (134;264;564;564a) w oparciu o informację renderowania (Mren), i obiektową informację kowariancji (E), informację o downmixie (D) dla uzyskania kanałów audio przetworzonej wersji drugiej informacji audio. The process of claim 10, wherein the processor (140;270;570;570a) of the audio signal is configured to perform pre-stereo processing of the second audio information (134;264;564;564a) based on rendering information (Mren), and object covariance information (E ), downmix information (D) to obtain audio channels of the processed version of the second audio information. 12. An audio signal decoder (100;200;500;590) according to claim 11, wherein the processor (140;270;570;570a) of the audio signal is configured to perform stereo processing for mapping the estimated input (ED * JX) of the audio object of the second audio information (134;264;564;564a) for multiple channels of upmix audio signal representation based on rendering information and covariance information. 12. Dekoder sygnału audio (100;200;500;590) według zastrzeżenia 11, w którym procesor (140;270;570;570a) sygnału audio jest skonfigurowany do realizacji przetwarzania stereo dla mapowania szacowanego wkładu (ED*JX) obiektu audio drugiej informacji audio (134;264;564;564a) na wiele kanałów reprezentacji sygnału audio upmixu w oparciu o informację renderowania i informację kowariancji. 13. An audio signal decoder according to claim 11 or claim 12, wherein the audio signal processor is configured to add a de-correlated audio signal input (P2Xd) derived from one or more audio channels of the second audio information to the second audio information, or information obtained from the second audio information based on the (R) upmix error information of the rendering and one or more values (wd1, wd2) for scaling the decrelated signal intensities. 13. Dekoder sygnału audio według zastrzeżenia 11 albo zastrzeżenia 12, w którym procesor sygnału audio jest skonfigurowany do dodania zdekorelowanego wkładu (P2Xd) sygnału audio, uzyskanego na bazie jednego lub większej liczby kanałów audio drugiej informacji audio, do drugiej informacji audio, lub informacji uzyskanej z drugiej informacji audio, w oparciu o informację (R) błędu upmixu renderowania i jednej lub większej liczby wartości (wd1, wd2) skalowania zdekorelowanych intensywności sygnału. 14. An audio signal decoder according to one of claims 1 to 10, wherein the processor (140;270;570;570a) of the audio signal is configured to perform post-processing of the second audio information (134;264;564;564a) based on the information (A ) about rendering, (E) object-oriented covariance, and (D) downmix information. 14. Dekoder sygnału audio według jednego z zastrzeżeń od 1 do 10, w którym procesor (140;270;570;570a) sygnału audio jest skonfigurowany do realizacji przetwarzania końcowego drugiej informacji audio (134;264;564;564a) w oparciu o informację (A) o renderowaniu, informację (E) o obiektowej kowariancji i informację o (D) downmixie. 15. The audio signal decoder according to claim 14, wherein the audio signal processor is configured to perform mono-to-binaural processing of the second audio information to map a single channel of the second audio information to two channels of the upmix signal representation, including the head transfer function. 15. Dekoder sygnału audio według zastrzeżenia 14, w którym procesor sygnału audio jest skonfigurowany do realizacji przetwarzania mono-do-binauralnego drugiej informacji audio, do mapowania pojedynczego kanału drugiej informacji audio na dwa kanały reprezentacji sygnału upmixu, uwzględniając funkcję transmitancji głowy. 16. The audio signal decoder according to claim 14, wherein the audio signal processor is configured to perform mono-to-stereo processing of the second audio information to map a single channel of the second audio information to two channels of the upmix signal representation. 16. Dekoder sygnału audio według zastrzeżenia 14, w którym procesor sygnału audio jest skonfigurowany do realizacji przetwarzania mono-do-stereo drugiej informacji audio, do mapowania pojedynczego kanału drugiej informacji audio na dwa kanały reprezentacji sygnału upmixu. 17. The audio signal decoder according to claim 14, wherein the audio signal processor is configured to perform stereo-to-binaural processing of the second audio information to map two channels of the second audio information to two channels of the upmix signal representation, including the head transfer function. 17. Dekoder sygnału audio według zastrzeżenia 14, w którym procesor sygnału audio jest skonfigurowany do realizacji przetwarzania stereo-do-binauralnego drugiej informacji audio, do mapowania dwóch kanałów drugiej informacji audio na dwa kanały reprezentacji sygnału upmixu, uwzględniając funkcję transmitancji głowy. 18. An audio signal decoder according to claim 14, wherein the audio signal processor is configured to perform stereo-to-stereo processing of the second audio information to map two channels of the second audio information to two channels of the upmix signal representation. 18. Dekoder sygnału audio według zastrzeżenia 14, w którym procesor sygnału audio jest skonfigurowany do realizacji przetwarzania stereo-do-stereo drugiej informacji audio, do mapowania dwóch kanałów drugiej informacji audio na dwa kanały reprezentacji sygnału upmixu. 19. An audio signal decoder according to one of claims 1 to 18, wherein the audio object separator is configured to treat audio objects of the second type of audio objects with which no residual information is associated as individual audio objects and in which the processor (140;270;570;570a) of the audio signal is configured to include rendering object parameters associated with the audio objects of the second type of audio objects to match the shares of the audio objects of the second type of audio objects in the upmix signal representation. 19. Dekoder sygnału audio według jednego z zastrzeżeń od 1 do 18, w którym separator obiektów audio jest skonfigurowany do traktowania obiektów audio drugiego typu obiektów audio, z którymi nie jest powiązana żadna informacja resztkowa, jako pojedynczych obiektów audio, i w którym procesor (140;270;570;570a) sygnału audio jest skonfigurowany do uwzględnienia obiektowych parametrów renderowania powiązanych z obiektami audio z drugiego typu obiektów audio dla dopasowania udziałów obiektów audio z drugiego typu obiektów audio w reprezentacji sygnału upmixu. 20. Dekoder sygnału audio według jednego z zastrzeżeń od 1 do 19, w którym separator obiektów audio jest skonfigurowany do uzyskania jednej lub dwóch wspólnych wartości (OLDL, OLDR) różnicy poziomów obiektów dla wielu obiektów audio z drugiego typu obiektów audio;i w którym separator obiektów audio jest skonfigurowany do użycia wspólnej wartości różnicy poziomów obiektów do obliczania współczynników (CPC) predykcji kanału;i w którym separator obiektów jest skonfigurowany do użycia współczynników predykcji kanału dla uzyskania jednego lub dwóch kanałów audio reprezentujących drugą informacje audio. twenty. An audio signal decoder according to one of claims 1 to 19, wherein the audio object separator is configured to obtain one or two common values (OLDL, OLDR) of the object level difference for multiple audio objects from the second type of audio objects;and wherein the audio object separator is configured to use a common object level difference value to calculate channel prediction coefficients (CPC);and wherein the object separator is configured to use channel prediction coefficients to obtain one or two audio channels representing the second audio information. 21. An audio signal decoder according to one of claims 1 to 20, wherein the audio object separator is configured to obtain one or two common values (OLDL, OLDR) of the object level difference for multiple audio objects from the second type of audio objects;and wherein the audio object separator is configured to use a common object level difference value to calculate matrix words (M);and wherein the audio object separator is configured to use a matrix (M) to obtain one or more audio channels representing the second audio information. 21. Dekoder sygnału audio według jednego z zastrzeżeń od 1 do 20, w którym separator obiektów audio jest skonfigurowany do uzyskania jednej lub dwóch wspólnych wartości (OLDL, OLDR) różnicy poziomów obiektów dla wielu obiektów audio z drugiego typu obiektów audio;i w którym separator obiektów audio jest skonfigurowany do użycia wspólnej wartości różnicy poziomów obiektów do obliczania wyrazów macierzy (M);i w którym separator obiektów audio jest skonfigurowany do użycia macierzy (M) dla uzyskania jednego lub większej liczby kanałów audio reprezentujących drugą informację audio. 22. The audio signal decoder according to one of claims 1 to 21, wherein the audio object separator is configured to selectively obtain a common value (IOCL, R) of the inter-object correlation associated with the audio object from the second type of audio objects based on the object parametric information if it is found, that there are two audio objects from the second type of audio objects, and to set the cross-object correlation value associated with the audio objects of the second type of audio objects to zero if it is determined that there are more or less than two audio objects from the second type of audio objects;and wherein the object separator is configured to use a common cross-object correlation value to calculate matrix words (M);and wherein the object separator is configured to use a common cross-object correlation value associated with the audio objects of the second type of audio objects to obtain one or more audio channels representing the second audio information. 22. Dekoder sygnału audio według jednego z zastrzeżeń od 1 do 21, w którym separator obiektów audio jest skonfigurowany do selektywnego uzyskiwania wspólnej wartości (IOCL,R) korelacji międzyobiektowej powiązanej z obiektem audio z drugiego typu obiektów audio w oparciu o obiektową informację parametryczną jeśli zostanie stwierdzone, że istnieją dwa obiekty audio z drugiego typu obiektów audio, i do ustawienia wartości korelacji międzyobiektowej powiązanej z obiektami audio z drugiego typu obiektów audio na zero, jeśli zostanie stwierdzone, że jest więcej lub mniej niż dwa obiekty audio z drugiego typu obiektów audio;i w którym separator obiektów jest skonfigurowany do użycia wspólnej wartości korelacji międzyobiektowej do obliczania wyrazów macierzy (M);i w którym separator obiektów jest skonfigurowany do użycia wspólnej wartości korelacji międzyobiektowej powiązanej z obiektami audio z drugiego typu obiektów audio dla uzyskania jednego lub większej liczby kanałów audio reprezentujących drugą informację audio. 23. The audio signal decoder according to one of claims 1 to 22, wherein the audio signal processor is configured to render the second audio information based on the object parametric information to obtain a rendered representation of the audio objects from the second type of audio objects as a processed version of the second audio information. 23. Dekoder sygnału audio według jednego z zastrzeżeń od 1 do 22, w którym procesor sygnału audio jest skonfigurowany do renderowania drugiej informacji audio w oparciu o obiektową informację parametryczną dla uzyskania renderowanej reprezentacji obiektów audio z drugiego typu obiektów audio, jako przetworzonej wersji drugiej informacji audio. 24. An audio signal decoder according to one of claims 1 to 23, wherein the audio object separator is configured to provide the second audio information in such a way that the second audio information describes more than two audio objects from the second type of audio objects. 24. Dekoder sygnału audio według jednego z zastrzeżeń od 1 do 23, w którym separator obiektów audio jest skonfigurowany do dostarczania drugiej informacji audio w taki sposób, że druga informacja audio opisuje więcej niż dwa obiekty audio z drugiego typu obiektów audio. 25. The audio signal decoder according to claim 24, wherein the audio object separator is configured to obtain, as the second audio information, a single channel representation of the audio signal or a two channel representation of the audio signal representing more than two audio objects from the second type of audio objects. 25. Dekoder sygnału audio według zastrzeżenia 24, w którym separator obiektów audio jest skonfigurowany do uzyskania, jako drugiej informacji audio, jednokanałowej reprezentacji sygnału audio lub dwukanałowej reprezentacji sygnału audio reprezentującej więcej niż dwa obiekty audio z drugiego typu obiektów audio. 26. An audio signal decoder according to one of claims 1 to 25, wherein the audio signal processor is configured to receive second audio information and to process the second audio information based on the object parametric information, taking into account the object parametric information associated with more than two audio objects from the second type of audio objects. 26. Dekoder sygnału audio według jednego z zastrzeżeń od 1 do 25, w którym procesor sygnału audio jest skonfigurowany do odbioru drugiej informacji audio i do przetwarzania drugiej informacji audio w oparciu o obiektową informację parametryczną, uwzględniając obiektową informację parametryczną powiązaną z więcej niż dwoma obiektami audio z drugiego typu obiektów audio. 27. Audio signal decoder according to one of claims 1 to 26, wherein the audio signal decoder is configured to obtain information (bsNumObjects) of the total number of objects and the number (bsNum-GroupsFGO) of the foreground objects from the configuration information (SAOCSpecificConfig) of the parametric information information and to be determined the number of audio objects from the second type of audio objects by creating a difference in information about the total number of objects and the number of foreground objects information. 27. Dekoder sygnału audio według jednego z zastrzeżeń od 1 do 26, w którym dekoder sygnału audio jest skonfigurowany do pozyskania informacji (bsNumObjects) całkowitej liczby obiektów i liczby (bsNum- GroupsFGO) obiektów przedniego planu z informacji konfiguracji (SAOCSpecificConfig) obiektowej informacji parametrycznej i do ustalenia liczby obiektów audio z drugiego typu obiektów audio, poprzez utworzenie różnicy informacji całkowitej liczby obiektów i informacji liczby obiektów przedniego planu. 28. Audio signal decoder according to one of claims 1 to 27, wherein the audio object separator is configured to use NEAO-related parametric information audio objects of the first type of audio objects to obtain, as the first audio information, NEAO (XEAO) audio signals representing NEAO audio objects from the first type of audio objects and to be obtained as second audio information, one or two audio signals (XOBJ) representing N-NEAO audio objects from the second type of audio objects, treating N-NEAO audio objects from the second type of audio objects as a single single-channel or two-channel audio object;and wherein the audio signal processor is configured to individually render N-NEAO audio objects represented by one or two audio signals from the second audio information using the object-oriented parametric information associated with N-NEAO audio objects from the second type of audio objects. 28. Dekoder sygnału audio według jednego z zastrzeżeń od 1 do 27, w którym separator obiektów audio jest skonfigurowany do użycia obiektowej informacji parametrycznej powiązanej z NEAO obiektami audio z pierwszego typu obiektów audio dla uzyskania, jako pierwszej informacji audio, NEAO (XEAO) sygnałów audio reprezentujących NEAO obiektów audio z pierwszego typu obiektów audio i do uzyskania, jako drugiej informacji audio, jednego lub dwóch sygnałów audio (XOBJ) reprezentujących N-NEAO obiektów audio z drugiego typu obiektów audio, traktując N-NEAO obiektów audio z drugiego typu obiektów audio, jako pojedynczy jednokanałowy lub dwukanałowy obiekt audio;i w którym procesor sygnału audio jest skonfigurowany do indywidualnego renderowania N-NEAO obiektów audio reprezentowanych przez jeden lub dwa sygnały audio z drugiej informacji audio z użyciem obiektowej informacji parametrycznej powiązanej z N-NEAO obiektami audio z drugiego typu obiektów audio. 29. A method of providing an upmix signal representation based on a downmix signal representation and object-oriented parametric information, the method comprising: 29. Sposób dostarczania reprezentacji sygnału upmixu w oparciu o reprezentację sygnału downmixu i obiektową informację parametryczną, przy czym sposób obejmuje: a distribution of the downmix signal representation for providing the first audio information describing the first set of one or more audio objects from the first type of audio objects and the second audio information describing the second set of one or more audio objects from the second type of audio objects based on the downmix signal representation and using at least part of the object parametric information, wherein the second audio information is audio information describing audio objects of the second type of audio objects in a combined manner;and processing the second audio information depending on the parametric object information to obtain a processed version of the second audio information;and combining the first audio information with the processed version of the second audio information to obtain an upmix signal representation;rozkład reprezentacji sygnału downmixu dla dostarczania pierwszej informacji audio opisującej pierwszy zestaw jednego lub większej liczby obiektów audio z pierwszego typu obiektów audio i drugiej informacji audio opisującej drugi zestaw jednego lub większej liczby obiektów audio z drugiego typu obiektów audio w oparciu o reprezentację sygnału downmixu i z użyciem przynajmniej części obiektowej informacji parametrycznej, przy czym druga informacja audio jest informacją audio opisującą obiekty audio drugiego typu obiektów audio w połączony sposób;i przetwarzanie drugiej informacji audio w zależności od obiektowej informacji parametrycznej, dla uzyskania przetworzonej wersji drugiej informacji audio;i łączenie pierwszej informacji audio z przetworzoną wersją drugiej informacji audio w celu uzyskania reprezentacji sygnału upmixu;przy czym reprezentacja sygnału upmixu jest dostarczana w oparciu o informację resztkową powiązaną z podzestawem obiektów audio reprezentowanym przez reprezentację sygnału downmixu;wherein the upmix signal representation is provided based on the residual information associated with the subset of audio objects represented by the downmix signal representation;przy czym reprezentacja sygnału downmixu jest rozkładana dla dostarczania pierwszej informacji audio opisującej pierwszy zestaw jednego lub większej liczby obiektów audio z pierwszego typu obiektów audio, z którym jest powiązana żadna informacja resztkowa i drugiej informacji audio opisującej drugi zestaw jednego lub większej liczby obiektów audio z drugiego typu obiektów audio, z którym nie jest powiązana żadna informacja resztkowa, w oparciu o reprezentację sygnału downmixu i z użyciem informacji resztkowej;wherein the downmix signal representation is broken down to provide the first audio information describing the first set of one or more audio objects from the first type of audio objects with which no residual information is associated and the second audio information describing the second set of one or more audio objects of the second type audio objects with no residual information associated with it, based on the downmix signal representation and using residual information;przy czym indywidualne obiektowe przetwarzanie obiektów audio z drugiego typu obiektów audio jest realizowane z użyciem obiektowej informacji parametrycznej powiązanej z więcej niż dwoma obiektami audio z drugiego typu obiektów audio;i przy czym informacja resztkowa opisuje zniekształcenie resztkowe, które zgodnie z oczekiwaniem pozostanie, jeśli obiekt audio z pierwszego typu obiektów audio jest wyizolowany wyłącznie z użyciem obiektowej informacji parametrycznej. wherein the individual object-oriented processing of audio objects from the second type of audio objects is performed using object-oriented parametric information associated with more than two audio objects from the second type of audio objects;and wherein the residual information describes a residual distortion that is expected to remain if the audio object of the first type of audio object is isolated using only parametric object information. 30. Program komputerowy do realizacji sposobu określonego w zastrzeżeniu 29, gdy program komputerowy jest wykonywany na komputerze. thirty. A computer program for performing the method of claim 29 when the computer program is executed on a computer. Fraunhofer-Gesellschaft zur Forderung der angewandten Forschung e.V., Niemcy Pełnomocnik: Fraunhofer-Gesellschaft zur Forderung der angewandten Forschung eV, Germany Representative: EP 2 446 435 B1 EP 2 446 435 B1 Z-10708/13 Z-10708/13 FIG 1 FIG 1 EP 2 446 435 B1 EP 2 446 435 B1 Z-10708/13 Z-10708/13 GENERAL STRUCTURE OF THE TRANSCODER / DECODER ARCHITECTURE SAOC OGÓLNA STRUKTURA ARCHITEKTURY TRANSKODERA/DEKODERA SAOC ARCHITEKTURA PROCESORA RESZTKOWEGO RESIDUAL PROCESSOR ARCHITECTURE FIG 3A FIG 3A ARCHITEKTURA PROCESORA RESZTKOWEGO RESIDUAL PROCESSOR ARCHITECTURE FIG 3B FIG 3B EP 2 446 435 B1 EP 2 446 435 B1 Z-10708/13 Z-10708/13 MO MO FJG 4A FJG 4A EP 2 446 435 B1 EP 2 446 435 B1 Z-10708/13 Z-10708/13 FIG 4Β FIG 4Β EP 2 446 435 B1 EP 2 446 435 B1 Z-10708/13 ιρι Z-10708/13 ιρι & DM. DM&. WCLD here. WCLD oto. IOC parameter regulator 'x-2x * 2 Lx-2 IOC regulator parametru 'x-2x*2 Łx-2 i) οπ tnl · i) οπ tnl· C, = M_FD'J C,= M_FD’J 1C | 1C| X X U-fDED '}) U-fDED’} ) G-Drra optional c G-Drrą opcjonalnie c ~ px t+ ł decorelat or ~px t+ł dekorelat or FIG 4C gain vector, n FIG 4C wektor wzmocnienia ,nł EP 2 446 435 B1 EP 2 446 435 B1 Z-10708/13 Z-10708/13 NI. NI. OMG OLD.IOC (&> OMG OLD.IOC (&> 0, = Μ "ΕΟΊ they you | G and ·" "" - optional gain vector 0,= Μ„ΕΟΊ oni ty |G i·"" „— wektor opcjonalnie wzmocnienia J. parameter regulator J. regulator parametru Ar decorator1* Γ Ar dekore lator r1* Γ FIG 4D FIG 4D EP 2 446 435 B1 EP 2 446 435 B1 Z-10708/13 downmix stereo (x-2-5) (tryb transkodowania) Z-10708/13 stereo downmix (x-2-5) (transcoding mode) DMC..DCLD OlOJOC DMC..DCLD OlOJOC M. M. C, -M, "E0 * 4 (J- (DED '|') parameter controller | C3, matched matched C,-M,„E0*4 (J-(DED'|') regulator parametru | C3,dopasowane dopasowane And optional I Ili opcjonalnie in.___ gain vector w.___ wektor wzmocnienia CPCl ??KCnICC COD CPCl??KCn.CŁD ICC Gum Gum X *1 , ή.Ι X*1 ,ή.Ι Φ - *. .p. - XX - X - X. - * XXX lpp;Φ--* . .p. - X. X. — X — X. — * X X X l p p;ri decorator ri dekore lator FIG 4E FIG 4E EP 2 446 435 B1 EP 2 446 435 B1 Z-10708/13 macierz renderowania Z-10708/13 rendering matrix FIG 4F macierz parametry renderowania HRTF FIG 4F matrix HRTF rendering parameters FIG 4G FIG 4G EP 2 446 435 B1 EP 2 446 435 B1 Z-10708/13 combined EKS SAOC system Z-10708/13 połączony system EKS SAOC 514 514 BASIC STRUCTURE OF THE CONNECTED SYSTEM EKS SAOC FIG SA PODSTAWOWA STRUKTURA POŁĄCZONEGO SYSTEMU EKS SAOC FIG SA 590 590 GENERAL STRUCTURE OF THE CONNECTED EKS SAOC SYSTEM UOGÓLNIONA STRUKTURA POŁĄCZONEGO SYSTEMU EKS SAOC FIG5B FIG5B EP 2 446 435 B1 EP 2 446 435 B1 Z-10708/13 Z-10708/13 FIG 6A FIG 6A FIG 6B FIG 6B EP 2 446 435 B1 EP 2 446 435 B1 Z-10708/13 Z-10708/13 EP 2 446 435 B1 EP 2 446 435 B1 Z-10708/13 Z-10708/13 EP 2 446 435 B1 EP 2 446 435 B1 Z-10708/13 Z-10708/13 FIG 6E FIG 6E EP 2 446 435 B1 EP 2 446 435 B1 Z-10708/13 Z-10708/13 700 700 710 710 720 720 -730 -730 FFG 7 FFG 7 EP 2 446 435 B1 EP 2 446 435 B1 Z-10708/13 Z-10708/13 1-1 ΓΝ 1-1 ΓΝ Φ- 'łt ·. "4 R = Φ- 'łt·. "4^= - From £? -O -θ '-OJ = and OOO CD C = | -£Z? -O -θ' -O J=i O O O CD C=| FIG 8 FIG 8 EP 2 446 435 B1 EP 2 446 435 B1 Z-10708/13 Z-10708/13 FIG 9A FIG 9A EP 2 446 435 B1 EP 2 446 435 B1 Z-10708/13 Z-10708/13 FIG 9B FIG 9B EP 2 446 435 B1 EP 2 446 435 B1 Z-10708/13 SAOC transcoder for MPEG Surround 986 Z-10708/13 transkoder SAOC do MPEG Surround 986 FIG 9C FIG 9C EP 2 446 435 B1 EP 2 446 435 B1 Z-10708/13 signal 1 enriched Z-10708/13 sygnał 1 wzbogaconych P P ?
552 paragraphs in 5 sections, as filed
Technical Field [0001] Embodiments of the invention relate to an audio signal decoder for providing an upmix signal representation based on a downmix signal representation and parametric object information.
[0002] Further embodiments of the invention relate to a method of providing an upmix signal representation based on the downmix signal representation and the parametric object information.
[0003] Further embodiments of the invention relate to a computer program.
[0004] Certain embodiments of the invention relate to the improved Karaoke / Solo SAOC system.
Background of the Invention [0005] In modern audio systems, it is desirable to transfer and store audio information in a streamlined manner. In addition, it is often desirable to play audio content using two or more speakers that are spatially positioned in the room. In such cases, it is desirable to use the capabilities of such multiple speaker systems to allow the user to spatially identify different audio content or components of a single audio content. This can be achieved by individually distributing different audio content to different speakers.
In other words, in the field of audio processing, audio transmission and audio storage, there is an increasing need to manipulate multi-channel content to improve the listening experience. The use of multi-channel audio content brings significant benefits to the user. For example, three-dimensional auditory sensations can be obtained, providing the user with greater satisfaction in entertainment applications. However, multi-channel content is also useful in a professional environment, such as in teleconferencing applications, because the speaker's intelligibility can be improved by using multi-channel playback.
[0007] However, a good balance between audio quality and bit rate requirements is also desirable to avoid excessive resource loads caused by multi-channel applications.
[0008] Recently, parametric techniques have been proposed for efficient bit rate and / or memory transmission of audio scenes containing multiple audio objects, for example Binaural Cue Coding (Type I) (see for example references [BCC]), Joint Source Coding (see for example references [JSC]) , and MPEG Spatial Audio Object Coding (SAOC) (see for example references [SAOC1], [SAOC2].
[0009] These techniques are aimed at the perceptual reconstruction of the desired audio output scene, rather than by means of waveform matching.
[0010] Fig. 8 shows an outline of such a system (in this case MPEG SAOC). The MPEG SAOC 800 system shown in Fig. 8 includes the SAOC 810 encoder and the SAOC 820 decoder. The SAOC 810 encoder receives a plurality of object signals from x1 to xN, which can be represented, for example, as time domain signals or frequency domain signals (e.g., in the form of set of Fourier transform coefficients or in the form of QMF subband signals). The SAOC 810 encoder also typically receives downmix d1 to dN coefficients that are associated with object signals x1 to xN. Separate sets of downmix coefficients may be available for each downmix signal channel. The SAOC 810 encoder is typically configured to obtain a downmix signal channel by combining object signals = x1 to xN in accordance with the associated downmix d1 to dN coefficients. Typically, there are fewer downmix channels than object signals x1 to xN. To enable (at least partially) the separation (or separate processing) of the object signals on the SAOC 820 decoder side, the SAOC 810 encoder provides both one or more downmix signals (labeled as downmix channels) 812 and additional information 814. Additional information 814 describes the signal properties object x1 to xN to enable object-oriented processing on the decoder side.
[0011] The SAOC decoder 820 is configured to receive both one or more downmix signals 812 and additional information 814. Also, the SAOC decoder 820 is typically configured to receive user interaction information and / or user control 822 that describes the desired rendering layout. For example, user interaction / user control information 822 may describe the speaker arrangement and the desired spatial arrangement of objects provided by object signals x1 to xN.
[0012] The SAOC decoder 820 is configured to provide, for example, many decoded upmix channel signals y<sub>1</sub> to y<sub>M</sub>. Upmix channel signals may for example be associated with individual speakers of a multi-speaker rendering system. The SAOC decoder 820 may, for example, include an object separator 820a that is configured to reconstruct, at least in part, object signals x1 to xN based on one or more downmix signals 812 and additional information 814, thereby obtaining reconstructed object signals 820b. However, the reconstructed object signals 820b may deviate slightly from the original object signals x1 to xN, for example because the additional information 814 is not completely sufficient for perfect reconstruction due to flow restrictions. The SAOC decoder 820 may further include a mixer 820c that can be configured to receive reconstructed object signals 820b and user interaction / user control information 822 and to provide y signals based thereon<sub>1</sub> to y<sub>M</sub> upmix channels. Mixer 820c may be configured to use user interaction / control information 822 to determine the contribution of individual reconstructed object signals 820b in signals y1 to yM of the upmix channels. For example, user interaction / user control information 822 may include rendering parameters (also referred to as rendering factors) that determine the contribution of individual reconstructed object signals 820b in y1 signals to yM upmix channels.
[0013] However, it should be noted that in many embodiments, object separation, which is indicated by the object separator 820a in Fig. 8, and mixing, which is indicated by the mixer 820c in Fig. 8, are carried out in one step. To this end, comprehensive parameters that describe the direct mapping of one or more downmix signals 812 to y1 to yM upmix channels can be calculated. These parameters can be calculated based on additional information 814 and user interaction / control information 822.
[0014] Referring now to Figs. 9a, 9b and 9c, various devices will be described for obtaining an upmix signal representation based on a downmix signal representation and object-oriented side information. Fig. 9a is a block diagram of an MPEG SAOC 900 system including a SAOC 920 decoder. The SAOC 920 decoder includes, as separate function blocks, an object decoder 922 and a mixing / rendering module 926. The object decoder 922 provides a plurality of reconstructed object signals 924 based on the downmix signal representation (e.g., in the form of one or more downmix signals represented in the time domain or frequency domain) and additional object information (e.g., in the form of object meta data). Mix / render module 926 receives the reconstructed object signals 924 associated with the number N of objects and provides, on their basis, one or more signals of the upmix channel 928. In the SAOC 920 decoder, the acquisition of object signals 924 is performed separately from mixing / rendering, which allows separation of the object decoding function from the mixing / rendering function, but entails relatively high computational complexity.
[0015] Referring now to Fig. 9b, another MPEG SAOC 930 system that includes the SAOC 950 decoder will be briefly described. The SAOC 950 decoder provides a plurality of upmix channel signals 958 based on a downmix channel representation (e.g. in the form of one or more) downmix signals) and additional object-oriented information (e.g. in the form of object-oriented meta data). The SAOC 950 decoder includes a combined object decoder and a mixing / rendering module that is configured to obtain 958 upmix channel signals in a combined mixing process without object decoding separation and mixing / rendering, wherein the parameters for said combined upmix process are dependent on both object-oriented additional information and rendering information. The combined upmix process also depends on the downmix information, which is considered part of the object-oriented additional information.
[0016] In summary of the above, providing signals 928, 958 of the upmix channels may be implemented in a one-step or two-step process.
[0017] Referring now to Fig. 9c, the MPEG SAOC 960 system will be described. The MPEG SAOC 960 system includes a 980 SAOC to MPEG Surround transcoder instead of a SAOC decoder.
[0018] The SAOC to MPEG Surround transcoder includes an additional information transcoder 982 that is configured to receive object-oriented additional information (e.g., in the form of object-oriented meta data) and optionally information about one or more downmix signals and rendering information. The additional information transcoder is also configured to provide 984 MPEG Surround additional information (e.g., as an MPEG Surround bit stream) based on the received data. Accordingly, the secondary information transcoder 982 is configured to transform the object (parametric) additional information that is received from the object encoder into the channel (parametric) additional information 984, including rendering information and optionally information about the content of one or more downmix signals.
[0019] Optionally, the SAOC 980 to MPEG Surround transcoder may be configured to manipulate one or more downmix signals described for example by a downmix signal representation to obtain a manipulated 988 downmix signal representation. However, the downmix signal manipulation module 986 may be omitted, so that the downmix output 988 of the SAOC to MPEG Surround transcoder 988 is identical to the downstream of the SAOC to MPEG Surround transcoder signal. The downmix signal manipulation module 986 can for example be used if the 984 MPEG Surround channel additional information does not provide the desired listening experience based on the downmix signal representation of the 980 SAOC to MPEG Surround transcoder, which may occur in some rendering arrangements.
[0020] Correspondingly, the SAOC 980 to MPEG Surround transcoder provides a 988 downmix signal representation and a 984 MPEG Surround bit stream so that multiple upmix channel signals can be generated that represent audio objects according to the rendering information input to the 980 SAOC to MPEG transcoder
Surround, using an MPEG Surround decoder that receives a 984 MPEG Surround bit stream and a representation of a 988 downmix signal.
[0021] In summary of the above, various methods of decoding SAOC-encoded audio signals can be used. In some cases, a SAOC decoder is used that provides upmix channel signals (e.g., signals 928, 958 upmix channels) based on the downmix signal representation and the object parametric additional information. Examples of this method are shown in Figs. 9a and 9b. Alternatively, the SAOC encoded audio information may be transcoded to obtain a downmix signal representation (e.g., as a 988 downmix signal representation) and additional channel information (e.g., MPEG Surround channel bit stream) that can be used by the MPEG Surround decoder to provide desired upmix channel signals.
[0022] In the 800 MPEG SAOC system, the general outline of which is shown in Fig. 8, the general processing is performed in a frequency selective manner and can be described as follows, in each frequency band:
* N input audio object signals x1 to xN are downmixed as part of the SAOC encoder processing. In the case of mono downmix, downmix coefficients are denoted by d1 to dN. In addition, the SAOC 810 encoder acquires additional information 814 describing the properties of the input audio objects. In the case of MPEG SAOC, the relationship of object power to each other is the most basic of such additional information.
* The downmix signal (or signal) 812 and additional information 814 are transmitted and / or stored. To this end, the audio downmix signal can be compressed using well-known perceptual audio encoders such as MPEG-1 Layer II or III (also known as ".mp3"), MPEG Advanced Audio Coding (AAC) or any other audio encoder.
* On the receiving side, the SAOC 820 decoder conceptually attempts to reconstruct the original object signal ("object separation") using the transmitted additional information 814 (and of course one or more downmix signals 812). These approximate object signals (also referred to as reconstructed object signals 820b) are then mixed into the target scene represented by M output audio channels (which may, for example, be represented by y signals<sub>1</sub> to y<sub>M</sub> upmix channels) using a rendering matrix. For mono output, the rendering matrix coefficients are given by r1 to rN.
* Effectively, object signal separation is rarely performed (or even never performed), because both the separation stage (indicated by the object 820a separator) and the mixing stage (indicated by the 820c mixer) are combined in one transcoding stage, which often allows a huge reduction in complexity computing.
[0023] This method has been found to be extremely efficient, both in terms of bit rate (only a few downmix channels need to be transmitted plus some additional information instead of N separate audio object signals or a discrete system) and computational complexity (processing complexity is mainly related to the number output channels rather than with the number of audio objects). Further benefits for the user on the receiving side include the freedom to choose the rendering layout according to his / her choice (mono, stereo, virtual headphone playback, etc.) and the user interaction function: the rendering matrix, and thus the output scene, can be set and changed interactively by the user according to his will, personal preferences or other criteria. For example, it is possible to locate speakers from one group together in one spatial area to maximize discrimination from other speakers. This interactivity is obtained by providing a decoder user interface.
[0024] For each transmitted sound object, its relative level and (for non-mono rendering) the spatial position of the rendering can be adjusted. This can occur in real time when the user changes the positions of the associated sliders of the graphical user interface (GUI) (for example: object level = + 5dB, object position = -30 degrees).
[0025] However, it has been found difficult to manipulate audio objects from different types of audio objects in such a system. In particular, it has been found that it is difficult to process audio objects from different types of audio objects, e.g., audio objects with which different additional information is associated, if the total number of audio objects for processing is not pre-determined.
[0026] In view of this situation, the object of the present invention is to create a concept that allows computationally efficient and flexible decoding of an audio signal comprising a downmix signal representation and parametric object information in which the object parametric information describes the audio objects of two or more different types of audio objects .
Summary of the Invention [0027] This object is achieved by an audio signal decoder for providing an upmix signal representation based on a downmix signal representation and parametric object information, a method of providing an upmix signal representation based on a downmix signal representation and object parametric information and a computer program as defined in independent reservations.
[0028] An embodiment of the invention provides an audio signal decoder for providing an upmix signal representation based on the downmix signal representation and the parametric object information. The audio signal decoder includes an object separator configured to decompose the downmix signal to provide the first audio information describing the first set of one or more audio objects of the first type of audio objects and the second audio information describing the second set of one or more audio objects of the second type of audio objects based for representing the downmix signal and using at least part of the object parametric information, wherein the second audio information is audio information describing audio objects of the second type of audio objects in a combined manner. The audio signal decoder also includes an audio signal processor configured to receive the second audio information and to process the second audio information depending on the parametric object information to obtain a processed version of the second audio information. The audio signal decoder also includes an audio signal combining module configured to combine the first audio information with the processed version of the second audio information to obtain an upmix signal representation.
[0029] The main idea of the present invention is that efficient processing of different types of audio objects can be achieved in a cascading structure that allows the separation of different types of audio objects using at least part of the object parametric information in the first processing step, implemented by the object separator and which enables additional spatial processing in the second processing stage implemented based on at least part of the object parametric information by the audio signal processor.
[0030] It has been found that acquiring a second audio information that includes audio objects of the second type of audio objects from the downmix signal representation can be performed with moderate complexity, even if there are more audio objects of the second type of audio objects. In addition, it has been found that the spatial processing of audio objects of the second type of audio objects can be performed efficiently after the second audio information is separated from the first audio information describing the audio objects of the first type of audio objects.
[0031] In addition, it has been found that the processing algorithm implemented by the object separator for separating the first audio information and the second audio information can be implemented with relatively low complexity if the individual object processing of the audio objects of the second type of audio objects is delayed for the audio signal processor and is not implemented simultaneously with the separation of the first audio information from the second audio information.
[0032] The audio signal decoder is therefore further configured to provide an upmix signal representation based on the downmix signal representation, parametric object information and residual information associated with the subset of audio objects represented by the downmix signal representation. In this case, the object separator is configured to distribute the downmix signal representation to provide the first audio information describing the first set of one or more audio objects (e.g., FGO foreground objects) of the first type of audio objects with which the residual information and the second audio information are associated describing a second set of one or more audio objects (e.g., BGO background objects) of a second type of audio object, with which no residual information is associated, based on the downmix signal representation and using at least part of the object parametric information and the residual information.
[0033] The audio signal processor is configured to receive the second audio information and to process the second audio information based (at least in part) on the parametric object information using the object parametric information associated with more than two audio objects of the second type of audio objects. Accordingly, individual object processing is carried out by the audio processor and such individual object processing is not carried out for audio objects of the second type of audio objects by the object separator.
[0034] This is based on the finding that a particularly accurate separation of the first audio information describing the first set of audio objects from the first type of audio objects and the second audio information describing the second set of audio objects from the second type of audio objects can be obtained using residual information in addition to object information parametric. It has been found that the use of only object-oriented parametric information will in many cases bring distortion that can be significantly reduced or even completely eliminated by using residual information. The residual information describes the residual distortion that is expected to remain if the audio object of the first type of audio object is isolated only using object-oriented parametric information. Residual information is typically estimated by the audio signal encoder. By using residual information, the separation between audio objects of the first type of audio objects and audio objects of the second type of audio objects can be improved.
[0035] This enables the first audio information and the second audio information to be obtained with particularly good separation between the audio objects of the first type of audio objects and the audio objects of the second type of audio objects, which in turn enables obtaining high quality spatial processing of the audio objects of the second type of audio objects processing the second audio information in an audio signal processor.
[0036] In a preferred embodiment, the object separator is thus configured to provide the first audio information in such a way that the audio objects of the first type of audio objects are highlighted relative to the audio objects of the second type of audio objects in the first audio information. The object separator is also configured to provide the second audio information in such a way that the audio objects of the second type of audio objects are enhanced relative to the audio objects of the first type of audio objects in the second audio information.
[0037] In a preferred embodiment, the audio signal decoder is configured to perform two-stage processing in such a way that the processing of the second audio information in the audio signal processor is performed after separation between the first audio information describing the first set of one or more audio objects of the first type of objects audio and second audio information describing a second set of one or more audio objects from the second type of audio objects.
[0038] In a preferred embodiment, the audio signal processor is configured to process the second audio information based on the object parametric information associated with the audio objects of the second type of audio objects and independent of the object parametric information associated with the audio objects of the first type of audio objects. Thus, separate processing of audio objects from the first type of audio objects and audio objects from the second type of audio objects can be obtained.
[0039] In a preferred embodiment, the object separator is configured to obtain the first audio information and the second audio information using a linear connection of one or more downmix channels and one or more residual channels. In this case, the object separator is configured to obtain connection parameters for performing a linear connection based on the downmix parameters associated with the audio objects of the first type of audio objects and based on the channel prediction coefficients of the audio objects from the first type of audio objects. The calculation of the channel prediction coefficients of the audio objects of the first type of audio objects may for example include audio objects of the second type of audio objects as a single common audio object. Accordingly, the separation process may be carried out with a sufficiently low computational complexity, which for example may be independent of the number of audio objects from the second type of audio objects.
[0040] In a preferred embodiment, the object separator is configured to use a rendering matrix for the first audio information to map object signals of the first audio information to the audio channels of the audio upmix signal. This can be done because the object separator may be able to obtain separate audio signals individually representing the audio objects from the first type of audio objects. Accordingly, it is possible to map object signals of the first audio information directly to audio channels from the audio upmix signal representation.
In a preferred embodiment, the audio processor is configured to perform stereo processing of the second audio information based on rendering information, object covariance information and downmix information to obtain audio channels of the audio upmix signal representation.
[0042] Thus, the stereo processing of audio objects of the second type of audio objects is separated from separating the audio objects of the first type of audio objects from the audio objects of the second type of audio objects. In this way, the efficient separation of audio objects from the first type of audio objects from the audio objects of the second type of audio objects is not disturbed (or degraded) by stereo processing, which typically leads to the distribution of audio objects on multiple audio channels without providing a high degree of separation of objects that can be obtained in an object separator, for example using residual information.
[0043] In a preferred embodiment, the audio signal processor is configured to perform post-processing of the second audio information based on rendering information, object covariance information and downmix information. This form of post-processing allows the spatial arrangement of audio objects from the second type of audio objects on the audio scene. Regardless, thanks to the cascade concept, the computational complexity of the audio signal processor can be kept low enough because the audio processor does not have to consider object-oriented parametric information associated with the audio objects of the first type of audio objects.
[0044] Additionally, various types of processing can be performed by an audio processor, such as, for example, mono to binaural processing, mono to stereo processing, stereo to binaural processing or stereo to stereo processing.
[0045] In a preferred embodiment, the object separator is configured to treat the audio objects of the second type of audio objects to which no residual information is associated as a single audio object. In addition, the audio signal processor is configured to include object-oriented rendering parameters to match object inputs from the second type of audio objects to the upmix signal representation. In this way, the audio objects of the second type of audio objects are treated by the object separator as a single audio object, which significantly reduces the complexity of the object separator as well as allows you to have unique residual information that is independent of the rendering parameters associated with the audio objects of the second type of audio objects.
In a preferred embodiment, the object separator is configured to obtain a common object value of the level difference for a plurality of audio objects from the second type of audio objects. The object separator is configured to use the level difference object value to calculate channel prediction coefficients. In addition, the object separator is configured to use channel prediction coefficients to obtain one or two audio channels representing the second audio information. To obtain the object value of the level difference, audio objects from the second type of audio objects can be efficiently treated by the object separator as a single audio object.
[0047] In a preferred embodiment, the object separator is configured to obtain a common object value of level difference for multiple audio objects of the second type of audio objects and the object separator is configured to use a common object value of level difference to calculate the words of the energy mode mapping matrix. The object separator is configured to use an energy mode mapping matrix to obtain one or more audio channels representing the second audio information. Again, the common value of the object level difference allows computationally efficient common treatment of audio objects from the second type of audio objects by the object separator.
[0048] In a preferred embodiment, the object separator is configured to selectively obtain a common cross-object correlation value associated with the audio objects of the second type of audio objects based on the parametric object information, if it is determined that there are two audio objects of the second type of audio objects and to be set cross-object correlation value associated with audio objects from the second type of audio objects to zero, if it is found that there are more or less than two audio objects from the second type of audio objects.
[0049] The object separator is configured to use a common cross-object correlation value associated with the audio objects of the second type of audio objects to obtain one or more audio channels representing the second audio information. In this approach, the inter-object correlation value is used if it is achievable with high computational efficiency, i.e. if there are two audio objects from the second type of audio objects. Otherwise, it would be computationally demanding to obtain cross-object correlation values. Therefore, it was found that a good compromise in terms of auditory experience and computational complexity is to set the cross-object correlation value associated with audio objects from the second type of audio objects to zero if there are more or less than two audio objects from the second type of audio objects.
[0050] In a preferred embodiment, the audio processor is configured to render the second audio information based (at least in part) on the parametric object information to obtain a rendered representation of the audio objects from the second type of audio objects as a processed version of the second audio information. In this case, rendering can be done independently of the audio objects from the first type of audio objects.
[0051] In a preferred embodiment, the object separator is configured to provide the second audio information in such a way that the second audio information describes more than two audio objects from the second type of audio objects. Embodiments of the invention allow flexible matching of the number of audio objects from the second type of audio objects, which is greatly facilitated by the cascade processing structure.
[0052] In a preferred embodiment, the object separator is configured to obtain, as the second audio information, a single-channel representation of the audio signal or a two-channel representation of audio representing more than two audio objects from the second type of audio objects. Acquiring one or two channels of audio signal can be carried out by an object separator with low computational complexity. In particular, the complexity of the object separator can be kept clearly lower compared to the case where the object separator would have to deal with more than two audio objects from the second type of audio objects. However, it has been independently found that a computationally efficient representation of audio objects from a second type of audio object is to use one or two audio signal channels.
[0053] In a preferred embodiment, the audio decoder is configured to obtain information about the total number of objects and information about the number of foreground objects from the configuration information associated with the object parametric information. The audio decoder is also configured to determine the number of audio objects of the second type of audio objects by creating a difference between information about the total number of objects and information about the number of foreground objects. In this way, efficient signaling of the number of audio objects of the second type of audio objects is obtained. In addition, this concept provides a high degree of flexibility in the number of audio objects of the second type of audio objects.
[0054] In a preferred embodiment, the object separator is configured to use object-oriented parametric information associated with Neao audio objects from the first type of audio objects to obtain, as first audio information, Neao, audio signals representing (preferably, individually) Neao audio objects of the first type audio objects and for obtaining second audio information, one or two audio signals representing N-Neao audio objects from the second type of audio objects, treating NNeao audio objects from the second type of audio objects as a single single-channel or two-channel audio object. The audio signal processor is configured to individually render N-Neao audio objects represented by one or two audio signals from the second audio information using object-oriented parametric information associated with N-Neao audio objects from the second type of audio objects. Accordingly, the separation of audio objects between the audio objects of the first type of audio objects and the audio objects of the second type of audio objects is separated from the subsequent processing of audio objects of the second type of audio objects.
[0055] An embodiment of the invention provides a method of providing an upmix signal representation based on the downmix signal representation and the parametric object information as defined in independent claim 29.
[0056] Another embodiment of the present invention provides a computer program for performing said method as defined in independent claim 30.
Brief description of the figures [0057] Embodiments of the invention will then be described with reference to the attached figures, in which:
Fig. 1 shows a block diagram of an audio decoder according to an embodiment of the invention;
Fig. 2 shows a block diagram of another audio decoder according to an embodiment of the invention;
Figures 3a and 3b are block diagrams of a residual processor that can be used as an object separator in an embodiment of the invention;
Figures 4a to 4e show block diagrams of audio signal processors that can be used in an audio signal decoder according to an embodiment of the invention
Fig. 4f is a block diagram of a SAOC transcoder processing mode;
Fig. 4g shows a block diagram of a SAOC decoder processing mode;
Fig. 5a is a block diagram of an audio decoder according to an embodiment of the invention;
Fig. 5b shows a block diagram of another audio decoder, according to an embodiment of the invention;
Fig. 6a is a Table representing a description of the listening test structure;
Fig. 6b is a Table representing the system being tested;
Fig. 6c is a Table representing the listening test elements and rendering matrices;
Fig. 6d shows a graphical representation of the average MUSHRA results for the Karaoke / Solo listening test;
Fig. 6e is a graphical representation of the average MUSHRA results for the classic listening test;
Fig. 7 is a flowchart of a method of providing an upmix signal representation according to an embodiment of the invention;
Fig. 8 is a block diagram of a reference MPEG SAOC system;
Fig. 9a is a block diagram of a reference SAOC system using a separate decoder and mixer;
Fig. 9b is a block diagram of a reference SAOC system using an integrated decoder and mixer;
and Fig. 9c is a block diagram of a reference SAOC system using a SAOC to MPEG transcoder.
A detailed description of the embodiments
1. Audio signal decoder according to Fig. 1 [0058] Fig. 1 is a block diagram of an audio signal decoder 100 according to an embodiment of the invention.
[0059] The audio signal decoder 100 is configured to receive object parametric information 110 and a downmix signal representation 112. The audio signal decoder 100 is configured to provide the upmix signal representation 120 based on the downmix signal representation and the object parametric information 110. The audio decoder 100 includes a 130 object separator, which is configured to decompose the downmix signal representation 112 to provide the first audio information 132 describing the first set of one or more audio objects of the first type of audio objects and the second audio information 134 describing the second set of one or more audio objects of the second type of audio objects based on representation of a 112 downmix signal and using at least part of the object parametric information 110. The audio signal decoder 100 also includes an audio signal processor 140 that is configured to receive the second audio information 134 and to process the second audio information based on at least a portion of the object parametric information 112 to obtain the processed version 142 of the second audio information 134. The audio signal decoder 100 also includes an audio combining module 150 configured to combine the first audio information 132 with the processed version 142 of the second audio information 134 to obtain an upmix signal representation 120.
[0060] The audio signal decoder 100 performs cascade processing of a downmix signal representation that represents audio objects from the first type of audio objects and audio objects from the second type of audio objects in a combined manner.
[0061] In a first processing step that is performed by the object separator 130, the second audio information describing the second set of audio objects of the second type of audio objects is separated from the first audio information 132 describing the set of audio objects of the first type of audio objects using the object parametric information 110. However, the second audio information 134 typically is audio information (e.g., a single-channel audio signal or a two-channel audio signal) describing audio objects of the second type of audio objects in a combined manner.
[0062] In a second processing step, the audio signal processor 140 processes the second audio information 134 based on the object parametric information. Thus, the audio signal processor 140 is capable of performing individual object processing or rendering of audio objects of a second type of audio objects that are described by the second audio information 134, which is typically not implemented by the object separator 130.
[0063] In this way, although audio objects from the second type of audio objects are preferably not processed in an individual object manner by the object separator 130, audio objects from the second type of audio objects are actually processed in an individual object way (e.g. rendered in an individual object way ) in a second processing step which is carried out by the audio signal processor 140. Thus, the separation between the audio objects of the first type of audio objects and the audio objects of the second type of audio objects, which is implemented by the object separator 130, is separated from the individual object-oriented processing of audio objects of the second type of audio objects, which is then implemented by the signal processor 140 audio. Accordingly, the processing that is carried out by the object separator 130 is substantially independent of the number of audio objects of the second type of audio objects. In addition, the format (e.g., single-channel audio or two-channel audio) of the second audio information 134 is typically independent of the number of audio objects of the second type of audio objects. Thus, the number of audio objects of the second type of audio objects can change without the need to modify the structure of the 130 object separator. In other words, audio objects from the second type of audio objects are treated as a single (e.g. single-channel or two-channel) audio object for which common object parametric information (e.g. a common value of object level difference) is obtained by the object separator 140.
[0064] Accordingly, the audio signal decoder 100 of Fig. 1 is capable of handling a variable number of audio objects of the second type of audio objects without structural modification of the object separator 130. Additionally, various algorithms for processing audio objects by the object separator 130 and the audio signal processor 140 may be used. Accordingly, for example, it is possible to perform audio object separation using residual information through the object separator 130, which enables particularly good separation of different audio objects, using residual information, which is additional information for improving the quality of object separation. In contrast, the audio signal processor 140 may perform individual object-oriented processing without using residual information. For example, the audio signal processor 140 may be configured to perform spatial-audio-object-coding (SAOC) audio signal processing to render various audio objects.
2. Audio signal decoder according to Fig. 2 [0065] In the following, an audio signal decoder 200 according to an embodiment of the invention will be described. A block diagram of this audio decoder 200 is shown in Fig. 2.
[0066] The audio signal decoder 200 is configured to receive a downmix signal 210, so-called SAOC bit stream 212, rendering matrix information 214 and optional head-related-transferfunction (HRTF) parameters 216. The audio decoder 200 is also configured to provide the downmix / MPS signal 220 and (optional) MPS 222 bit stream
2.1. Input signals and output signals of the audio decoder 200 [0067] Various details regarding the input and output signals of the audio decoder 200 will be described below.
[0068] The downmix signal 200 may for example be a single-channel audio signal or a two-channel audio signal. The downmix signal 210 may, for example, be obtained from an encoded representation of the downmix signal.
[0069] The bit stream SAOC 212 may for example contain object-oriented parametric information. For example, the SAOC bit stream 212 may contain object level difference information, for example, in the form of object level difference OLD parameters, inter-object correlation information, for example, in the form of inter-object correlation IOC parameters.
[0070] In addition, the SAOC bit stream 212 may include downmix information describing how the downmix signals have been provided based on the signals of the audio objects using the downmix process. For example, the SAOC bit stream may include DMG downmix gain parameters and (optional) DCLD parameters of downmix channel level differences.
[0071] Rendering matrix information 214 may, for example, describe how different audio objects should be rendered by an audio decoder. For example, rendering matrix information 214 may describe the allocation of audio objects to one or more output / MPS downmix signal channels 220.
[0072] Optional information 216 head transfer function (HRTF) may further describe the transfer function to obtain a binaural headphone signal.
[0073] The output / MPEG Surround downmix signal 220 (also briefly referred to as "the output downmix signal / MPS") represents one or more audio channels, for example in the form of a time domain representation of a audio signal or a frequency domain representation of an audio signal. Alone or in combination with the optional 222 MPEG Surround bit stream (MPS bit stream), which includes MPEG Surround parameters describing the mapping of the 220 downmix output / MPS signal to multiple audio channels, an upmix signal representation is created.
2.2. Structure and functions of the audio decoder 200 [0074] In the following, the structure of the audio signal decoder 200 will be described in more detail, which may act as a SAOC transcoder or a SAOC decoder.
[0075] The audio signal decoder 200 comprises a downmix processor 230 that is configured to receive the downmix signal 210 and to provide an output / MPS downmix signal 220 thereof. The downmix processor 230 is also configured to receive at least part of the SAOC bit stream information 212 and at least part of the information 214 of the rendering matrix. In addition, the downmix processor 230 can also receive 240 processed SAOC parameter information from the 250 parameter processor.
[0076] The parameter processor 250 is configured to receive SAOC bit stream information 212, rendering matrix information 214 and optional head transfer function information information 260, and to provide MPEG Surround 222 bit stream based thereon containing MPEG Surround parameters (if MPEG Surround parameters are required , for example, in the transcoding mode of operation). In addition, the 250 parameter processor provides processed SAOC information 240 (if processed SAOC information is required).
[0071] The structure and functions of the downmix processor 230 will be described in more detail below.
[0078] The downmix processor 230 comprises a residual processor 260 that is configured to receive the downmix signal 210 and to provide on its basis a first signal 262 of an audio object describing so-called enhanced audio objects (EAOs) that can be referred to as first-type audio objects audio. The first audio object signal may contain one or more audio channels and may be considered as the first audio information. The residual processor 260 is also configured to provide a second signal 264 of audio objects that describes the audio objects of the second type of audio objects and can be considered as the second audio information. The second audio object signal 264 may include one or more channels and typically may include one or more audio channels describing multiple audio objects. Typically, the second audio object signal may describe even more than two audio objects from the second type of audio objects.
[0079] Downmix processor 230 also includes a SAOC downmix pre-processor 270 that is configured to receive the second audio object signal 264 and to provide a processed version thereof 272 of the second audio object signal 264, which may be considered as the processed version of the second audio information.
[0080] The downmix processor 230 also includes an audio signal combining module 280 that is configured to receive the first audio object signal 262 and the processed version 272 of the second audio object signal 264 and to provide an output / MPS downmix signal 220 thereof that can be recognized. , alone or together with the (optional) MPEG Surround 222 bit stream, for representing the upmix signal.
[0081] The functions of the individual downmix processor 230 units will be discussed in more detail below.
[0082] The residual processor 260 is configured to separately provide the first audio object signal 262 and the second audio object signal 264. To this end, the residual processor 260 may be configured to use at least part of the SAOC bit stream information 212. For example, the residual processor 260 may be configured to evaluate the parametric object information associated with the audio objects of the first type of audio objects, i.e. EAO so-called "enhanced audio objects". In addition, the residual processor 260 may be configured to obtain comprehensive information describing the audio objects from the second type of audio objects, for example the so-called colloquially "enriched audio objects". The residual processor 260 can also be configured to evaluate the residual information, which is provided in the SAOC bit stream information 212, for separation between enriched audio objects (audio objects from the first type of audio objects) and non-enriched audio objects (audio objects from the second type of audio objects) . The residual information may, for example, encode a time domain residual signal that is used to obtain particularly pure separation between enriched audio objects and unenriched audio objects. In addition, the residual processor 260 may optionally evaluate at least a portion of the information 214 of the rendering matrix, for example to determine the distribution of enriched audio objects in the audio channels of the first signal 262 of audio objects, [0083] The SAOC downmix pre-processor 270 includes a channel redistribution module 274 which is configured to receiving one or more audio channels from the other 264 audio objects and for delivery based thereon, one or more (typically two) audio channels of the processed second signal 272 of the audio object. In addition, the SAOC downmix pre-processor 270 includes a de-correlated signal provider 276, which is configured to receive one or more audio channels of a second signal 264 audio objects and to provide one or more de-correlated signals 278a, 278b based thereon which are added to the signals supplied by the channel redistribution module 274 to obtain the processed version 272 of the second audio object signal 264.
[0084] Further details regarding the SAOC downmix processor will be discussed below.
[0085] The audio signal combining module 280 combines the first audio object signal 262 with the processed version 272 of the second audio object signal. To this end, channel bonding can be carried out. Accordingly, an output / MPS downmix signal 220 is obtained.
[0086] The parameter processor 250 is configured to obtain (optional) MPEG Surround parameters, which are the MPEG Surround bit stream 222 of the upmix signal representation based on the SAOC bit stream, including rendering matrix information 214 and optional HRTF parameter information 216. In other words, the SAOC parameter processor 252 is configured to transform object parameter information, which is described by SAOC bit stream information 212 into channel parametric information, which is described by the MPEG Surround 222 bit stream.
[0087] The following is a brief overview of the structure of the SAOC transcoder / decoder architecture shown in Fig. 2. Spatial audio object coding (SAOC) is a parametric multi-object coding technique. It is designed to transmit a number of audio objects in an audio signal (e.g., an audio downmix 210 signal) that contains M channels. Together with the backwards compatible downmix signal, object parameters (e.g. using SAOC bit stream information 212) that are able to reconstruct and manipulate the original object signals are transmitted. The SAOC encoder (not shown here) produces a downmix of object signals at its input and acquires these object parameters. The number of objects that can be handled is, in principle, unlimited. Object parameters are quantized and encoded efficiently into the SAOC 212 bit stream. Downmix 210 signal can be compressed and sent without the need to update existing encoders and infrastructures. The object parameters or SAOC decoder information are transmitted in an additional low bit rate channel, e.g., in the auxiliary part of the downmix bit stream data.
[0088] On the decoder side, input objects are reconstructed and rendered into a number of playback channels. Rendering information containing the reconstruction level and panoramic position for each object is provided by the user or can be obtained from the SAOC bit stream (for example, as preset information). The rendering information may be variable over time. Output circuits can range from mono to multi-channel (e.g. 5.1) and are independent of both the number of input objects and the number of downmix channels. Binaural rendering of objects is possible, including azimuth and elevation of virtual objects. The optional effects interface allows advanced object signal manipulation, in addition to level and panorama modifications.
[0089] The objects themselves may be mono signals, stereo signals as well as multi-channel signals (e.g., 5.1 channels). Typical downmix configurations are mono and stereo.
[0090] The basic structure of the SAOC transcoder / decoder will be explained below, which is shown in Fig. 2. The SAOC transcoder / decoder module discussed here can act either as a stand-alone decoder or as a transcoder from SAOC to the MPEG Surround bit stream, depending on the target configuration output channel. In the first mode of operation, the output signal configuration is mono, stereo or binaural configuration and two output channels are used. In the first case, the SAOC module can operate in the decoder mode, and the output of the SAOC module is the output in pulse-code modulation (PCM output). In the first case, no MPEG Surround decoder is required. Instead, the upmix signal representation may contain only output signal 220 and provision of the MPEG Surround 222 bit stream may be omitted. In the second case, the output configuration is a multi-channel configuration with more than two output channels. The SAOC module can operate in transcoder mode. The SAOC module output signal can include both downmix signal 220 and MPEG Surround 222 bit stream, as shown in Fig. 2. Accordingly, an MPEG Surround decoder is necessary to obtain the final representation of the audio signal for reproduction through the speakers.
[0091] Fig. 2 shows the basic structure of the SAOC transcoder / decoder architecture. The residual processor 216 obtains enriched audio objects from the downmix input 210 using the residual information contained in the SAOC 212 bit stream. The downmix pre-processor 270 processes regular audio objects (which are, for example, non-enriched audio objects, i.e. no audio objects for which no residual information in the bit stream SAOC 212). Enriched audio objects (represented by the first audio object signal 262) and processed regular audio objects (represented by, for example, the processed version 272 of the second audio object signal 264) are combined into output signal 220 for SAOC decoder mode or as a signal
220 MPEG Surround downmix for SAOC transcoder mode. A detailed description of the processing blocks is given below.
3. Architecture and functions of the residual processor and the energy mode processor [0092] The details of the residual processor which, for example, can take over the function of the separator 130 of the audio decoder objects 100 or the residual processor 260 of the audio decoder 200 will be discussed below. To this end, Figures 3a and 3b show block diagrams of such a residual processor 300 that can take the place of the object separator 130 or the residual processor 260. 3a shows less details than Fig. 3b. However, the following description applies to the residual processor 300 according to Fig. 3a as well as the residual processor 380 according to Fig. 3b.
[0093] The residual processor 300 is configured to receive the SAOC downmix signal 300, which may be equivalent to the downmix signal representation 112 of Figure 1 or the downmix signal representation 210 of Figure 2. The residual processor 300 is configured to provide first information based thereon. audio 320 describing one or more enhanced audio objects, which may be, for example, equivalent to the first audio information 132 or the second signal 262 of the audio object. Also, the residual processor 300 may provide the second audio information 322 describing one or more other audio objects (e.g., non-enriched audio objects for which no residual information is available), wherein the second audio information 322 may be equivalent to the second audio information 134 or with the second signal 264 of the audio object.
[0094] The residual processor 300 includes a 1-to-N / 2-to-N 330 unit (OTN / TTN unit) that receives the SAOC downmix signal 310 and which also receives SAOC 332 data and debris. The 330 1-to-N unit / 2-to-N also provides a 334 enriched audio object signal that describes the enriched audio objects (EAOs) contained in the SAOC downmix 310 signal. Also, the 330 1-to-N / 2-to-N unit provides second audio information 322. Residual processor 300 also includes a rendering unit 340 that receives the signal 334 of enhanced audio objects and rendering matrix information 342 and provides the first audio information 320 thereon.
[0095] The following will describe in more detail the processing of enhanced audio objects (EAO processing) that is performed by the residual processor 300.
3.1. Introduction to operation of the residual processor 300 [0096] When considering the functionality of the residual processor 300, it should be noted that the SAOC technology allows individual manipulation of a number of audio objects in terms of their gain / attenuation without significantly compromising the final sound quality, only in a very limited way. The special "karaoke" application mode requires complete (or almost total) suppression of certain objects, typically the lead vocal, while maintaining undisturbed perceptual quality of the background sound.
[0097] A typical use case includes up to four Enriched Audio Object (EAO) signals, which may, for example, represent two independent stereo objects (e.g., two independent stereo objects that are prepared for their removal on the decoder side).
[0098] It should be noted that one or more qualitatively enriched audio objects (or more precisely, audio signal inputs associated with the enriched audio objects) is included in the SAOC downmix signal 310. Typically, audio signal cartridges associated with one or more enriched audio objects are mixed in the downmix processing implemented by the audio signal encoder with the audio signal cartridges of other audio objects that are not enriched audio objects. Also, it should be noted that audio signal inputs from many enriched audio objects are also typically superimposed or mixed by downmix processing implemented by the audio signal encoder.
3.2 SAOC architecture supporting enriched audio objects [0099] The details of the residual processor 300 will be discussed below. Enriched audio object processing includes 1-to-N or 2-to-N units depending on the SAOC downmix mode. The 1-to-N processing unit is dedicated to the mono downmix signal and the 2-to-N processing unit is dedicated to the stereo downmix signal 310. Both of these units represent the generalized and enriched modification of the 2-to-2 block (TTT block) known from the ISO / IEC 23003-1: 2007 standard. In the encoder, regular signals and EAO signals are combined into a downmix signal. OTN processing units<sup>-1</sup>/ TTN<sup>-1</sup> (which are inverted 1-to-N processing units or inverted 2-to-N processing units) are used to generate and encode the corresponding residual signals.
[0100] EAO signals and regular signals are reconstructed from the downmix signal 310 by 330 OTN / TTN units using the SAOC additional information and residual signals contained. The reconstructed EAOs (which are described by the signal 334 of the enriched audio objects) are provided to the rendering unit 340 which represents (or provides) the product of the respective rendering matrix (described by the rendering matrix information 342) and the final output signal from the OTN / TTN unit. Regular audio objects (which are described by the second audio information 322) are provided to the SAOC downmix pre-processor, e.g., SAOC downmix pre-processor 270, for further processing. Figures 3a and 3b show the general structure of the residual processor, i.e. the architecture of the residual processor.
[0101] The residual processor output signals 320, 322 are calculated as ^ OBJ <sup>=</sup> X<sub>paradise</sub> ,
<img file="PL2446435T3_D0001.tif" />
where XOBJ represents the downmix signal of regular audio objects (i.e. not EAO) and XEAO is the rendered EAO output signal for SAOC decoding mode or the corresponding EAO downmix signal for SAOC transcoding mode.
[0102] The residual processor may operate in a prediction mode (using residual information) or in an energy mode (without residual information). The Xres extended input signal is defined as follows:
<img file="PL2446435T3_D0002.tif" />
For prediction mode \ /
X, For energy mode [0103] In this case, X may, for example, represent one or more channels of the downmix signal representation 310, which may be carried in a bit stream representing multi-channel audio content. res can mean one or more residual signals that can be described by a bit stream representing multi-channel audio content.
[0104] OTN / TTN processing is represented by the M matrix and the EAO processor by the AEAO matrix.
[0105] The OTN / TTN processing M matrix is defined according to the EAO mode of operation (i.e. prediction or energy) as
· Fredietiort <sub>(</sub> ^ Frediefion »For prediction mode
M
L s 'w' For energy mode [0106] The OTN / TTN processing M matrix is represented as where the MOBJ matrix relates to regular audio objects (ie not EAO) and the MEAO matrix relates to enriched audio objects (EAO).
[0107] In some embodiments, one or more multi-channel background objects (MBOs) can be treated in the same way by the residual processor 300.
[0108] The Multi-channel Background Object (MBO) is a mono or stereo MPS downmix that is part of the SAOC downmix. Unlike the use of individual SAOC objects for each channel in a multi-channel signal, MBO can be used enabling SAOC to more efficiently manipulate a multi-channel object. For MBO, the SAOC overhead decreases when the SAOC parameters of the MBO are associated only with downmix channels rather than all upmix channels.
3.3 Additional definitions
3.3.1 Dimension of signals and parameters [0109] The dimensions of signals and parameters will be briefly discussed below to explain how often various calculations are performed.
[0110] Audio signals are defined for each time interval n and each hybrid subband k (which may be a frequency subband). The corresponding SAOC parameters are defined for each time interval of 1 parameter and processing bandwidth m. The following mapping between hybrid and parameter domain is specified in Table A. 31 ISO / IEC 23003-1: 2007. Hence, all calculations are performed for certain time / band indices and the corresponding dimensions are implied for each variable entered.
[0111] However, below the time and frequency band indices will sometimes be omitted to maintain notation compactness.
3.3.2 Calculation of the AEAO matrix [0112] The AEAO pre-rendering matrix of EAO objects is defined according to the number of output channels (i.e. mono, stereo or binaural) as
<img file="PL2446435T3_D0003.tif" />
in the mono case in other cases.
[0113] Matrices defined as
<img file="PL2446435T3_D0004.tif" />
with size 1 x N<sub>so</sub> and + size 2 x N<sub>so</sub> are
A - η<sup>ε</sup>ΑθΚΛ<sup>ε</sup>*°
I ”^ 16 '
IN,
EAO ,,, ΕΑΟ ... EAO. "£ 40 EAO
IN,
EDA
).
<img file="PL2446435T3_D0005.tif" />
Λ / where the rendering sub-matrix · '· -<sup>1</sup> corresponds to EAO rendering (and describes the desirable mapping of enhanced audio objects to upmix signal representation channels).
, m><sup>EA0</sup> , [0114] The values · are calculated depending on the rendering information associated with the enriched audio objects using the appropriate EAO elements and using the equations in section 4.2.2.1.
[0115] For binaural rendering, the matrix is defined by the equations given in section 4.1.2, for which the respective target binaural rendering matrix only contains elements associated with EAOs.
3.4 Calculation of OTN / TTN elements in residual mode [0116] Hereinafter, it will be discussed how the SAOC downmix signal 310, which typically contains one or two audio channels, is mapped to a 334 enriched audio object signal, which typically contains one or more enriched object channels audio and second audio information 322, which typically includes one or two channels of regular audio objects.
[0117] The function of the 1-to-N or 2-to-N unit 330 may, for example, be performed using a matrix vector product such that a vector describing both signal channels 334 of enhanced audio objects and channels of second audio information 322 is obtained by multiplying the vector describing the SAOC downmix signal channels 310 and (optional) one or more residual signals by an MPrediction or MEnergy matrix. Accordingly, the determination of the MPrediction or MEnergy matrix is an important step in obtaining the first audio information 320 and the second audio information 322 from the SAOC 310 downmix.
[0118] In summary, the OTN / TTN upmix process is represented either by the MPrediction matrix for the prediction mode or by the MEnergy matrix for the energy mode.
[0119] The energy based coding / decoding procedure is designed for encoding a downmix signal not maintaining a waveform. Therefore, the upmix matrix
OTN / TTN for the appropriate energy mode is not based on specific waveforms, but only describes the relative energy distribution of the input audio objects, as will be described in more detail below.
3.4.1 Prediction mode [0120] In prediction mode, the MPrediction matrix is defined using the downmix information contained in matrix D '<sup>1</sup> and CPC data from matrix C:
<img file="PL2446435T3_D0006.tif" />
[0121] For several SAOC modes, the extended downmix D matrix and CPC C matrix have the following dimensions and structures:
3.4.1.1 Stereo downmix (TTN) modes:
[0122] For stereo downmix (TTN) modes (for example, in the case of a stereo downmix based on two channels of regular audio objects and NEAO channels of enhanced audio objects), the (extended) downmix D matrix and CPC C matrix can be obtained as follows:
<td></td><td> ' 1 0</td><td> 0 1</td><td>! - 1 ! n<sub>0</sub>AND_______</td><td></td>
<td>D =</td><td></td><td> «0</td><td>! -i ...</td><td> 0</td>
<td></td><td> + +</td><td> *</td><td>! about</td><td></td>
<td></td><td></td><td></td><td> ' 0</td><td>J</td>
<td></td><td>ri</td><td> 0</td><td> 1 0 ...</td><td> 0></td>
<td></td><td> 0</td><td>l</td><td>! about ...</td><td> 0</td>
<td></td><td> _____</td><td></td><td>.L _ _</td><td></td>
<td>c-</td><td><sup>C</sup>0.0</td><td></td><td>! l ...</td><td> 0</td>
<td></td><td></td><td></td><td></td><td><sub>+</sub></td>
<td></td><td></td><td> »</td><td></td><td></td>
<td></td><td></td><td></td><td></td><td> *</td>
<td></td><td></td><td></td><td> |0 ...</td><td>J</td>
[0123] In the case of a stereo downmix, each EAOj contains two CPC cj, 0 and cj, 1 giving a matrix C.
[0124] The residual processor output signals are calculated as
<img file="PL2446435T3_D0007.tif" />
<img file="PL2446435T3_D0008.tif" />
[0125] Thus, two signals yL, yR (which are represented by XOBJ) are obtained, which represent one or two or even more than two regular audio objects (also referred to as non-expanded audio objects). Also, NEAO signals (represented by XEAO) are obtained representing the enriched NEAO audio objects. These signals are obtained on the basis of two SAOC downmix signals l0, r0 and NEAO res0 signals res0 to resNEAO-1, which will be encoded in the SAOC additional information, for example, as part of the object parametric information.
[0126] It should be noted that the yL and yR signals can be equivalent to signal 322 and that the Y0, EAO to YNEAO-1, EAO (which represent XEAO) signals can be equivalent to 320 signals.
[0127] The AEAO matrix is a rendering matrix. Matrix words<sup>EDA</sup> they can for example describe mapping of enhanced audio objects to 334 enhanced audio object (XEAO) signal channels.
[0128] Accordingly, the correct selection of the AEAO matrix allows the optional integration of the rendering unit function 340 in such a way that the multiplication of the vector describing the channels (10, r0) of the SAOC downmix signal 310 and one or more <sub>r</sub> and EAO> <Pttdiciian <sub>r</sub> residual signals (res0, ..., resNEAO-1) through the matrix can directly create a XEAO representation of the first audio information 320.
3.4.1.2 Mono downmix modes (OTN):
[0129] The following will acquire the signals 320 of enhanced audio objects (or alternatively signals 334 of enhanced audio objects) and the signal 322 of regular audio objects for the case in which the SAOC downmix signal 310 contains only the signal channel.
[0130] For mono downmix (OTN) modes (e.g., mono downmix based on one channel of regular audio objects and channels of enriched NEAO audio objects) (extended) downmix D matrix and CPC matrix C can be obtained as follows:
<img file="PL2446435T3_D0009.tif" />
<td></td><td>'and! about</td><td> ...</td>
<td>c =</td><td>! and</td><td>~ & Apos;</td>
<td></td><td> : !<sup>0</sup></td><td></td>
<td></td><td>! θ</td><td> ... 1<sub>;</sub></td>
[0131] In the case of mono downmix, one EAOj is predicted by one factor cj giving matrix C. All elements of the matrix cj are obtained, for example, from SAOC parameters (for example from SAOC 322 data) according to the relationship given below (section 3.4.1.4 ).
[0132] The residual processor output signals are calculated as
<img file="PL2446435T3_D0010.tif" />
<img file="PL2446435T3_D0011.tif" />
[0133] The XOBJ output signal includes, for example, one channel describing regular audio objects (non-enriched audio objects). The XEAO output signal includes, for example, one, two or even more channels describing enhanced audio signals (preferably NEAo channels describing enhanced audio objects). Again, said signals are equivalent to signals 320, 322.
3.4.1.3 calculation of the inverse extended downmix matrix [0134] Matrix D<sup>-1</sup> is the inversion of the downmix D matrix, and C implies CPC.
[0135] Matrix D<sup>-1</sup> is the inversion of the downmix matrix D and can be calculated as
<img file="PL2446435T3_D0012.tif" />
[0136] Elements di, j (e.g., inversion D<sup>-1</sup> 6x6 extended downmix matrix D) are obtained using the following values:
<img file="PL2446435T3_D0013.tif" />
<img file="PL2446435T3_D0014.tif" />
d<sub>l?</sub> = m, +, ~ + ηι<sub>2</sub>η<sub>4</sub> ^ i, 5 ~ + myi: +, rf, <sub>6</sub> = W<sub>4</sub> + <sup>+ m</sup>4<sup>IN</sup>2 + <sup>m</sup>4<sup>n</sup>l "ΎψΒ, and
<img file="PL2446435T3_D0015.tif" />
Aj = «, + Υ +«, wj - m, -τη, ηι, η, - m, m, n<sub>AND</sub>,
A < <sup>= +</sup> Ύ Y + Y + <sup>n</sup>and<sup>m</sup>t ~ - iWjJW ^ Wj -ηΊ<sub>2</sub>η ^ η<sub>Λ</sub>, d<sub>2 and</sub> = / ¾ + n<sub>3</sub>mf + w ^ mj + n<sub>3</sub> Y - -m<sub>3</sub>m<sub>AND</sub>n<sub>4</sub>,
A s = 4 «+ λ« Y + n<sub>4</sub> Y + n<sub>AND</sub>rr ^ - rn ^ rty - m<sub>4</sub>m<sub>2</sub>n<sub>2</sub> - m ^ n ,,. *<sup>4 +</sup> rfjj = -1 - £ «, '" Σ Υ' "W" ~ "" Y * ~<sup>m</sup>2<sup>n</sup>* ~<sup>+</sup> 2 ^ 2 ^ 3 ^ 4) + 2m, m<sub>4</sub>«^»<sub>4</sub> + 2 / «, / η<sub>4</sub>η, η<sub>4</sub> / «» J = J, t /<sub>4</sub> = / 0, // ¾ + 0, / ¾ + // 1 ^ 05 + // ^ 0, / ¾ + // ^ // ¾ ^ + / ^ / Oj / i<sub>4</sub> - / η ^ ζη, ^ η, -ηγη, η, η, -η ^ / η, η, η ^ -ζη / η, η, η ,.
d<sub>3i</sub> = / o, / o, + n, n, + // ¾<sup>2</sup> o, o, + / 0 ^ 0.0, + / 0, / 0.0<sup>2</sup> + η ^ ηι ^ ηΐ - m ^ m ^ n ^ nj - / η, ηι, η, η, -m<sub>3</sub>m, n, n, - / η, πι, η, η ,, d<sub>3b</sub> = / o, / o<sub>4</sub> + o, o<sub>4</sub> + mfan, 4- mfan<sub>t</sub> + / ο, / η<sub>4</sub>ο<sup>2</sup> + m, m, nl - / η ^, η, η, -ιη, / η, η, η, - / η / η ^ η, -m ^ / η, η, η ,,
4 d ,, = -Ι-Σ ”ί“ Σ "/ ~<sup>m2</sup>^ - "fó -" fri - "^<sup>n</sup>4 - "fó + 2m<sub>and</sub>m<sub>and</sub>n, n<sub>3</sub>2m +<sub>l</sub>m<sub>t</sub>n, n, + 2 / n, / n<sub>4</sub>o, o<sub>4</sub>>
/ ··> -ij »i / ed ,, - m, m, + / ij / i, + rnn, n, + / π'ο, ο, + / 05 / 0.0,<sup>2</sup> + m<sub>:</sub>m, n<sup>2</sup>, - ιη, / η, ηη, - / η, / ο, ο, ο, - / η, / η, / ι ^, - / η, / η, η, η ,, d, <sub>b</sub> = m, m<sub>t</sub> + η ^ η, + mj<sup>2</sup>mercury B4 + mfan, + / η, / η, ο,<sup>2</sup> + / N / n4o,<sup>J</sup>- / π, / η, η, / ^ - η, / π, η, η, - / η, / η ^ η, η, -πι ^ / η, η, η ,, _ 4 4
Ϋ = "1" ΣΎΣ "/ ~<sup>m</sup>4<sup>n</sup>y -'Φ'ϊ -<sup>m</sup>4 ^ - »ί" ί ~ "^<sup>η</sup>ί + + 2m, / o<sub>4</sub>Well<sub>4</sub> + 2 / ^ / o<sub>4</sub>/ yi<sub>4</sub>.
d<sub>}and</sub> = / nm, + η, η, + nifan, + rnny, + / ο, / η<sub>4</sub>ο,<sup>2</sup> + / η, / η, η ^ - / o, / n<sub>4</sub>n, o, - η ^ / η, / ^ η, - ntm ^ n, - η ^ / η, η, η ,,
X <sup>e</sup>-> - 2Z *<sup>m</sup>J "ΣΥ -" &! ~> Ή- "Μ-" fó ~ "tó -o ^ o,<sup>2</sup> + 2o ^ / OjO, Oj + 2 / η / 0,0,0, + 2 // ^ / 1 / ¾.
/ * Jl <sup>4 4</sup> ć & n = l + Σ '·; + Σ · ί + / n / t + / nj / ę + // ζο, + m? Nf + / njnf + rr ^ / y + / n? / I ^ + / n ^ nf + // 1 ^ / 1 , + / n? n<sup>2</sup> + m ^ oj + jt / = i + / o,<sup>2</sup>about<sup>2</sup> -2 / 0, / 05/1, / ¾ -2 / ο, / ο, π, ο, -2nmr, n -2 // \ m, n, n, -2 // ^ // 1, / ij / i -2m<sub>and</sub>m, n, n ,.
[0137] Coefficients m<sub>j</sub> in the extended downmix matrix means the downmix values for each EAOj for the right and left downmix channels as
<img file="PL2446435T3_D0016.tif" />
[0138] The downmix di, j elements of the D downmix matrix are obtained using the DMG downmix gain information and (optional) DCLD information of the downmix channel level difference, which is included in the SAOC 332 information, which is represented, for example, by parameter object information 110 or information 212 SAOC bit stream.
[0139] In the case of stereo downmix, the D downmix matrix with dimensions 2xN with elements di, j (i = 0,1; j = 0, ..., N -1) is obtained from the parameters DMG and DCLD as
<img file="PL2446435T3_D0017.tif" />
o.osiwc,, 0 1DCLD.
VI + 10, <hdcld, <sup>L</sup>
<img file="PL2446435T3_D0018.tif" />
[0140] For the downmix mono D downmix matrix with 1xN dimensions with di, j elements (i = 0; j = 0, ..., N -1) is obtained from DMG parameters as
<img file="PL2446435T3_D0019.tif" />
[0141] In this case, the dequantized DMGj and DCLDJ parameters of the downmix are obtained, for example, from parametric side information 110 or from the bit stream
SAOC 212.
[0142] The EAO (j) function calculates the mapping between the input channel indexes of the audio objects and the EAO signals:
<img file="PL2446435T3_D0020.tif" />
<img file="PL2446435T3_D0021.tif" />
3.4.1.4 Calculation of matrix C [0143] Matrix C implies CPCs and is obtained from the sent SAOC parameters (e.g. oLDs, IOCS, DMGs and DCLDS) as
X o - (1<sup>—</sup>*
[0144] In other words, limited CPCs are obtained according to the above equations that can be considered as a constraint algorithm. However, limited CPCs can also be derived from the values ". ··· '¥ using a different limitation approach (limitation algorithm) or can be set equal to the values<sup>C</sup>/.ABOUT· <sup>c</sup>/.1.
[0145] It should be noted that the words cj, 1 matrix (and direct quantities on the basis of which the words cj, 1 matrix are calculated) are typically only required if the downmix signal is a stereo downmix signal.
[0146] CPCs are limited by another limiting function:
<img file="PL2446435T3_D0022.tif" />
with a weighting factor λ determined as
<img file="PL2446435T3_D0023.tif" />
[0147] For one specific EAO channel j = 0 ... NEAO -1, unlimited CPCs are estimated by p P -P p <sup>J</sup> LoCoj<sup>1</sup> ft> * ftoCaj * LeRc
PP -P<sup>1</sup><sup>1</sup> lt><sup>{</sup> fto <sup>1</sup> lophine
Y.<sup>=</sup>
PP - P<sup>1</sup> * La'Re> <sup>L</sup> LetRa [0148] Energy quantities PLo, PRo, PLoRo, PLoCo, ji and PRoCo, j are calculated as <a name="caption1"></a>And «OLD<sub>l</sub> + £ Σ · yo * -o
<img file="PL2446435T3_D0024.tif" />
<img file="PL2446435T3_D0025.tif" />
<img file="PL2446435T3_D0026.tif" />
[0149] The covariance matrix ei, j is defined as follows: The covariance matrix E with the dimension NxN with elements ei, j represents the approximation of the covariance matrix E χ SS * of the original signal and is obtained from OLD and IOC as
<img file="PL2446435T3_D0027.tif" />
[0150] In this case, the dequantized parameters of the OLDi, IOCij objects are obtained, for example, from parametric additional information 110 or from the SAOC 212 bit stream.
[0151] In addition, eL, R can for example be obtained as
<img file="PL2446435T3_D0028.tif" />
[0152] The OLDL, OLDR and IOCL, R parameters correspond to regular (audio) objects and can be obtained using downmix information:
<img file="PL2446435T3_D0029.tif" />
<img file="PL2446435T3_D0030.tif" />
[0153] As can be seen, two common values for the difference in the level of OLDL and OLDR objects are calculated for regular audio objects in the case of a stereo downmix signal (which preferably implies a two-channel signal of regular audio objects). In contrast, only one common OLDL of the object level difference is calculated for regular audio objects in the case of a single-channel (mono) downmix signal (which preferably implies a single-channel signal of regular audio objects).
[0154] As can be seen, the first (in the case of a two-channel downmix signal) or the only (in the case of a single-channel downmix signal) the common OLDL value of the object level difference is obtained by adding the contributions of regular audio objects having the index (or indexes) of the audio object and, to the left channel (or only channel) SAOC downmix signal 310.
[0155] A second common OLDR value of the object level difference (which is used in the case of a two-channel downmix signal) is obtained by summing the contributions of regular audio objects having the index (or indexes) of the audio object and, to the right channel of the SAOC downmix signal 310.
[0156] OLDL input of regular audio objects (having audio object indexes i = 0 to i = N-NEAO -1) to the left channel signal (or only channel signal) of the SAOC downmix, signal 710 is calculated, for example, taking into account d0 gain, and downmix describing the downmix gain used for a regular audio object having an audio object index and when the left channel signal 310 of the SAOC downmix signal is obtained as well as the level of the regular object of the audio object having the index of the audio object i, which is represented by the OLDi value.
[0157] Similarly, a common OLDR value of the object level difference is obtained using downmix d1 coefficients, and describing the downmix gain that is used for a regular audio object having an audio object index and when the right downmix signal channel SAOC 310 and level information is created OLDi associated with a regular audio object having an audio object index i.
[0158] As can be seen, the equations for calculating PLo, PRo, PLoRo, PLoCo, and J and PRoCo, j do not distinguish between individual regular audio objects, but only use common OLDL, OLDR values of the object level difference, thus treating regular audio objects (having indexes audio object i) as a single audio object.
[0159] Also, the IOCL, R value of the inter-object correlation, which is associated with regular audio objects, is set to 0 unless there are two regular audio objects.
[0160] The matrix ei, j (and eL, R) covariance is defined as follows:
[0161] The N-N covariance matrix E with elements ei, j represents the approximation of the E χ SS * matrix of the original signal covariance and is obtained from the OLD and IOC parameters as
- jOLflOLDJOC ,,.
[0162] For example <sup>e</sup>t ,, R <sup>=</sup> yjOLDfOLDfiIOC<sub>ts</sub>, where OLDL and OLDR and IOCL, R are calculated as described above.
[0163] In this case, the dequantized object parameters are obtained as
<img file="PL2446435T3_D0031.tif" />
where DOLD and DIOC are matrices containing object level difference parameters and inter-object correlation parameters.
3.4.2 Energy mode [0164] A different concept will be described below that can be used to separate the signals 320 of extended audio objects and signals 322 of regular audio objects (non-extended audio objects), and which can be used in conjunction with non-waveform audio coding SAOC downmix 310 channels.
[0165] In other words, the energy-based coding / decoding procedure is designed for a non-observing waveform coding of a downmix signal. Thus, the OTN / TTN upmix matrix for the respective energy mode does not depend on specific waveforms, but only describes the relative energy distribution in the input audio objects.
[0166] Also, the concept discussed here, which is designed as an "energy mode" concept, can be used without transmitting the residual signal information. Again, regular audio objects (non-enriched audio objects) are treated as a single single-channel or two-channel audio object having one or two common OLDL, OLDR object level difference values.
[0167] For the energy mode, the MEnergy matrix is defined using downmix and OLD information as will be described below.
3.4.2.1 Energy mode for stereo downmix (TTN) modes [0168] For stereo (for example, a stereo downmix based on two channels,, Μ® * · ® 'regular objects and N channels<sub>EA</sub>enhanced audio objects) matrices - and are obtained from the corresponding OLDs according to
<img file="PL2446435T3_D0032.tif" />
<img file="PL2446435T3_D0033.tif" />
[0169] The residual processor output signals are calculated as
<img file="PL2446435T3_D0034.tif" />
<img file="PL2446435T3_D0035.tif" />
[0170] The yL, yR signals, which are represented by the XOBJ signal, describe regular audio objects (and can be equivalent to 322 signals), and the signals Y0, EAO to YNEAO-1, EAO, which are described by the XEAO signal, describe the enriched audio objects (and can be equivalent to 334 signals or 320 signals).
[0171] If a mono upmix signal is desired for a stereo downmix signal case, 2-to-1 processing may be performed, for example, by a pre-processor 270 based on a two-channel XOBJ signal.
3.4.2.2 Energy mode for mono downmix (OTN) modes [0172] For mono (for example mono downmix based on one channel of regular audio objects and NEAO channels of enhanced audio objects) matrices
Did he burn ΰπΓΓχ? * jj Enerjy and are obtained from the corresponding OLDs in accordance with
<img file="PL2446435T3_D0036.tif" />
OLD<sub>l</sub>
OID<sub>t</sub> + £ m / OLA 'no)
<img file="PL2446435T3_D0037.tif" />
[0173] The residual processor output signals are calculated as
<img file="PL2446435T3_D0038.tif" />
[0174] A single channel 322 of regular audio objects (represented by XOBJ) and NEAO channels 320 of enriched audio objects (represented by XEAO) can be obtained by using a matrix and for representing a single channel signal 310 of the SAoC downmix (represented here by).
[0175] If a two-channel (stereo) upmix signal is desired for a single-channel (mono) downmix signal, 1-to-2 processing may be performed, e.g., by a pre-processor 270 based on a single-channel XoBJ signal.
4. Architecture and operation of the SAoC downmix pre-processor [0176] In the following, the operation of the SAoC downmix pre-processor 270 will be described for both certain decoding modes and for certain transcoding modes.
4.1 Operation in decoding modes
4.1.1 Introduction [0177] The following describes how to obtain an output signal using the SAoC parameters and panorama information (or rendering information) associated with each audio object. The SAoC 495 decoder is shown in Fig. 4g and consists of a SAoC processor 495 and a downmix processor 497.
[0178] It should be noted that the SAoC decoder 494 can be used to process regular audio objects and therefore can receive, as downmix signal 497a, a second audio object signal 264 or a regular audio object signal 322 or a second audio information 134. processor 497 accordingly downmix may provide, as output signals 497b, the processed version 272 of the second signal 264 audio objects or the processed version 142 of the second audio information 134. Thus, the downmix processor 497 may act as a SAoC downmix pre-processor 270, or as an audio signal processor 140.
[0179] The SAoC parameters processor 496 may act as the SAoC parameters processor 252 and as a result provide downmix information 496a.
4.1.2 Downmix processor [0180] In the following, the downmix processor will be described in more detail, which is part of the audio signal processor 140 and which is referred to as "SAoC downmix pre-processor 270" in the embodiment of Fig. 2 and which is designated 497 in the SAoC decoder 495.
[0181] For the SAoC system decoder mode, the downmix processor output signal 142, 272, 497b (represented in the QMF hybrid domain) is provided to the appropriate synthesis filter bank (not shown in Figs. 1 and 2) as described in the ISo / IEC 23003 standard -1: 2007 generating the final PCM output signal. Regardless, the output signal 142, 272, 497b of the downmix processor is typically coupled to one or more audio signals 132, 262 representing enhanced audio objects. This connection can be made before the respective synthesis filter bank (so that the combined signal combining the downmix processor output signal and one or more signals representing enriched audio objects is fed into the synthesis filter bank).
Alternatively, the downmix processor output may be combined with one or more audio signals representing enriched audio objects only after processing in the synthesis filter bank. Accordingly, the upmix signal representation 120, 220 may be either a GMF domain representation or a PCM domain representation (or any other suitable representation). Downmix processing includes, for example, mono processing, stereo processing and, if desired, subsequent binaural processing.
[0182] Output signal X of the downmix pre-processor 270, 497 (designated
142 also 142, 272, 497b) is calculated from the mono downmix X signal (also denoted 134, 262, 497a) and the de-correlated Xd mono downmix signal as
<img file="PL2446435T3_D0039.tif" />
[0183] The de-correlated mono downmix Xd signal is calculated as
X<sub>d</sub> = decorrFunć.
[0184] The de-correlated Xd signals are created in the decorrelator described in ISO / IEC 23003-1: 2007, subclause 6.6.2. According to this scheme, the configuration bsDecorrConfig == 0 with the decorrelator index X = 8 in accordance with Table A.26 to Table A.29 in ISO / IEC 23003-1: 2007 should be used. Hence decorrFunc () means the process of decorrelation:
V -ί<sup>Χ |</sup>* Ί.ΓdecorrFune ((] 0) Ρ, ΧΓ j ^ Ac (wtFuw ((0 l) P, X) [0185] For the binaural output signal, upmix G and P2 parameters obtained from SAOC data, rendering information and HRTF parameters are applied
Λ in the downmix X signal (and Xd) for the binaural output signal X, see Fig. 2, reference numeral 270, where the basic structure of the downmix processor is shown.
[0186] Target matrix A<sup>l, m</sup> 2xN binaural rendering consists
Ijj. and<sup>1</sup>'™. Each element is obtained from HRTF parameters and from a rendering matrix with elements' λ ·· for example by the SAOC parameter processor. Target matrix A<sup>l, m</sup> binaural rendering represents the relationship between all audio input objects y and the desired binaural output signal.
<img file="PL2446435T3_D0040.tif" />
jr τ Ih f Γ JW 2 [0187] HRTF parameters are given by 'i for each processing band m. Spatial positions for which HRTF parameters are available are marked with index i. These parameters are described in ISO / IEC 23003-1: 2007.
4.1.2.1 General Outline [0188] Below, an overview of the downmix processing will be provided with reference to Figs. 4a and 4b, which show a block diagram of the downmix processing that can be implemented by the audio signal processor 140 or by the combination of processor 252 SAOC parameters and pre-processor 270 SAOC downmix, or by a combination of a SAOC downmix 496 processor and a 497 downmix processor.
[0189] Referring now to Fig. 4a, the downmix processing receives a rendering M matrix, object level difference OLD information, cross object correlation IOC information, downmix gain DMG information and (optional) DCLD level difference information of downmix channel objects. The downmix processing 400 of Fig. 4a obtains a rendering matrix A based on the rendering matrix M, for example using a parameter adjustment module and M-to-A mapping. Also, the words of the matrix E covariance are obtained depending on the OLD information of the object level difference and the IOC information of inter-object correlation, for example as discussed above. Similarly, the downmix D matrix words are obtained depending on the downmix gain DMG information and the DCLD information of the downmix channel level differences.
[0190] The f words of the desired F covariance matrix are obtained depending on the rendering matrix A and the covariance matrix E. Also, the scalar value v is obtained depending on the matrix E of covariance and the matrix D of the downmix (or depending on their words).
[0191] The PL, PR gain values for the two channels are obtained depending on the words of the desired matrix F covariance and the scalar value v. Also the value of φ<sub>ε </sub>cross-object phase difference is obtained depending on the words f of the desired matrix F covariance. The angle of rotation α is also obtained depending on the words f of the desired matrix F covariance, taking into account, for example, the constant c. In addition, a second angle of rotation β is obtained, for example, depending on the gains PL, PR of the channels and the first angle of rotation α. The words of the G matrix are obtained, for example, depending on the two-channel values of PL, PR of the gains as well as depending on the inter-object phase difference φC and optionally, on the rotation angles α, β. Similarly, the words in the P2 matrix are determined depending on some or all of the P values listed<sub>L</sub>, P<sub>R</sub>, φε, α, β.
[0192] Hereinafter, it will be discussed how the G and / or P2 matrix (or words thereof) that can be used by a downmix processor as discussed above can be obtained for different processing modes.
4.1.2.2 Mono to binaural "x-1-b" processing mode [0193] The following will discuss processing modes in which regular audio objects are represented by a single-channel downmix 134, 264, 322, 497a signal in which binaural rendering is desirable.
[0194] Upmix parameters G<sup>l, m</sup> and P2<sup>l, m</sup> are calculated as
<img file="PL2446435T3_D0041.tif" />
<img file="PL2446435T3_D0042.tif" />
[0195] '· and' · reinforcements
/.mi for left and right output channels are
<img file="PL2446435T3_D0043.tif" />
fK * 1 [0196] Desired matrix F<sup>l, m</sup> 2x2 covariance with elements is given as
<img file="PL2446435T3_D0044.tif" />
[0197] Scalar v<sup>l, m</sup> is calculated as
<img file="PL2446435T3_D0045.tif" />
[0198] The inter-channel phase difference is given as
i. in another case [0199] Inter-channel coherence is calculated as
<img file="PL2446435T3_D0046.tif" />
[0200] Angles of rotation a<sup>l, m</sup> e<sup>l, m</sup> are given as yarccos cos (arg (/ Jf) |), 0 <at £ 11, p '<sup>1</sup>'<0.6, - arccos 2 (rf ·).
otherwise
<img file="PL2446435T3_D0047.tif" />
4.1.2.3 "x-1-2" mono-to-stereo processing mode [0201] The following will discuss the processing mode in which regular audio objects are represented by a single-channel signal 134, 264, 222 and in which stereo rendering is desired.
[0202] For the stereo output signal, the "x-1-b" processing mode can be used without using HRTF information. This can be done by obtaining f7 '<sup>m</sup> all elements - rendering matrix A, obtaining:
<img file="PL2446435T3_D0048.tif" />
4.1.2.4 "x-1-1" mono-to-mono processing mode [0203] The following describes the processing mode in which regular audio objects are represented by signal channel 134, 264, 322, 497a and in which regular two-channel rendering is desirable audio objects.
[0204] For the mono output signal, the "x-1-2" processing mode can be used with the following words:
4.1.2.5 Stereo-to-binaural "x-2-b" processing mode [0205] Below, the processing mode in which regular audio objects are represented by a two-channel signal 134, 264, 322, 497a and in which binaural rendering is desirable will be described regular audio objects.
<sub>lm</sub> Ρ<sup>ί</sup>·"
[0206] Parameters G<sup>l, m</sup> and upmix are calculated as
<img file="PL2446435T3_D0049.tif" />
<img file="PL2446435T3_D0050.tif" />
D<sup>1</sup>·"·<sup>1</sup> pi, · », * ηή * ι pl, m [0207] Suitable gains for the left and right output channels are
<img file="PL2446435T3_D0051.tif" />
fl.mj [0208] Desired matrix F<sup>l, m, x</sup> 2x2 covariance with the elements · *, * is given as
<img file="PL2446435T3_D0052.tif" />
[0209] Matrix C<sup>l, m</sup> 2x2 covariance with elements of "dry" binaural signal is estimated as
where
<img file="PL2446435T3_D0053.tif" />
[0210] Suitable scalars v<sup>l, m, x</sup> iv<sup>l, m</sup> are calculated as
<img file="PL2446435T3_D0054.tif" />
[0211] Downmix matrix D<sup>l, x</sup> 1xN size with elements obtained as
<img file="PL2446435T3_D0055.tif" />
may be
<img file="PL2446435T3_D0056.tif" />
ji [0212] D 2xmi downmix matrix with elements' can be obtained as [0213] Matrix E<sup>l, m, x</sup> with elements is obtained from the following dependence
<img file="PL2446435T3_D0057.tif" />
[0214] Inter-channel phase differences' are given as
<img file="PL2446435T3_D0058.tif" />
I.-Λ ™ i '<sup>L</sup>'are calculated as i *
Pt = mm
<td>f Ιλή</td><td>and |</td><td>and '.-and 1</td>
<td>IAj and r '7' ί *</td><td>"PY = mtn</td><td></td>
<td></td><td></td><td>JmaA «Γ4 '. £<sup>2</sup>) j</td>
[0216] Rotation angles a<sup>l, m</sup> e<sup>l, m</sup> are given as a<sup>fm</sup> Y (arccos (A0 -arccos (p £ ") j, β<sup>1</sup>"= Arctan
4.1.2.6 Stereo-to-stereo "x-2-2" processing mode [0217] The following describes the processing mode in which regular audio objects are described by a two-channel (stereo) signal 134, 264, 322, 497a and in which it is desirable two-channel rendering (stereo).
[0218] For the stereo output signal, the stereo pre-processing is directly applied, which will be described below in section 4.2.2.3.
4.1.2.7 Stereo-to-mono "x-2-1" processing mode [0219] In the following, a processing mode will be described in which regular audio objects are represented by a two-channel (stereo) signal 134, 264, 322, 497a and in which Single-channel (mono) rendering is desirable.
[0220] For the mono output signal, stereo pre-processing with a single active matrix expression is rendered, as described below in section 4.2.2.3.
4.1.2.8 Conclusions [0221] Referring again to Figs. 4a and 4b, processing is shown here that can be used for single-channel or two-channel signal 134, 264, 322, 497a representing regular audio objects after separation between extended audio objects and regular objects audio. Figures 4a and 4b show the processing, wherein the processing in Figs. 4a and 4b differ in that the optional parameter adjustment is introduced at various stages of processing.
4.2 Operation in transcoding mode
4.2.1 Introduction [0222] The following explains how to combine SAOC parameters and panorama information (or rendering information) associated with each audio object (or preferably any regular audio object) in a standard compliant MPEG Surround bit stream (MPS bit stream) .
[0223] The SAOC 490 transcoder is shown in Fig. 4f and consists of a SAOC parameter processor 491 and a downmix processor 492 in use for a stereo downmix.
[0224] The SAOC 490 transcoder may for example take over the role of an audio signal processor 140. Alternatively, the SAOC 490 transcoder can take over the role of the SAOC downmix pre-processor 270 in conjunction with the SA2 252 processor.
[0225] For example, the SAOC parameter processor 491 may receive the SAOC bit stream 491a, which is equivalent to the object parametric information 110 or the SAOC bit stream 212. Also, the SAOC parameter processor 491 may receive the rendering matrix information 491b, which may be included in the object parametric information 110 , or it can be equivalent to 214 rendering matrix information. The SAOC parameter processor 491 may also provide downmix processing information 491c to the downmix processor 492, which may be equivalent to information 240. In addition, the SAOC parameter processor 491 may include an MPEG Surround 491d bit stream (or MPEG Surround parameter bit stream) that includes parametric Surround information, which is compatible with the MPEG Surround standard. The MPEG Surround 491d bit stream may for example be part of the processed version 142 of the second audio information, or may for example be part of the MPS 222 bit stream or replace it.
[0226] Downmix processor 492 is configured to receive downmix signal 492a, which is preferably a single-channel downmix signal or a two-channel downmix signal and which is preferably equivalent to second audio information 134, or second signal 264, 322 of audio objects. The downmix processor 492 may also provide an MPEG Surround downmix signal 492b that is equivalent to the processed version 272 (or part thereof) of the second signal 264 of audio objects.
[0227] However, there are various ways to combine the MPEG Surround downmix signal 492b with the signal 132, 262 of enhanced audio objects. Joining can be done in the MPEG Surround domain.
[0228] Alternatively, however, the MPEG Surround representation, comprising the bit stream 491d of the MPEG Surround parameters and the MPEG Surround downmix signal 492b of regular audio objects may be converted back to a multi-domain time domain representation or a multi-channel frequency domain representation (individually representing different audio channels) by MPEG Surround decoder and can then be combined with signals of enhanced audio objects.
[0229] It should be noted that the transcoding modes include both one or more mono downmix processing modes and one or more stereo downmix processing modes. However, only stereo processing mode will be discussed below, because the processing of regular audio objects signals is more complex in stereo downmix processing mode.
4.2.2 Downmix processing in stereo downmix processing mode ("x-2-5")
4.2.2.1 Introduction [0230] The following section will describe the SAOC transcoding mode for the stereo downmix case.
[0231] Object parameters (difference of OLD object levels, inter-object IOC correlation, DMG downmix gain and DCMD downmix channel difference) from the SAOC bit stream are transcoded into spatial parameters (preferably channel, CLD channel level difference, inter-object correlation ICC, channel prediction coefficients) CPC) for the MPEG Surround bit stream according to the rendering information. Downmix is modified according to object parameters and the rendering matrix.
[0232] Referring now to Figs. 4c, 4d and 4e, an overview of the processing, and in particular of the downmix modification, will be given.
[0233] Fig. 4c shows a processing block diagram that is implemented for modifying a downmix signal, e.g., a downmix signal 134, 264, 322,492a describing one or preferably more regular audio objects. As can be seen in Figs. 4c, 4d and 4e, the processing receives a rendering Mren matrix, downmix gain DMG information, DCLD downmix channel level difference information, object level difference OLD information and inter-object correlation IOC information. The rendering matrix can optionally be modified by adjusting the parameter as shown in Fig. 4c. The downmix D matrix words are obtained depending on the DMG downmix gain information and DCLD information of the downmix channel level differences. The words of the E coherence matrix are obtained depending on the OLD information of the object level difference and the IOC information of inter-object correlation. Additionally, the matrix J can be obtained depending on the downmix matrix D and the coherence matrix E, or depending on their words. Then the C3 matrix can be obtained depending on the rendering Mren matrix, downmix matrix D, coherence matrix E and matrix J. The matrix G can be obtained depending on the DTTT matrix, which can be a matrix having predefined words, and also depending on the matrix C3. Matrix G may optionally be modified to obtain a modified Gmod matrix. Matrix G or its modified version of Gmod can be used to obtain the processed version 142, 272, 492b of the second audio information 134, 264 from the second audio information 134, 264, 492a (with the second audio information 134,
Λ
264 is marked with X and its processed version 142, 272 is marked with X.
[0234] The energy rendering of the object which is performed to obtain MPEG Surround parameters will be discussed below. Also, the stereo pre-processing that is implemented to obtain the processed version 142, 272, 492b of the second audio information 134, 264, 492a representing regular audio objects will be described.
4.2.2.2 Rendering energy of objects [0235] The transcoder sets parameters for the MPS decoder according to the target rendering described by the rendering matrix Mren. The six-channel target covariance is designated F and given by
<img file="PL2446435T3_D0059.tif" />
[0236] The transcoding process may be conceptually divided into two parts. In one part, three-channel rendering is performed to the left, right and center channel. At this stage, the parameters for downmix modification are obtained as well as the prediction parameters for the TTT block for the MPS decoder. In the rest, CLD and ICC parameters are determined for rendering between front and surround channels (OTT parameters, left - front left, surround left, right front - surround right).
4.2.2.2.1 Rendering of the left, right and center channel [0237] At this stage, spatial parameters are determined that control the rendering to the left and right channel consisting of front and surround signals. These parameters describe the TTT block prediction matrix for MPS CTTT decoding (CPC parameters for the MPS decoder) and the G matrix of the downmix converter.
[0238] CTTT is a prediction matrix for obtaining the target rendering from
Λ modified downmix X = GX:
<img file="PL2446435T3_D0060.tif" />
[0239] A3 is a 3xN reduced rendering matrix describing rendering for left, right and center channels respectively. It is obtained as A3 = D36Mren with a D36 partial downmix matrix 6 to 3 determined by
<img file="PL2446435T3_D0061.tif" />
[0240] The weights wp, p = 1,2,3 of the partial downmix are adjusted in such a way that the energy wp (y2p-1 + y2p) is equal to the sum of energy || y2p-1 ||<sup>2</sup>|| + || y2p<sup>2</sup> up to the limiting factor
<img file="PL2446435T3_D0062.tif" />
where fi, j are elements F.
[0241] For the estimation of the desired CTTT prediction matrix and the G downmix pre-processing matrix, we define a 3x2 size prediction matrix C3, which leads to target rendering
<img file="PL2446435T3_D0063.tif" />
[0242] Such a matrix is obtained by taking into account the normal equations c (ded) ^ a<sub>3</sub>ed \ [0243] The solution of normal equations gives the best possible fit of the waveform to the target output signal for a given object covariance model. G and CTTT are now obtained by solving the system of equations
<img file="PL2446435T3_D0064.tif" />
[0244] To avoid numerical problems when calculating J = (DED *)<sup>-1</sup>, J is modified. First, the proprietary properties λ1,2 J are calculated, solving det (J - A1,2I) = 0.
[0245] Eigenvalues are sorted in descending order (λ<sub>1</sub>> λ<sub>2</sub>) and the eigenvector corresponding to the greater eigenvalue is calculated according to the equation above. Its position in the positive x plane is ensured (the first element must be positive). The second eigenvector is obtained from the first by rotation through an angle of -90 degrees:
<img file="PL2446435T3_D0065.tif" />
[0246] The weighing matrix is calculated from the downmix D matrix and the C3 prediction matrix, W = (D diag (C3)).
[0247] Since CTTT is a function of the MPS prediction parameters c1 and c2 (described in ISO / IEC 23003-1: 2007), CTTT G = C3 is rewritten as follows to find the point or fixed points of the function
Γ
<img file="PL2446435T3_D0066.tif" />
with Γ = (DTTT C3) W (DTTT C3) * ib = GWC3V, where
<img file="PL2446435T3_D0067.tif" />
iv = (1 1 -1).
[0248] If Γ does not provide a unique solution (det (r) <10<sup>-3</sup>), the point closest to the point giving the TTT transition is selected. In the first stage, the row is selected and for Γ, γ = [γΐ1γΐ2] where the elements contain the most energy, so η, 1<sup>2</sup> + γ, .2<sup>2</sup> -+1<sup>2</sup> + Yj, 2<sup>2</sup>, j = 1.2.
[0249] Next, a solution is established such that
<img file="PL2446435T3_D0068.tif" />
[0250] If the solution obtained for i<sub>2</sub> lies outside the permitted range for the prediction coefficients, which is defined as -2 <ćj <3 (as defined in ISO / IEC 23003-1: 2007), Ćj should be calculated as below.
[0251] First, define a set of points, xp as:
<img file="PL2446435T3_D0069.tif" />
and the distance function dis / Func (^) = - 2bx<sub>p</sub> [0252] Next, prediction parameters are defined according to:
<img file="PL2446435T3_D0070.tif" />
[0253] Prediction parameters are limited according to:
Ci - 0 - + Ay Cj = 0 -, where λ, γ1 and γ2 are defined as
<img file="PL2446435T3_D0071.tif" />
"Λ.3 + Λ, 3 <sup>+</sup>Λϊ
<img file="PL2446435T3_D0072.tif" />
<img file="PL2446435T3_D0073.tif" />
[0254] For the MPS decoder, CPC and corresponding ICCTTT are provided as follows
<img file="PL2446435T3_D0074.tif" />
4.2.2.2.2 Rendering between front and surround channels [0255] Parameters determining the rendering between front and surround channels can be estimated directly from the target matrix F covariance
<img file="PL2446435T3_D0075.tif" />
ICC '* =
<img file="PL2446435T3_D0076.tif" />
with (a, b) = (1,2) and (3,4).
[0256] MPS parameters are provided in the form
CLDf = 0 ^, 1, m) and 1CC'P = D<sub>) cc</sub>(h, lm), for each OTT block h.
4.2.2.3 Stereo processing [0257] The stereo processing of the signal 134 to 64, 322 regular audio objects will be discussed below. Stereo processing is used to obtain the process for general representation 142, 272 based on a two-channel representation of regular audio objects.
[0258] Stereo downmix X, which is represented by signals 134, 264, 492a
Λ regular audio objects are converted into a modified downmix X signal, which is represented by processed signals 142, 272 of regular audio objects:
<img file="PL2446435T3_D0077.tif" />
where
<img file="PL2446435T3_D0078.tif" />
[0259] The final output signal from the SAOC transcoder, X is formed by mixing X with the decorrelated signal component according to:
<img file="PL2446435T3_D0079.tif" />
where the de-correlated Xd signal is calculated as described above, and the GMod and P2 mix matrices as below.
[0260] First, define the rendering upmix error matrix as
<img file="PL2446435T3_D0080.tif" />
where
Ajiit - - GD, and also define the covariance matrix of the predicted R signal as
<img file="PL2446435T3_D0081.tif" />
[0261] The gvec gain vector can then be calculated as:
<img file="PL2446435T3_D0082.tif" />
and the GMod mix matrix is given as:
ftlag (g ^) G, η <sub>2</sub> > Oh,
G, in another case [0262] Similarly, the P2 mix matrix is given as:
<img file="PL2446435T3_D0083.tif" />
v<sub>R</sub>^ (W<sub>e</sub>) otherwise [0263] To obtain vR and Wd, the characteristic equation R must be solved:
det (R-2,<sub>2</sub>I) = 0 "giving the eigenvałues, Ą and Ą [0264] The respective eigenvectors R, vR1 and vR2 can be calculated by solving the system of equations:
(R Λ, <sub>J</sub>l) Y<sub>Fti R2</sub> - ABOUT.
[0265] Eigenvalues are sorted in descending order (λ1> λ2), and the eigenvector corresponding to the greater eigenvalue is calculated according to the equation above. It shall be positioned in the positive x plane (the first element must be positive). The second eigenvector is obtained from the first by rotation through an angle of -90 degrees:
<img file="PL2446435T3_D0084.tif" />
[0266] By entering P1 = (1 1) G, Rd can be calculated according to:
<img file="PL2446435T3_D0085.tif" />
which gives
<img file="PL2446435T3_D0086.tif" />
and finally the mix matrix
<img file="PL2446435T3_D0087.tif" />
4.2.2.4 Dual mode [0267] The SAOC transcoder may enable the calculation of the mix matrix P1, 2 and the prediction matrix C3 according to an alternative method for the higher frequency range. This alternative method is particularly useful for downmix signals where the higher frequency range is encoded by a non-waveform coding algorithm, e.g. SBR in high-performance AAC.
[0268] For the upper parameter bands, defined by bsTttBandsLow <pb <numBands, P1, P2 and C3 should be calculated according to the alternative method described below:
<img file="PL2446435T3_D0088.tif" />
o) [0269] Define the energy downmix and target energy vectors, respectively:
<img file="PL2446435T3_D0089.tif" />
and an auxiliary matrix
<img file="PL2446435T3_D0090.tif" />
[0270] Then calculate the gain vector
<img file="PL2446435T3_D0091.tif" />
which ultimately gives a new prediction matrix
<img file="PL2446435T3_D0092.tif" />
g / l, l ϊι ^ ι.ι ίΑ «£ / li ίΐ ^ ϊ, Ι <sup>1 </sup>S & 2 J
5. Combined EKS SAOC decoding / transcoding mode, encoder according to Fig. 10 and systems according to Figs. 5a and 5b [0271] Below is a brief description of the method of combined EKS SAOC processing. A preferred "EKS SAOC processing" method is proposed, in which EKS processing is integrated into the regular SAOC decoding / transcoding chain in a cascade system.
5.1 Audio encoder according to Fig. 5 [0272] In a first step, the objects dedicated to EKS processing (enhanced Karaoke / solo processing) are identified as foreground objects (FGO) and their number NFGO (also denoted NEAO) is determined by the variable " bsNumGroupsFGO "bit stream. Said bit stream variable may for example be included in the SAOC bit stream described above.
[0273] For generating the bit stream (in the audio signal encoder), the parameters of all Nobj input objects are arranged in such a way that FGO foreground objects contain the last NFGO parameters (or alternatively NEAO) in each case, for example, OLDi for [Nobj - Nfgo and Nobj - 1].
[0274] From other objects, which are, for example, BGO background objects or unenriched audio objects, a "regular SAOC style" downmix signal is generated, which also acts as a BGO background object. Then, the background object and foreground objects are mixed down in the "EKS processing style" and residual information is obtained from each foreground object. In this way, there is no need to enter any additional processing steps. In this way, the bit stream syntax was changed.
[0275] In other words, on the encoder side, un-enriched audio objects are distinguished from enriched audio objects. A single-channel or two-channel downmix signal of regular audio objects is provided, which represents regular audio objects (un-enriched audio objects), with one, two or even more regular audio objects (un-enriched audio objects). The single-channel or two-channel downmix signal of regular audio objects is then combined with one or more signals of the enriched audio objects (which may be, for example, a single-channel downmix signal or a two-channel downmix signal) to obtain a common downmix signal (which may be, for example, a single-channel downmix signal or two-channel downmix signal) combining audio signals of enriched audio objects and a downmix signal of regular audio objects.
[0276] In the following, the basic structure of such a cascade encoder will be briefly described with reference to Fig. 10, which shows a block diagram of the SAOC 1000 encoder according to an embodiment of the invention. The SAOC 1000 encoder includes a SAOC downmix module 1010, which is typically a SAOC downmix module that does not provide residual information. The SAOC downmix module 1010 is configured to receive multiple signals of 1012 NBGO audio objects from regular (non-enriched) audio objects. Also, SAOC downmix module 1010 is configured to provide the regular audio object downmix signal 1014 based on regular audio objects 1012 in such a way that the regular audio object signal 1014 combines the 1012 regular audio object signals according to the downmix parameters. The SAOC downmix module 1010 also provides SAOC 1016 regular audio object information that describes the signals and downmix of regular audio objects. For example, SAOC information 1016 of regular audio objects may include DMG downmix gain information and DCLD information of the downmix channel level difference describing the downmix made by the SAOC downmix module 1010. In addition, SAOC 1016 regular audio object information may include object level difference information and inter-object correlation information describing the relationship between regular audio objects described by the regular audio object signal 1012.
[0277] Encoder 1000 also includes a second SAOC downmix module 1020 that is typically configured to provide residual information. Second SAOC downmix module 1020 is preferably configured to receive one or more signals 1022 of enhanced audio objects as well as to receive downmix signal 1014 of regular audio objects.
[0278] The second SAOC downmix module 1020 is also configured to provide a common SAOC downmix signal 1024 based on the signals 1022 of enhanced audio objects and the signal 1014 of regular audio objects. When providing a common SAOC downmix signal, the second SAOC downmix module 1020 typically treats the regular audio object downmix signal 1014 as a single channel or two channel signal.
[0279] Second module 1020 is also configured to provide SAOC information of enriched audio objects, which describes, for example, DCLD values of the downmix channel level differences associated with enriched audio objects and OLD values of the object level difference associated with enriched audio objects and IOC values of inter-object correlation associated with enriched audio objects. In addition, the second SAOC downmix module 1020 is preferably configured to provide residual information associated with each of the enriched audio objects in such a way that the residual information associated with the enriched audio objects describes the difference between the original signal of the individual enriched audio objects and the expected signal of the individual enriched audio objects that can be obtained from a downmix signal using DMG downmix information, DCLD and OLD object information, IOC.
[0280] The audio encoder 1000 is well adapted to interact with the audio decoder described herein.
5.2 Audio signal decoder according to Fig. 5a [0281] The basic structure of the EKS SAOC 500 hybrid decoder will be described below, the block diagram of which is shown in Fig. 5a.
[0282] The audio decoder 500 of Fig. 5a is configured to receive the downmix signal 510, SAOC bit stream information 512 and rendering matrix information 514. The audio decoder 500 includes Karaoke / solo processing and rendering of 520 foreground objects that is configured to provide a first signal 562 of audio objects that describes the rendered foreground objects and a second signal 564 of audio objects that describes the background objects. Foreground objects may, for example, be called "enhanced audio objects" and background objects may, for example, be called "regular audio objects" or "non-enriched audio objects". The audio decoder 500 also includes regular SAOC 570 decoding, which is configured to receive a second signal 562 of audio objects and to provide a processed version 572 of the second signal 564 audio objects based thereon. The audio decoder 500 also includes a combining module 580 that is configured to combine the first signal 562 of audio objects and the processed version 572 of the second signal 564 audio objects to obtain output signal 520.
[0283] The functionality of the audio decoder 500 will be discussed below with some additional details. On the SAOC decoding / transcoding side, the upmix process results in a cascade system that first includes enhanced Karaoke / solo processing (EKS processing) to break down the downmix signal into a background object (BGO) and foreground objects (FGO). The required object level differences (OLD) and inter-object correlations (IOC) for the background object are derived from the object and downmix information (both of which are object-oriented parametric information and which both are typically contained in the SAOC bit stream):
OLD<sub>l</sub>= Υ d ^ .OLD ι-Ο
W-l-jko
OLD<sub>r</sub>^ £
1-0
<img file="PL2446435T3_D0093.tif" />
1OC ", NN<sub>m</sub>-2,
Oh, otherwise.
[0284] In addition, this step (which is typically performed by EKS processing and rendering of 520 foreground objects) includes mapping the foreground objects to final output channels (in such a way that, for example, the first signal 562 of the audio objects is a multi-channel signal in which each of foreground objects are mapped to one or more channels). The foreground object (which typically contains many so-called "regular audio objects") is rendered to the corresponding output channels by the regular SAOC decoding process (or alternatively in some cases by the SAOC transcoding process). For example, this process can be accomplished by regularly decoding SAOC 570. The final mixing step (e.g., combining module 580) provides the output with the desired combination of signals of the rendered foreground and background objects.
[0285] This hybrid EKS SAOC system represents a combination of all the beneficial properties of a regular SAOC system and its EKS mode. This approach allows obtaining appropriate results using the proposed system with the same bit stream for both classic (moderate rendering) and Karaoke / solo (extreme rendering) scenario.
5.3 Generalized structure according to Fig. 5b [0286] The generalized structure of the combined EKS SAOC 590 system will be discussed below with reference to Fig. 5b, which shows a block diagram of such a generalized combined EKS SAOC system. The combined EKS SAOC 590 system of Fig. 5b can also be seen as an audio decoder.
[0287] The combined EKS SAOC 590 system is configured to receive the downmix signal 510a, SAOC bit stream information 512a and rendering matrix information 514a. Also, the combined EKS SAOC 590 system is configured to provide output 520a based on them.
[0288] The combined EKS SAOC 590 system includes a SAOC type I 520a processing member that receives the downmix signal 510a, SAOC bit stream information 512a (or at least a portion thereof), and rendering matrix information 514a (or at least a portion thereof). In particular, the SAOC type I 520a processing member receives values (OLDs) of the level difference of the first member objects. The SAOC processing member I 520a provides one or more signals 562a describing the first set of objects (e.g., audio objects of the first type of audio objects). The SAOC processing member I 520a also provides one or more signals 564a describing a second set of objects.
[0289] The combined EKS SAOC system also includes a SAOC type II 570a processing member that is configured to receive one or more signals 564a describing a second set of objects and to provide on its basis one or more signals 572a describing a third set of objects using differences second-member object levels that are included in SAOC bit stream information 512a, as well as at least part of the rendering matrix information 514. The combined EKS SAOC system also includes a combining module 580a, which may, for example, be an adder to provide output signals 520a by combining one or more signals 570a describing a third set of objects (wherein the third set of objects may be a processed version of the second set of objects).
[0290] In summary of the above, Fig. 5b shows a generalized form of the basic structure described with reference to Fig. 5a above, in a preferred embodiment of the invention.
6. Perceptual evaluation of the way EKS SAOC combined processing
6.1 Test methodology, design and components [0291] These subjective listening tests were carried out in an acoustically insulated listening room, which is adapted for high-quality listening. playback was done using headphones (STAX SR Lambda Pro with D / A converter Lake-People and STAX SRM monitor). The test method followed standard procedures used in spatial audio verification tests based on the MUSHRA method (multiple stimulus with hidden reference and anchors) for subjective assessment of intermediate audio quality (see references [7]).
[0292] Eight listeners participated in the conducted test. According to the MUSHRA methodology, listeners were asked to compare all test conditions against reference. Test conditions were automatically randomized for each test element and for each monitor. Subjective responses were recorded by the MUSHRA computer program on a scale of 0 to 100. Instant switching between test items was enabled. The MUSHRA test was carried out to assess the perceptual results of the SAoC modes under consideration and the proposed system described in the table of Fig. 6a, which provides a description of the listening test structure.
[0293] The corresponding downmix signals were encoded using an AAC core encoder at 128 kbps. To assess the perceptual quality of the proposed combined EKS SAoC system, it is compared with the regular SAoC RM system (model reference system SAoC) and the current EKS mode (enriched Karaoke / solo mode) for two different rendering test scenarios described in the table in Fig. 6b, which describes the systems tested.
[0294] 20 kbps residual encoding was used for the current EKS mode and the proposed combined EKS SAoC system. Note that for the current EKS mode it is necessary to generate a stereo background object (BGo) before the actual encoding / decoding procedure, because this mode has limitations on the number of input object types.
[0295] The listening test material and the appropriate downmix and rendering parameters used in the tests were selected from the set of CfP "call-for-proposals" audio elements described in document [2]. relevant data for the "Karaoke" and "Classic" rendering application scenarios can be found in the table of Fig. 6c, which describes the listening test elements and rendering matrices.
6.2 Listening test results [0296] A brief outline in the form of graphs showing the results of the listening test can be found in Figs. 6d and 6e, where Fig. 6d shows the MUSHRA results for the Karaoke / solo rendering type listening test, and Fig. 6e shows the average scores
MUSHRA listening test of classic rendering. The charts show the average MUSHRA scores for the items for all listeners and the statistical average for all items assessed together with associated ranges of 95% confidence intervals.
[0297] Based on the results of the listening tests, the following conclusions can be made:
* Fig. 6d shows a comparison for the current EKS mode with a connected EKS SAOC system for Karaoke applications. For all test elements, no significant differences in quality (in a statistical sense) are observed between the two systems. From this observation it can be concluded that the combined EKS SAOC system is capable of efficiently using residual information, achieving the quality of the EKS mode. You will also notice that the quality of the regular SAOC system (no residue) is lower than both other systems.
* Fig. 6e shows a comparison for the current SAOC regular mode with the connected EKS SAOC system for classic rendering scenarios. For all tested components, the quality of these two systems is statistically the same. This demonstrates the proper operation of the combined EKS SAOC system for the classic rendering scenario.
[0298] Hence, it can be concluded that the proposed unified system combining EKS mode with regular SAOC retains the advantages of subjective audio quality in the respective types of rendering.
[0299] Given that the proposed combined EKS SAOC system has no restrictions on the BGO object but has fully flexible rendering capabilities for regular SAOC mode and can use the same bit stream for all rendering types, it seems beneficial to incorporate it into the MPEG SAOC standard .
7. Method of Fig. 7 [0300] Hereinafter, a method of providing an upmix signal representation based on a downmix signal representation and object-oriented parametric information will be described with reference to Fig. 7, which shows a flowchart of such a method.
[0301] Method 700 includes step 710 of decomposing the downmix signal representation for providing the first audio information describing the first set of one or more audio objects from the first type of audio objects and the second audio information describing the second set of one or more audio objects from the second type of audio objects, based on the downmix signal representation and at least part of the object parametric information. Method 700 also includes step 720 of processing the second audio information based on the object parametric information to obtain a processed version of the second audio information.
[0302] Method 700 also includes the step 730 of combining the first audio information with the processed version of the second audio information to obtain an upmix signal representation.
[0303] The method 700 of Fig. 7 can be supplemented by any of the features and functions discussed herein in relation to the device of the invention. Also, method 700 provides the benefits discussed with respect to the device of the invention.
8. Alternative implementations [0304] Although some aspects have been described in the context of the device, it is clear that these aspects also represent a description of the respective method, where the block or device corresponds to the method step or the properties of the method step. Similarly, aspects described in the context of a method step also represent a description of the respective block or position or properties of the respective device. Some or all of the method steps may be carried out by hardware devices (or with their help), such as a microprocessor, programmable computer or electronic circuit. In some embodiments, one or more of the most important process steps may be carried out by such a device.
[0305] The audio signal encoded according to the invention may be stored on a digital storage medium or may be transmitted by means of transmission such as wireless transmission means or wired transmission means such as the Internet.
[0306] Depending on some implementation requirements, embodiments of the invention can be implemented in hardware or in software. The implementation can be implemented using digital storage media, e.g. floppy disks, Blue-Ray discs, DVDs, CDs, ROMs, PROMs, EPROMs, EEPROMs or FLASHs containing electronically readable control signals that cooperate with them (or are capable of such interaction with the programmed computer system so that the appropriate method is implemented. Therefore, the digital storage medium can be computer readable.
[0307] Some embodiments of the invention include a data carrier containing electronically readable control signals that interact with a programmable computer system such that one of the methods described herein is implemented.
[0308] In general, embodiments of the present invention may be implemented as a computer program product with a program code, which program code may operate to implement one method of the invention when the computer program product is running on a computer. For example, the program code can be saved on a machine-readable medium.
[0309] Other embodiments include a computer program for performing one of the methods described herein, stored on a machine readable medium.
[0310] In other words, an embodiment of the method of the invention is thus a computer program containing program code for implementing one of the methods described herein when the computer program product is running on a computer.
[0311] A further embodiment of the methods of the invention is thus a data carrier (or digital storage medium, or a computer readable medium) comprising the computer program stored therein for carrying out one of the methods described herein. The data carrier, or digital storage medium, or recorded medium, is typically tangible and / or non-transmitting.
[0312] A further embodiment of the method of the invention is thus a data stream or a sequence of signals representing a computer program for carrying out one of the methods described herein. For example, the data stream or signal sequence may be configured to be sent over a data link, e.g., via the Internet.
[0313] Another embodiment of the method of the invention comprises processing means, e.g. a computer, or programmable logic device, configured or adapted to implement one of the methods described herein.
[0314] Another embodiment of the method of the invention comprises a computer in which a computer program is installed to perform one of the methods described herein.
[0315] In some embodiments, the programmable logic device (e.g., a user programmable logic table) may be used to perform some or all of the functions of the methods described herein. In some embodiments, the user programmable logic table may interact with a microprocessor to implement one of the methods described herein. Generally, the methods are preferably carried out by any hardware device.
[0316] The above described embodiments are merely illustrative for the principles of the present invention. It should be understood that modifications and variants of the systems and details described herein are obvious to those skilled in the art. It is therefore intended that the restrictions arise only from the scope of the following claims and not from the specific details provided for the description and explanation of the present variants of the invention.
9. Conclusions [0317] Some aspects and benefits of the combined EKS SAOC system of the present invention will be summarized below. In Karaoke and Solo playback situations, the SAOC EKS processing mode supports both the reproduction of only background objects / foreground objects and a free mix (determined by the rendering matrix) of these groups of objects.
[0318] Also, the first mode is considered the main purpose of EKS processing, while the second mode provides additional flexibility.
[0319] It has been found that generalization of EKS functionality consequently involves efforts to combine EKS with the regular SAOC processing mode to obtain one unified system. The capabilities of such a unified system include:
* One single transparent SAOC decoding / transcoding structure;
* One bit stream for both EKS and SAOC regular mode;
* No restrictions on the number of input objects containing a background object (BGO), so that there is no need to generate a background object before the SAOC coding stage; and * Residual coding support for foreground objects providing enhanced perceptual quality in demanding Karaoke / Solo playback situations.
[0320] These benefits can be achieved using the unified system described herein.
Bibliography [0321] [1] ISO / IEC JTC1 / SC29 / WG11 (MPEG), Document N8853, "Call for Proposals on Spatial Audio Object Coding", 79th MPEG Meeting, Marrakech, January 2007.
[2] ISO / IEC JTC1 / SC29 / WG11 (MPEG), Document N9099, "Final Spatial Audio Object Coding Evaluation Procedures and Criterion", 80th MPEG Meeting, San Jose, April 2007.
[3] ISO / IEC JTC1 / SC29 / WG11 (MPEG), Document N9250, "Report on Spatial Audio Object Coding RM0 Selection", 8 1st MPEG Meeting, Lausanne, July 2007.
[4] ISO / IEC JTC1 / SC29 / WG11 (MPEG), Document M15123, "Information and Verification Results for CE on Karaoke / Solo system improving the performance of MPEG SAOC RM0", 83rd MPEG Meeting, Antalya, Turkey, January 2008.
[5] ISO / IEC JTC1 / SC29 / WG11 (MPEG), Document N10659, "Study on ISO / IEC 23003-2: 200x Spatial Audio Object Coding (SAOC)", 88th MPEG Meeting, Maui, USA, April 2009.
[6] ISO / IEC JTC1 / SC29 / WG11 (MPEG), Document M10660, "Status and Workplan on SAOC Core Experiments", 88th MPEG Meeting, Maui, USA, April 2009.
[7] EBU Technical recommendation: "MUSHRA-EBU Method for Subjective Listening Tests of Intermediate Audio Quality", Doc. B / AIM022, October 1999.
[8] ISO / IEC 23003-1: 2007, Information technology - MPEG audio technologies - Part 1: MPEG Surround.
Fraunhofer-Gesellschaft zur Forderung der angewandten Forschung eV, Germany Representative:
EP 2 446 435 B1 Z-10708/13
Contents5
41 members in 20 offices
Priority claims6
| Document | Office | Kind | Date |
|---|---|---|---|
| 22004209 | United States of America | P | |
| 10727721 | European Patent Office (EPO) | A | |
| 2010058906 | European Patent Office (EPO) | W | |
| EP20100727721 | – | – | – |
| US20090220042P | – | – | – |
| WO2010EP58906 | – | – | – |
Members41
| Document | Office | Kind | |
|---|---|---|---|
| CA2766727A1 | Canada | A1 | |
| CA2855479A1 | Canada | A1 | |
| WO2010149700A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW201108204A | Taiwan Province of China | A | |
| AR077226A1 | Argentina | A1 | |
| AU2010264736A1 | Australia | A1 | |
| SG177277A1 | Singapore | A1 | |
| MX2011013829A | Mexico | A | |
| KR20120023826A | Republic of Korea | A | |
| EP2446435A1 | European Patent Office (EPO) | A1 | |
| CN102460573A | China | A | |
| US2012177204A1 | United States of America | A1 | |
| CO6480949A2 | Colombia | A2 | |
| ZA201109112B | South Africa | B | |
| JP2012530952A | Japan | A | |
| EP2535892A1 | European Patent Office (EPO) | A1 | |
| HK1170329A1 | Hong Kong, China | A1 | |
| EP2446435B1 | European Patent Office (EPO) | B1 | |
| RU2012101652A | Russian Federation | A | |
| HK1180100A1 | Hong Kong, China | A1 | |
| ES2426677T3 | Spain | T3 | |
| PL2446435T3This record | Poland | T3 | |
| CN103474077A | China | A | |
| CN103489449A | China | A | |
| AU2010264736B2 | Australia | B2 | |
| KR101388901B1 | Republic of Korea | B1 | |
| TWI441164B | Taiwan Province of China | B | |
| CN102460573B | China | B | |
| EP2535892B1 | European Patent Office (EPO) | B1 | |
| ES2524428T3 | Spain | T3 | |
| US8958566B2 | United States of America | B2 | |
| JP5678048B2 | Japan | B2 | |
| PL2535892T3 | Poland | T3 | |
| MY154078A | Malaysia | A | |
| RU2558612C2 | Russian Federation | C2 | |
| BRPI1009648A2 | Brazil | A2 | |
| CA2766727C | Canada | C | |
| CN103474077B | China | B | |
| CA2855479C | Canada | C | |
| CN103489449B | China | B | |
| BRPI1009648B1 | Brazil | B1 |
Numbers
- Publication, DOCDB
- 2446435
- Publication, EPODOC
- PL2446435T
- Application
- 727721
- Application, DOCDB
- 10727721
- Application, EPODOC
- PL20100727721T
Titles2
- English
- AUDIO SIGNAL DECODER, METHOD FOR DECODING AN AUDIO SIGNAL AND COMPUTER PROGRAM USING CASCADED AUDIO OBJECT PROCESSING STAGES
- Polish
- Dekoder sygnału audio, sposób dekodowania sygnału audio i program komputerowy wykorzystujący kaskadowe etapy przetwarzania obiektów audio
Classification
- IPC, 2
- G10L19 008
- G10L19 20