Audio signal decoder, method for decoding an audio signal and computer program using cascaded audio object processing stages
Abstract
This record has no abstract on file.
Term
3.7 yearsto projected expiry
Projected expiry 23 June 2030, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
1 claim: 1 independent, 0 dependent
- 1Patent claims Zastrzeżenia patentowe 1. Audio signal decoder (100; 200; 500; 590) for providing an upmix signal representation based on a downmix signal representation (112; 210; 510; 510a) and parametric object information (110; 212; 512; 512a), the signal decoder audio includes:1. Dekoder sygnału audio (100;200;500;590) do dostarczania reprezentacji sygnału upmixu w oparciu o reprezentację sygnału downmixu (112;210;510;510a) i obiektową informację parametryczną (110;212;512;512a), przy czym dekoder sygnału audio zawiera: an object separator (130;260;520;520a), configured to distribute the downmix signal representation, for providing the first audio information (132;262;562;562a) describing the first set of one or more audio objects from the first type of audio objects and the second audio information (134;264;564;564a), describing a second set of one or more audio objects of the second type of audio objects, based on the representation of the downmix signal and using at least part of the object parametric information;separator obiektów (130;260;520;520a), skonfigurowany do rozkładu reprezentacji sygnału downmixu, dla dostarczania pierwszej informacji audio (132;262;562;562a), opisującej pierwszy zestaw jednego lub większej liczby obiektów audio z pierwszego typu obiektów audio i drugiej informacji audio (134;264;564;564a), opisującej drugi zestaw jednego lub większej liczby obiektów audio z drugiego typu obiektów audio, w oparciu o reprezentację sygnału downmixu i z użyciem przynajmniej części obiektowej informacji parametrycznej;an audio signal processor configured to receive the second audio information (134;264;564;564a) and to process the second audio information depending on the parametric object information to obtain a processed version (142;272;572;572a) of the second audio information;and a combining module (150;280;580;580a) of the audio signal, configured to combine the first audio information with the processed version of the second audio information to obtain an upmix signal representation;procesor sygnału audio skonfigurowany do odbioru drugiej informacji audio (134;264;564;564a) i do przetwarzania drugiej informacji audio w zależności od obiektowej informacji parametrycznej, dla uzyskania przetworzonej wersji (142;272;572;572a) drugiej informacji audio;oraz moduł łączenia (150;280;580;580a) sygnału audio, skonfigurowany do łączenia pierwszej informacji audio z przetworzoną wersją drugiej informacji audio, dla uzyskania reprezentacji sygnału upmixu;przy czym separator obiektów jest skonfigurowany do uzyskania pierwszej informacji audio i drugiej informacji audio zgodnie z gdzie ^^Prediction _ jj-i£ gdzie gdzie XoBJ reprezentuje kanały drugiej informacji audio;wherein the object separator is configured to obtain first audio information and second audio information according to where ^^ Prediction _ jj-i £ where XoBJ represents the channels of the second audio information;gdzie X;Ao reprezentuje sygnały obiektów pierwszej informacji audio;where X;Ao represents object signals of the first audio information;gdzie £> reprezentuje macierz, która jest odwróceniem rozszerzonej macierzy downmixu;where £> represents the matrix, which is the inversion of the extended downmix matrix;Ci Q, Ci ή gdzie C opisuje macierz reprezentującą wiele współczynników j"" ' predykcji kanału;gdzie l0 i r0 reprezentują kanały reprezentacji sygnału downmixu;gdzie res0 do resNEAO-1 reprezentują kanały resztkowe;i gdzie AEAO jest macierzą EAO wstępnego renderowania, której wyrazy opisują mapowanie wzbogaconych obiektów audio na kanały sygnału X;Ao wzbogaconych obiektów audio;przy czym separator obiektów jest skonfigurowany do uzyskania odwróconej macierzy D" downmixu, jako odwrócenia rozszerzonej macierzy E downmixu, która jest zdefiniowana jako przy czym separator obiektów jest skonfigurowany do uzyskania macierzy c jako gdzie m0 do "Χαο-ι są wartościami downmixu powiązanymi z obiektami audio pierwszego typu obiektów audio;Ci Q, Ci ή where C describes a matrix representing many coefficients j"" 'channel prediction;where l0 and r0 represent downmix signal representation channels;where res0 to resNEAO-1 represent residual channels;and where A.EDA is an EAO pre-rendering matrix whose words describe the mapping of enhanced audio objects to X signal channels;Ao enhanced audio objects;wherein the object separator is configured to obtain an inverted D 'downmix matrix as inverse of the extended downmix E matrix, which is defined as wherein the object separator is configured to obtain a matrix c as where m0 to "Χαο-ι are downmix values associated with the audio objects of the first type of audio objects;gdzie n0 do 1Xeac-i są wartościami downmixu powiązanymi z obiektami audio pierwszego typu obiektów audio;where n0 down 1Xeac-i are downmix values associated with audio objects of the first type of audio objects;przy czym separator obiektów jest skonfigurowany do obliczania współczynników predykcji i J 1 jako oraz przy czym separator obiektów jest skonfigurowany do uzyskania ograniczonych wherein the object separator is configured to calculate the prediction coefficients and J 1 as and where the object separator is configured to obtain limited C;n . Oj współczynników predykcji j i cj-,1 ze współczynników predykcji J·0 i 7,1 z użyciem algorytmu ograniczenia, lub do użycia współczynników predykcji ' ' i J ’ jako współczynników predykcji cj,0 i cj,1;C;n. Oh prediction coefficients jicj-,1 from prediction coefficients J0 and 7,1 using a constraint algorithm, or to use prediction coefficients '' and J 'as prediction coefficients cj, 0 and cj, 1;przy czym wielkości energii PLo, PRo, PLoRo, PLoCoj i PRoCoj są zdefiniowane jako = ^f^OLD, + ^ef ft - m.OLD} - £ i=0 + mfa* - nflLDj - £ netj i=0 /*/ gdzie parametry OLDL, OLDR i IOCL,R odpowiadają obiektom audio drugiego typu obiektów audio i są zdefiniowane zgodnie z where the energy quantities PLo, PRo, PLoRo, PLoCoj and PRoCoj are defined as = ^ f ^ OLD, + ^ ef ft - m.OLD} - £ i = 0 + mfa * - nflLDj - £ neie i = 0 / * / where parameters OLDL, OLDR and IOCL, R correspond to audio objects of the second type of audio objects and are defined according to OLDl= £ <O £ Ą, i-0 OLDl= £ <O£Ą, i-0 In-WGP-I W-Wgp-I OLDh^ £ ĄOLD ,, ł-0 OLDh^ £ ĄOLD,, ł-0 IOC ,,, NNE4O = 2, θ 'otherwise. IOC,,, N-NE4O = 2, θ’ w innym przypadku. gdzie d0,i i d1,i są wartościami downmixu powiązanymi z obiektami audio z drugiego typu obiektów audio;where d0, ii d1, i are downmix values associated with the audio objects of the second type of audio objects;gdzie OLDi są wartościami różnicy poziomów obiektów powiązanymi z obiektami audio z drugiego typu obiektów audio;where OLDi are the object level difference values associated with the audio objects of the second type of audio objects;gdzie N jest całkowitą liczbą obiektów audio;where N is the total number of audio objects;gdzie NEAO jest liczbą obiektów audio pierwszego typu obiektów audio;where NEAO is the number of audio objects of the first type of audio objects;gdzie IOC0,1 jest wartością korelacji międzyobiektowej powiązaną z parą obiektów audio drugiego typu obiektów audio;where IOC0,1 is the cross-object correlation value associated with the pair of audio objects of the second type of audio objects;gdzie eij i eL,R są wartościami kowariancji uzyskanymi z parametrów różnicy poziomów obiektów i parametrów korelacji międzyobiektowej;i gdzie eij są powiązane z parą obiektów audio z pierwszego typu obiektów audio, a eL,R jest powiązany z parą obiektów audio z drugiego typu obiektów audio. where eij and eL, R are covariance values obtained from object level difference parameters and inter-object correlation parameters;and where e and j are associated with a pair of audio objects from the first type of audio objects and eL, R is associated with a pair of audio objects from the second type of audio objects. 2. Audio signal decoder (100;200;500;590) for providing an upmix signal representation depending on the downmix signal representation (112;210;510;510a) and object-oriented parametric information (110;212;512;512a), the signal decoder audio includes: 2. Dekoder sygnału audio (100;200;500;590) do dostarczania reprezentacji sygnału upmixu w zależności od reprezentacji sygnału downmixu (112;210;510;510a) i obiektowej informacji parametrycznej (110;212;512;512a), przy czym dekoder sygnału audio zawiera: an object separator (130;260;520;520a) configured to distribute the downmix signal representation, to provide the first audio information (132;262;562;562a) describing the first set of one or more audio objects from the first type of audio objects, and the second audio information (134;264;564;564a), describing a second set of one or more audio objects from the second type of audio objects, depending on the downmix signal representation and using at least part of the object parametric information;separator (130;260;520;520a) obiektów skonfigurowany do rozkładu reprezentacji sygnału downmixu, do dostarczania pierwszej informacji audio (132;262;562;562a), opisującej pierwszy zestaw jednego lub większej liczby obiektów audio z pierwszego typu obiektów audio, oraz drugiej informacji audio (134;264;564;564a), opisującej drugi zestaw jednego lub większej liczby obiektów audio z drugiego typu obiektów audio, w zależności od reprezentacji sygnału downmixu i z użyciem co najmniej części obiektowej informacji parametrycznej;an audio signal processor configured to receive the second audio information (134;264;564;564a) and to process the second audio information depending on the parametric object information to obtain a processed version (142;272;572;572a) of the second audio information;and module (150;280;580;580a) combining the audio signal, configured to combine the first audio information with the processed version of the second audio information to obtain the upmix signal representation, the object separator being configured to obtain the first audio information and the second audio information according to \ ο 7 γ - A Εη «Ε > procesor sygnału audio skonfigurowany do odbioru drugiej informacji audio (134;264;564;564a) i do przetwarzania drugiej informacji audio w zależności od obiektowej informacji parametrycznej, dla uzyskania przetworzonej wersji (142;272;572;572a) drugiej informacji audio;oraz moduł (150;280;580;580a) łączenia sygnału audio, skonfigurowany do łączenia pierwszej informacji audio z przetworzoną wersją drugiej informacji audio dla uzyskania reprezentacji sygnału upmixu, przy czym separator obiektów jest skonfigurowany do uzyskania pierwszej informacji audio i drugiej informacji audio zgodnie z \ ο 7 γ — A Εη«Ε> ΛΕ ^ Ο ~ Λ ™ ΕΛΟ d> and where XOBJ represents the channels of the second audio information;ΛΕ^Ο ~ Λ ™ ΕΛΟ d> i gdzie XOBJ reprezentuje kanały drugiej informacji audio;gdzie XEAO reprezentuje sygnały obiektów pierwszej informacji audio;gdzie gdzie m0 to mNEAO-1 są wartościami downmixu powiązanymi z obiektami audio z pierwszego typu obiektów audio;where XEAO represents the signals of the first audio information objects;where m0 is mNEAO-1 are downmix values associated with the audio objects of the first type of audio objects;gdzie n0 do nNEAO-1 są wartościami downmixu powiązanymi z obiektami audio z pierwszego typu obiektów audio;where n0 to nNEAO-1 are downmix values associated with the audio objects of the first type of audio objects;gdzie OLDi są wartościami różnicy poziomów obiektów powiązanymi z obiektami audio z pierwszego typu obiektów audio;where OLDi are the object level difference values associated with the audio objects of the first type of audio objects;gdzie OLDL i OLDR są wspólnymi wartościami różnicy poziomów obiektów powiązanymi z obiektami audio drugiego typu obiektów audio;i gdzie AEAO jest macierzą EAO wstępnego renderowania, której wyrazy opisują mapowanie wzbogaconych obiektów audio na kanały sygnału XEAO wzbogaconych obiektów audio. where OLDL and OLDR are common object level difference values associated with audio objects of the second type of audio objects;and where A.EDA is an EAO pre-rendering matrix whose words describe the mapping of enriched audio objects to the XEAO signal channels of enriched audio objects. 3. An audio decoder (100;200;500;590) for providing an upmix signal representation depending on the downmix signal representation (112;210;510;510a) and object-oriented parametric information (110;212;512;512a), the signal decoder audio includes: 3. Dekoder (100;200;500;590) sygnału audio do dostarczania reprezentacji sygnału upmixu w zależności od reprezentacji sygnału downmixu (112;210;510;510a) i obiektowej informacji parametrycznej (110;212;512;512a), przy czym dekoder sygnału audio zawiera: a separator (130;260;520;520a) of audio objects configured to distribute the downmix signal representation for providing the first audio information (132;262;562;562a) describing the first set of one or more audio objects from the first type of audio objects, and the second audio information (134;264;564;564a) describing a second set of one or more audio objects from the second type of audio objects, depending on the representation of the downmix signal and using at least part of the object parametric information;separator (130;260;520;520a) obiektów audio skonfigurowany do rozkładu reprezentacji sygnału downmixu, dla dostarczania pierwszej informacji audio (132;262;562;562a) opisującej pierwszy zestaw jednego lub większej liczby obiektów audio z pierwszego typu obiektów audio, oraz drugiej informacji audio (134;264;564;564a) opisującej drugi zestaw jednego lub większej liczby obiektów audio z drugiego typu obiektów audio, w zależności od reprezentacji sygnału downmixu i z użyciem co najmniej części obiektowej informacji parametrycznej;an audio signal processor configured to receive the second audio information (134;264;564;564a) and to process the second audio information depending on the parametric object information to obtain a processed version (142;272;572;572a) of the second audio information;and a module (150;280;580;580a) for combining audio signals, configured to combine the first audio information with the processed version of the second audio information to obtain an upmix signal representation;procesor sygnału audio skonfigurowany do odbioru drugiej informacji audio (134;264;564;564a) i do przetwarzania drugiej informacji audio w zależności od obiektowej informacji parametrycznej, dla uzyskania przetworzonej wersji (142;272;572;572a) drugiej informacji audio;oraz moduł (150;280;580;580a) łączenia sygnałów audio, skonfigurowany do łączenia pierwszej informacji audio z przetworzoną wersją drugiej informacji audio, dla uzyskania reprezentacji sygnału upmixu;przy czym separator obiektów jest skonfigurowany do uzyskania pierwszej informacji audio i drugiej informacji audio zgodnie z wherein the object separator is configured to obtain first audio information and second audio information in accordance with Χ "" = Μ ^ <ί. Χ„„ =Μ^<ί. ^£40 = gdzie XOBJ reprezentuje kanał drugiej informacji audio;^£40 = where XOBJ represents the channel of the second audio information;gdzie XEAO reprezentuje sygnały obiektów pierwszej informacji audio;gdzie gdzie m0 do mNEAO-1 są wartościami downmixu powiązanymi z obiektami audio z pierwszego typu obiektów audio;where XEAO represents the signals of the first audio information objects;where m0 to mNEAO-1 are downmix values associated with the audio objects of the first type of audio objects;gdzie OLDi są wartościami różnicy poziomów obiektów powiązanymi z obiektami audio z pierwszego typu obiektów audio;where OLDi are the object level difference values associated with the audio objects of the first type of audio objects;gdzie OLDL jest wspólną wartością różnicy poziomów obiektów powiązanych z obiektami audio drugiego typu obiektów audio;i gdzie AEAO jest macierzą EAO wstępnego renderowania, której wyrazy opisują mapowanie wzbogaconych obiektów audio na kanały sygnału XEAO wzbogaconych obiektów audio;where OLDL is the common value of the difference in the level of objects associated with the audio objects of the second type of audio objects;and where A.EDA is an EAO pre-rendering matrix whose words describe the mapping of enriched audio objects to XEAO enriched audio object channels;are used for the d representation0 whereby the matrix and single SAOC downmix signal. są zastosowane dla reprezentacji d0 przy czym macierze i pojedynczego sygnału downmixu SAOC. 4. A method of providing an upmix signal representation depending on the downmix signal representation and object-oriented parametric information, the method comprising: a distribution of the downmix signal representation for providing the first audio information describing the first set of one or more audio objects from the first type of audio objects and the second audio information describing the second set of one or more audio objects from the second type of audio objects based on the downmix signal representation from using at least part of the object parametric information;and processing the second audio information depending on the parametric object information to obtain a processed version of the second audio information;and combining the first audio information with the processed version of the second audio information to obtain an upmix signal representation;4. Sposób dostarczania reprezentacji sygnału upmixu w zależności od reprezentacji sygnału downmixu i obiektowej informacji parametrycznej, przy czym sposób obejmuje: rozkład reprezentacji sygnału downmixu, dla dostarczania pierwszej informacji audio opisującej pierwszy zestaw jednego lub większej liczby obiektów audio z pierwszego typu obiektów audio i drugiej informacji audio opisującej drugi zestaw jednego lub większej liczby obiektów audio z drugiego typu obiektów audio, w oparciu o reprezentację sygnału downmixu i z użyciem przynajmniej części obiektowej informacji parametrycznej;oraz przetwarzanie drugiej informacji audio w zależności od obiektowej informacji parametrycznej, dla uzyskania przetworzonej wersji drugiej informacji audio;oraz łączenie pierwszej informacji audio z przetworzoną wersją drugiej informacji audio dla uzyskania reprezentacji sygnału upmixu;przy czym pierwsza informacja audio i druga informacja audio są uzyskane zgodnie z gdzie wherein the first audio information and the second audio information are obtained according to where MPrediction D C f gdzie łitł/ł MPrediction DC f where łitł / ł - - ™ 4ar -i— *> From Pr ^ lo ^lf, ao where XOBJ represents the channels of the second audio information;- — ™ 4ar -i— * >Z Pr^ięlim ^lf,ao gdzie XOBJ reprezentuje kanały drugiej informacji audio;gdzie XEAO reprezentuje sygnały obiektów pierwszej informacji audio;where XEAO represents the signals of the first audio information objects;gdzie D”reprezentuje macierz, która jest odwróceniem rozszerzonej macierzy downmixu;where D "represents the matrix, which is the inversion of the extended downmix matrix;Ci q, Cs ~ and where C describes a matrix representing many coefficients j channel prediction;Ci q, Cs ~i gdzie C opisuje macierz reprezentującą wiele współczynników j predykcji kanału;gdzie l0 i r0 reprezentują kanały reprezentacji sygnału downmixu;gdzie res0 do resNEAO-1 reprezentują kanały resztkowe;i gdzie AEAO jest macierzą EAO wstępnego renderowania, której wyrazy opisują mapowanie wzbogaconych obiektów audio na kanały sygnału XEAO wzbogaconych obiektów audio;przy czym odwrócona macierz D" downmixu, jako odwrócenie rozszerzonej macierzy E downmixu, która jest zdefiniowana jako where l0 and r0 represent downmix signal representation channels;where res0 to resNEAO-1 represent residual channels;and where A.EDA is an EAO pre-rendering matrix whose words describe the mapping of enriched audio objects to XEAO enriched audio object channels;wherein the inverted D "downmix matrix, as the inversion of the extended downmix E matrix, which is defined as m n me NEjtO ' ó" NEjtO 'ó " przy czym macierz C jest uzyskana jako gdzie m0 do mNEAO-1 są wartościami downmixu powiązanymi z obiektami audio pierwszego typu obiektów audio;wherein matrix C is obtained as where m0 to mNEAO-1 are downmix values associated with the audio objects of the first type of audio objects;gdzie n0 do nNEAO-1 są wartościami downmixu powiązanymi z obiektami audio pierwszego typu obiektów audio;where n0 to nNEAO-1 are downmix values associated with the audio objects of the first type of audio objects;C L' 1 przy czym współczynniki predykcji '' i ;1 są obliczone jako oraz ograniczone współczynniki predykcji cj,0 i cj,1 są pozyskane ze współczynników predykcji cy.o i ^J·1 z użyciem algorytmu ograniczenia, lub przy czym współczynniki predykcji C/.° i Cy·1 są użyte jako współczynniki predykcji F/.O i C/·1;CL '1 with prediction coefficients'' and ;1 are calculated as and limited prediction coefficients cj, 0 and cj, 1 are obtained from prediction coefficients cyo i ^ J ·1 using a constraint algorithm, or with prediction coefficients C/.Cho i Cy·1 are used as the prediction coefficients F / .O and C/·1;przy czym wielkości energii PLo, PRo, PLoRo, PLoCoj i PRoCoj są zdefiniowane jako ^£.0 "I fffclg-1 where the energy quantities PLo, PRo, PLoRo, PLoCoj and PRoCoj are defined as ^ £ .0 "I fffclg-1 Λ. = oldl + Σ Σ / = 0 i-0 ^ = α + Σ Σ w> * Λ. = oldl + Σ Σ /=0 i-0 ^=α+Σ Σ w>* J-0 fc-0 gdzie parametry OLDL, OLDR i IOCL,R odpowiadają obiektom audio drugiego typu obiektów audio i są zdefiniowane zgodnie z owt= Σ <%,oLD(t i-0 J-0 fc-0 where OLDL, OLDR and IOCL, R parameters correspond to audio objects of the second type of audio objects and are defined according tot= Σ <%, oLD(t i-0 W - ^, - 1 oldh^ X ąold, ł-o where d0, ii d1, i are downmix values associated with audio objects from the second type of audio objects;W-^,-1 oldh^ X Ąold, ł-o gdzie d0,i i d1,i są wartościami downmixu powiązanymi z obiektami audio z drugiego typu obiektów audio;gdzie OLDi są wartościami różnicy poziomów obiektów powiązanymi z obiektami audio z drugiego typu obiektów audio;where OLDi are the object level difference values associated with the audio objects of the second type of audio objects;gdzie N jest całkowitą liczbą obiektów audio;where N is the total number of audio objects;gdzie NEAO jest liczbą obiektów audio pierwszego typu obiektów audio;where NEAO is the number of audio objects of the first type of audio objects;gdzie IOC0,1 jest wartością korelacji międzyobiektowej powiązaną z parą obiektów audio drugiego typu obiektów audio;where IOC0,1 is the cross-object correlation value associated with the pair of audio objects of the second type of audio objects;gdzie eij i eL,R są wartościami kowariancji uzyskanymi z parametrów różnicy poziomów obiektów i parametrów korelacji międzyobiektowej;i gdzie eij są powiązane z parą obiektów audio z pierwszego typu obiektów audio, a eL,R jest powiązany z parą obiektów audio z drugiego typu obiektów audio. where eij and eL, R are covariance values obtained from object level difference parameters and inter-object correlation parameters;and where e and j are associated with a pair of audio objects from the first type of audio objects and eL, R is associated with a pair of audio objects from the second type of audio objects. 5. A method of providing an upmix signal representation depending on the downmix signal representation and object-oriented parametric information, the method comprising: distribution of the downmix signal representation for providing the first audio information describing the first set of one or more audio objects from the first type of audio objects, and the second audio information describing the second set of one or more audio objects from the second type of audio objects, depending on the downmix signal representation and using at least part of the object parametric information;5. Sposób dostarczania reprezentacji sygnału upmixu w zależności od reprezentacji sygnału downmixu i obiektowej informacji parametrycznej, przy czym sposób obejmuje: rozkład reprezentacji sygnału downmixu, dla dostarczania pierwszej informacji audio opisującej pierwszy zestaw jednego lub większej liczby obiektów audio z pierwszego typu obiektów audio, oraz drugiej informacji audio opisującej drugi zestaw jednego lub większej liczby obiektów audio z drugiego typu obiektów audio, w zależności od reprezentacji sygnału downmixu i z użyciem co najmniej części obiektowej informacji parametrycznej;processing the second audio information depending on the parametric object information to obtain a processed version of the second audio information;and combining the first audio information with the processed version of the second audio information to obtain an upmix signal representation, wherein the first audio information and the second audio information are obtained in accordance with \rai where XOBJ represents the channels of the second audio information;przetwarzanie drugiej informacji audio w zależności od obiektowej informacji parametrycznej dla uzyskania przetworzonej wersji drugiej informacji audio;oraz łączenie pierwszej informacji audio z przetworzoną wersją drugiej informacji audio dla uzyskania reprezentacji sygnału upmixu, przy czym pierwsza informacja audio i druga informacja audio są uzyskane zgodnie z \ra i gdzie XOBJ reprezentuje kanały drugiej informacji audio;gdzie XEAO reprezentuje sygnały obiektów pierwszej informacji audio;where XEAO represents the signals of the first audio information objects;gdzie gdzie m0 do mNEAO-1 są wartościami downmixu powiązanymi z obiektami audio z pierwszego typu obiektów audio;where m0 to mNEAO-1 are downmix values associated with the audio objects of the first type of audio objects;gdzie n0 do nNEAO-1 są wartościami downmixu powiązanymi z obiektami audio z pierwszego typu obiektów audio;where n0 to nNEAO-1 are downmix values associated with the audio objects of the first type of audio objects;gdzie OLDi są wartościami różnicy poziomów obiektów powiązanymi z obiektami audio z pierwszego typu obiektów audio;where OLDi are the object level difference values associated with the audio objects of the first type of audio objects;gdzie OLDL i OLDR są wspólnymi wartościami różnicy poziomów obiektów powiązanymi z obiektami audio drugiego typu obiektów audio;i gdzie AEAO jest macierzą EAO wstępnego renderowania, której wyrazy opisują mapowanie wzbogaconych obiektów audio na kanały sygnału XEAO wzbogaconych obiektów audio. where OLDL and OLDR are common object level difference values associated with audio objects of the second type of audio objects;and where A.EDA is an EAO pre-rendering matrix whose words describe the mapping of enriched audio objects to the XEAO signal channels of enriched audio objects. 6. A method of providing an upmix signal representation based on a downmix signal representation and object-oriented parametric information, the method comprising: a distribution of the downmix signal representation for providing the first audio information describing the first set of one or more audio objects from the first type of audio objects and the second audio information describing the second set of one or more audio objects from the second type of audio objects based on the downmix signal representation from using at least part of the object parametric information;and processing the second audio information depending on the parametric object information to obtain a processed version of the second audio information;and combining the first audio information with the processed version of the second audio information to obtain an upmix signal representation;6. Sposób dostarczania reprezentacji sygnału upmixu w oparciu o reprezentację sygnału downmixu i obiektową informację parametryczną, przy czym sposób obejmuje: rozkład reprezentacji sygnału downmixu, dla dostarczania pierwszej informacji audio opisującej pierwszy zestaw jednego lub większej liczby obiektów audio z pierwszego typu obiektów audio i drugiej informacji audio opisującej drugi zestaw jednego lub większej liczby obiektów audio z drugiego typu obiektów audio, w oparciu o reprezentację sygnału downmixu i z użyciem przynajmniej części obiektowej informacji parametrycznej;oraz przetwarzanie drugiej informacji audio w zależności od obiektowej informacji parametrycznej, dla uzyskania przetworzonej wersji drugiej informacji audio;i łączenie pierwszej informacji audio z przetworzoną wersją drugiej informacji audio dla uzyskania reprezentacji sygnału upmixu;przy czym pierwsza informacja audio i druga informacja audio sa uzyskane zgodnie z ^£40 = gdzie XOBJ reprezentuje kanał drugiej informacji audio;wherein the first audio information and the second audio information are obtained in accordance with ^ £ 40 = where XOBJ represents the channel of the second audio information;gdzie XEAO reprezentuje sygnały obiektów pierwszej informacji audio;gdzie gdzie m0 do mNEAO-1 są wartościami downmixu powiązanymi z obiektami audio z pierwszego typu obiektów audio;where XEAO represents the signals of the first audio information objects;where m0 to mNEAO-1 are downmix values associated with the audio objects of the first type of audio objects;gdzie OLDi są wartościami różnicy poziomów obiektów powiązanymi z obiektami audio z pierwszego typu obiektów audio;where OLDi are the object level difference values associated with the audio objects of the first type of audio objects;gdzie OLDL jest wspólną wartością różnicy poziomów obiektów powiązanych z obiektami audio drugiego typu obiektów audio;i gdzie AEAO jest macierzą EAO wstępnego renderowania, której wyrazy opisują mapowanie wzbogaconych obiektów audio na kanały sygnału XEAO wzbogaconych obiektów audio;where OLDL is the common value of the difference in the level of objects associated with the audio objects of the second type of audio objects;and where A.EDA is an EAO pre-rendering matrix whose words describe the mapping of enriched audio objects to XEAO enriched audio object channels;£ |£ k*n»'Are used for the representation d0 whereby the matrix and single SAOC downmix signal. Ł|£k*n»' są zastosowane dla reprezentacji d0 przy czym macierze i pojedynczego sygnału downmixu SAOC. 7. A computer program for carrying out the method defined in one of claims 4 to 6, when the computer program is executed on a computer. 7. Program komputerowy do realizacji sposobu określonego w jednym z zastrzeżeń od 4 do 6, gdy program komputerowy jest wykonywany na komputerze. Fraunhofer-Gesellschaft zur Forderung der angewandten Forschung e .V . , Niemcy Pełnomocnik : Fraunhofer-Gesellschaft zur Forderung der angewandten Forschung e .V. , Germany Representative: EP 2 535 892 B1 EP 2 535 892 B1 Z-12693 Z-12693 FIG 1 FIG 1 EP 2 535 892 B1 EP 2 535 892 B1 Z-12693 Z-12693 300 ____ / _ * {—342 matrix :: SAOC downmix render SAOC data remains 300 ____/_ *{ —342 macierz :: renderowania downmix SAOC resztki danych SAOC 3i0 l — i. · 332 -) 3i0 l—i. 332·--) ARCHITEKTURA PROCESORA RESZTKOWEGO RESIDUAL PROCESSOR ARCHITECTURE FIG 3A FIG 3A MD procesor resztkowy | dane SAOC MD residual processor SAOC data X 31D 1 /330 X 31D 1/330 3Ξ2 SAOC downmix 3Ξ2 downmix SAOC residual data saoc · '·· Ί X »4. "_ _ .___________ L. resztki danych saoc ·'·· Ί X» 4 . „ _ _ .___________L. V "" * V""* 3Ż0 3Ż0 ARCHITEKTURA PROCESORA RESZTKOWEGO RESIDUAL PROCESSOR ARCHITECTURE FIG 3B FIG 3B EP 2 535 892 B1 EP 2 535 892 B1 Z-12693 Z-12693 FIG 4A FIG 4A EP 2 535 892 B1 EP 2 535 892 B1 Z-12693 Z-12693 FIG 4 ES FIG 4 ES EP 2 535 892 B1 EP 2 535 892 B1 Z-12693 Z-12693 EP 2 535 892 B1 EP 2 535 892 B1 Z-12693 Z-12693 IBL IBl 0Π 0Π DMG OLD.IOC '4 τ DMG OLD.IOC '4 τ C3= M "eED'J these?) C3=M„eED'J te?) G— 0πιϋ, tL · τ vector optional strengthen And parameter controller G— 0πιϋ, tL· τ wektor opcjonalnie wzmocnienia I regulator parametru * Λ ' Λ*' A* AND* Φ-P, ήΛ decorator Φ-P, ήΛ dekore lator Bjt bjt Ji.l Ji.l FIG 4D FIG 4D 100 100 EP 2 535 892 B1 EP 2 535 892 B1 Z-12693 Z-12693 FIG 4E FIG 4E 101 101 EP 2 535 892 B1 EP 2 535 892 B1 Z-12693 macierz renderowania Z-12693 rendering matrix FIG 4F FIG 4F 497 ^ SAOC decoder 497^ dekoder SAOC Vdownmix Vdownmix P do rocesor ownmixu wyjście P to ownmix processor output 495 495 497B 497b FROM Z Wa bit stream Wa strumień bitów SAOC SAOC 496a parameter processor 496a procesor parametru SAOC ' “i-I-T\π _ - T — macierz parametry renderowania HRTF SAOC '' iIT \ π _ - T - matrix HRTF rendering parameters 496 496 FIG 4G FIG 4G 102 102 EP 2 535 892 B1 EP 2 535 892 B1 Z-12693 Z-12693 514 514 BASIC STRUCTURE OF THE CONNECTED EKS SAOC FIG5A SYSTEM PODSTAWOWA STRUKTURA POŁĄCZONEGO SYSTEMU EKS SAOC FIG5A 5l4fl 593 5l4fl 593 GENERAL STRUCTURE OF THE CONNECTED EKS SAOC SYSTEM UOGÓLNIONA STRUKTURA POŁĄCZONEGO SYSTEMU EKS SAOC FfG 5B FfG 5B 103 103 EP 2 535 892 B1 EP 2 535 892 B1 Z-12693 Z-12693 FIG 6A FIG 6A FIG6B FIG6B 104 104 EP 2 535 892 B1 EP 2 535 892 B1 Z-12693 Z-12693 ο co ο ο what ο 105 105 EP 2 535 892 B1 EP 2 535 892 B1 Z-12693 Z-12693 106 106 EP 2 535 892 B1 EP 2 535 892 B1 Z-12693 Z-12693 FIG 6E FIG 6E 107 107 EP 2 535 892 B1 EP 2 535 892 B1 Z-12693 Z-12693 720 720 730 730 FIG 7 FIG 7 108 108 EP 2 535 892 B1 EP 2 535 892 B1 Z-12693 Z-12693 -¼ 4 * 1 | t -¼ 4* 1|t _. and. _. i. ju .o' l—.·' ~j' ju. o 'l—. · '~ J' O O O CD C=| OOO CD C = | FIG 8 FIG 8 109 109 EP 2 535 892 B1 EP 2 535 892 B1 Z-12693 Z-12693 CD O O O O CD OOOO LU LU Q Q CO WHAT ABOUT O LO LO ABOUT O Oi Oi LU LU Q Q ABOUT O FIG 9A FIG 9A 110 110 EP 2 535 892 B1 EP 2 535 892 B1 Z-12693 Z-12693 FIG 9B FIG 9B 111 SAOC transcoder to MPEG Surround 986 111 transkoder SAOC do MPEG Surround 986 EP 2 535 892 B1 Z-12693 EP 2 535 892 B1 Z-12693 Qtoj. # 1 Qtoj.#1 Ohi. # 3 Ohi.#3, Obl ·. # ^ Ot> j. * Hand object encoder downmix signal (s) optional: downmix signal manipulator Obl·.#^ ot>j.*rę koder obiektów sygnał(y) downmixu opcjonalnie: manipulator sygnału ' downmixu 98;98;downmix signal (s) sygnał(y) downmixu And object metadata I metadane obiektowe And additional information transcoder I transkoder informacji dodatkowej 984 MPEG Surround bit stream 984 strumień bitów MPEG Surround 9fi0 9fi0 982 rendering information 982 informacja renderowania FIG 9C FIG 9C 112 112 EP 2 535 892 B1 EP 2 535 892 B1 Z-12693 u 'C O (U < Z-12693 u 'CO (U < Έ m 5 | Έ m 5 | ŁD ϋ signal 1 enriched ŁD ϋ sygnał 1 wzbogaconych > "O> * u O 3. -Ω '-Ί θ' N 2? § ro (U c Έ σι fU c >"O >* u O 3. -Ω ’-Ί θ’ N 2? § ro (U c Έ σι fU c ϊ_ ro ϊ_ ro Zi σι (U · - S ro U '1 £ OŹ E <<uo to s what's this Zi σι (U ·- S ro U '1 £ O ŹŹ E < <u o to s c o ta o T3 = 5 T3 =5 ΓΌ ΓΌ
538 paragraphs in 9 sections, as filed
Technical Field [0001] Embodiments of the invention relate to an audio signal decoder for providing an upmix signal representation based on a downmix signal representation and parametric object information.
[0002] Further embodiments of the invention relate to a method of providing an upmix signal representation based on the downmix signal representation and the parametric object information.
[0003] Further embodiments of the invention relate to a computer program.
[0004] Certain embodiments of the invention relate to the improved Karaoke / Solo SAOC system.
Background of the Invention [0005] In modern audio systems, it is desirable to transfer and store audio information in a streamlined manner. In addition, it is often desirable to play audio content using two or even more speakers that are spatially arranged in a room. In such cases, it is desirable to use the capabilities of such multi-speaker systems to enable the user to spatially identify different audio content or components of a single audio content. This can be achieved by individually distributing different audio content to different speakers.
[0006] In other words, in the field of audio processing, audio transmission and audio storage, there is an increasing need to manipulate multi-channel content to improve the listening experience. The use of multi-channel audio content brings significant benefits to the user. For example, three-dimensional auditory sensations can be obtained, providing the user with greater satisfaction in entertainment applications. However, multi-channel content is also useful in a professional environment, such as in teleconferencing applications, because the speaker's intelligibility can be improved by using multi-channel playback.
[0007] However, a good balance between audio quality and bit rate requirements is also desirable to avoid excessive resource loads caused by multi-channel applications.
[0008] Recently, parametric techniques have been proposed for efficient bit rate and / or memory transmission of audio scenes containing multiple audio objects, for example Binaural Cue Coding (Type I) (see for example references [BCC]), Joint Source Coding (see for example references [JSC]) , and MPEG Spatial Audio Object Coding (SAOC) (see for example references [SAOC1], [SAOC2].
[0009] These techniques are aimed at the perceptual reconstruction of the desired audio output scene, rather than by means of waveform matching.
[0010] Fig. 8 shows an outline of such a system (in this case MPEG SAOC). The MPEG SAOC 800 system shown in Fig. 8 includes the SAOC 810 encoder and SAOC 820 decoder. The SAOC 810 encoder receives a plurality of object signals from x1 to xN, which can be represented, for example, as time domain signals or frequency domain signals (e.g., in the form of set of Fourier transform coefficients or in the form of QMF subband signals). The SAOC 810 encoder also typically receives downmix d1 to dN coefficients that are associated with x1 to xN object signals. Separate sets of downmix coefficients may be available for each downmix signal channel. The SAOC 810 encoder is typically configured to obtain a downmix signal channel by combining object signals = x1 to xN according to associated downmix d1 to dN coefficients. Typically, there are fewer downmix channels than object signals x1 to xN. To allow (at least partially) the separation (or separate processing) of the object signals on the SAOC 820 decoder side, the SAOC 810 encoder provides both one or more downmix signals (designated as downmix channels) 812 and additional information 814.
Additional information 814 describes the properties of object signals x1 to xN to allow object processing on the decoder side.
[0011] The SAOC decoder 820 is configured to receive both one or more downmix signals 812 and additional information 814. Also, the SAOC decoder 820 is typically configured to receive user interaction information and / or user control 822 that describes the desired rendering system. For example, user interaction / user control information 822 may describe the speaker arrangement and the desired spatial arrangement of objects provided by object signals x1 to xN.
[0012] The SAOC decoder 820 is configured to provide, for example, a plurality of * Λ decoded upmix channel signals' to -<sup>1</sup> . Upmix channel signals may for example be associated with individual speakers of a multi-speaker rendering system. The SAOC decoder 820 may, for example, include an object separator 820a that is configured to reconstruct, at least in part, object signals x1 to xN based on one or more downmix signals 812 and additional information 814, thereby obtaining reconstructed object signals 820b. However, the reconstructed object signals 820b may differ slightly from the original object signals x1 to xN, for example because the additional information 814 is not completely sufficient for perfect reconstruction due to flow restrictions. The SAOC decoder 820 may further include a mixer 820c that can be configured to receive reconstructed object signals 820b and user interaction / control information 822 and to provide signals based on * 'Λ thereof<sup>1</sup>- 'upmix channels. Mixer 820c may be configured to use user interaction / control information 822 to determine the contribution of individual reconstructed object signals 820b in the 'to - signals<sup>1</sup>upmix channels. For example, user interaction / user control information 822 may include rendering parameters (also referred to as rendering factors) that determine the contribution of the individual reconstructed object signals 820b to the 'to -''-' signals of the upmix channels.
[0013] However, it should be noted that in many embodiments, object separation, which is indicated by the object separator 820a in Fig. 8, and mixing, which is indicated by the mixer 820c in Fig. 8, are carried out in one step. To this end, comprehensive parameters that describe the direct mapping * of the one or more downmix signals 812 to the 'to -' upmix channels can be calculated. These parameters can be calculated based on additional information 814 and user interaction / control information 822.
[0014] Referring now to Figs. 9a, 9b and 9c, various devices will be described for obtaining an upmix signal representation based on a downmix signal representation and object-oriented side information. Fig. 9a is a block diagram of an MPEG SAOC 900 system comprising a SAOC 920 decoder. The SAOC 920 decoder includes, as separate function blocks, an object decoder 922 and a mixing / rendering module 926. The object decoder 922 provides a plurality of reconstructed object signals 924 based on the representation of the downmix signal (e.g., in the form of one or more downmix signals represented in the time domain or frequency domain) and additional object information (e.g., in the form of object meta data). Mix / render module 926 receives the reconstructed object signals 924 associated with the number N of objects and provides, on their basis, one or more signals of the upmix channel 928. In the SAOC 920 decoder, the acquisition of object signals 924 is performed separately from mixing / rendering, which allows separation of the object decoding function from the mixing / rendering function, but entails relatively high computational complexity.
[0015] Referring now to Fig. 9b, another MPEG SAOC 930 system will be briefly described that includes a SAOC 950 decoder. The SAOC 950 decoder provides a plurality of upmix channel signals 958 based on a downmix channel representation (e.g. in the form of one or more) downmix signals) and additional object-oriented information (e.g. in the form of object-oriented meta data). The SAOC 950 decoder includes a combined object decoder and a mixing / rendering module that is configured to obtain 958 upmix channel signals in a combined mixing process without object decoding separation and mixing / rendering, wherein the parameters for said combined upmix process are dependent on both object-oriented additional information and rendering information. The combined upmix process also depends on the downmix information, which is considered part of the object-oriented additional information.
[0016] In summary of the above, the upmix channel signals 928, 958 may be provided in a one-step or two-step process.
[0017] Referring now to Fig. 9c, the MPEG SAOC 960 system will be described. The MPEG SAOC 960 system includes a 980 SAOC to MPEG Surround transcoder instead of a SAOC decoder.
[0018] The SAOC to MPEG Surround transcoder includes an additional information transcoder 982 that is configured to receive object-oriented additional information (e.g., as object-oriented metadata) and optionally information about one or more downmix signals and rendering information. The additional information transcoder is also configured to provide 984 MPEG Surround additional information (e.g., as an MPEG Surround bit stream) based on the received data. Accordingly, the secondary information transcoder 982 is configured to transform the object (parametric) additional information that is received from the object encoder into the channel (parametric) additional information 984, including rendering information and optionally information about the content of one or more downmix signals.
[0019] Optionally, the SAOC 980 to MPEG Surround transcoder may be configured to manipulate one or more downmix signals described, for example, by a downmix signal representation to obtain a manipulated 988 downmix signal representation. However, the downmix signal manipulation module 986 may be omitted, so that the downmix output 988 of the SAOC to MPEG Surround transcoder 988 is identical to the downstream of the SAOC to MPEG Surround transcoder signal. The downmix signal manipulation module 986 may for example be used if the 984 MPEG Surround channel additional information does not provide the desired listening experience based on the downmix signal input of the transducer 980 SAOC to MPEG Surround, which may occur in some rendering systems.
[0020] Correspondingly, the 980 SAOC to MPEG Surround transcoder provides a 988 downmix signal representation and a 984 MPEG Surround bit stream so that many upmix channel signals can be generated that represent audio objects according to the rendering information fed into the 980 SAOC transcoder to MPEG Surround , using an MPEG Surround decoder that receives a 984 MPEG Surround bit stream and a representation of a 988 downmix signal.
[0021] In summary of the above, various methods of decoding SAOC-encoded audio signals can be used. In some cases, a SAOC decoder is used that provides upmix channel signals (e.g., signals 928, 958 upmix channels) based on the downmix signal representation and the object parametric additional information. Examples of this method are shown in Figs. 9a and 9b. Alternatively, the SAOC encoded audio information may be transcoded to obtain a downmix signal representation (e.g., a 988 downmix signal representation) and additional channel information (e.g., MPEG Surround channel bit stream) that can be used by the MPEG Surround decoder to provide desired upmix channel signals.
[0022] In the 800 MPEG SAOC system, the general outline of which is shown in Fig. 8, the general processing is performed in a frequency selective manner and can be described as follows, in each frequency band:
* N input object audio signals x1 to xN are downmixed as part of the SAOC encoder processing. In the case of mono downmix, downmix coefficients are denoted by d1 to dN. In addition, the SAOC 810 encoder acquires additional information 814 describing the properties of the input audio objects. In the case of MPEG SAOC, the relationship of object power to each other is the most basic of such additional information.
* The downmix signal (or signals) 812 and additional information 814 are transmitted and / or stored. To this end, the audio downmix signal can be compressed using well-known perceptual audio coders such as MPEG-1 Layer II or III (also known as ".mp3"), MPEG Advanced Audio Coding (AAC) or any other audio encoder.
* On the receiving side, the SAOC 820 decoder conceptually attempts to reconstruct the original object signal ("object separation") using additional information 814 (and of course one or more downmix signals 812). These approximate object signals (also referred to as reconstructed object signals 820b) are then mixed into the target scene represented by M output audio channels (which can, for example, be represented by * 'Λ' signals to -<sup>1</sup>^ upmix channels) using a rendering matrix. For mono output, the rendering matrix coefficients are given by r1 to rN.
* Effectively, object signal separation is rarely performed (or even never performed), because both the separation stage (indicated by the object 820a separator) and the mixing stage (indicated by the 820c mixer) are combined in one transcoding stage, which often allows a huge reduction in complexity computing.
[0023] Such a method has been found to be extremely efficient, both in terms of bit rate (only a few downmix channels need to be transmitted plus some additional information instead of N separate audio object signals or a discrete system) and computational complexity (processing complexity is mainly related to the number of output channels rather than with the number of audio objects). Further benefits for the user on the receiving side include the freedom to choose the rendering layout according to his / her choice (mono, stereo, virtual headphone playback, etc.) and the user interaction function: the rendering matrix, and thus the output scene, can be set and changed interactively by the user according to his will, personal preferences or other criteria. For example, it is possible to locate speakers from one group together in one spatial area to maximize discrimination from other speakers. This interactivity is obtained by providing a decoder user interface.
[0024] For each transmitted sound object, its relative level and (for non-mono rendering) the spatial position of the rendering can be adjusted. This can occur in real time when the user changes the positions of the associated GUI sliders (graphical user interface, graphical user interface) (for example: object level = + 5dB, object position = -30 degrees).
[0025] However, it has been found difficult to manipulate audio objects from different types of audio objects in such a system. In particular, it has been found that it is difficult to process audio objects from different types of audio objects, e.g., audio objects with which different additional information is associated, if the total number of audio objects for processing is not predetermined.
[0026] In view of this situation, the object of the present invention is to create a concept that allows computationally efficient and flexible decoding of an audio signal comprising a downmix signal representation and parametric object information in which the object parametric information describes the audio objects of two or more different types of audio objects .
Summary of the Invention [0027] This object is achieved by audio decoders for providing an upmix signal representation based on a downmix signal representation and parametric object information, methods for providing an upmix signal representation based on a downmix signal representation and object parametric information and a computer program as defined in independent reservations.
[0028] Embodiments of the invention as set forth in independent claims 1 to 3 create audio signal decoders for providing an upmix signal representation based on the downmix signal representation and the parametric object information. Audio signal decoders include an object separator configured to distribute the downmix signal representation to provide the first audio information describing the first set of one or more audio objects of the first type of audio objects and the second audio information describing the second set of one or more audio objects of the second type of audio objects based for representation of the downmix signal and using at least part of the object parametric information. The audio signal decoders also include an audio signal processor configured to receive the second audio information and to process the second audio information depending on the object parametric information to obtain a processed version of the second audio information. Audio signal decoders also include an audio signal combining module configured to combine the first audio information with the processed version of the second audio information to obtain an upmix signal representation.
[0029] The main idea of the present invention is that efficient processing of different types of audio objects can be obtained in a cascading structure that allows the separation of different types of audio objects using at least part of the object parametric information in the first processing step, implemented by the object separator and which enables additional spatial processing in the second processing stage implemented based on at least part of the object parametric information by the audio signal processor. It has been found that acquiring a second audio information that includes audio objects of the second type of audio objects from the downmix signal representation can be performed with moderate complexity, even if there are more audio objects of the second type of audio objects. In addition, it has been found that the spatial processing of audio objects of the second type of audio objects can be performed efficiently after the second audio information is separated from the first audio information describing the audio objects of the first type of audio objects.
[0030] In addition, it has been found that the processing algorithm implemented by the object separator for separating the first audio information and the second audio information can be implemented with relatively low complexity if the individual object processing of the audio objects of the second type of audio objects is delayed for the audio signal processor and is not implemented simultaneously with the separation of the first audio information from the second audio information.
[0031] For example, an audio signal decoder may be configured to provide an upmix signal representation based on a downmix signal representation, parametric object information and residual information associated with a subset of audio objects represented by a downmix signal representation. In this case, the object separator may be configured to distribute the downmix signal representation to provide the first audio information describing the first set of one or more audio objects (e.g., FGO foreground objects) of the first type of audio objects with which the residual information and the second information are associated audio describing a second set of one or more audio objects (e.g., BGO background objects) of a second type of audio object, with which no residual information is associated, based on the downmix signal representation and using at least part of the object parametric information and the residual information.
[0032] This implementation is based on the finding that a particularly accurate separation of the first audio information describing the first set of audio objects from the first type of audio objects and the second audio information describing the second set of audio objects from the second type of audio objects can be obtained using residual information in addition to object information parametric information. It has been found that the use of only object-oriented parametric information will in many cases bring distortion that can be significantly reduced or even completely eliminated by using residual information. For example, residual information describes the residual distortion that is expected to remain if the audio object of the first type of audio object is isolated only with the use of object-oriented parametric information. Residual information is typically estimated by the audio signal encoder. By using residual information, the separation between audio objects of the first type of audio objects and audio objects of the second type of audio objects can be improved.
[0033] This enables the first audio information and the second audio information to be obtained with particularly good separation between the audio objects of the first type of audio objects and the audio objects of the second type of audio objects, which in turn allows obtaining high quality spatial processing of the audio objects of the second type of audio objects during processing the second audio information in an audio signal processor.
[0034] In an embodiment, the object separator may thus be configured to provide the first audio information in such a way that the audio objects of the first type of audio objects are highlighted relative to the audio objects of the second type of audio objects in the first audio information. The object separator is also configured to provide the second audio information in such a way that audio objects from the second type of audio objects are enhanced relative to the audio objects of the first type of audio objects in the second audio information.
[0035] In implementation, the audio signal decoder may be configured to perform two-stage processing in such a way that the processing of the second audio information in the audio signal processor is performed after separation between the first audio information describing the first set of one or more audio objects of the first type of audio objects and second audio information describing a second set of one or more audio objects from the second type of audio objects.
[0036] In an embodiment, the audio signal processor may be configured to process the second audio information based on the object-related parametric information associated with the audio objects of the second type of audio objects and independent of the object-related parametric information associated with the audio objects of the first type of audio objects. Thus, separate processing of audio objects from the first type of audio objects and audio objects from the second type of audio objects can be obtained.
[0037] In an embodiment, the object separator may be configured to obtain the first audio information and the second audio information using a linear connection of one or more downmix channels and one or more residual channels. In this case, the object separator can be configured to obtain connection parameters for performing a linear connection based on the downmix parameters associated with the audio objects of the first type of audio objects and based on the channel prediction coefficients of the audio objects from the first type of audio objects. The calculation of the channel prediction coefficients of the audio objects from the first type of audio objects may, for example, include audio objects from the second type of audio objects as a single common audio object. Accordingly, the separation process may be performed with a sufficiently low computational complexity, which for example may be independent of the number of audio objects from the second type of audio objects.
[0038] In an embodiment, the object separator may be configured to use a rendering matrix for the first audio information to map object signals of the first audio information to the audio channels of the audio upmix signal representation. This can be done because the object separator may be able to obtain separate audio signals individually representing the audio objects from the first type of audio objects. Accordingly, it is possible to map object signals of the first audio information directly to audio channels from the audio upmix signal representation.
[0039] In an implementation, the audio processor may be configured to perform stereo processing of the second audio information based on rendering information, object covariance information and downmix information to obtain audio channels of the audio upmix signal representation.
[0040] Thus, the stereo processing of audio objects of the second type of audio objects may be separated from separating the audio objects of the first type of audio objects from the audio objects of the second type of audio objects. In this way, the efficient separation of audio objects from the first type of audio objects from the audio objects from the second type of audio objects is not disturbed (or degraded) by stereo processing, which typically leads to the distribution of audio objects on multiple audio channels without providing a high degree of separation of objects that can be obtained in an object separator, for example using residual information.
[0041] In implementation, the signal processor may be configured to perform post-processing of the second audio information based on rendering information, object covariance information and downmix information. This form of post-processing enables the spatial arrangement of audio objects from the second type of audio objects on the audio scene. Regardless, thanks to the cascade concept, the computational complexity of the audio signal processor can be kept low enough because the audio processor does not have to consider object-oriented parametric information associated with the audio objects of the first type of audio objects.
[0042] In addition, various types of processing may be performed by an audio processor, such as, for example, mono to binaural processing, mono to stereo processing, stereo to binaural processing or stereo to stereo processing.
[0043] In an implementation, the object separator may be configured to treat the audio objects of the second type of audio objects to which no residual information is associated as a single audio object. In addition, the audio signal processor may be configured to include object-oriented rendering parameters to match object inputs from the second type of audio objects to the upmix signal representation. In this way, audio objects from the second type of audio objects are treated by the object separator as a single audio object, which significantly reduces the complexity of the object separator as well as allows you to have unique residual information that is independent of the rendering parameters associated with the audio objects of the second type of audio objects.
[0044] In an implementation, the object separator may be configured to obtain a common object value of the level difference for a plurality of audio objects from the second type of audio objects. The object separator can be configured to use the level difference object value to calculate channel prediction coefficients. In addition, the object separator can be configured to use channel prediction coefficients to obtain one or two audio channels representing the second audio information. To obtain the object value of the level difference, audio objects from the second type of audio objects can be efficiently treated by the object separator as a single audio object.
[0045] In an embodiment, the object separator may be configured to obtain a common object value of level difference for multiple audio objects of the second type of audio objects and the object separator may be configured to use a common object value of level difference to calculate the words of the energy mode mapping matrix. The object separator may be configured to use an energy mode mapping matrix to obtain one or more audio channels representing the second audio information. Again, the common value of object level difference allows computationally efficient common treatment of audio objects from the second type of audio objects by the object separator.
[0046] In implementation, the object separator may be configured to selectively obtain a common cross-object correlation value associated with audio objects from the second type of audio objects based on the parametric object information, if it is determined that there are two audio objects of the second type of audio objects and to set the value cross-object correlation associated with audio objects from the second type of audio objects to zero, if found that there are more or less than two audio objects from the second type of audio objects. The object separator may be configured to use a common cross-object correlation value associated with the audio objects of the second type of audio objects to obtain one or more audio channels representing the second audio information. In this approach, the inter-object correlation value is used if it is achievable with high computational efficiency, i.e. if there are two audio objects from the second type of audio objects. Otherwise, it would be computationally demanding to obtain cross-object correlation values. Therefore, it was found that a good compromise in terms of auditory experience and computational complexity is to set the cross-object correlation value associated with audio objects from the second type of audio objects to zero if there are more or less than two audio objects from the second type of audio objects.
[0047] In an embodiment, the audio processor may be configured to render the second audio information based (at least in part) on the parametric object information to obtain a rendered representation of the audio objects from the second type of audio objects as a processed version of the second audio information. In this case, rendering can be done independently of the audio objects from the first type of audio objects.
[0048] In an embodiment, the object separator may be configured to provide the second audio information in such a way that the second audio information describes more than two audio objects from the second type of audio objects. The implementations allow flexible adjustment of the number of audio objects from the second type of audio objects, which is greatly facilitated by the cascade processing structure.
[0049] In an embodiment, the object separator may be configured to obtain, as the second audio information, a single-channel representation of the audio signal or a two-channel representation of audio representing more than two audio objects from the second type of audio objects. Acquiring one or two channels of audio signal can be carried out by an object separator with low computational complexity. In particular, the complexity of the object separator can be kept clearly lower compared to the case where the object separator would have to deal with more than two audio objects from the second type of audio objects. However, it has been independently found that a computationally efficient representation of audio objects from a second type of audio object is to use one or two audio signal channels.
In implementation, the audio signal processor may be configured to receive the second audio information and to process the second audio information depending on (at least partially) the parametric object information, including the object parametric information associated with more than two audio objects of the second type of audio objects. Accordingly, the processing of individual audio objects is carried out by the audio processor, while such processing of individual audio objects is not carried out for the audio objects of the second type of audio objects by the object separator.
[0051] In implementation, the audio decoder may be configured to obtain information about the total number of objects and information about the number of foreground objects from the configuration information associated with the object parametric information. The audio decoder can also be configured to determine the number of audio objects of the second type of audio objects by creating a difference between information about the total number of objects and information about the number of foreground objects. In this way, efficient signaling of the number of audio objects of the second type of audio objects is obtained. In addition, this concept provides a high degree of flexibility in the number of audio objects of the second type of audio objects.
[0052] In implementation, the object separator may be configured to use object-oriented parametric information associated with NEAO audio objects from the first type of audio objects to obtain, as the first audio information, NEAO, audio signals representing (preferably, individually) NEAO audio objects from the first type of objects audio and to obtain, as a second audio information, one or two audio signals representing N-N<sub>EDA</sub> audio objects from the second type of audio objects, treating N-N<sub>EDA</sub> audio objects from the second type of audio object as a single single-channel or two-channel audio object. The audio signal processor is configured to individually render N-NEAO audio objects represented by one or two audio signals from the second audio information using object-oriented parametric information associated with N-NEAO audio objects from the second type of audio objects. Accordingly, the separation of audio objects between audio objects from the first type of audio objects and audio objects from the second type of audio objects is separated from the subsequent processing of audio objects from the second type of audio objects.
[0053] Embodiments of the invention create methods, as set out in independent claims 4 to 6, for providing an upmix signal representation based on the downmix signal representation and the parametric object information.
[0054] Another embodiment of the present invention provides a computer program for performing said methods, as defined in independent claim 7.
Brief description of the figures [0055] Embodiments of the invention will then be described with reference to the attached figures, in which:
Fig. 1 shows a block diagram of an audio decoder according to an embodiment of the invention;
Fig. 2 shows a block diagram of another audio decoder according to an embodiment of the invention;
Figures 3a and 3b are block diagrams of a residual processor that can be used as an object separator in an embodiment of the invention;
Figures 4a to 4e show block diagrams of audio signal processors that can be used in an audio signal decoder according to an embodiment of the invention. Figure 4f shows a block diagram of a SAOC transcoder processing mode;
Fig. 4g shows a block diagram of a SAOC decoder processing mode;
Fig. 5a shows a block diagram of an audio signal decoder, according to an embodiment of the invention;
Fig. 5b shows a block diagram of another audio decoder, according to an embodiment of the invention;
Fig. 6a is a Table representing a description of the listening test structure;
Fig. 6b is a Table representing the system being tested;
Fig. 6c is a Table representing the listening test elements and rendering matrices;
Fig. 6d shows a graphical representation of the average MUSHRA results for the Karaoke / Solo listening test;
Fig. 6e is a graphical representation of the average MUSHRA results for the classic listening test;
Fig. 7 is a flowchart of a method of providing an upmix signal representation according to an embodiment of the invention;
Fig. 8 is a block diagram of a reference MPEG SAOC system;
Fig. 9a is a block diagram of a reference SAOC system using a separate decoder and mixer;
Fig. 9b is a block diagram of a reference SAOC system using an integrated decoder and mixer;
and Fig. 9c is a block diagram of a reference SAOC system using a SAOC to MPEG transcoder.
A detailed description of the embodiments
1. Audio signal decoder according to Fig. 1 [0056] Fig. 1 is a block diagram of an audio signal decoder 100 according to an embodiment of the invention.
[0057] The audio signal decoder 100 is configured to receive object parametric information 110 and a downmix signal representation 112. The audio signal decoder 100 is configured to provide the upmix signal representation 120 based on the downmix signal representation and the object parametric information 110. The audio signal decoder 100 includes a 130 object separator, which is configured to decompose the downmix signal representation 112 to provide the first audio information 132 describing the first set of one or more audio objects of the first type of audio objects and the second audio information 134 describing the second set of one or more audio objects of the second type of audio objects based on representation of a 112 downmix signal and using at least part of the object parametric information 110. The audio signal decoder 100 also includes an audio signal processor 140 that is configured to receive the second audio information 134 and to process the second audio information based on at least a portion of the object parametric information 112 to obtain the processed version 142 of the second audio information 134. The audio signal decoder 100 also includes an audio signal combining module 150 configured to combine the first audio information 132 with the processed version 142 of the second audio information 134 to obtain an upmix signal representation 120.
[0058] The audio signal decoder 100 performs cascade processing of a downmix signal representation that represents audio objects from the first type of audio objects and audio objects from the second type of audio objects in a combined manner.
[0059] In a first processing step that is performed by the object separator 130, the second audio information describing the second set of audio objects of the second type of audio objects is separated from the first audio information 132 describing the set of audio objects of the first type of audio objects using the object parametric information 110. However, the second audio information 134 typically is audio information (e.g., a single-channel audio signal or a two-channel audio signal) describing audio objects from the second type of audio objects in a combined manner.
[0060] In a second processing step, the audio signal processor 140 processes the second audio information 134 based on the object parametric information. Thus, the audio signal processor 140 is capable of implementing individual object processing or rendering of audio objects of the second type of audio objects that are described by the second audio information 134, which is typically not implemented by the object separator 130.
[0061] In this way, although audio objects from the second type of audio objects are preferably not processed in an individual object way by the object separator 130, audio objects from the second type of audio objects are actually processed in an individual object way (e.g. rendered in an individual object way ) in the second processing step, which is carried out by the audio signal processor 140. Thus, the separation between the audio objects of the first type of audio objects and the audio objects of the second type of audio objects, which is implemented by the object separator 130, is separated from the individual object-oriented processing of audio objects of the second type of audio objects, which is then implemented by the signal processor 140 audio. Accordingly, the processing that is carried out by the object separator 130 is substantially independent of the number of audio objects of the second type of audio objects. In addition, the format (e.g., single-channel audio or two-channel audio) of the second audio information 134 is typically independent of the number of audio objects of the second type of audio objects. Thus, the number of audio objects of the second type of audio objects can change without the need to modify the structure of the 130 object separator. In other words, audio objects from the second type of audio objects are treated as a single (e.g. single-channel or two-channel) audio object for which common object parametric information (e.g., common value of object level difference) is obtained by the object separator 140.
[0062] Accordingly, the audio signal decoder 100 of Fig. 1 is capable of handling a variable number of audio objects of the second type of audio objects without structural modification of the object separator 130. Additionally, different algorithms for processing audio objects by the object separator 130 and the audio signal processor 140 may be used. Accordingly, for example, it is possible to perform audio object separation using residual information through the object separator 130, which enables particularly good separation of different audio objects, using residual information, which is additional information for improving the quality of object separation. In contrast, the audio signal processor 140 may perform individual object-oriented processing without using residual information. For example, the audio signal processor 140 may be configured to perform spatial-audioobject-coding (SAOC) audio signal processing to render various audio objects.
2. Audio signal decoder according to Fig. 2 [0063] Hereinafter, the audio signal decoder 200 according to an embodiment of the invention will be described. A block diagram of this audio decoder 200 is shown in Fig.
2.
[0064] The audio signal decoder 200 is configured to receive downmix signal 210, so-called SAOC bit stream 212, rendering matrix information 214 and optional headrelated-transfer-function parameters (HRTF) 216. The audio signal decoder 200 is also configured to provide the downmix / MPS signal 220 and (optionally) the MPS 222 2.1 bit stream. Input signals and output signals of the audio decoder 200 [0065] Various details regarding the input signals and output signals of the audio decoder 200 will be described below.
[0066] The downmix signal 200 may for example be a single-channel audio signal or a two-channel audio signal. The downmix signal 210 may, for example, be obtained from an encoded representation of the downmix signal.
[0067] The SAOC bit stream 212 may for example contain object-oriented parametric information. For example, the SAOC bit stream 212 may include object level difference information, for example, object OLD parameters, object level difference information, for example, inter object object correlation IOC parameters.
[0068] In addition, the SAOC bit stream 212 may include downmix information describing how the downmix signals have been provided based on the signals of the audio objects using the downmix process. For example, the SAOC bit stream may include DMG downmix gain parameters and (optional) DCLD parameters of downmix channel level differences.
[0069] Rendering matrix information 214 may, for example, describe how different audio objects should be rendered by the audio decoder. For example, rendering matrix information 214 may describe the allocation of audio objects to one or more output / MPS downmix signal channels 220.
[0070] Optional information 216 head transfer function (HRTF) may further describe the transfer function to obtain a binaural headphone signal.
[0071] The output / MPEG Surround downmix signal 220 (also briefly referred to as "the output downmix signal / MPS") represents one or more audio channels, for example in the form of a time domain representation of a audio signal or a frequency domain representation of an audio signal. Alone or in combination with the optional 222 MPEG Surround bit stream (MPS bit stream), which includes MPEG Surround parameters describing the mapping of the 220 downmix output / MPS signal to multiple audio channels, an upmix signal representation is created.
2.2. Structure and functions of the audio decoder 200 [0072] In the following, the structure of the audio signal decoder 200 will be described in more detail, which may act as a SAOC transcoder or a SAOC decoder.
[0073] The audio signal decoder 200 comprises a downmix processor 230 that is configured to receive the downmix signal 210 and to provide an output / MPS downmix signal 220 thereof. The downmix processor 230 is also configured to receive at least part of the SAOC bit stream information 212 and at least part of the information 214 of the rendering matrix. In addition, the downmix processor 230 may also receive 240 processed SAOC parameter information from the 250 parameter processor.
[0074] The parameter processor 250 is configured to receive SAOC bit stream information 212, rendering matrix information 214 and optional head transfer function information information 260 and to provide an MPEG Surround bit stream 222 based thereon containing MPEG Surround parameters (if MPEG Surround parameters are required , for example, in the transcoding mode of operation). In addition, the 250 parameter processor provides processed SAOC information 240 (if processed SAOC information is required).
[0075] The structure and functions of the downmix processor 230 will be described in more detail below.
[0076] The downmix processor 230 includes a residual processor 260 that is configured to receive the downmix signal 210 and to provide on its basis a first signal 262 of an audio object describing so-called enhanced audio objects (EAOs) that can be referred to as first-type audio objects audio. The first audio object signal may contain one or more audio channels and may be considered as the first audio information. The residual processor 260 is also configured to provide a second signal 264 of audio objects, which describes the audio objects of the second type of audio objects and can be considered as the second audio information. The second audio object signal 264 may include one or more channels and typically may include one or more audio channels describing a plurality of audio objects. Typically, the second audio object signal may describe even more than two audio objects from the second type of audio objects.
[0077] Downmix processor 230 also includes a SAOC downmix pre-processor 270 that is configured to receive the second audio object signal 264 and to provide a processed version thereof 272 of the second audio object signal 264 which may be considered the processed version of the second audio information.
[0078] The downmix processor 230 also includes an audio signal combining module 280 that is configured to receive the first audio object signal 262 and the processed version 272 of the second audio object signal 264 and to provide an output / MPS downmix signal 220 thereof that can be recognized. , alone or together with the (optional) MPEG Surround 222 bit stream, for representing the upmix signal.
[0079] The functions of the individual downmix processor 230 units will be discussed in more detail below.
[0080] The residual processor 260 is configured to separately provide the first audio object signal 262 and the second audio object signal 264. To this end, the residual processor 260 may be configured to use at least part of the SAOC bit stream information 212. For example, the residual processor 260 may be configured to evaluate the parametric object information associated with the audio objects of the first type of audio objects, i.e. EAO so-called "enhanced audio objects". In addition, the residual processor 260 may be configured to obtain comprehensive information describing the audio objects from the second type of audio objects, for example the so-called colloquially "enriched audio objects". The residual processor 260 can also be configured to evaluate the residual information, which is provided in the SAOC bit stream information 212, for separation between enriched audio objects (audio objects from the first type of audio objects) and non-enriched audio objects (audio objects from the second type of audio objects) . The residual information may, for example, encode a time domain residual signal that is used to obtain particularly pure separation between enriched audio objects and unenriched audio objects. In addition, the residual processor 260 may optionally evaluate at least a portion of the information 214 of the rendering matrix, for example, to determine the distribution of enriched audio objects in the audio channels of the first signal 262 of audio objects, [0081] The SAOC downmix pre-processor 270 includes a channel redistribution module 274 which is configured to receiving one or more audio channels from the other 264 audio objects and for delivery based thereon, one or more (typically two) audio channels of the processed second signal 272 of the audio object. In addition, the SAOC downmix pre-processor 270 includes a de-correlated signal provider 276, which is configured to receive one or more audio channels of a second signal 264 audio objects and to provide one or more de-correlated signals 278a, 278b based thereon which are added to the signals supplied through the channel redistribution module 274 to obtain the processed version 272 of the second audio signal 264.
[0082] Further details regarding the SAOC downmix processor will be discussed below. [0083] The audio signal combining module 280 combines the first audio object signal 262 with the processed version 272 of the second audio object signal. To this end, channel bonding can be carried out. Accordingly, an output / MPS downmix signal 220 is obtained.
[0084] The parameter processor 250 is configured to obtain (optional) MPEG Surround parameters, which are the MPEG Surround bit stream 222 of the upmix signal representation based on the SAOC bit stream, including rendering matrix information 214 and optional HRTF parameter information 216. In other words, the SAOC parameter processor 252 is configured to transform object parameter information, which is described by SAOC bit stream information 212 into channel parametric information, which is described by the MPEG Surround 222 bit stream.
[0085] The following is a brief overview of the structure of the SAOC transcoder / decoder architecture shown in Fig. 2. Spatial audio object coding (SAOC) is a parametric multi-object coding technique. It is designed to transmit a number of audio objects in an audio signal (e.g., audio downmix 210 signal) that contains M channels. Together with the backward compatible downmix signal, object parameters (e.g., using SAOC bit stream information 212) that are able to reconstruct and manipulate the original object signals are transmitted. The SAOC encoder (not shown here) produces a downmix of object signals at its input and acquires these object parameters. The number of objects that can be served is, in principle, unlimited. Object parameters are quantized and coded efficiently into the SAOC 212 bit stream. Downmix signal 210 can be compressed and sent without the need to update existing encoders and infrastructures. The object parameters or SAOC decoder information are transmitted in an additional low bit rate channel, e.g., in the auxiliary part of the downmix bit stream data.
[0086] On the decoder side, the input objects are reconstructed and rendered into a number of playback channels. Rendering information containing the reconstruction level and panoramic position for each object is provided by the user or can be obtained from the SAOC bit stream (for example, as pre-determined information). The rendering information may be variable over time. Output circuits can range from mono to multi-channel (e.g. 5.1) and are independent of both the number of input objects and the number of downmix channels. Binaural rendering of objects is possible, including azimuth and elevation of virtual objects. The optional effects interface allows advanced manipulation of object signals, in addition to level and panorama modifications.
[0087] The objects themselves may be mono signals, stereo signals as well as multi-channel signals (e.g., 5.1 channels). Typical downmix configurations are mono and stereo. [0088] The basic structure of the SAOC transcoder / decoder will be explained below, which is shown in Fig. 2. The SAOC transcoder / decoder module discussed here can act either as a stand-alone decoder or as a transcoder from SAOC to the MPEG Surround bit stream, depending on the target configuration output channel. In the first mode of operation, the output signal configuration is mono, stereo or binaural configuration and two output channels are used. In the first case, the SAOC module can operate in the decoder mode, and the output of the SAOC module is the output in pulse-code modulation (PCM output). In the first case, no MPEG Surround decoder is required. Instead, the upmix signal representation may contain only the output signal 220 and provision of the MPEG Surround 222 bit stream may be omitted. In the second case, the output configuration is a multi-channel configuration with more than two output channels. The SAOC module can operate in transcoder mode. The SAOC module output signal may include both downmix signal 220 and MPEG Surround 222 bit stream as shown in Fig. 2. Accordingly, an MPEG Surround decoder is necessary to obtain the final representation of the audio signal for playback through the speakers.
[0089] Fig. 2 shows the basic structure of the SAOC transcoder / decoder architecture. Residual processor 216 obtains enriched audio objects from the downmix input 210 using the residual information contained in the SAOC 212 bit stream. The downmix pre-processor 270 processes regular audio objects (which are, for example, non-enriched audio objects, i.e. no audio objects for which no residual information in the bit stream SAOC 212). Enriched audio objects (represented by the first audio object signal 262) and processed regular audio objects (represented for example by the processed version 272 of the second audio object signal 264) are combined into output signal 220 for SAOC decoder mode or 220 MPEG Surround downmix signal for mode SAOC transcoder. A detailed description of the processing blocks is given below.
3. Architecture and functions of the residual processor and energy mode processor [0090] The details of the residual processor that, for example, can take over the function of the separator 130 of the audio decoder objects 100 or the residual processor 260 of the audio decoder 200 will be discussed below. To this end, Figures 3a and 3b show block diagrams of such a residual processor 300 that can take the place of the object separator 130 or the residual processor 260. 3a shows less details than Fig. 3b. However, the following description applies to the residual processor 300 according to Fig. 3a as well as the residual processor 380 according to Fig. 3b.
[0091] The residual processor 300 is configured to receive the SAOC downmix signal 300, which may be equivalent to the downmix signal representation 112 of Fig. 1 or the downmix signal representation 210 of Fig. 2. The residual processor 300 is configured to provide first information based thereon. audio 320 describing one or more enhanced audio objects, which may be, for example, equivalent to the first audio information 132 or the second signal 262 of the audio object. Also, the residual processor 300 may provide the second audio information 322 describing one or more other audio objects (e.g., non-enriched audio objects for which no residual information is available), wherein the second audio information 322 may be equivalent to the second audio information 134 or with the second signal 264 of the audio object.
[0092] The residual processor 300 includes a 1-to-N / 2-to-N 330 unit (OTN / TTN unit) that receives the SAOC downmix signal 310 and which also receives SAOC 332 data and debris. The 330 1-to-N unit / 2-to-N also provides a 334 enriched audio object signal that describes the enriched audio objects (EAOs) contained in the SAOC downmix 310 signal. Also, the 330 1-to-N / 2-to-N unit provides second audio information 322. The residual processor 300 also includes a rendering unit 340 that receives the signal 334 of the enhanced audio objects and the rendering matrix information 342 and provides the first audio information 320 thereon.
[0093] In the following, the processing of enriched audio objects (EAO processing) that is implemented by the residual processor 300 will be described in more detail.
3.1. Introduction to operation of the residual processor 300 [0094] When considering the functionality of the residual processor 300, it should be noted that SAOC technology allows the individual manipulation of a number of audio objects in terms of their gain / attenuation without significantly compromising the final sound quality, only in a very limited way. The special "karaoke" application mode requires complete (or almost total) suppression of certain objects, typically the lead vocal, while maintaining undisturbed perceptual quality of the background sound.
[0095] A typical use case includes up to four Enriched Audio Object (EAO) signals, which may, for example, represent two independent stereo objects (e.g. two independent stereo objects that are prepared for their removal on the decoder side).
[0096] It should be noted that one or more qualitatively enriched audio objects (or more precisely, audio signal inputs associated with the enriched audio objects) is included in the SAOC downmix signal 310. Typically, audio signal cartridges associated with one or more enriched audio objects are mixed in downmix processing implemented by the audio signal encoder with the audio signal cartridges of other audio objects that are not enriched audio objects. Also, it should be noted that audio signal inputs from many enriched audio objects are also typically superimposed or mixed by downmix processing implemented by the audio signal encoder.
3.2 SAOC architecture supporting enriched audio objects [0097] Details of the residual processor 300 will be discussed below. Enriched audio object processing includes 1-to-N or 2-to-N units depending on the SAOC downmix mode. The 1-to-N processing unit is dedicated to the mono downmix signal and the 2-to-N processing unit is dedicated to the signal
310 stereo downmix. Both of these units represent the generalized and enriched modification of the 2-to-2 block (TTT block) known from the ISO / IEC 23003-1: 2007 standard. In the encoder, regular signals and EAO signals are combined into a downmix signal. OTN processing units<sup>-1</sup>/ TTN<sup>-1</sup> (which are inverted 1-to-N processing units or inverted 2-to-N processing units) are used to generate and encode the corresponding residual signals.
[0098] EAO signals and regular signals are reconstructed from the downmix signal 310 by 330 OTN / TTN units using the SAOC additional information and the residual signals contained. The reconstructed EAOs (which are described by the signal 334 of the enriched audio objects) are provided to the rendering unit 340 which represents (or provides) the product of the respective rendering matrix (described by the rendering matrix information 342) and the final output signal from the OTN / TTN unit. Regular audio objects (which are described by the second audio information 322) are provided to the SAOC downmix pre-processor, e.g., SAOC downmix pre-processor 270, for further processing. Figures 3a and 3b show the general structure of the residual processor, i.e. the architecture of the residual processor.
[0099] Output signals 320, 322 of the residual processor are calculated as ^ os ~ ^ EAO ^ -EAO ^ EAO ^ on [0100] where XOBJ represents the downmix signal of regular audio objects (i.e. not EAO) and XEAO is the rendered EAO output signal for SAOC decoding mode or the appropriate EAO downmix signal for SAOC transcoding mode.
[0101] The residual processor may operate in a prediction mode (using residual information) or in an energy mode (without residual information). The Xres extended input signal is defined as follows:
x ~ =
Ϊ<sup>Χ</sup>Ί <a name="caption1"></a>Μ
χ.
for the prediction mode for the energy mode.
[0102] In this case, X may, for example, represent one or more channels of a downmix signal representation 310, which may be carried in a bit stream representing multi-channel audio content. res can mean one or more residual signals that can be described by a bit stream representing multi-channel audio content.
[0103] OTN / TTN processing is represented by the M matrix and the EAO processor by the AEAO matrix.
[0104] The OTN / TTN processing M matrix is defined according to the EAO mode of operation (i.e. prediction or energy) as <sup>for trVbu predVkcj</sup>'
M
L Fhety 'for energy mode.
[0105] The OTN / TTN processing M matrix is represented as
<img file="PL2535892T3_D0001.tif" />
where the MOBJ matrix applies to regular audio objects (i.e. not EAO) and the MEAO matrix relates to enriched audio objects (EAO).
[0106] In some embodiments, one or more multi-channel background objects (MBOs) can be treated in the same way by the residual processor 300.
[0107] The Multi-channel Background Object (MBO) is a MPS mono or stereo downmix that is part of the SAOC downmix. Unlike the use of individual SAOC objects for each channel in a multi-channel signal, MBO can be used enabling SAOC to more efficiently manipulate a multi-channel object. In the case of MBO, the SAOC overhead decreases when the SAOC parameters of the MBO are associated only with downmix channels rather than all upmix channels.
3.3 Additional definitions
3.3.1 Dimensioning of signals and parameters [0108] The dimensions of the signals and parameters will be briefly discussed below to explain how often various calculations are performed.
[0109] Audio signals are defined for each time interval n and each hybrid subband k (which may be a frequency subband). The corresponding SAOC parameters are defined for each time interval of 1 parameter and m processing band. The following mapping between hybrid and parameter domain is specified in Table A. 31 ISO / IEC 23003-1: 2007. Hence, all calculations are performed for certain time / band indices and the corresponding dimensions are implied for each variable entered.
[0110] However, below time and frequency band indices will sometimes be omitted to maintain notation compactness.
3.3.2 Calculation of the AEAO matrix [0111] The AEAO pre-rendering matrix of EAO objects is defined according to the number of output channels (i.e. mono, stereo or binaural) as for mono case, A / °, for other cases
ILO EAO Α<sup>Ελ0</sup> with size 1 x N ^ and with size 2 x N; - are defined as »A-> fOiur £ AO
-(
IN,
GAC in:
, £ XO in, £ 4ϋ
IN,
EAO vp:
EAO) ·
AEAO <sub>n</sub>EDA EDA% *
J = <sup>D</sup>26 <sup>M</sup>™
<img file="PL2535892T3_D0002.tif" />
Λ / where the rendering sub matrix corresponds to the rendering of the EAO (and describes the desirable mapping of enhanced audio objects to upmix signal representation channels).
[0113] Values <sup>υ</sup>. ' are calculated depending on the rendering information associated with the enriched audio objects using the appropriate EAO elements and using the equations in section 4.2.2.1.
[0114] For binaural rendering, the matrix is defined by the equations given in section 4.1.2 for which the respective target binaural rendering matrix only contains elements associated with EAOs.
3.4 Calculation of OTN / TTN elements in residual mode [0115] Below we will discuss how the SAOC downmix signal 310, which typically contains one or two audio channels, is mapped to a 334 enriched audio object signal, which typically contains one or more enriched audio object channels and to the second audio information 322, which typically includes one or two channels of regular audio objects.
[0116] The function of the 1-to-N or 2-to-N unit 330 may, for example, be performed using a matrix vector product such that a vector describing both the signal channels 334 of the enriched audio objects and the channels of the second audio information 322 is obtained by multiplying the vector describing the SAOC downmix signal channels 310 and (optional) one or more residual signals by an MPrediction or MEnergy matrix. Accordingly, the determination of the MPrediction or MEnergy matrix is an important step in obtaining the first audio information 320 and the second audio information 322 from the SAOC 310 downmix.
[0117] In summary, the OTN / TTN upmix process is represented either by the MPrediction matrix for the prediction mode or by the MEnergy matrix for the energy mode.
[0118] The energy based coding / decoding procedure is designed for encoding a downmix signal not maintaining a waveform. Therefore, the OTN / TTN upmix matrix for the respective energy mode is not based on specific waveforms, but only describes the relative energy distribution of the input audio objects, as will be described in more detail below.
3.4.1 Prediction mode [0119] In prediction mode, the MPrediction matrix is defined using the downmix information contained in the t matrix<sup>1</sup>'and CPC data from matrix C:
<img file="PL2535892T3_D0003.tif" />
[0120] For several SAOC modes, the extended downmix D0 matrix and CPC C matrix have the following dimensions and structures:
3.4.1.1 Stereo downmix (TTN) modes:
[0121] For stereo downmix (TTN) modes (for example, in the case of a stereo downmix based on two channels of regular audio objects and NEAO channels of enhanced audio objects), (extended) matrix l<sup>J</sup> downmix and C CPC matrix can be obtained as follows:
<img file="PL2535892T3_D0004.tif" />
<td></td><td>fl 0</td><td> 0 1</td><td>oo 1</td><td>1 oo _</td>
<td>c =</td><td><sup>C</sup>OJO</td><td>au</td><td>! and ... 1</td><td> 0</td>
<td></td><td></td><td></td><td><sup>1</sup> : *„ 1 * -</td><td> *</td>
<td></td><td></td><td>"Łrtj-h</td><td>JO ...</td><td><sup>1</sup></td>
[0122] In the case of a stereo downmix, each EAOj contains two CPC cj, 0 and cj, 1 giving a matrix C.
[0123] The residual processor output signals are calculated as
<img file="PL2535892T3_D0005.tif" />
<img file="PL2535892T3_D0006.tif" />
[0124] Thus, two signals yL, yR (which are represented by XOBJ) are obtained that represent one or two or even more than two regular audio objects (also referred to as non-extended audio objects). Also, NEAO signals (represented by XEAO) are obtained representing the enriched NEAO audio objects. These signals are obtained on the basis of two SAOC downmix signals l0, r0 and NEAO residual signals res0 to resNEAO-1, which will be encoded in the SAOC additional information, for example, as part of the object parametric information.
[0125] It should be noted that signals yL and yR may be equivalent to signal 322 and that signals Y0, EAO to YNEAO-1, EAO (which represent XEAO) may be equivalent to signals
320.
[0126] Matrix A<sup>EDA</sup> is the rendering matrix. Matrix words<sup>EDA</sup> they may describe, for example, the mapping of enhanced audio objects to 334 enhanced audio object (XEAO) signal channels.
[0127] Accordingly, the correct choice of matrix A<sup>EDA</sup> enables the optional integration of the rendering unit function 340 in such a way that the multiplication of the vector describing the channels (l0, r0) of the SAOC downmix signal 310 and one or more A £ xO M / ł Pr Edfc / łtfn residual signals (res<sub>0</sub>, ..., res<sub>NEAO-1</sub>) by the matrix can directly create an XEAO representation of the first audio information 320.
3.4.1.2 Mono downmix modes (OTN):
[0128] The following will acquire the signals 320 of enriched audio objects (or alternatively signals 334 of enriched audio objects) and the signal 322 of regular audio objects for the case in which the SAOC downmix signal 310 contains only the signal channel.
[0129] For mono downmix (OTN) modes (e.g. mono downmix based on one channel of regular audio objects and channels of enhanced NEAO audio objects) (extended) matrix I<sup>J</sup> downmix and C CPC matrix can be obtained as follows:
<img file="PL2535892T3_D0007.tif" />
<img file="PL2535892T3_D0008.tif" />
ο ϊ
ο [0130] For mono downmix, one EAOj is predicted by one factor cj giving matrix C. All elements of the matrix cj are obtained, for example, from SAOC parameters (for example from SAOC 322 data) according to the relationship given below (section 3.4.1.4 ).
[0131] The residual processor output signals are calculated as
<img file="PL2535892T3_D0009.tif" />
<img file="PL2535892T3_D0010.tif" />
[0132] The XOBJ output signal includes, for example, one channel describing regular audio objects (non-enriched audio objects). The XEAO output signal includes, for example, one, two or even more channels describing enriched audio signals (preferably NEAO channels describing enriched audio objects). Again, said signals are equivalent to signals 320, 322.
3.4.1.3 Calculation of the inverse extended downmix matrix [0133] Matrix I> is the inversion of matrix I <sup>J</sup>downmix, and C implies CPC.
[0134] Matrix IJ is an inversion of matrix I <sup>J</sup>downmix and can be calculated as
<img file="PL2535892T3_D0011.tif" />
[0135] Elements. (for example, inversions b of a 6x6 extended b downmix matrix) are obtained using the following values:
/=>
<img file="PL2535892T3_D0012.tif" />
Σ> Λ ·
U * 'J d<sub>l?</sub> = m, + + m, n<sub>4</sub> - m<sub>2</sub>NN-<sub>2</sub> - <sub>s</sub> d<sub>14</sub> - wij + m<sub>2</sub>r ^ + m<sub>2</sub>n<sub>2</sub> + m<sub>2</sub>n<sup>2</sup>4 -m ^ n, - - tn ^ n,., ^ 1.5 ~ <sup>m</sup>s<sup>+</sup> + <sup>m</sup>3<sup>n</sup>2 '+ <sup>m</sup>3<sup>n</sup>4 ~ W | «3« l “WjBjMj - Ι«<sub>4</sub>^ / Ϊ<sub>4</sub> , = W<sub>4</sub> + m<sub>4</sub>rf + m<sub>4</sub>iTj + m<sub>4</sub>n<sup>2</sup> , <*<sub>in</sub> = i + 2X '
7-1 j <sup>=</sup> «Ι + n, ml +«, 7n<sub>4</sub> - γμ, λϊ ^ -0ϊ, ηί<sub>4</sub>π<sub>4ι</sub> c /<sub>3 4</sub> = Hj + n ^ tnj + / jj / Kj + «jj / ę - ^ 01,«, - -m, fn.n. <sub>f</sub> d<sub>it</sub> -n<sub>}</sub>+ n<sub>3</sub>m? + njiiij + n<sub>3</sub>m<sup>2</sup>a - - m<sub>3</sub>m<sub>4</sub>n<sub>4</sub>, ^ 2, t = <sup>n</sup>A + <sup>rt</sup>and<sup>tn</sup>\ <sup>+</sup> «4 ^ 2 + W<sup>1</sup>! - - «W».<sup>44</sup> rf<sub>M</sub> = -1 - £ m<sup>2</sup> - £ / »'-« X - mfó -nfó -<sup>m</sup>4 ^ -m & i - <sup>m</sup>and'<sup>n</sup>l + 2 / η, / ιζ, π, π, + 2 / β, / β, β, ιι, + ζπ ^ η, η, / -i 7, dj<sub>4</sub> =/^/^+/1,^+^/1,/7, + //7(77,77, +//^/¾<sup>2</sup> +/^7/1,/7,<sup>2</sup> -τ / τ, / ττ, / τ, η, - / η / Η, Η, η, -m / n ^ n, -m / n ^ n ,,
4<sub>3S</sub> = / «, /«, + / 7, ^ + / 17 (/ 7, / ¾ + Zn<sub>4</sub>w, / i, + / n, / n, / ij + / 7i, 7w, 7i (-τ / ι, / η, / ι, / ι, -τη, / π, / ι, η, -τ / ι, / τ /, / ι, / ι, - // 1, ///, / 1, / 1 ,,
J <sub>6</sub> =771,/71, + 71,71, +/71(/1,71, + /7^71,7¾ + 7/1,7/7,71( + 771,771,7¾<sup>2</sup> - TT!, / », /!, / !, - / ^ 771, / 1, ^ - / 71.7 / 1, / 1, / 1, - 7/1, // ¾ / 1, / ¾, _ <sup>4 4</sup> d<sub>t</sub>, = -1 - £ m (- Y / i (-mtf - // 1 (/ 1 (-771 (7¾<sup>2 _m</sup>X - ^ (+ ^ η, ηπ ^ + 2 / η, / π, / ι, / ι, + 2 / η, / π, / ι, / ι ,,
J ,, = m<sub>2</sub>m<sub>J</sub>+ n<sub>2</sub>n<sub>l</sub> + m<sub>l</sub><sup>2</sup>n<sub>2</sub>n<sub>s</sub>+ mtn<sub>J</sub>n<sub>)</sub>+ m<sub>2</sub>m<sub>l</sub>n (+ / ^ / ^ / 1 (-ιη, ιιι, ηπ, - / ąmtfn, - / η, / η, / ι, / ι, - / η, / η, η, / ι ,, d<sub>t6</sub>-m<sub>2</sub>m<sub>4</sub> + R ^ n<sub>4</sub> +/11(/1,/1,+/7^71,7¾ + /11,//1,/¾<sup>2</sup> +//1,//1,/¾<sup>2</sup> - // 1, // 1, / 1, / 3, - / η, τ / ι, / ι, η, - / η, / η, / ι, / ι, -τη, / η, η, / ι ,,
J<sub>ss</sub> = -l- £ // l (~ Σ<sup>η</sup>* ~<sup>m</sup>2<sup>n</sup>^~<sup>m</sup>and<sup>n</sup>and<sup>2 _ / N</sup>|<sup>2</sup>”5 -<sup>m</sup>4^<sup>-Λ</sup>ί<sup>η</sup>«~ * & Ί + 2 // 1, // 1, / 11/1, + 2 // 1,771, / 1, / 1, + 2/71, ^ 1,71,7¾.
<Z <sub>6</sub> = / n, w, + «,«, + rfn, n<sub>t</sub> + nĄiyi, + m / nrf + τη, / η, η,<sup>2</sup> - τη, / π, π, / ¾ - τη, / η, η, η, - τη, τη, η, η, - τ / ι, / η, / ι, / ι ,, d<sub>66</sub> = —1 - £ // ι,<sup>2</sup> - Υ / ι (- »& -" tó ~ "^<sup>η</sup>3 -μ-ιΊ] <sup>+</sup> ^<sup>m</sup>and<sup>m</sup>and<sup>n</sup>and<sup>n</sup>and<sup>+</sup> + 2 / π, / η, π, η ,,> 1 j-1
4 den = i + £ m<sup>2</sup> + £ η<sup>2</sup> + / η (/ ι (+ / η (/ ι (+ / ^ / ¾<sup>2</sup> + τη (/ ι (+ +171 (7¾<sup>2</sup> + // 1 (/ 1 (+ / n (n (+ m} / ¾<sup>2</sup> + hand + // 1 (/ 1 (+ i'ij = i + / w £ / 7<sup>2</sup> -2/71,^,/1,7¾ -2//1,//1,/7,/¾ -2//1,7/7,^7¾ -2/7^771,11,71, -2/«,/«,«,/7, - 2/^//1,/¾^.
[0136] Coefficients mj and nj of the extended matrix l<sup>J</sup> downmix means the downmix values for each EAOj for the right and left downmix channels as [0137] The downmix di, j matrix elements D are obtained using DMG downmix gain information and (optional) DCLD information of the downmix channel level information that is included in the SAOC 332 information which is represented, for example, by parameter object information 110 or SAOC bit stream information 212.
[0138] In the case of a stereo downmix, a 2xN downmix D matrix with elements d<sub>and</sub>, j (i = 0,1; j = 0, ..., N -1) is obtained from the parameters DMG and DCLD as
<img file="PL2535892T3_D0013.tif" />
<img file="PL2535892T3_D0014.tif" />
[0139] For the downmix mono D downmix matrix with 1xN dimensions with di, j elements (i = 0; j = 0, ..., N -1) is obtained from DMG parameters as
<img file="PL2535892T3_D0015.tif" />
O. OJJSUK [0140] In this case, the dequantized DMGj and DCLDj parameters of the downmix are obtained, for example, from parametric additional information 110 or from the SAOC 212 bit stream.
[0141] The EAO (j) function calculates the mapping between the input channel indexes of the audio objects and the EAO signals:
EAO (j) = N-1-j, j = 0, ^, N<sub>E</sub>AND<sub>ABOUT</sub>-1
3.4.1.4 Calculation of matrix C [0142] Matrix C implies CPCs and is obtained from transmitted SAOC parameters (e.g. oLDs, IOCS, DMGs and DCLDS) as
<img file="PL2535892T3_D0016.tif" />
<img file="PL2535892T3_D0017.tif" />
[0143] In other words, limited CPCs are obtained according to the above equations, which can be considered as a limitation algorithm. However, limited CPCs can also be derived from values.<sup>:</sup> ... and with a different containment approach (containment algorithm) or can be set equal to the values'<sup>:</sup> '.
[0144] It should be noted that the words cj, 1 matrix (and direct quantities on the basis of which the words cj, 1 matrix are calculated) are typically only required if the downmix signal is a stereo downmix signal.
[0145] CPCs are limited by another limiting function:
<img file="PL2535892T3_D0018.tif" />
<img file="PL2535892T3_D0019.tif" />
with a weighting factor λ determined as
<img file="PL2535892T3_D0020.tif" />
[0146] For one specific EAO channel j = 0 ... NEAO -1, unlimited CPCs are estimated by <sup>c</sup>yo /> P - PP
Ro HoCoJ<sup>1</sup> loro
PP - P<sup>2</sup><sup>£</sup> Lr> Ro <sup>r</sup>loro
<img file="PL2535892T3_D0021.tif" />
[0147] The energy amounts PLo, PRo, PLoRo, PLoCo, j and PRoCo, j are calculated as ^ = 010, + £ £ <sup>m</sup>and<sup>m</sup>t<sup>e</sup>n rt-eO <sup>n</sup>fao ~ \ / «O <sup>=</sup> OLDr + z> <sup>n</sup>j<sup>n</sup>k<sup>e</sup>j, k * j-0 Ł = 0 ^ Ł10 "<sup>1 IN</sup>C4O<sup>-</sup>’
PloRo <sup>= e</sup>L, n <sup>+</sup> /. from,<sup>m</sup>j<sup>n</sup>ifijjt> y = o * ° o = ^<sup>£</sup>AND<sup>+</sup>V '.<sub>n</sub>^<sup>m</sup>,<sup>O £ D</sup>j "Σ / 0
<img file="PL2535892T3_D0022.tif" />
[0148] The covariance matrix ei, j is defined as follows: The covariance matrix E with the dimension NxN with elements ei, j represents the approximation of the covariance matrix E ~ SS * of the original signal and is obtained from OLD and IOC as
<img file="PL2535892T3_D0023.tif" />
[0149] In this case, the dequantized parameters of the OLDi, IOCij objects are obtained, for example, from parametric additional information 110 or from the SAOC 212 bit stream. [0150] In addition, eL, R can, for example, be obtained as
<img file="PL2535892T3_D0024.tif" />
[0151] The OLDL, OLDR and IOCL, R parameters correspond to regular (audio) objects and can be obtained using downmix information:
OLD<sub>l</sub> = £ dl, OLD, f = 0
<img file="PL2535892T3_D0025.tif" />
roc<sub>L</sub>"= / OC", "NN<sub>M</sub> = 2, otherwise.
[0152] As can be seen, two common values for the difference in the level of OLDL and OLDR objects are calculated for regular audio objects in the case of a stereo downmix signal (which preferably implies a two-channel signal of regular audio objects). In contrast, only one common OLDL of the object level difference is calculated for regular audio objects in the case of a single-channel (mono) downmix signal (which preferably implies a single-channel signal of regular audio objects).
[0153] As can be seen, the first (in the case of a two-channel downmix signal) or the only (in the case of a single-channel downmix signal) the common OLDL value of the object level difference is obtained by adding the contributions of regular audio objects having the index (or indexes) of the audio object and, to the left channel (or only channel) of the SAOC downmix signal 310.
[0154] A second common OLDR value of the object level difference (which is used in the case of a two-channel downmix signal) is obtained by summing the contributions of regular audio objects having the index (or indexes) of the audio object and, to the right channel of the SAOC downmix signal 310.
[0155] OLDL input of regular audio objects (having audio object indexes i = 0 to i = N-NEAO-1) to the left channel signal (or only channel signal) of the SAOC downmix, signal 710 is calculated, for example, taking into account d0 gain, and downmix describing the downmix gain used for a regular audio object having an audio object index and when the left channel signal 310 of the SAOC downmix signal is obtained as well as the level of a regular audio object having an audio object index and which is represented by the OLDi value.
[0156] Similarly, a common OLDR value of the object level difference is obtained using downmix d1 coefficients, and describing the downmix gain that is applied to a regular audio object having an audio object index and when the right downmix signal channel SAOC 310 and level information is created OLDi associated with a regular audio object having an audio object index i.
[0157] As can be seen, the equations for calculating PLo, PRo, PLoRo, PLoCo, and J and PRoCo, j do not distinguish between individual regular audio objects, but only use common OLDL, OLDR values of the object level difference, thus treating regular audio objects (having indexes audio object i) as a single audio object.
[0158] Also, the IOCL, R value of the inter-object correlation, which is associated with regular audio objects, is set to 0 unless there are two regular audio objects.
[0159] The matrix ei, j (and eL, R) covariance is defined as follows:
[0160] N-N covariance matrix E with elements ei, j represents the approximation of the E ~ SS * matrix of the original signal covariance and is obtained from OLD and IOC as
<img file="PL2535892T3_D0026.tif" />
[0161] For example
<img file="PL2535892T3_D0027.tif" />
where OLDL and OLDR and IOCL, R are calculated as described above.
[0162] In this case, the dequantized object parameters are obtained as
<img file="PL2535892T3_D0028.tif" />
<img file="PL2535892T3_D0029.tif" />
where DOLD and DIOC are matrices containing object level difference parameters and inter-object correlation parameters.
3.4.2 Energy mode [0163] A different concept will be described below that can be used to separate the signals 320 of extended audio objects and signals 322 of regular audio objects (non-extended audio objects), and which can be used in conjunction with non-waveform audio coding SAOC downmix 310 channels.
[0164] In other words, the energy-based coding / decoding procedure is designed for a non-observing waveform coding of a downmix signal. Thus, the OTN / TTN upmix matrix for the appropriate energy mode does not depend on specific waveforms, but only describes the relative energy distribution in the input audio objects.
[0165] Also, the concept discussed here, which is designed as an "energy mode" concept, can be used without transmitting the residual signal information. Again, regular audio objects (non-enriched audio objects) are treated as a single single-channel or two-channel audio object having one or two common values OLDL, OLDR of object level difference.
[0166] For the energy mode, the MEnergy matrix is defined using downmix and OLD information, as will be described below.
3.4.2.1 Energy mode for stereo downmix (TTN) modes [0167] For stereo (e.g. stereo downmix based on two regular object channels and NEAO channels of enriched audio objects) matrices are obtained from the corresponding OLD according to
<img file="PL2535892T3_D0030.tif" />
<img file="PL2535892T3_D0031.tif" />
[0168] The residual processor output signals are calculated as
OBJ and '
<img file="PL2535892T3_D0032.tif" />
[0169] The signals yL, yR, which are represented by the XOBJ signal, describe regular audio objects (and can be equivalent to 322 signals), and the signals y0, EAO to yNEAO-1, EAO, which are described by the XEAO signal, enhanced audio objects (and can be equivalent to 334 signals or 320 signals).
[0170] If a mono upmix signal is desired for a stereo downmix signal case, 2-to-1 processing may be performed, for example, by a pre-processor 270 based on a two-channel XOBJ signal.
3.4.2.2 Energy mode for mono downmix (OTN) modes [0171] For mono (e.g. mono downmix based on one channel of regular audio objects and NEAO channels of enriched audio objects) matrices and are obtained from the corresponding OLD<sub>S</sub> according to
<img file="PL2535892T3_D0033.tif" />
<img file="PL2535892T3_D0034.tif" />
[0172] The residual processor output signals are calculated as [0173] A single channel 322 of regular audio objects (represented by XOBJ) and 10 NEAO channels 320 of enriched audio objects (represented by XEAO) may
BACKGROUND & iffEy be obtained by using a matrix and for the representation of the single channel 310 SAOC downmix signal (represented here by to).
[0174] If a two-channel (stereo) upmix signal is desired for a single-channel (mono) downmix signal, 1-to-2 processing may be performed, e.g., by a pre-processor 270 based on a single-channel XOBJ signal.
4. Architecture and operation of the SAOC downmix pre-processor [0175] In the following, the operation of the SAOC downmix pre-processor 270 will be described for both certain decoding modes and for certain transcoding modes.
4.1 Operation in decoding modes
4.1.1 Introduction [0176] The following describes how to obtain an output using SAOC parameters and panorama information (or rendering information) associated with each audio object. The SAOC 495 decoder is shown in Fig. 4g and consists of a SAOC parameter processor 495 and a downmix processor 497.
[0177] It should be noted that the SAOC decoder 494 can be used to process regular audio objects and therefore can receive, as downmix signal 497a, a second audio object signal 264 or a regular audio object signal 322 or a second audio information 134. Accordingly, processor 497 downmix may provide, as output signals 497b, the processed version 272 of the second signal 264 audio objects or the processed version 142 of the second audio information 134. Thus, the downmix processor 497 may act as a SAOC downmix pre-processor 270, or as an audio signal processor 140.
[0178] The SAOC parameters processor 496 may act as the SAOC parameters processor 252 and as a result provide downmix information 496a.
4.1.2 Downmix processor [0179] Below, the downmix processor which is part of the audio signal processor 140 and which is referred to as "SAOC downmix pre-processor 270" in the embodiment of Fig. 2 and which is designated 497 in the SAOC decoder will be described in more detail below. 495.
[0180] For the SAOC system decoder mode, the output signal 142, 272, 497b of the downmix processor (represented in the hybrid QMF domain) is provided to the appropriate synthesis filter bank (not shown in Figs. 1 and 2) as described in the ISO / IEC 23003 standard -1: 2007 generating the final PCM output signal. Regardless, the output signal 142, 272, 497b of the downmix processor is typically combined with one or more audio signals 132, 262 representing enhanced audio objects. This connection can be made before the appropriate synthesis filter bank (so that the combined signal combining the downmix processor output signal and one or more signals representing enriched audio objects is fed into the synthesis filter bank). Alternatively, the downmix processor output may be combined with one or more audio signals representing enriched audio objects only after processing in the synthesis filter bank. Accordingly, the upmix signal representation 120, 220 may be either a QMF domain representation or a PCM domain representation (or any other suitable representation). Downmix processing includes, for example, mono processing, stereo processing and, if desired, subsequent binaural processing.
[0181] The output X signal of the downmix pre-processor 270, 497 (also denoted 142, 272, 497b) is calculated from the X signal of the mono downmix (also denoted 134, 262, 497a) and the de-correlated Xd signal of the downmix as
<img file="PL2535892T3_D0035.tif" />
[0182] The de-correlated mono downmix signal Xd is calculated as
X<sub>d</sub> = decorrFuncllC} [0183] The de-correlated Xd signals are created in the decorrelator described in ISO / IEC 23003-1: 2007, subclause 6.6.2. According to this scheme, the configuration bsDecorrConfig == 0 with the decorrelator index X = 8 in accordance with Table A.26 to Table A.29 in ISO / IEC 23003-1: 2007 should be used. Hence decorrFunc () means the process of decorrelation:
(x<sub>d</sub> =
<img file="PL2535892T3_D0036.tif" />
θ) Ρ, Χ / [0184] For binaural output signal, upmix G and P2 parameters obtained from SAOC data, rendering information and HRTF parameters are used in the downmix X signal (and X<sub>d</sub>) giving the binaural output signal X, see Fig. 2, reference numeral 270, where the basic structure of the downmix processor is shown.
[0185] Target matrix A<sup>l, m</sup> 2xN binaural rendering consists of
And a '' *. Each element is obtained from HRTF parameters and from the matrix
M '· ™ m<sup>m</sup> rendering with elements for example by the SAOC parameter processor.
Target matrix A<sup>l, m</sup> binaural rendering represents the relationship between all audio input objects y and the desired binaural output signal.
<img file="PL2535892T3_D0037.tif" />
<img file="PL2535892T3_D0038.tif" />
ττβτ [0186] The HRTF parameters are given by and for each processing band m. The spatial positions for which HRTF parameters are available are indicated by the index i. These parameters are described in ISO / IEC 23003-1: 2007.
4.1.2.1 Overview [0187] The following will outline the downmix processing in general with reference to Figs. 4a and 4b, which are a block diagram of the downmix processing that can be implemented by an audio signal processor 140 or by a combination of a processor 252 SAOC parameters and a pre-processor 270 SAOC downmix, or by a combination of a SAOC downmix 496 processor and a 497 downmix processor. [0188] Referring now to Fig. 4a, downmix processing receives M rendering matrix, object level difference OLD information, cross object correlation IOC information, downmix gain DMG information and (optional) DCLD downmix channel level difference information. The downmix processing 400 of Fig. 4a obtains a rendering matrix A based on the rendering matrix M, for example using a parameter adjustment module and M-to-A mapping. Also, the words of the matrix E covariance are obtained depending on the OLD information of the object level difference and the IOC information of inter-object correlation, for example as discussed above. Similarly, the downmix D matrix words are obtained depending on the downmix gain DMG information and the DCLD information of the downmix channel level differences.
[0189] The f words of the desired F covariance matrix are obtained depending on the rendering matrix A and the covariance matrix E. Also, the scalar value v is obtained depending on the matrix E of the covariance and the matrix D of the downmix (or depending on their words).
[0190] The PL, PR gain values for the two channels are obtained depending on the words of the desired matrix F covariance and the scalar value v. Also the value of φ<sub>ε </sub>cross-object phase difference is obtained depending on the words f of the desired matrix F covariance. The angle of rotation α is also obtained depending on the words f of the desired matrix F covariance, taking into account, for example, the constant c. In addition, a second angle of rotation β is obtained, for example, depending on the gains PL, PR of the channels and the first angle of rotation α. The words of the G matrix are obtained, for example, depending on the two-channel values of PL, PR of the gains as well as depending on the inter-object phase difference φ<sub>ε</sub> and optionally, from rotation angles α, β. Similarly, the words in the matrix P<sub>2</sub> are determined depending on some or all of the P values listed<sub>L</sub>, P<sub>R</sub>, φ<sub>ε</sub>, α, β.
[0191] Hereinafter, it will be discussed how the G and / or P2 matrix (or words thereof) that can be used by a downmix processor as discussed above can be obtained for different processing modes.
4.1.2.2 Mono to binaural "x-1-b" processing mode [0192] The following will discuss processing modes in which regular audio objects are represented by a single-channel downmix 134, 264, 322, 497a signal in which binaural rendering is desirable.
[0193] G upmix parameters<sup>l, m</sup> and P.<sub>2</sub><sup>l, m</sup> are calculated as
<img file="PL2535892T3_D0039.tif" />
r
<img file="PL2535892T3_D0040.tif" />
pi.tii p /, tt [0194] Gains i for the left and right output channels are
<img file="PL2535892T3_D0041.tif" />
<img file="PL2535892T3_D0042.tif" />
[0195] Desired matrix F<sup>l, m</sup> 2x2 covariance with elements given as
Im
[0196] Scalar v<sup>l, m</sup> is calculated as
<img file="PL2535892T3_D0043.tif" />
[0197] The inter-channel phase difference is given as j /, - πϊ
9c arg ^ f), 0Sm <11, ^ "> 0.6, and \ otherwise.
[0198] Inter-channel coherence
<img file="PL2535892T3_D0044.tif" />
is calculated as
<img file="PL2535892T3_D0045.tif" />
[0199] Angles of rotation a<sup>l, m</sup> e<sup>l, m</sup> are given as arccos (/ ^ '"cos (arg0 <m <11, p' ^ <0.6, | arccos (p'-<sup>in</sup>), otherwise.
<img file="PL2535892T3_D0046.tif" />
4.1.2.3 "x-1-2" mono-to-stereo processing mode [0200] The processing mode in which regular audio objects are represented by single-channel signal 134, 264, 222 and in which stereo rendering is desired is discussed below.
[0201] For the stereo output signal, the "x-1-b" processing mode can be used without using HRTF information. This can be accomplished by acquiring all elements of the A matrix of rendering, obtaining:
<img file="PL2535892T3_D0047.tif" />
<img file="PL2535892T3_D0048.tif" />
4.1.2.4 "x-1-1" mono-to-mono processing mode [0202] The following describes the processing mode in which regular audio objects are represented by signal channel 134, 264, 322, 497a and in which regular two-channel rendering is desirable audio objects.
[0203] For the mono output, the "x-1-2" processing mode can be used with the following words:
<img file="PL2535892T3_D0049.tif" />
4.1.2.5 Stereo-to-binaural "x-2-b" processing mode [0204] Below, the processing mode in which regular audio objects are represented by a two-channel signal 134, 264, 322, 497a and in which binaural rendering is desirable will be described regular audio objects.
[0205] The G 'and upmix parameters are calculated as
<img file="PL2535892T3_D0050.tif" />
<img file="PL2535892T3_D0051.tif" />
nl, v ρΙ, ι »[0206] Appropriate gains for the left and right output channels are
<img file="PL2535892T3_D0052.tif" />
<sub>l</sub> / Ίαι, -γ [0207] Desired F matrix<sup>l, m, x</sup> 2x2 covariance with elements is given as
JW [0208] Matrix C<sup>l, m</sup> 2x2 covariance with elements of "dry" binaural signal is estimated as where
<img file="PL2535892T3_D0053.tif" />
[0209] Suitable scalars v<sup>l, m, x</sup> iv<sup>l, m</sup> are calculated as
λ # [0210] Downmix matrix D<sup>l, x</sup> with 1xN size with elements can be obtained as
<img file="PL2535892T3_D0054.tif" />
<img file="PL2535892T3_D0055.tif" />
d '[0211] 2xN downmix D' matrix with elements can be obtained as [0212] Matrix E<sup>l, m, x</sup> with the elements' · '··' is obtained from the following dependence. Λ
<img file="PL2535892T3_D0056.tif" />
f.iw
Jf.
[0213] Inter-channel phase differences are given as
<img file="PL2535892T3_D0057.tif" />
[0214] ICC and <sup>L</sup>'· Are calculated as
<img file="PL2535892T3_D0058.tif" />
[0215] Angles of rotation a<sup>l, m</sup> e<sup>l, m</sup> are given as
<img file="PL2535892T3_D0059.tif" />
<img file="PL2535892T3_D0060.tif" />
4.1.2.6 Stereo-to-stereo "x-2-2" processing mode [0216] The following describes the processing mode in which regular audio objects are described by a two-channel (stereo) signal 134, 264, 322, 497a and in which it is desirable two-channel rendering (stereo).
[0217] For the stereo output signal, the stereo pre-processing is directly applied, which will be described below in section 4.2.2.3.
4.1.2.7 Stereo-to-mono "x-2-1" processing mode [0218] In the following, a processing mode is described in which regular audio objects are represented by a two-channel (stereo) signal 134 , 264, 322, 497a and in which is single channel (mono) rendering.
[0219] For the mono output signal, stereo pre-processing with a single active matrix expression is rendered, as described below in section 4.2.2.3.
4.1.2.8 Conclusions [0220] Referring again to Figs. 4a and 4b, processing is shown that can be applied to single-channel or two-channel signal 134, 264, 322, 497a representing regular audio objects after separation between extended audio objects and regular objects audio. Figures 4a and 4b show the processing, wherein the processing in Figs. 4a and 4b differ in that the optional parameter control is introduced at various stages of processing.
4.2 Operation in transcoding mode
4.2.1 Introduction [0221] The following explains how to combine SAOC parameters and panorama information (or rendering information) associated with each audio object (or preferably with any regular audio object) in a standard compliant MPEG Surround bit stream (MPS bit stream).
[0222] The SAOC 490 transcoder is shown in Fig. 4f and consists of a SAOC parameter processor 491 and a downmix processor 492 in use for a stereo downmix.
[0223] The SAOC transcoder 490 may for example take over the role of an audio signal processor 140. Alternatively, the SAOC 490 transcoder may take over the role of the SAOC downmix pre-processor 270 in conjunction with the SA2 252 processor.
[0224] For example, the SAOC parameter processor 491 may receive the SAOC bit stream 491a, which is equivalent to the object parametric information 110 or the SAOC bit stream 212. Also, the SAOC parameter processor 491 may receive the rendering matrix information 491b, which may be included in the object parametric information 110 , or it can be equivalent to 214 rendering matrix information. The SAOC parameter processor 491 may also provide downmix processing information 491c to the downmix processor 492, which may be equivalent to information 240. In addition, the SAOC parameter processor 491 may include an MPEG Surround 491d bit stream (or MPEG Surround parameter bit stream) that includes parametric Surround information, which is compatible with the MPEG Surround standard. The MPEG Surround 491d bit stream may for example be part of the processed version 142 of the second audio information, or may for example be part of the MPS 222 bit stream or replace it.
[0225] Downmix processor 492 is configured to receive downmix signal 492a, which is preferably a single-channel downmix signal or a two-channel downmix signal and which is preferably equivalent to second audio information 134, or second signal 264, 322 of audio objects. The downmix processor 492 may also provide an MPEG Surround downmix signal 492b that is equivalent to the processed version 272 (or part thereof) of the second signal 264 of audio objects.
[0226] However, there are various ways to combine the MPEG Surround downmix signal 492b with the signal 132, 262 of enhanced audio objects. Joining can be done in the MPEG Surround domain.
[0227] Alternatively, however, the MPEG Surround representation, comprising the bit stream 491d of the MPEG Surround parameters and the MPEG Surround downmix signal 492b of regular audio objects can be converted back to a multi-domain time domain representation or a multi-channel frequency domain representation (individually representing different audio channels) by MPEG Surround decoder and can then be combined with signals of enhanced audio objects. [0228] It should be noted that the transcoding modes include both one or more mono downmix processing modes and one or more stereo downmix processing modes. However, only stereo processing mode will be discussed below, because the processing of regular audio objects is more complex in stereo downmix processing mode.
4.2.2 Downmix processing in stereo downmix processing mode ("x-2-5")
4.2.2.1 Introduction [0229] The following section will describe the SAOC transcoding mode for the stereo downmix case.
[0230] Object parameters (OLD object level difference, IOC inter-object correlation, DMG downmix gain and DCMD downmix channel difference) from the SAOC bit stream are transcoded into spatial parameters (preferably channel, CLD channel level difference, inter-object correlation ICC, channel prediction coefficients) CPC) for the MPEG bit stream
Surround according to the rendering information. Downmix is modified according to object parameters and the rendering matrix.
[0231] Referring now to Figs. 4c, 4d and 4e, an overview of the processing, and in particular of the downmix modification, will be given.
[0232] Fig. 4c shows a processing block diagram that is implemented for modifying a downmix signal, e.g., a downmix signal 134, 264, 322,492a describing one or preferably more regular audio objects. As can be seen in Figs. 4c, 4d and 4e, the processing receives a rendering Mren matrix, downmix gain DMG information, DCLD downmix channel level difference information, object level difference OLD information and inter-object correlation IOC information. The rendering matrix can optionally be modified by adjusting the parameter as shown in Fig. 4c. The downmix D matrix words are obtained depending on the DMG downmix gain information and DCLD information of the downmix channel level differences. The words of the E coherence matrix are obtained depending on the OLD information of the object level difference and the IOC information of inter-object correlation. Additionally, the matrix J can be obtained depending on the downmix matrix D and the coherence matrix E, or depending on their words. Then the C3 matrix can be obtained depending on the rendering Mren matrix, downmix matrix D, coherence matrix E and matrix J. The matrix G can be obtained depending on the DTTT matrix, which can be a matrix having predefined words, and also depending on the matrix C3. Matrix G may optionally be modified to obtain a modified Gmod matrix. Matrix G or its modified version Gmod can be used to obtain the processed version 142, 272, 492b of the second audio information 134, 264 from the second audio information 134, 264, 492a (where the second audio information 134, 264 is marked X, and its
And the processed version 142, 272 is marked with X.
[0233] The energy rendering of the object that is performed to obtain MPEG Surround parameters will be discussed below. Also, stereo processing that is implemented to obtain the processed version 142, 272, 492b of the second audio information 134, 264, 492a representing regular audio objects will be described.
4.2.2.2 Rendering energy of objects [0234] The transcoder sets parameters for the MPS decoder according to the target rendering described by the rendering matrix Mren. The six-channel target covariance is designated F and given by
<img file="PL2535892T3_D0061.tif" />
[0235] The transcoding process may be conceptually divided into two parts. In one part, three-channel rendering is performed to the left, right and center channel. At this stage, the parameters for downmix modification are obtained as well as the prediction parameters for the TTT block for the MPS decoder. In the rest, CLD and ICC parameters are determined for rendering between the front and surround channels (OTT parameters, left - front left, surround left, right front - right surround).
4.2.2.2.1 Rendering of the left, right and center channel [0236] At this stage, spatial parameters are determined that control the rendering to the left and right channel consisting of front and surround signals. These parameters describe the TTT block prediction matrix for MPS CTTT decoding (CPC parameters for the MPS decoder) and the downmix converter G matrix.
[0237] CTTT is a prediction matrix for obtaining target rendering from
AND
V-CY modified downmix "'' ':
C<sub>TTT</sub>X = C.<sub>m</sub>GX ^ A<sub>J</sub>S.
[0238] A3 is a 3xN reduced rendering matrix describing rendering for left, right and center channels respectively. It is obtained as A3 = D36 Mren with a D36 partial downmix matrix 6 to 3 determined by
<img file="PL2535892T3_D0062.tif" />
[0239] The weights wp, p = 1,2,3 of the partial downmix are adjusted in such a way that the energy w<sub>p</sub>(s<sub>2p1</sub> + y<sub>2p</sub>) is equal to the sum of energy || y<sub>2p</sub>_<sub>1</sub>1|<sup>2</sup> + || y<sub>2p</sub>||<sup>2</sup> up to the limiting factor
<img file="PL2535892T3_D0063.tif" />
<img file="PL2535892T3_D0064.tif" />
w3 = 0.5, where fi, j are elements F.
[0240] For the estimation of the desired CTTT prediction matrix and downmix G processing matrix, we define a 3x2 prediction matrix C3, which leads to target rendering
CX ^ AS [0241] Such matrix is obtained by taking into account the normal equations ą (DED ') «A, ED \ [0242] The solution of normal equations gives the best possible fit of the wave shape to the target output signal for a given object covariance model. G and CTTT are now obtained by solving the system of equations
C<sub>TT</sub>tG = C<sub>3</sub>.
[0243] To avoid numerical problems when calculating J = (DED *)<sup>-1</sup>, J is modified. First, the proprietary properties λ1,2 J are calculated, solving det (J λ1,2ΐ) = 0.
[0244] Eigenvalues are sorted in descending order (λι> λ<sub>2</sub>) and the eigenvector corresponding to the greater eigenvalue is calculated according to the equation above.
Its position in the positive x plane is ensured (the first element must be positive). The second eigenvector is obtained from the first by rotation through an angle of -90 degrees:
f J fi \ '=, W.
[0245] The weighing matrix is calculated from the downmix D matrix and the C3 prediction matrix, W = 5 (D diag (C3)).
[0246] Because CTTT is a function of the MPS prediction parameters c1 and c2 (described in ISO / IEC 23003-1: 2007), CTTT G = C3 is rewritten as follows to find the point or fixed points of the function
<img file="PL2535892T3_D0065.tif" />
with Γ = (Dttt C3) in (Dttt C3) * and b = GWC 37, where
<img file="PL2535892T3_D0066.tif" />
[0247] If r does not provide a unique solution (det (r) <10<sup>-3</sup>), the point closest to the point giving the TTT transition is selected. In the first stage, the row is selected and for ΒΓ, γ = γ γ<sub>2</sub>] where the elements contain the most energy, so Y<sub>and</sub>,<sub>1</sub><sup>2</sup> + Yi, 2<sup>2</sup> You and,<sub>1</sub><sup>2</sup> + Yj,<sub>2</sub><sup>2</sup>, j = 1.2.
[0248] Next, a solution is determined such that / Y \<sup>C</sup>2j
-3y
JJ
'UJ'
V-L2 [0249] If the solution obtained for "i": lies outside the allowable range for the prediction coefficients, which is defined as -—<sup>:</sup> (as specified in ISO / IEC
23003-1: 2007), · should be calculated as below.
[0250] First, define a set of points, xp as:
<img file="PL2535892T3_D0067.tif" />
and distance function) = 1 ^, -2bip.
[0251] Next, prediction parameters are defined according to:
<img file="PL2535892T3_D0068.tif" />
[0252] Prediction parameters are limited according to:
c, = (l - A) ć, + Λ / ,, c<sub>2</sub> = (1-λ) £<sub>2</sub>+ λχ<sub>21</sub> where λ, γ1 and γ2 are defined as
<img file="PL2535892T3_D0069.tif" />
<img file="PL2535892T3_D0070.tif" />
<img file="PL2535892T3_D0071.tif" />
<img file="PL2535892T3_D0072.tif" />
[0253] For the MPS decoder, CPC and corresponding ICCTTT are provided as follows
4.2.2.2.2 Rendering between front and surround channels [0254] Parameters determining the rendering between front and surround channels can be estimated directly from the target matrix F covariance
<img file="PL2535892T3_D0073.tif" />
<img file="PL2535892T3_D0074.tif" />
[0255] MPS parameters are provided in the form
<img file="PL2535892T3_D0075.tif" />
for each OTT block h.
4.2.2.3 Stereo processing [0256] The stereo processing of the signal 134 to 64, 322 regular audio objects will be discussed below. Stereo processing is used to obtain the process for general representation 142, 272 based on a two-channel representation of regular audio objects.
[0257] The stereo downmix X, which is represented by the signals 134, 264, 492a of regular audio objects, is converted into the signal X of the modified downmix that is represented by the processed signals 142, 272 of regular audio objects:
<img file="PL2535892T3_D0076.tif" />
where
<img file="PL2535892T3_D0077.tif" />
0 [0258] The final output signal from the SAOC transcoder, X is created by mixing X with the de-correlated signal component according to:
Λ
Χ = β "" Χ + Ρ<sub>2</sub>Χ ", where the de-correlated Xd signal is calculated as described above, and the mix matrixes Gmod and P2 as shown below.
[0259] First, define the rendering upmix error matrix as
<img file="PL2535892T3_D0078.tif" />
where
<img file="PL2535892T3_D0079.tif" />
and further define the covariance matrix of the predicted signal ii as
* GDED'G.
'1,1 [0260] The gvec gain vector can then be calculated as:
<img file="PL2535892T3_D0080.tif" />
and the GMod mix matrix is given as:
<img file="PL2535892T3_D0081.tif" />
^, (G<sub>in <</sub>,) G
G,,> 0, otherwise.
[0261] Similarly, the P2 mix matrix is given as:
Ρ<sub>2</sub> 'ο ο'
»* 1,2> 0, otherwise.
[0262] To obtain vR and Wd, the characteristic equation R: det (R-Xi,<sub>2</sub>I) = 0, having eigenvalues λι and λ<sub>2</sub> [0263] The respective eigenvectors R, vR1 and vR2 can be calculated by solving the system of equations:
(B— ą ^ I) v<sub>RrR2</sub> = 0.
[0264] Eigenvalues are sorted in descending order (λι> λ<sub>2</sub>), and the eigenvector corresponding to the greater eigenvalue is calculated according to the equation above. It shall be positioned in the positive x plane (the first element must be positive). The second eigenvector is obtained from the first by rotation through an angle of -90 degrees:
R = (WrJ <sub>n</sub> , Ow < <sup>AT</sup> 4 J [0265] By entering P1 = (1 1) G, Rd can be calculated according to:
<img file="PL2535892T3_D0082.tif" />
which gives
<img file="PL2535892T3_D0083.tif" />
and finally the mix matrix <sup>P</sup>2 = (? X
IN<sub>dL</sub> Oh
4.2.2.4 Dual mode [0266] The SAOC transcoder may enable the calculation of the mix matrix P1, P2 and the prediction matrix C3 according to an alternative method for the higher frequency range. This alternative method is particularly useful for downmix signals where the higher frequency range is encoded by a non-waveform coding algorithm, e.g. SBR in high-performance AAC.
[0267] For the upper parameter bands defined by bsTttBandsLow <pb <numBands, P1, P2 and C3 should be calculated according to the alternative method described below:
P-
<img file="PL2535892T3_D0084.tif" />
<img file="PL2535892T3_D0085.tif" />
[0268] Define the energy downmix and target energy vectors, respectively:
<img file="PL2535892T3_D0086.tif" />
= diag (DED ') + eI,
<img file="PL2535892T3_D0087.tif" />
and an auxiliary matrix
T =
<img file="PL2535892T3_D0088.tif" />
= AjD * +.
[0269] Then calculate the gain vector
<img file="PL2535892T3_D0089.tif" />
which ultimately gives a new prediction matrix
<td></td><td>^ | I | J</td><td> 5/1.1</td>
<td>C<sub>3</sub> =</td><td>ffAl</td><td>5l £ 2 ^</td>
<td></td><td></td><td>5 / w></td>
[0270] 5. Combined EKS SAOC decoding / transcoding mode, encoder according to Fig. 10 and systems according to Figs. 5a and 5b [0271] A brief description of the method of combined EKS SAOC processing will be given below. A preferred "EKS SAOC processing" method is proposed in which the EKS processing is integrated into the regular SAOC decoding / transcoding chain in a cascade system.
5.1 Audio encoder according to Fig. 5 [0272] In the first step, the objects dedicated to EKS processing (enhanced Karaoke / solo processing) are identified as foreground objects (FGO) and their number NFGO (also denoted NEAO) is determined by the variable " bsNumGroupsFGO "bit stream. Said bit stream variable may for example be included in the SAOC bit stream described above.
[0273] For generating the bit stream (in the audio signal encoder), the parameters of all Nobj input objects are arranged in such a way that FGO foreground objects contain the last NFGO parameters (or alternatively NEAO) in each case, for example, OLD<sub>and</sub> for N<sub>vol</sub> - N<sub>FGO</sub> and N<sub>vol</sub> - 1].
[0274] From the other objects, which are, for example, BGO background objects or unenriched audio objects, a "regular SAOC style" downmix signal is generated, which also acts as a BGO background object. Then, the background object and foreground objects are mixed down in the "EKS processing style" and residual information is obtained from each foreground object. In this way, there is no need to enter any additional processing steps. In this way, the bit stream syntax was changed.
[0275] In other words, on the encoder side, un-enriched audio objects are distinguished from enriched audio objects. A single-channel or two-channel downmix signal of regular audio objects is provided, which represents regular audio objects (un-enriched audio objects), with one, two or even more regular audio objects (un-enriched audio objects). The single-channel or two-channel downmix signal of regular audio objects is then combined with one or more signals of the enriched audio objects (which may be e.g. a single-channel downmix signal or two-channel downmix signal) to obtain a common downmix signal (which may be e.g. a single-channel downmix signal or two-channel downmix signal) combining audio signals of enriched audio objects and a downmix signal of regular audio objects.
[0276] In the following, the basic structure of such a cascade encoder will be briefly described with reference to Fig. 10, which shows a block diagram of the SAOC 1000 encoder according to an embodiment of the invention. The SAOC 1000 encoder includes a SAOC downmix module 1010, which is typically a SAOC downmix module that does not provide residual information. The SAOC downmix module 1010 is configured to receive multiple signals of 1012 NBGO audio objects from regular (non-enriched) audio objects. Also, SAOC downmix module 1010 is configured to provide the regular audio object downmix signal 1014 based on regular audio objects 1012 in such a way that the regular audio object signal 1014 combines the 1012 regular audio object signals according to the downmix parameters. The SAOC downmix module 1010 also provides SAOC 1016 regular audio object information that describes the signals and downmix of regular audio objects. For example, SAOC information 1016 of regular audio objects may include DMG downmix gain information and DCLD information about the downmix channel level difference describing the downmix made by the SAOC downmix module 1010. In addition, SAOC 1016 regular audio object information may include object level difference information and inter-object correlation information describing the relationship between regular audio objects described by the regular audio object signal 1012.
[0277] Encoder 1000 also includes a second SAOC downmix module 1020 that is typically configured to provide residual information. The second SAOC downmix module 1020 is preferably configured to receive one or more signals 1022 of enhanced audio objects as well as to receive the downmix signal 1014 of regular audio objects.
[0278] The second SAOC downmix module 1020 is also configured to provide a common SAOC downmix signal 1024 based on the signals 1022 of enhanced audio objects and the signal 1014 of regular audio objects. When providing a common SAOC downmix signal, the second SAOC downmix module 1020 typically treats the regular audio object downmix signal 1014 as a single channel or two channel signal.
[0279] Second module 1020 is also configured to provide SAOC information of enriched audio objects, which describes, for example, DCLD values of downmix channel level differences associated with enriched audio objects and OLD values of object level difference associated with enriched audio objects and IOC values of inter-object correlation associated with enriched audio objects. Additionally, the second SAOC downmix module 1020 is preferably configured to provide residual information associated with each of the enriched audio objects in such a way that the residual information associated with the enriched audio objects describes the difference between the original signal of the individual enriched audio objects and the expected signal of the individual enriched audio objects that can be obtained from a downmix signal using DMG downmix information, DCLD and OLD object information, IOC.
[0280] The audio encoder 1000 is well adapted to interact with the audio decoder described herein.
5.2 Audio signal decoder according to Fig. 5a [0281] The basic structure of the EKS SAOC 500 hybrid decoder will be described below, the block diagram of which is shown in Fig. 5a.
[0282] The audio decoder 500 of Fig. 5a is configured to receive downmix signal 510, SAOC bit stream information 512 and rendering matrix information 514. The audio decoder 500 includes Karaoke / solo processing and rendering of 520 foreground objects that is configured to provide a first signal 562 of audio objects that describes rendered foreground objects and a second signal 564 of audio objects that describes background objects. Foreground objects may, for example, be called "enhanced audio objects" and background objects may, for example, be called "regular audio objects" or "non-enriched audio objects". The audio decoder 500 also includes regular SAOC 570 decoding, which is configured to receive a second signal 562 of audio objects and to provide a processed version 572 of the second signal 564 audio objects based thereon. The audio decoder 500 also includes a combining module 580 that is configured to combine the first signal 562 of audio objects and the processed version 572 of the second signal 564 audio objects to obtain output signal 520.
[0283] The functionality of the audio decoder 500 will be discussed below with some additional details. On the SAOC decoding / transcoding side, the upmix process results in a cascade system that first includes enriched Karaoke / solo processing (EKS processing) to decompose the downmix signal into a background object (BGO) and foreground objects (FGO). The required object level differences (OLD) and inter-object correlations (IOC) for the background object are obtained from the object and downmix information (both of which are object-oriented parametric information and which both are typically contained in the SAOC bit stream):
OLD<sub>l</sub>= £ · d ^ OLD,
<img file="PL2535892T3_D0090.tif" />
<img file="PL2535892T3_D0091.tif" />
<img file="PL2535892T3_D0092.tif" />
[0284] In addition, this step (which is typically performed by EKS processing and rendering of 520 foreground objects) includes mapping the foreground objects to final output channels (in such a way that, for example, the first signal 562 of the audio objects is a multi-channel signal in which each of foreground objects are mapped to one or more channels). The foreground object (which typically contains many so-called "regular audio objects") is rendered to the corresponding output channels by the regular SAOC decoding process (or alternatively in some cases by the SAOC transcoding process). For example, this process can be accomplished by regularly decoding SAOC 570. The final mixing step (e.g., combining module 580) provides the output with the desired combination of signals of the rendered foreground and background objects.
[0285] This hybrid EKS SAOC system represents a combination of all the beneficial properties of a regular SAOC system and its EKS mode. This approach allows obtaining appropriate results using the proposed system with the same bit stream for both classic (moderate rendering) and Karaoke / solo (extreme rendering) scenario.
5.3 Generalized structure according to Fig. 5b [0286] The generalized structure of the combined EKS SAOC 590 system with reference to Fig. 5b will be discussed below, which shows a block diagram of such a generalized combined EKS SAOC system. The combined EKS SAOC 590 system of Fig. 5b can also be seen as an audio decoder.
[0287] The combined EKS SAOC 590 system is configured to receive the downmix signal 510a, SAOC bit stream information 512a and rendering matrix information 514a. Also, the combined EKS SAOC 590 system is configured to provide output 520a based on them.
[0288] The combined EKS SAOC 590 system includes a SAOC type I processing member 520a that receives the downmix signal 510a, SAOC bit stream information 512a (or at least a portion thereof), and rendering matrix information 514a (or at least a portion thereof). In particular, the SAOC processing member I 520a receives values (OLDs) of the difference in object level of the first member. The SAOC processing member I 520a provides one or more signals 562a describing the first set of objects (e.g., audio objects of the first type of audio objects). The SAOC processing member I 520a also provides one or more signals 564a describing a second set of objects.
[0289] The combined EKS SAOC system also includes a SAOC type II 570a processing member that is configured to receive one or more signals 564a describing a second set of objects and to provide based on it one or more signals 572a describing a third set of objects using differences second-member object levels that are included in SAOC bit stream information 512a, as well as at least part of the rendering matrix information 514. The combined EKS SAOC system also includes a combining module 580a, which may, for example, be an adder to provide output signals 520a by combining one or more signals 570a describing a third set of objects (wherein the third set of objects may be a processed version of the second set of objects).
[0290] In summary of the above, Fig. 5b shows a generalized form of the basic structure described with reference to Fig. 5a above, in a preferred embodiment of the invention.
6. Perceptual evaluation of the way EKS SAOC combined processing
6.1 Test methodology, design and components [0291] These subjective listening tests were carried out in an acoustically insulated listening room, which is adapted for high-quality listening. Playback was done using headphones (SAX SR Lambda Pro with D / A converter Lake-People and monitor STAX SRM). The test method followed standard procedures used in spatial audio verification tests based on the MUSHRA method (multiple stimulus with hidden reference and anchors) for subjective evaluation of intermediate audio quality (see references [7]).
[0292] Eight listeners participated in the conducted test. According to the MUSHRA methodology, listeners were asked to compare all test conditions against reference. Test conditions were automatically selected randomly for each element of the test and for each listener. Subjective responses were recorded by the MUSHRA computer program on a scale of 0 to 100. Instant switching between test items was enabled. The MUSHRA test was carried out to assess the perceptual results of the SAOC modes under consideration and the proposed system described in the table of Fig. 6a, which provides a description of the listening test structure.
[0293] Corresponding downmix signals were encoded using an AAC core encoder at 128 kbps. To assess the perceptual quality of the proposed combined EKS SAOC system, it is compared with the regular SAOC RM system (SAOC reference model system) and the current EKS mode (enriched Karaoke / solo mode) for two different rendering test scenarios described in the table in Fig. 6b, which describes the systems tested.
[0294] 20 kbps residual encoding was used for the current EKS mode and the proposed combined EKS SAOC system. It should be noted that for the current EKS mode it is necessary to generate a stereo background object (BGO) before the actual encoding / decoding procedure, because this mode has limitations on the number of input object types.
[0295] The listening test material and the appropriate downmix and rendering parameters used in the tests were selected from the set of CfP "call-for-proposals" audio elements described in document [2]. The corresponding data for the "Karaoke" and "Classic" rendering application scenarios can be found in the table of Fig. 6c, which describes the listening test elements and rendering matrices.
6.2 Listening test results [0296] A brief outline of the graphs showing the results of the listening test can be found in Figs. 6d and 6e, where Fig. 6d shows the MUSHRA results for the Karaoke / solo rendering type listening test, and Fig. 6e shows the average scores MUSHRA listening test of classic rendering. The graphs show the average MUSHRA scores for the items for all listeners and the statistical average for all items assessed together with associated ranges of 95% confidence intervals.
[0297] Based on the results of the listening tests, the following conclusions can be made:
* Fig. 6d shows a comparison for the current EKS mode with a combined EKS SAOC system for Karaoke applications. For all test elements, no significant differences in quality (in a statistical sense) are observed between the two systems. From this observation it can be concluded that the combined EKS SAOC system is capable of efficiently using residual information, achieving the quality of the EKS mode. It can also be seen that the quality of the regular SAOC system (no residue) is lower than both other systems.
* Fig. 6e shows a comparison for the current SAOC regular mode with the connected EKS SAOC system for classic rendering scenarios. For all tested components, the quality of these two systems is statistically the same. This demonstrates the proper operation of the combined EKS SAOC system for the classic rendering scenario.
[0298] Hence, it can be concluded that the proposed unified system combining EKS mode with regular SAOC retains the advantages of subjective audio quality in the respective types of rendering.
[0299] Given that the proposed combined EKS SAOC system has no restrictions on the BGO object, but has fully flexible rendering capabilities for regular SAOC mode and can use the same bit stream for all rendering types, it seems beneficial to incorporate it into the MPEG SAOC standard .
7. Method of Fig. 7 [0300] Hereinafter, a method of providing an upmix signal representation based on a downmix signal representation and object-oriented parametric information will be described with reference to Fig. 7, which shows a flowchart of such a method.
[0301] The method 700 includes step 710 of decomposing the downmix signal representation for providing the first audio information describing the first set of one or more audio objects from the first type of audio objects and the second audio information describing the second set of one or more audio objects from the second type of audio objects, based on the downmix signal representation and at least part of the object parametric information. Method 700 also includes step 720 of processing the second audio information based on the object parametric information to obtain a processed version of the second audio information.
[0302] The method 700 also includes the step 730 of combining the first audio information with the processed version of the second audio information to obtain an upmix signal representation.
[0303] The method 700 according to Fig. 7 can be supplemented by any of the features and functions discussed herein in relation to the device according to the invention. Also, method 700 provides the benefits discussed with respect to the device of the invention.
8. Alternative implementations [0304] Although some aspects have been described in the context of the device, it is clear that these aspects also represent a description of the respective method, where the block or device corresponds to the method step or the properties of the method step. Similarly, the aspects described in the context of the method step also represent a description of the respective block or position or properties of the respective device. Some or all of the method steps may be carried out by hardware devices (or with their help), such as a microprocessor, programmable computer or electronic circuit. In some embodiments, one or more of the most important method steps may be carried out by such a device.
[0305] The audio signal encoded according to the invention may be stored on a digital storage medium or may be transmitted by means of transmission such as wireless transmission means or wired transmission means such as the Internet. [0306] Depending on some implementation requirements, embodiments of the invention can be implemented in hardware or in software. The implementation can be implemented using digital storage media, e.g. floppy disks, Blue-Ray discs, DVDs, CDs, ROMs, PROMs, EPROMs, EEPROMs or FLASHs containing electronically readable control signals that cooperate with them (or are capable of such interaction with the programmed computer system so that the appropriate method is implemented. Therefore, the digital storage medium can be computer readable.
[0307] Some embodiments of the invention include a data carrier containing electronically readable control signals that interact with a programmable computer system such that one of the methods described herein is implemented. [0308] In general, embodiments of the present invention may be implemented as a computer program product with a program code, which program code may operate to implement one method of the invention when the computer program product is running on a computer. For example, the program code can be saved on a machine-readable medium.
[0309] Other embodiments comprise a computer program for performing one of the methods described herein, stored on a machine readable carrier.
In other words, an embodiment of a method of the invention is thus a computer program containing program code for implementing one of the methods described herein when the computer program product is running on a computer.
[0311] A further embodiment of the methods of the invention is thus a data carrier (or digital storage medium, or a computer readable medium) comprising the computer program stored therein for carrying out one of the methods described herein. The data carrier, or digital storage medium, or recorded medium, is typically tangible and / or non-transmitting.
[0312] A further embodiment of the method of the invention is thus a data stream or a sequence of signals representing a computer program for carrying out one of the methods described herein. The data stream or signal sequence may, for example, be configured to be sent over a data link, e.g., via the Internet.
[0313] Another embodiment of the method of the invention comprises processing means, e.g. a computer, or a programmable logic device, configured or adapted to implement one of the methods described herein.
[0314] Another embodiment of the method of the invention comprises a computer in which a computer program is installed to perform one of the methods described herein.
[0315] In some embodiments, a programmable logic device (e.g., a user programmable logic table) may be used to perform some or all of the functions of the methods described herein. In some embodiments, the user programmable logic array may interact with a microprocessor to implement one of the methods described herein. Generally, the methods are preferably carried out by any hardware device.
[0316] The above described embodiments are merely illustrative for the principles of the present invention. It should be understood that modifications and variants of the systems and details described herein are obvious to those skilled in the art. It is therefore intended that the restrictions arise only from the scope of the following claims, and not from the specific details provided for the purposes of describing and explaining the present variants of the invention.
9. Conclusions [0317] Some aspects and benefits of the combined EKS 5 SAOC system according to the present invention will be summarized below. In Karaoke and Solo playback situations, the SAOC EKS processing mode supports both the reproduction of only background objects / foreground objects and a free mix (determined by the rendering matrix) of these groups of objects.
[0318] Also, the first mode is considered the main purpose of EKS processing, while the second mode provides additional flexibility.
[0319] It has been found that the generalization of EKS functionality consequently involves the effort of combining EKS with the regular SAOC processing mode to obtain one unified system. The capabilities of such a unified system include:
* One single transparent SAOC decoding / transcoding structure;
* One bit stream for both EKS and SAOC regular mode;
* No restrictions on the number of input objects containing a background object (BGO), so that there is no need to generate a background object before the SAOC coding stage; and * Residual coding support for foreground objects providing enhanced perceptual quality in demanding Karaoke / Solo playback situations.
[0320] These benefits can be achieved using the unified system described herein.
Fraunhofer-Gesellschaft zur Forderung der angewandten Forschung eV, Germany
Proxy:
Bibliography [0321] [1] ISO / IEC JTC1 / SC29 / WG11 (MPEG), Document N8853, "Call for Proposals on Spatial Audio Object Coding", 79th MPEG Meeting, Marrakech, January 2007.
[2] ISO / IEC JTC1 / SC29 / WG11 (MPEG), Document N9099, "Final Spatial Audio Object
Coding Evaluation Procedures and Criterion ", 80th MPEG Meeting, San Jose, April 2007.
[3] ISO / IEC JTC1 / SC29 / WG11 (MPEG), Document N9250, "Report on Spatial Audio Object Coding RM0 Selection", 8 1st MPEG Meeting, Lausanne, July 2007.
[4] ISO / IEC JTC1 / SC29 / WG11 (MPEG), Document M15123, "Information and Verification
Results for CE on Karaoke / Solo system improving the performance of MPEG SAOC RM0 ", 83rd MPEG Meeting, Antalya, Turkey, January 2008.
[5] ISO / IEC JTC1 / SC29 / WG11 (MPEG), Document N10659, "Study on ISO / IEC 230032: 200x Spatial Audio Object Coding (SAOC)", 88th MPEG Meeting, Maui, USA, April
2009.
[6] ISO / IEC JTC1 / SC29 / WG11 (MPEG), Document M10660, "Status and Workplan on SAOC Core Experiments", 88th MPEG Meeting, Maui, USA, April 2009.
[7] EBU Technical recommendation: "MUSHRA-EBU Method for Subjective Listening Tests of Intermediate Audio Quality", Doc. B / AIM022, October 1999.
[8] ISO / IEC 23003-1: 2007, Information technology - MPEG audio technologies - Part 1:
MPEG Surround.
Fraunhofer-Gesellschaft zur Forderung der angewandten Forschung eV, Germany
Proxy:
EP 2 535 892 B1
Z-12693
Contents9
41 members in 20 offices
Priority claims6
| Document | Office | Kind | Date |
|---|---|---|---|
| 22004209 | United States of America | P | |
| 10727721 | European Patent Office (EPO) | A | |
| 12183562 | European Patent Office (EPO) | A | |
| EP20100727721 | – | – | – |
| EP20120183562 | – | – | – |
| US20090220042P | – | – | – |
Members41
| Document | Office | Kind | |
|---|---|---|---|
| CA2766727A1 | Canada | A1 | |
| CA2855479A1 | Canada | A1 | |
| WO2010149700A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW201108204A | Taiwan Province of China | A | |
| AR077226A1 | Argentina | A1 | |
| AU2010264736A1 | Australia | A1 | |
| SG177277A1 | Singapore | A1 | |
| MX2011013829A | Mexico | A | |
| KR20120023826A | Republic of Korea | A | |
| EP2446435A1 | European Patent Office (EPO) | A1 | |
| CN102460573A | China | A | |
| US2012177204A1 | United States of America | A1 | |
| CO6480949A2 | Colombia | A2 | |
| ZA201109112B | South Africa | B | |
| JP2012530952A | Japan | A | |
| EP2535892A1 | European Patent Office (EPO) | A1 | |
| HK1170329A1 | Hong Kong, China | A1 | |
| EP2446435B1 | European Patent Office (EPO) | B1 | |
| RU2012101652A | Russian Federation | A | |
| HK1180100A1 | Hong Kong, China | A1 | |
| ES2426677T3 | Spain | T3 | |
| PL2446435T3 | Poland | T3 | |
| CN103474077A | China | A | |
| CN103489449A | China | A | |
| AU2010264736B2 | Australia | B2 | |
| KR101388901B1 | Republic of Korea | B1 | |
| TWI441164B | Taiwan Province of China | B | |
| CN102460573B | China | B | |
| EP2535892B1 | European Patent Office (EPO) | B1 | |
| ES2524428T3 | Spain | T3 | |
| US8958566B2 | United States of America | B2 | |
| JP5678048B2 | Japan | B2 | |
| PL2535892T3This record | Poland | T3 | |
| MY154078A | Malaysia | A | |
| RU2558612C2 | Russian Federation | C2 | |
| BRPI1009648A2 | Brazil | A2 | |
| CA2766727C | Canada | C | |
| CN103474077B | China | B | |
| CA2855479C | Canada | C | |
| CN103489449B | China | B | |
| BRPI1009648B1 | Brazil | B1 |
Numbers
- Publication, DOCDB
- 2535892
- Publication, EPODOC
- PL2535892T
- Application
- 20120183562
- Application, DOCDB
- 12183562
- Application, EPODOC
- PL20120183562T
Titles2
- English
- Audio signal decoder, method for decoding an audio signal and computer program using cascaded audio object processing stages
- Polish
- Dekoder sygnału audio, sposób dekodowania sygnału audio i program komputerowy wykorzystujący kaskadowe etapy przetwarzania obiektów audio
Classification
- IPC, 4
- G10L19 008
- G10H1 36
- G10L19 20
- H04S7 00