Decoder, encoder and method for informed loudness estimation in object-based audio coding systems
Abstract
A decoder for generating an audio output signal having one or more audio output channels is provided, having a receiving interface for receiving an audio input signal having a plurality of audio object signals, for receiving loudness information on the audio object signals, and for receiving rendering information indicating whether one or more of the audio object signals shall be amplified or attenuated, further having a signal processor for generating the one or more audio output channels of the audio output signal, configured to determine a loudness compensation value depending on the loudness information and depending on the rendering information, and configured to generate the one or more audio output channels of the audio output signal from the audio input signal depending on the rendering information and depending on the loudness compensation value. One or more by-pass audio object signals are employed for generating the audio output signal. Moreover, an encoder is provided.

Term
8.2 yearsto projected expiry
Projected expiry 27 November 2034, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
22 claims: 17 independent, 5 dependent
- 1Zastrzeżenia patentowe 1. Dekoder do generowania wyjściowego sygnału audio zawierającego jeden lub większą liczbę wyjściowych kanałów audio, przy czym dekoder zawiera:interfejs odbiorczy (110) do odbioru wejściowego sygnału audio zawierającego wiele sygnałów obiektów audio, do odbioru informacji głośności sygnałów obiektów audio i do odbioru informacji renderowania wskazującej, jak jeden lub większa liczba sygnałów obiektów audio powinna zostać wzmocniona czy stłumiona, oraz procesor (120) sygnału do generowania jednego lub większej liczby wyjściowych kanałów audio wyjściowego sygnału audio, przy czym procesor (120) sygnału jest skonfigurowany do wyznaczania wartości kompensacji głośności w zależności od informacji głośności i w zależności od informacji renderowania, przy czym procesor (120) sygnału jest skonfigurowany do generowania jednego lub większej liczby wyjściowych kanałów audio wyjściowego sygnału audio z wejściowego sygnału audio w zależności od informacji renderowania i w zależności od wartości kompensacji głośności, przy czym procesor (120) sygnału jest skonfigurowany do generowania jednego lub większej liczby wyjściowych kanałów audio wyjściowego sygnału audio z wejściowego sygnału audio w zależności od informacji renderowania i w zależności od wartości kompensacji głośności tak, że głośność wyjściowego sygnału audio jest równa głośności wejściowego sygnału audio lub tak, że głośność wyjściowego sygnału audio jest bliższa głośności wejściowego sygnału audio niż głośność zmodyfikowanego sygnału audio, jaki byłby wynikiem modyfikacji wejściowego sygnału audio poprzez wzmocnienie lub stłumienie sygnałów obiektów audio wejściowego sygnału audio zgodnie z informacją renderowania.
- 2Dekoder według zastrz. 1, w którym procesor (120) sygnału jest skonfigurowany do generowania zmodyfikowanego sygnału audio poprzez modyfikację wejściowego sygnału audio poprzez wzmacnianie lub tłumienie sygnałów obiektów audio wejściowego sygnału audio zgodnie z informacją renderowania i w którym procesor (120) sygnału jest skonfigurowany do generowania wyjściowego sygnału audio poprzez zastosowanie wartości kompensacji głośności na zmodyfikowanym sygnale audio tak, że głośność wyjściowego sygnału audio jest równa głośności wejściowego sygnału audio, albo tak, że głośność wyjściowego sygnału audio jest bliższa głośności wejściowego sygnału audio niż głośność zmodyfikowanego sygnału audio.
- 3Dekoder według zastrz. 1 albo 2, w którym każdy z sygnałów obiektów audio wejściowego sygnału audio jest przydzielony do dokładnie jednej grupy z dwóch lub większej liczby grup, przy czym każda z dwóch lub większej liczby grup zawiera jeden lub większą liczbę sygnałów obiektów audio wejściowego sygnału audio, przy czym interfejs odbiorczy (110) jest skonfigurowany do odbierania wartości głośności dla każdej grupy z dwóch lub większej liczby grup jako informacji głośności, przy czym procesor (120) sygnału jest skonfigurowany do wyznaczania wartości kompensacji głośności w zależności od wartości głośności każdej z dwóch lub większej liczby grup, oraz przy czym procesor (120) sygnału jest skonfigurowany do generowania jednego lub większej liczby wyjściowych kanałów audio wyjściowego sygnału audio z wejściowego sygnału audio w zależności od wartości kompensacji głośności.
- 4Dekoder według zastrz. 3, w którym co najmniej jedna grupa z dwóch lub większej liczby grup zawiera dwa lub większą liczbę sygnałów obiektów audio.
- 5Dekoder według zastrz. 1 albo 2, w którym każdy z sygnałów obiektów audio wejściowego sygnału audio jest przypisany do dokładnie jednej grupy z więcej niż dwóch grup, przy czym każda z więcej niż dwóch grup zawiera jeden lub większą liczbę sygnałów obiektów audio wejściowego sygnału audio, przy czym interfejs odbiorczy (110) jest skonfigurowany do odbierania wartości głośności dla każdej grupy z więcej niż dwóch grup jako informacji głośności, przy czym procesor (120) sygnału jest skonfigurowany do wyznaczania wartości kompensacji głośności w zależności od wartości głośności każdej z więcej niż dwóch grup i przy czym procesor (120) sygnału jest skonfigurowany do generowania jednego lub większej liczby wyjściowych kanałów audio wyjściowego sygnału audio z wejściowego sygnału audio w zależności od wartości kompensacji głośności.
- 6Dekoder według zastrz. 5, w którym co najmniej jedna grupa z więcej niż dwóch grup zawiera dwa lub większą liczbę sygnałów obiektów audio.
- 7Dekoder według jednego z zastrz. 3 do 6, w którym procesor (120) sygnału jest skonfigurowany do wyznaczania wartości kompensacji głośności według wzoru i»' 1 '* SL = 101og 10 yyy Σ% (1 /=1 lub według wzoru N A£= 101og 10 ^—— /=1 przy czym Δί jest wartością kompensacji głośności. przy czym i oznacza i-ty sygnał obiektów audio z sygnałów obiektów audio, przy czym Li oznacza głośność i-tego sygnału obiektów audio, przy czym gi oznacza pierwszą wagę miksowania dla i-tego sygnału obiektów audio, przy czym hi oznacza drugą wagę miksowania dla i-tego sygnału obiektów audio, przy czym c oznacza wartość stałą, i przy czym N oznacza liczbę.
- 8Dekoder według jednego z zastrz. 3 do 6, w którym procesor (120) sygnału jest skonfigurowany do wyznaczania wartości kompensacji głośności według wzoru N <' K;/W 10 /=1 przy czym Δί oznacza wartość kompensacji głośności, przy czym i oznacza i-ty sygnał obiektów audio sygnałów obiektów audio, przy czym gi oznacza pierwszą wagę miksowania dla i-tego sygnału obiektów audio, przy czym hi oznacza drugą wagę miksowania dla i-tego sygnału obiektów audio, przy czym N oznacza liczbę, oraz przy czym Ki jest zdefiniowane według przy czym Li oznacza głośność i-tego sygnału obiektów audio i przy czym Lref oznacza głośność obiektu referencyjnego.
- 9Dekoder według zastrz. 3 albo 4, w którym każdy z sygnałów obiektów audio wejściowego sygnału audio jest przypisany do dokładnie jednej grupy z dokładnie dwóch grup jako dwie lub większa liczba grup, przy czym każdy z sygnałów obiektów audio wejściowego sygnału audio jest przydzielony albo do grupy obiektów planu pierwszego z dokładnie dwóch grup albo do grupy obiektów tła dokładnie dwóch grup, przy czym interfejs odbiorczy (110) jest skonfigurowany do odbierania wartości głośności grupy obiektów planu pierwszego, przy czym interfejs odbiorczy (110) jest skonfigurowany do odbierania wartości głośności grupy obiektów tła, przy czym procesor (120) sygnału jest skonfigurowany do generowania jednego lub większej liczby wyjściowych kanałów audio wyjściowego sygnału audio z wejściowego sygnału audio w zależności od wartości kompensacji głośności.
- 10Dekoder według zastrz. 9, w którym procesor (120) sygnału jest skonfigurowany do wyznaczania wartości kompensacji głośności według wzoru λ r , „ , , J <· । l^i ;। 0 |l ' AL = 10 log H) ----[Π''n 4 |(| w przy czym ΔL oznacza wartość kompensacji głośności, przy czym Kfgo oznacza wartość głośności grupy obiektu planu pierwszego, przy czym Kbgo oznacza wartość głośności grupy obiektu tła, przy czym mFGo oznacza wzmocnienie renderowania grupy obiektu planu pierwszego, przy czym mBGo oznacza wzmocnienie renderowania grupy obiektu tła.
- 11Dekoder według zastrz. 9, w którym procesor (120) sygnału jest skonfigurowany do wyznaczania wartości kompensacji głośności według wzoru przy czym ΔL oznacza wartość kompensacji głośności, przy czym Lfgo oznacza wartość głośności grupy obiektu planu pierwszego, przy czym Lbgo oznacza wartość głośności grupy obiektu tła przy czym qfgo oznacza wzmocnienie renderowania grupy obiektu planu pierwszego, a przy czym qbgo oznacza wzmocnienie renderowania grupy obiektu tła.
- 12Dekoder według jednego z poprzednich zastrzeżeń, w którym interfejs odbiorczy (110) jest skonfigurowany do odbierania sygnału downmixu zawierającego jeden lub większą liczbę kanałów downmixu jako wejściowego sygnału audio, przy czym jeden lub większa liczba kanałów downmixu zawiera sygnały obiektów audio i przy czym liczba jednego lub większej liczby kanałów downmixu jest mniejsza od liczby sygnałów obiektów audio, przy czym interfejs odbiorczy (110) jest skonfigurowany do odbierania informacji downmixu wskazującej, jak sygnały obiektów audio są miksowane w jednym lub większej liczbie kanałów downmixu, i przy czym procesor (120) sygnału jest skonfigurowany do generowania jednego lub większej liczby wyjściowych kanałów audio wyjściowego sygnału audio z wejściowego sygnału audio w zależności od informacji downmixu, w zależności od informacji renderowania i w zależności od wartości kompensacji głośności.
- 13Dekoder według zastrz. 12, w którym interfejs odbiorczy (110) jest skonfigurowany do odbierania jednego lub większej liczby kolejnych sygnałów obejścia obiektów audio, przy czym jeden lub większa liczba kolejnych sygnałów obejścia obiektów audio nie jest zmiksowana w sygnale downmixu, przy czym interfejs odbiorczy (110) jest skonfigurowany do odbierania informacji głośności wskazującej informację o głośności sygnałów obiektów audio, które są zmiksowane w sygnale downmixu i wskazującej informację o głośności jednego lub większej liczby kolejnych sygnałów obejścia obiektów audio, które nie są zmiksowane w sygnale downmixu, i przy czym procesor (120) sygnału jest skonfigurowany do wyznaczania wartości kompensacji głośności w zależności od informacji o głośności sygnałów obiektów audio, które są zmiksowane w sygnale downmixu i w zależności od informacji o głośności jednego lub większej liczby kolejnych sygnałów obejścia obiektów audio, które nie są zmiksowane w sygnale downmixu.
- 14Dekoder do generowania wyjściowego sygnału audio zawierającego jeden lub większą liczbę wyjściowych kanałów audio, przy czym dekoder zawiera:interfejs odbiorczy (110) do odbierania wejściowego sygnału audio zawierającego wiele sygnałów obiektów audio, do odbierania informacji głośności o sygnałach obiektów audio i do odbierania informacji renderowania wskazującej, czy jeden lub większa liczba sygnałów obiektów audio powinna być wzmocniona czy stłumiona, oraz procesor (120) sygnału do generowania jednego lub większej liczby wyjściowych kanałów audio wyjściowego sygnału audio, przy czym procesor (120) sygnału jest skonfigurowany do wyznaczania wartości kompensacji głośności w zależności od informacji głośności i w zależności od informacji renderowania, i przy czym procesor (120) sygnału jest skonfigurowany do generowania jednego lub większej liczby wyjściowych kanałów audio wyjściowego sygnału audio z wejściowego sygnału audio w zależności od informacji renderowania i w zależności od wartości kompensacji głośności, przy czym interfejs odbiorczy (110) jest skonfigurowany do odbierania sygnału downmixu zawierającego jeden lub większą liczbę kanałów downmixu jako wejściowego sygnału audio, przy czym jeden lub większa liczba kanałów downmixu zawiera sygnały obiektów audio, i przy czym liczba jednego lub większej liczby kanałów downmixu jest mniejsza od liczby sygnałów obiektów audio, przy czym interfejs odbiorczy (110) jest skonfigurowany do odbierania informacji downmixu wskazującej jak sygnały obiektów audio są miksowane w jednym lub większej liczbie kanałów downmixu, oraz przy czym procesor (120) sygnału jest skonfigurowany do generowania jednego lub większej liczby wyjściowych kanałów audio wyjściowego sygnału audio z wejściowego sygnału audio w zależności od informacji downmixu, w zależności od informacji renderowania i w zależności od wartości kompensacji głośności, przy czym interfejs odbiorczy (110) jest skonfigurowany do odbierania jednego lub większej liczby kolejnych sygnałów obejścia obiektów audio, przy czym jeden lub większa liczba kolejnych sygnałów obejścia obiektów audio nie jest zmiksowana w sygnale downmixu, przy czym interfejs odbiorczy (110) jest skonfigurowany do odbierania informacji głośności wskazującej informację o głośności sygnałów obiektów audio, które są zmiksowane w sygnale downmixu i wskazującej informację o głośności jednego lub większej liczby kolejnych sygnałów obejścia obiektów audio, które nie są zmiksowane w sygnale downmixu, oraz przy czym procesor (120) sygnału jest skonfigurowany do wyznaczania wartości kompensacji głośności w zależności od informacji o głośności sygnałów obiektów audio, które są zmiksowane w sygnale downmixu i w zależności od informacji o głośności jednego lub większej liczby kolejnych sygnałów obejścia obiektów audio, które nie są zmiksowane w sygnale downmixu.
- 15Koder, zawierający:jednostkę (210;710) kodowania na bazie obiektów do kodowania wielu sygnałów obiektów audio w celu uzyskania zakodowanego sygnału audio zawierającego wiele sygnałów obiektów audio, oraz jednostkę (220;720;820) kodowania głośności obiektów do kodowania informacji głośności o sygnałach obiektów audio, przy czym informacja głośności zawiera jedną lub większą liczbę wartości głośności, przy czym każda z jednej lub większej liczby wartości głośności zależy od jednego lub większej liczby sygnałów obiektów audio, przy czym każdy z sygnałów obiektów audio zakodowanego sygnału audio jest przydzielony do dokładnie jednej grupy z dwóch lub większej liczby grup, przy czym każda z dwóch lub większej liczby grup zawiera jeden lub większą liczbę sygnałów obiektów audio zakodowanego sygnału audio, przy czym co najmniej jedna grupa z dwóch lub większej liczby grup zawiera dwa lub większą liczbę sygnałów obiektów audio, przy czym jednostka (220;720;820) kodowania głośności obiektów jest skonfigurowana do wyznaczania jednej lub większej liczby wartości głośności informacji głośności poprzez wyznaczanie wartości głośności dla każdej grupy dwóch lub większej liczby grup, przy czym wspomniana wartość głośności wspomnianej grupy wskazuje głośność całkowitą jednego lub większej liczby sygnałów obiektów audio wspomnianej grupy.
- 16Koder, zawierający:jednostkę (210;710) kodowania na bazie obiektów do kodowania wielu sygnałów obiektów audio w celu uzyskania zakodowanego sygnału audio zawierającego wiele sygnałów obiektów audio, oraz jednostkę (220;720;820) kodowania głośności obiektów do kodowania informacji głośności o sygnałach obiektów audio, przy czym informacja głośności zawiera jedną lub większą liczbę wartości głośności, przy czym każda z jednej lub większej liczby wartości głośności zależy od jednego lub większej liczby sygnałów obiektów audio, przy czym jednostka (210;710) kodowania na bazie obiektów jest skonfigurowana do odbierania sygnałów obiektów audio, przy czym każdy z sygnałów obiektów audio jest przypisany do dokładnie jednej z dwóch grup, przy czym każda z dokładnie dwóch grup zawiera jeden lub większą liczbę sygnałów obiektów audio, przy czym co najmniej jedna grupa z dokładnie dwóch grup zawiera dwa lub większą liczbę sygnałów obiektów audio, przy czym jednostka (210;710) kodowania na bazie obiektów jest skonfigurowana do downmixowania sygnałów obiektów audio, zawartych w dokładnie dwóch grupach, w celu uzyskania sygnału downmixu zawierającego jeden lub większą liczbę kanałów downmixu audio jako zakodowanego sygnału audio, przy czym liczba jednego lub większej liczby kanałów downmixu jest mniejsza od liczby sygnałów obiektów audio zawartych w dokładnie dwóch grupach., przy czym jednostka (220;720;820) kodowania głośności obiektów jest skonfigurowana do odbierania jednego lub większej liczby kolejnych sygnałów obejścia obiektów audio, przy czym każdy z jednego lub większej liczby kolejnych sygnałów obejścia obiektów audio jest przydzielony do trzeciej grupy, przy czym każdy z jednego lub większej liczby kolejnych sygnałów obejścia obiektów audio nie jest zawarty w pierwszej grupie i nie jest zawarty w drugiej grupie, przy czym jednostka (210;710) kodowania na bazie obiektów jest skonfigurowana do nie downmixowania jednego lub większej liczby kolejnych sygnałów obejścia obiektów audio w sygnale downmixu, oraz przy czym jednostka (220;720;820) kodowania głośności obiektów jest skonfigurowana do wyznaczania pierwszej wartości głośności, drugiej wartości głośności i trzeciej wartości głośności informacji głośności, przy czym pierwsza wartość głośności wskazuje głośność całkowitą jednego lub większej liczby sygnałów obiektów audio pierwszej grupy, druga wartość głośności wskazuje głośność całkowitą jednego lub większej liczby sygnałów obiektów audio drugiej grupy a trzecia wartość głośności wskazuje głośność całkowitą jednego lub większej liczby kolejnych sygnałów obejścia obiektów audio trzeciej grupy, lub jest skonfigurowana do wyznaczania pierwszej wartości głośności i drugiej wartości głośności informacji głośności, przy czym pierwsza wartość głośności wskazuje głośność całkowitą jednego lub większej liczby sygnałów obiektów audio pierwszej grupy a druga wartość głośności wskazuje głośność całkowitą jednego lub większej liczby sygnałów obiektów audio drugiej grupy i jednego lub większej liczby kolejnych sygnałów obejścia obiektów audio trzeciej grupy.
- 17System zawierający:koder (310) zawierający: jednostkę (210;710) kodowania na bazie obiektów do kodowania wielu sygnałów obiektów audio w celu uzyskania zakodowanego sygnału audio zawierającego wiele sygnałów obiektów audio;oraz jednostkę (220;720;820) kodowania głośności obiektów do kodowania informacji głośności sygnałów obiektów audio, przy czym informacja głośności zawiera jedną lub większą liczbę wartości głośności, przy czym każda z jednej lub większej liczby wartości głośności zależy od jednego lub większej liczby sygnałów obiektów audio, dekoder (320) według jednego z zastrz. 1 do 14, do generowania wyjściowego sygnału audio zawierającego jeden lub większą liczbę wyjściowych kanałów audio, przy czym dekoder (320) jest skonfigurowany do odbierania zakodowanego sygnału audio wejściowego sygnału audio i do odbierania informacji głośności, przy czym dekoder (320) jest skonfigurowany do następnie odbierania informacji renderowania, przy czym dekoder (320) jest skonfigurowany do wyznaczania wartości kompensacji głośności w zależności od informacji głośności i w zależności od informacji renderowania, oraz przy czym dekoder (320) jest skonfigurowany do generowania jednego lub większej liczby wyjściowych kanałów audio wyjściowego sygnału audio z wejściowego sygnału audio w zależności od informacji renderowania i w zależności od wartości kompensacji głośności.
- 18Sposób generowania sygnału obiektów audio zawierającego jeden lub większą liczbę wyjściowych kanałów audio, przy czym sposób obejmuje:odbieranie wejściowego sygnału audio zawierającego wiele sygnałów obiektów audio, odbieranie informacji głośności o sygnałach obiektów audio, odbieranie informacji renderowania wskazującej jak jeden lub większa liczba sygnałów obiektów audio powinna być wzmocniona lub stłumiona, wyznaczanie wartości kompensacji głośności w zależności od informacji głośności i w zależności od informacji renderowania, oraz generowanie jednego lub większej liczby wyjściowych kanałów audio wyjściowego sygnału audio z wejściowego sygnału audio w zależności od informacji renderowania i w zależności od wartości kompensacji głośności. przy czym generowanie jednego lub większej liczby wyjściowych kanałów audio wyjściowego sygnału audio z wejściowego sygnału audio jest przeprowadzane w zależności od informacji renderowania i w zależności od wartości kompensacji głośności tak, że głośność wyjściowego sygnału audio jest równa głośności wejściowego sygnału audio lub tak, że głośność wyjściowego sygnału audio jest bliższa głośności wejściowego sygnału audio niż głośność zmodyfikowanego sygnału audio jaki byłby wynikiem modyfikacji wejściowego sygnału audio poprzez wzmocnienie lub stłumienie sygnałów obiektów audio wejściowego sygnału audio zgodnie z informacją renderowania.
- 19Sposób generowania wyjściowego sygnału audio zawierającego jeden lub większą liczbę wyjściowych kanałów audio przy czym sposób obejmuje:odbieranie wejściowego sygnału audio zawierającego wiele sygnałów obiektów audio, przy czym odbieranie wejściowego sygnału audio jest przeprowadzany poprzez odbieranie sygnału downmixu zawierającego jeden lub większą liczbę kanałów downmixu jako wejściowy sygnał audio, przy czym jeden lub większa liczba kanałów downmixu zawiera sygnały obiektów audio i przy czym liczba jednego lub większej liczby kanałów downmixu jest mniejsza od liczby sygnałów obiektów audio, odbieranie informacji renderowania wskazującej, czy jeden lub większa liczba sygnałów obiektów audio powinna być wzmocniona, czy stłumiona, odbieranie informacji downmixu wskazującej jak sygnały obiektów audio są miksowane w jednym lub większej liczbie kanałów downmixu, odbieranie jednego lub większej liczby kolejnych obejściowych sygnałów obiektów audio, przy czym jeden lub większa liczba kolejnych obejściowych sygnałów obiektów audio nie jest zmiksowana w sygnale downmixu. odbieranie informacji głośności o sygnałach obiektów audio, przy czym informacja głośności wskazuje informację o głośności sygnałów obiektów audio, które są zmiksowane w sygnale downmixu i wskazuje informację o głośności jednego lub większej liczby kolejnych obejściowych sygnałów obiektów audio, które nie są zmiksowane w sygnale downmixu, oraz wyznaczanie wartości kompensacji głośności w zależności od informacji głośności i w zależności od informacji renderowania, przy czym wyznaczanie wartości kompensacji głośności jest przeprowadzane w zależności od informacji o głośności sygnałów obiektów audio, które są miksowane w sygnale downmixu i w zależności od informacji o głośności jednego lub większej liczby kolejnych obejściowych sygnałów obiektów audio, które nie są zmiksowane w sygnale downmixu, oraz generowanie jednego lub większej liczby wyjściowych kanałów audio wyjściowego sygnału audio z wejściowego sygnału audio w zależności od informacji downmixu, w zależności od informacji renderowania i w zależności od wartości kompensacji głośności.
- 20Sposób kodowania, obejmujący:kodowanie wielu sygnałów obiektów audio w celu uzyskania zakodowanego sygnału audio zawierającego wiele sygnałów obiektów audio, i wyznaczanie informacji głośności o sygnałach obiektów audio, przy czym informacja głośności zawiera jedną lub większą liczbę wartości głośności, przy czym każda z jednej lub większej liczby wartości głośności zależy od jednego lub większej liczby sygnałów obiektów audio, przy czym wyznaczanie jednej lub większej liczby wartości głośności informacji głośności jest przeprowadzane poprzez wyznaczanie wartości głośności dla każdej grupy z dwóch lub większej liczby grup, przy czym wspomniana wartość głośności wspomnianej grupy wskazuje głośność całkowitą jednego lub większej liczby sygnałów obiektów audio wspomnianej grupy, kodowanie informacji głośności o sygnałach obiektów audio, przy czym każdy z sygnałów obiektów audio zakodowanego sygnału audio jest przypisany do dokładnie jednej grupy z dwóch lub większej liczby grup, przy czym każda z dwóch lub większej liczby grup zawiera jeden lub większą liczbę sygnałów obiektów audio zakodowanego sygnału audio, przy czym co najmniej jedna grupa z dwóch lub większej liczby grup zawiera dwa lub większą liczbę sygnałów obiektów audio.
- 21Sposób kodowania, obejmujący:odbieranie sygnałów obiektów audio, przy czym każdy z sygnałów obiektów audio jest przypisany do dokładnie jednej z dokładnie dwóch grup, przy czym każda z dokładnie dwóch grup zawiera jeden lub większą liczbę sygnałów obiektów audio, przy czym co najmniej jedna grupa z dokładnie dwóch grup zawiera dwa lub większą liczbę sygnałów obiektów audio, kodowanie wielu sygnałów obiektów audio w celu uzyskania zakodowanego sygnału audio zawierającego wiele sygnałów obiektów audio poprzez downmixowanie sygnałów obiektów audio, zawartych w dokładnie dwóch grupach, w celu uzyskania sygnału downmixu zawierającego jeden lub większą liczbę kanałów audio downmixu jako zakodowanego sygnału audio, przy czym liczba jednego lub większej liczby kanałów downmixu jest mniejsza od liczby sygnałów obiektów audio zawartych w dokładnie dwóch grupach, wyznaczanie informacji głośności o sygnałach obiektów audio, przy czym informacja głośności zawiera jedną lub większą liczbę wartości głośności, przy czym każda z jednej lub większej liczby wartości głośności zależy od jednego lub większej liczby sygnałów obiektów audio, poprzez wyznaczanie pierwszej wartości głośności, drugiej wartości głośności i trzeciej wartości głośności informacji głośności, przy czym pierwsza wartość głośności wskazuje głośność całkowitą jednego lub większej liczby sygnałów obiektów audio pierwszej grupy, druga wartość głośności wskazuje głośność całkowitą jednego lub większej liczby sygnałów obiektów audio drugiej grupy a trzecia wartość głośności wskazuje głośność całkowitą jednego lub większej liczby kolejnych sygnałów obejścia obiektów audio trzeciej grupy, lub poprzez wyznaczanie pierwszej wartości głośności i drugiej wartości głośności informacji głośności, przy czym pierwsza wartość głośności wskazuje głośność całkowitą jednego lub większej liczby sygnałów obiektów audio pierwszej grupy a druga wartość głośności wskazuje głośność całkowitą jednego lub większej liczby sygnałów obiektów audio drugiej grupy i jednego lub większej liczby kolejnych sygnałów obejścia obiektów audio trzeciej grupy, kodowanie informacji głośności sygnałów obiektów audio, odbieranie jednego lub większej liczby kolejnych sygnałów obejścia obiektów audio, przy czym każdy z jednego lub większej liczby kolejnych sygnałów obejścia obiektów audio jest przydzielony do trzeciej grupy, przy czym każdy z jednego lub większej liczby kolejnych sygnałów obejścia obiektów audio nie jest zawarty w pierwszej grupie i nie jest zawarty w drugiej grupie, oraz nie downmixowanie jednego lub większej liczby kolejnych sygnałów obejścia obiektów audio w sygnale downmixu.
- 22Program komputerowy do realizacji sposobu określonego z jednym z zastrz. 18 do 21, gdy jest wykonywany w komputerze lub procesorze sygnału. Fraunhofer-Gesellschaft zur Forderung der angewandten Forschung e. V., Niemcy 1/12 Pełnomocnik:Eliza Stypińska, rzecznik patentowy EP 3 074 971 B1 Z-17157 EP 3 074 971 B1 Z-17157 2/12 FIG 2 EP 3 074 971 B1 Z-17157 3/12 EP 3 074 971 B1 Z-17157 4/12 EP 3 074 971 B1 Z-17157 5/12 ω 5 ο EP 3 074 971 B1 Z-17157 6/12 J ingerencja użytkownika --------------------------i----------------------------► ! czas prawdziwy poziom i sygnału wyjściowego ---------------------estymacja głośności sygnału wyjściowego na bazie sygnału delay czas świadoma estymacja głośności sygnału wyjściowego czas czas FIG 6 EP 3 074 971 B1 Z-17157 7/12 EP 3 074 971 B1 Z-17157 8/12 EP 3 074 971 B1 Z-17157 9/12 EP 3 074 971 B1 Z-17157 10/12 EP 3 074 971 B1 Z-17157 11/12 tłumienie z downmixu (LU) EP 3 074 971 B1 Z-17157 12/12
Independent claims22
317 paragraphs in 5 sections, as filed
Description
The present invention relates to the encoding, processing and decoding of an audio signal, more particularly a decoder, an encoder and a method for conscious loudness estimation in object-based audio coding systems.
[0002] Recently, in the field of audio coding, parametric techniques for bit rate efficient transmission / storage of audio scenes containing audio object signals [BCC, JSC, SAOC, SAOC1, SAOC2] and conscious source separation [ISS1, ISS2, ISS3, ISS4, ISS5, ISS6]. These techniques aim to reconstruct a desired output audio scene or audio scene object based on additional auxiliary information describing the transmitted / stored audio scene and / or source objects in the audio scene. This reconstruction takes place in the decoder using the deliberate source separation method. The reconstructed objects can be combined to create the output audio scene. Depending on how the objects are combined, the perceptual loudness of the output scene may vary.
[0003] In TV and radio broadcasting, the loudness levels of the audio tracks of different programs can be normalized based on various aspects such as peak signal level or loudness level. Depending on the dynamic properties of the signal, two signals with the same peak level may have very different perceived loudness levels. In such a situation, when switching between programs or channels, the variations in signal volume are very annoying and are the main source of end-user complaints about broadcasts.
[0004] It has been proposed in the prior art to normalize all programs on all channels similar to a common reference level using a measure based on the perceptual loudness of the signal. One of such recommendations in Europe is EBU Recommendation R128 [EBU] (hereinafter referred to as R128).
[0005] The recommendation states that the "program loudness", eg, the average loudness in one program (or one advertisement or some other significant program unit) should be equal to a certain level (with little allowed variation). As more and more broadcasters follow this recommendation and the required standardization, the difference in average loudness between programs and channels should be minimized.
[0006] Estimation can be performed in several ways. There are several mathematical models for estimating the perceptual loudness of an audio signal. The EBU Recommendation R128 is based on the model presented in (hereinafter referred to as BS.1770) (see [ITU]) loudness estimation.
[0007] As previously stated, e.g. according to the EBU recommendation R128, the program loudness, e.g. the average loudness in one program, should be equal to a certain level with small allowed variations. However, this leads to significant problems when performing audio rendering, not solved so far in the prior art. Performing audio rendering on the decoder side has a significant impact on the overall / total loudness of the received input audio signal. However, even though the scene is rendered, the overall volume of the received audio signal should remain the same.
[0008] Currently, there is no specific solution to this problem on the decoder side.
[0009] EP 2 146 522 A1 ([EP]), relates to the concept of generating output audio signals using object-based metadata. At least one audio output signal representing a composite of at least two different audio object signals is generated, but this does not offer a solution to this problem.
[0010] WO 2008/035275 A2 ([BRE]) describes an audio system including an encoder that encodes audio objects in a coding unit that generates a downmix audio signal and parametric data representing a plurality of audio objects. The downmix audio signal and the parametric data are sent to a decoder which includes a decoding unit that generates approximate replicas of the audio objects and a rendering unit that generates an output from the audio objects. The decoder further comprises a processor for generating the encoding modification data that is sent to the encoder. The encoder then modifies the encoding of the audio objects, and in particular modifies the parametric data in response to the encoding modification data. This approach makes it possible to control the manipulation of the audio objects by the decoder but performed fully or partially by the encoder. In this way, the manipulation can be performed on actual independent audio objects, rather than rough replicas, thus providing better results.
[0011] EP 2 146 522 A1 ([SCH]) discloses an apparatus for generating at least one audio output signal representing a composite of at least two different audio objects, which comprises a processor for processing the input audio signal to provide an object representation of the input audio signal, wherein this object representation may be generated by parametrically controlled approximation of the original objects using the objects downmix signal. The object manipulator individually manipulates the objects using metadata based on audio objects, pertaining to the individual audio objects, in order to extract the manipulated audio objects. The manipulated audio objects are mixed in the object mixer to ultimately obtain an audio output containing one or more channel signals depending on the particular rendering system.
WO 2008/046531 A1 ([ENG]) describes an audio object encoder for generating an encoded object signal using a plurality of audio objects, which includes a downmix information generator for generating downmix information indicative of the distribution of the plurality of audio objects into at least two downmix channels. an audio object parameter generator for generating object parameters for the audio objects; and an output interface for generating an imported audio output signal using the downmix information and the object parameters. The audio synthesizer uses the downmix information to generate output data useful in creating multiple output channels of a predetermined audio output configuration.
[0013] WO 2008/035275 A2 relates to an audio system comprising an encoder which encodes the audio objects in a coding unit that generates the downmix audio signal into parametric data representing a plurality of audio objects. The downmix audio signal and the parametric data are sent to a decoder which includes a decoding unit that generates approximate replicas of the audio objects and a rendering unit that generates the output from the audio objects.
EP 2 146 522 A1 relates to an apparatus for generating at least one audio output signal representing a combination of at least two different audio objects. The object manipulator individually manipulates the objects using metadata based on the audio objects relating to the individual audio objects in order to extract the manipulated audio objects. The manipulated audio objects are mixed using an object mixer to ultimately obtain an audio output signal containing one or more channel signals depending on the particular rendering system.
WO 2008/046531 A1 relates to an audio object encoder for generating an encoded object signal using a plurality of audio objects, comprising a downmix information generator for generating downmix information indicative of the distribution of the plurality of audio objects into at least two downmix channels. an audio object parameter generator for generating object parameters for the audio objects; and an output interface for generating an imported audio output signal using the downmix information and the object parameters. The audio synthesizer uses the downmix information to generate output data useful for creating multiple output channels of a predetermined audio output configuration. A Guide to Dolby Metadata, 2005, Pages 1 - 28, httD: //www.dolbv.com/uDloadedFiles/
Assets / US / Doc / Professional / 18 Metadata.Guide.pdf refers to metadata providing content producers with the ability to deliver audio to consumers across a range of listening environments. It also offers a choice that allows consumers to tailor their settings to suit their listening environment.
It would be desirable to have an accurate estimate of the mean loudness output or change in mean loudness without delay, and when the program is not changing or the rendering scene does not change, the mean loudness estimate should also remain static.
It is an object of the present invention to provide improved concepts for encoding, processing and decoding an audio signal. The object of the present invention is achieved by a decoder as defined in claim 1, an encoder as defined in claim 15, a system as defined in claim 17, a method as defined in claim 18, a method as defined in claim 19 and a computer program as defined in claim 22. A deliberate method for estimating the loudness of an output signal in an object-based audio coding system is provided. The concepts provided are based on the loudness information of the objects in the audio mix delivered to the decoder. The decoder uses this information together with the rendering information to estimate the loudness of the output signal. This then enables, for example, an estimation of the loudness difference between the default downmix and the rendered output signal. It is then possible to compensate for the difference to obtain an approximately constant loudness in the output signal independent of the rendering information. Loudness estimation in the decoder is fully parametric and computationally light-weight and accurate compared to signal-based loudness estimation concepts.
[0016] Concepts for acquiring loudness information of a specific output scene using purely parametric concepts are provided, which then allow loudness processing without explicit signal-based loudness estimation at the decoder. Moreover, a specific technique of spatial coding of audio objects (SAOC) is described. Spatial Audio Object Coding) standardized by MPEG [SAOC], but the concepts provided can also be used in conjunction with other audio object coding techniques.
[0017] A decoder for generating the audio output signal, including one or more output audio channels, is provided. The decoder comprises a receiving interface for receiving an input audio signal comprising a plurality of audio object signals, for receiving loudness information of the audio object signals, and for receiving rendering information indicative of whether one or more of the audio object signals should be amplified or suppressed. Moreover, the decoder comprises a signal processor for generating one or more audio output channels of the audio output signal. The signal processor is configured to derive a loudness compensation value based on the loudness information and on the rendering information. Moreover, the signal processor is configured to generate one or more output audio channels of the audio output signal from the input audio signal depending on the rendering information and depending on the loudness compensation value.
[0018] According to an embodiment, the signal processor may be configured to generate one or more output audio channels of the audio output signal from the input audio signal depending on the rendering information and depending on the loudness compensation value, such that the volume of the output audio signal is equal to the volume of the audio input signal, or so that the volume of the output audio signal is closer to the volume of the input audio signal than the volume of a modified audio signal that would result from modifying the input audio signal by amplifying or attenuating the audio object signals of the input audio signal according to the rendering information.
According to another embodiment, each of the audio object signals of the input audio signal may be allocated to exactly one group of two or more groups, each of the two or more groups may include one or more audio object signals of the input signal. audio. In such an embodiment, the receiving interface may be configured to receive a loudness value for each group of two or more groups as loudness information, said loudness value indicating the original total loudness of one or more audio object signals of said group. Moreover, the receiving interface may be configured to receive rendering information indicating for at least one group of the two or more groups whether one or more of the audio object signals of said group should be amplified or suppressed by indicating a modified total volume of one or more groups. the audio object signals of said group. Moreover, in such an embodiment, the signal processor may be configured to determine a loudness compensation value depending on the modified total loudness of each of said at least one group of two or more groups and depending on the original total loudness of each of the two or more groups. Moreover, the signal processor may be configured to generate one or more output audio channels of the audio output signal from the input audio signal depending on the modified total loudness of each of the at least one group of the two or more groups and depending on the loudness compensation value.
[0020] In particular embodiments, at least one group of two or more groups may include two or more audio object signals.
[0021] Furthermore, an encoder is provided. The encoder comprises an object-based coding unit for encoding a plurality of the audio object signals to obtain an encoded audio signal comprising the plurality of audio object signals. Further, the encoder comprises an object loudness coding unit for encoding loudness information about the audio object signals. The loudness information comprises one or more loudness values, each of the one or more loudness values depending on one or more audio object signals.
[0022] According to an embodiment, each of the audio object signals of the encoded audio signal may be allocated to exactly one group of two or more groups, each of the two or more groups comprising one or more audio object signals of the encoded audio signal. The object loudness encoding unit may be configured to determine one or more loudness values of the loudness information by determining a loudness value for each group of two or more groups, said loudness value of said group indicating the original total loudness of one or more audio object signals of said object. groups.
[0023] Moreover, a system is provided. The system comprises an encoder according to one of the above-described embodiments, for encoding the plurality of audio object signals to obtain an encoded audio signal comprising the plurality of audio object signals, and for encoding loudness information about the audio object signals. Moreover, the system comprises a decoder according to one of the above-described embodiments for generating an audio output signal including one or more output audio channels. The decoder is configured to receive the encoded audio signal as input audio signal and loudness information. Moreover, the decoder is further configured to receive the rendering information. Moreover, the decoder is configured to determine the loudness compensation value depending on the loudness information and depending on the rendering information. Moreover, the decoder is configured to generate one or more output audio channels of the audio object signal from the input audio signal depending on the rendering information and depending on the loudness compensation value.
[0024] Furthermore, a method for generating an audio output signal including one or more output audio channels is provided. The method includes:
- Receiving an input audio signal containing multiple audio object signals.
- Receiving information on the loudness of signals from audio objects.
- Receiving rendering information indicating whether one or more of the audio object signals should be amplified or suppressed.
- Determining the value of loudness compensation depending on the loudness information and depending on the rendering information. And:
- Generating one or more output audio channels of the audio output signal from the input audio signal depending on the rendering information and depending on a loudness compensation value.
[0025] Moreover, an encoding method is provided. The method includes:
- Encoding of the input audio signal containing multiple audio object signals. And:
- Encoding loudness information of the audio object signals, the loudness information comprising one or more loudness values, the one or more loudness values depending on the one or more audio object signals.
[0026] Furthermore, a computer program is provided for performing the above-described method when it is executed in a computer or a signal processor.
[0027] Advantageous embodiments are disclosed in the dependent claims.
[0028] Embodiments of the present invention are described in more detail below with reference to the figures in which:
Fig. 1 illustrates a decoder for generating an audio output signal including one or more output audio channels according to an embodiment,
Fig. 2 illustrates an encoder according to an embodiment,
Fig. 3 illustrates a system according to an embodiment,
Fig. 4 illustrates a spatial audio object coding system including an SAOC encoder and a SAOC decoder,
Fig. 5 illustrates an SAOC decoder including an auxiliary information decoder, an object separator and a rendering material,
Fig. 6 illustrates the behavior of the output loudness estimate with a loudness change,
Fig. 7 shows a deliberate loudness estimation according to an embodiment, illustrating the components of an encoder and decoder according to an embodiment,
Fig. 8 illustrates an encoder according to another embodiment,
Fig. 9 illustrates an encoder and decoder according to an embodiment relating to SAOC-Dialog Enhancement which includes bypass channels,
Fig. 10 shows a first illustration of a measured change in loudness as a result of applying the provided concepts of parametric estimation of the loudness change in a parametric manner,
Fig. 11 shows a second illustration of the measured change in loudness as a result of applying the provided concepts of parametric estimation of the loudness change in a parametric manner, and
Fig. 12 illustrates another embodiment for performing loudness compensation.
[0029] Before the preferred embodiments are described in more detail, loudness estimation, spatial audio object coding (SAOC) and Dialogue Enhancement (DE) will be described.
[0030] First, loudness estimation will be described.
[0031] As previously stated, the EBU Recommendation R128 is based on the model presented in ITU-R BS.1770 for loudness estimation. This measure will be used as an example, but the concepts described can also be applied to other loudness measures.
[0032] The operation of the loudness estimation according to to BS.1770 is relatively simple and is based on the following main steps [ITU]:
- The input signal xi (or signals in the case of a multi-channel signal) is filtered with a K filter (combination of a shelving filter and a high-pass filter) to obtain the signal (s) yi,
- The mean square energy z of the signal yi is calculated.
- In the case of a multi-channel signal, the weighting Gi of the channel is applied and the weighted signals are summed up. The signal loudness is then defined as
Z = c + 101og<sub>10</sub>J ^, i
- with a constant value æ = -0.691. The output signal is then expressed in the units of "LKFS (Loudness, K-weighted, relative to Full Scale, Loudness, K-weighted, relative to full scale).
[0033] In the above formula, Gi may, for example, be 1 for some channels and Gi may, for example, be 1.41 for some other channels. For example, when considering the left channel, right channel, center channel, left surround channel, and right surround channel, the corresponding Gi weights may, for example, be 1 for the left, right and center channels and may be, for example, 1.41 for the left surround channel. and the right surround channel, see [ITU].
[0034] It can be seen that the loudness value L is closely related to the logarithm of the signal energy.
[0035] In the following, spatial coding of audio objects will be described.
Object-based audio coding concepts provide great flexibility on the path decoder side. An example of the concept of object-based audio coding is Spatial Audio Object Coding (SAOC).
[0037] Fig. 4 illustrates a Spatial Audio Object Coding (SAOC) system including the SAOC encoder 410 and the SAOC decoder 420.
[0038] The SAOC encoder 410 receives N audio object signals Si, .., Sn as input. In addition, the SAOC encoder 410 further receives the "Mixing information D" instruction regarding a method of combining these objects to obtain a downmix signal including the M downmix channels Xi,., Xm. The SAOC encoder 410 extracts some assistance information from the objects and from the mixing operation, and this assistance information is transmitted and / or stored along with the downmix signals.
[0039] The main feature of the SAOC system is that the X downmix signal comprising the Xi, .., Xm downmix channels. creates a semantically meaningful signal. In other words, it is possible to listen to the downmix signal. If, for example, the receiver does not have SAOC decoder functionality, the receiver can still provide the downmix signal at the output anyway.
[0040] Fig. 5 illustrates an SAOC decoder including an auxiliary information decoder 510, an object separator 520 and a render module 530. The SAOC decoder shown in Fig. 5 receives, e.g., from the SAOC encoder, the downmix signal and the side information. The downmix signal can be considered an input audio signal comprising the audio object signals since the audio object signals are mixed in the downmix signal (the audio object signals are mixed in one or more channels of the downmix signal).
[0041] The SAOC decoder may, e.g., try to then (virtually) reconstruct the original objects, e.g., by using the object separator 520, e.g., using decoded side information. These (virtual) object reconstructions S ^ i, .., S ^ n, e.g., the reconstructed audio object signals, are then combined based on rendering information, e.g., the rendering matrix R, to form K output audio channels Yi , .., Yk of the output audio signal Y.
[0042] In SAOC, often audio object signals are, for example, reconstructed, e.g., by using covariance information, e.g., a signal covariance matrix E, which is transmitted from the SAOC encoder to the SAOC decoder.
[0043] For example, the following formula can be used to reconstruct the audio object signals on the decoder side:
S = GX with G ED ED (DEDV where
N is the number of audio object signals,
Nsampies is the number of samples of audio object signals to consider,
M is the number of downmix channels
X is the downmix audio signal, size M x Nsampies,
D is the downmix matrix, size M x N
E is the signal covariance matrix, size N x N defined as E = XX<sup>H.</sup>
S are parametrically reconstructed N audio object signals, size N x Nsampies (·) <sup>H.</sup> is a self-adjoint (Hermitian) operator that represents the conjugate transposition (·).
[0044] Then, the rendering matrix R may be applied on the reconstructed audio object signals S to obtain the audio output channels of the audio output signal Y, e.g. according to the formula:
Y = RS where
K is the number of output audio channels Yi, ..., Yx signals of the audio objects Y,
R is a rendering matrix of size K x N
Y is the output audio signal including K output audio channels, size is Kx Nsampies
[0045] In Fig. 5, an operation for reconstructing objects, e.g., performed by the object separator 520, is specified with the phrase "virtual" or "optional" because it does not necessarily take place, but the desired functionality can be obtained by combining the reconstruction steps. and rendering in the parametric domain (i.e., combining equations).
In other words, instead of first reconstructing the audio object signals using the mix information D and the covariance information E and then applying the rendering information R on the reconstructed audio object signals to obtain the output audio channels Yi, .., Yk, both steps may be performed in one step such that the output audio channels Yi, .., Yk are directly generated from the downmix channels.
[0047] For example, the following formula may be used:
¥ = RGX with G «ED * (DE DV
[0048] In principle, the rendering information R may require any combination of original audio object signals. However, in practice, object reconstructions may include reconstruction errors and the desired output scene may not necessarily be obtained. As a rough rule of thumb, covering many practical cases, the more the desired output scene differs from the downmix signal, the more reconstruction errors will be heard.
[0049] The dialogue improvement (DE) is described below. For example, SAOC technology can be used to implement this scenario. It should be noted that while the name "dialogue enhancement" implies the focus on dialogue-oriented signals, the same principle can also be used with other types of signals.
[0050] In the DE situation, the degrees of freedom in the system are limited with respect to the general case.
[0051] For example, the audio object signals Si,., Sn = S are grouped (and potentially mixed) in two foreground meta objects (FGO) Sfgo and background object (BGO) Sbgo.
[0052] Moreover, the output scene Yi, .., Yk = Y resembles the downmix signal Xi, .., Xm = X. In particular, both signals have the same dimensions, i.e., K = M and the end user can only control the relative mixing levels of two meta objects FGO and BGO. More specifically, the downmix signal is obtained by mixing FGO and BGO with some scalar weights
X h<sub>FGO</sub>S.<sub>FGO</sub> + h<sub>BGC</sub>S.<sub>BG0</sub>and the output scene is acquired similarly with some FGO scalar weighting i
BGO:
Y ~ Sfgo ^ fgo <sup>4</sup> Sbgo ^ bgo
[0053] Depending on the relative values of the mixing weights, the balance between FGO and BGO can vary. For example, in the settings
SfctO <sup>></sup> K.
Sbgo <sup>=</sup>
TGO
[0054] it is possible to increase the relative level of FGO in the mixture. If the FGO is a dialogue, this setting provides dialogue enhancement functionality.
[0055] As a practical example, the BGO may be stadium noise and other ambient noise at sporting events and the FGO is the commentator's voice. The functionality of DE allows the end user to enhance or suppress the commentator level relative to the background.
[0056] The embodiments are based on the finding that using SAOC (or the like) technology in a broadcast situation makes it possible to offer extended signal manipulation functionality to the end-user. More functionality is provided than just channel change and playback level adjustment.
[0057] One possibility of using DE technology is briefly discussed above. If the broadcast signal which is the downmix signal for SAOC has a normalized level, e.g. according to R128, different programs have similar loudness when no SAOC processing is applied (or the rendering description is the same as the downmix description). However, when some processing (SAOC) is applied, the output signal differs from the default downmix signal and the output volume may be different from the default downmix volume. From the end-user point of view, this may lead to a situation where the output signal between channels or programs, again, may contain undesirable jumps or differences. In other words, the benefits of the standardization applied by the broadcaster are partially lost.
[0058] This problem is not only specific to the SAOC or the DE situation, but may also arise in other audio coding concepts that allow the end-user to interfere with the content. However, in many cases it does not cause any harm if the output signal is at a different volume than the default downmix.
[0059] As previously stated, the total program loudness of the audio input signal should be equal to a predetermined level with little variation allowed. However, as already stated, this leads to significant problems when audio rendering is performed as the rendering can have a significant influence on the overall / total loudness of the received input audio signal. However, despite the rendering of the scene, the total volume of the received audio signal should remain the same.
[0060] One approach will be to estimate the loudness of a signal during its reproduction, and with a proper time integration concept, the estimation may converge to the true mean loudness after some time. However, the time required for convergence is problematic from the end-user point of view. When the loudness estimation changes, even when no changes to the signal are applied, the compensation for the change in loudness should also be responsive and change its behavior. This would lead to an output signal with a temporarily varying average loudness which can be perceived as quite annoying.
[0061] Fig. 6 illustrates the operation of the estimation of the output loudness when changing the loudness. Among other things, an estimation of the loudness of the output signal based on the signal is presented, which illustrates the effect of the solution just described. Estimation approaches the correct estimation quite slowly. Preferably, instead of estimating the loudness of the output from the signal, it would be a deliberate estimation of the loudness of the output which immediately and correctly determines the loudness of the output.
[0062] Specifically, in Fig. 6, user intervention, e.g. the dialog item level, changes at a point in time T by increasing the value. The true level of the output signal, and accordingly the volume, varies at the same moment in time. When an output loudness estimation is made from an output signal with a certain time integration time, the estimate will change gradually and will be correct after a certain delay. During this time, the estimation values change and cannot be reliably used for further processing of the output signal, e.g., for loudness level correction.
[0063] As already stated, it would be desirable to have an accurate estimate of the initial loudness average or change of the mean loudness without delay and when the program is not changing or the rendering scene is unchanged, the mean loudness estimation should also remain static. In other words, when some loudness change compensation is applied, the compensation parameter should only change when either the program changes or there is some user interaction.
[0064] The desired behavior is illustrated in the bottom illustration of Fig. 6 (conscious loudness estimation of the output signal). The estimation of the loudness of the output signal should change as soon as the user input changes.
[0065] Fig. 2 illustrates an encoder according to an embodiment.
[0066] The encoder includes an object-based coding unit 210 for encoding the plurality of audio object signals to obtain an encoded audio signal including the plurality of audio object signals.
[0067] Further, the encoder comprises an object loudness coding unit 220 for coding loudness information about the audio object signals. The loudness information comprises one or more loudness values, each of the one or more loudness values depending on one or more audio object signals.
[0068] According to an embodiment, each of the audio object signals of the encoded audio signal is allocated to exactly one group of two or more groups, each of the two or more groups comprising one or more audio object signals of the encoded audio signal. The object loudness coding unit 220 is configured to determine one or more loudness values of the loudness information by determining a loudness value for each group of two or more groups, said loudness value of said group indicating the original total loudness of the one or more audio object signals from the group. said group.
[0069] Fig. 1 illustrates a decoder for generating an audio output signal including one or more output audio channels according to an embodiment.
The decoder comprises a receiving interface 110 for receiving an input audio signal including a plurality of audio object signals, for receiving loudness information of the audio object signals, and for receiving rendering information indicative of whether one or more of the audio object signals should be amplified or suppressed.
[0071] Moreover, the decoder comprises a signal processor 120 for generating one or more audio output channels of the audio output signal. Signal processor 120 is configured to determine a loudness compensation value based on the loudness information and depending on the rendering information. Moreover, the signal processor 120 is configured to generate one or more output audio channels of the audio output signal from the input audio signal depending on the rendering information and depending on the loudness compensation value.
[0072] According to an embodiment, the signal processor 120 is configured to generate one or more output audio channels of the output audio signal from the input audio signal depending on the rendering information and depending on the loudness compensation value, such that the loudness of the output audio signal is equal to the loudness of the input. audio signal, or so that the volume of the output audio signal is closer to the volume of the input audio signal, than the loudness of the modified audio signal that would result from modifying the input audio signal by amplifying or attenuating the audio object signals of the input audio signal according to the rendering information.
[0073] According to another embodiment, each of the audio object signals of the input audio signal is assigned to exactly one group of two or more groups, each of the two or more groups including one or more audio object signals of the input audio signal.
[0074] In such an embodiment, the receiving interface 110 is configured to receive a loudness value for each group of two or more groups as loudness information, said loudness value indicating the original total loudness of one or more audio object signals of said group. Moreover, the receiving interface 110 is configured to receive rendering information indicating for at least one group of the two or more groups whether one or more of the audio object signals of said group should be amplified or suppressed by indicating a modified total volume of one or more. the number of audio object signals of said group. Moreover, in such an embodiment, the signal processor 120 is configured to determine a loudness compensation value depending on the modified total loudness of said at least one group of two or more groups and according to the original total loudness of each of the two or more groups. Moreover, the signal processor 120 is configured to generate one or more output audio channels of the audio output from the input audio signal depending on the modified total loudness of each of the at least one of the two or more groups and depending on the loudness compensation value.
[0075] In specific embodiments, at least one group of two or more groups comprises two or more audio object signals.
[0076] There is a direct relationship between the energy ei of the audio object signal and the loudness Li of the audio object signal and according to the formulas:
L. = c +10 log<sub>10</sub> e<sub>and</sub>, e<sub>f</sub> = 1O ^<sup>-</sup>^<sup>10</sup> where c is a constant value.
[0077] The embodiments are based on the following findings: The different audio object signals of the input audio signal may have different loudness and therefore different energies. For example, if a user wants to increase the loudness of one of the audio object signals, the rendering information can be suitably adjusted and increasing the loudness of this audio object signal increases the energy of that audio object. This would lead to an increase in the volume of the audio output signal. In order to keep the overall loudness constant, loudness compensation must be performed. In other words, the modified audio signal that would result from applying rendering information to the input audio signal would have to be matched. However, the exact amplification effect of one of the audio object signals on the total loudness of the modified audio signal depends on the original loudness of the amplified audio object signal, e.g., an audio object signal, the loudness of which is increased. If the original loudness of this object corresponds to an energy that was quite low, the effect on the overall loudness of the input audio signal will be small. However, if the original loudness of this object corresponds to an energy that was quite high, the effect on the overall loudness of the input audio signal will be significant.
[0078] Two examples may be considered. In both examples, the input audio signal comprises two audio object signals and in both examples, by using the rendering information, the energy of one of the audio object signals is increased by 50%.
[0079] In a first example, the first audio object signal is 20% and the second audio object signal is 80% of the total energy of the input audio signal. However, in the second example, the first audio object, the first audio object signal is 40% and the second audio object signal is 60% of the total energy of the audio object signal. In both examples, these contributions can be obtained from the loudness information about the signals of the audio objects, since there is a direct relationship between loudness and energy.
[0080] In the first example, a 50% increase in the energy of the first audio object leads to a modified audio signal that is generated by amplifying rendering information about the audio object signal has a total energy of 1.5 x 20% + 80% = 110 % of the energy of the input audio signal.
[0081] In a second example, an increase of 50% of the energy of the first audio object leads to a modified audio signal that is generated by applying the rendering information on the input audio signal has a total energy of 1.5 x 40% + 60% = 120 % of the energy of the input audio signal.
[0082] Thus, after applying the rendering information to the input audio signal, in the first example, the total energy of the modified audio signal only needs to be reduced by 9% (10/110) in order to have the same energy in both the input audio signal and the output audio signal. the audio signal, while in the second example, the total energy of the modified audio signal must be reduced by 17% (20/120). For this, a loudness compensation value can be calculated.
[0083] For example, the loudness compensation value may be a scalar that is applied to all audio output channels of the audio output signal.
[0084] According to an embodiment, the signal processor is configured to generate a modified audio signal by modifying the input audio signal by amplifying or attenuating the audio object signals of the input audio signal according to the rendering information. Moreover, the signal processor is configured to generate the output audio signal by applying a loudness compensation value on the modified audio signal such that the loudness of the output audio signal is equal to the loudness of the input audio signal, or such that the loudness of the output audio signal is closer to the loudness of the input audio signal than the loudness of the output audio signal. modified audio signal.
For example, in the first example above, the loudness compensation value lcv may for example be set to lcv = 10/11 and a multiplication factor of 10/11 may be applied to all channels resulting from rendering the input audio channels according to information rendering.
[0086] Accordingly, for example, in the second example above, the loudness compensation value may, for example, be set to = 10/12 = 5/6, and a multiplication factor of 5/6 may be applied to all channels resulting from the rendering. audio input channels according to the render information.
[0087] In other embodiments, each of the audio object signals may be allocated to one of a plurality of groups and a loudness value may be transmitted for each of the groups, indicating the total loudness value of the audio object signals of said group. If the rendering information indicates that the energy of one of the groups is suppressed or enhanced, e.g., 50% boost as above, an increase in total energy can be calculated and a loudness compensation value determined as described above.
[0088] For example, according to an embodiment, each of the audio object signals of an input audio signal is assigned to exactly one group of exactly two groups as two or more groups. Each of the audio object signals of the input audio signal is assigned either to a foreground object group of exactly two groups or to a background object group of exactly two groups. Receiving interface 110 is configured to receive the original total loudness of one or more audio object signals of the foreground object group. In addition, the receiving interface 110 is configured to receive the original total loudness of one or more audio object signals of the background object group. Moreover, the receiving interface 110 is configured to receive rendering information indicating for at least one group of exactly two groups whether one or more of the audio object signals of each of said at least one group should be amplified or suppressed by indicating a modified total loudness. one or more audio object signals of said group.
[0089] In such an embodiment, the signal processor 120 is configured to determine a loudness compensation value in dependence on the modified total loudness of each of said at least one group as a function of the original total loudness of one or more foreground group audio object signals and in dependence of from the original total loudness of one or more audio object signals of the background object group. Moreover, the signal processor 120 is configured to generate one or more output audio channels of the audio output signal from the input audio signal depending on the modified total loudness of each of the at least one group and depending on the loudness compensation value.
[0090] According to some embodiments, each of the audio object signals is allocated to one of three or more groups and the receiving interface may be configured to receive a loudness value for each of the three or more groups, indicating the total loudness of the audio object signals of said group. .
[0091] According to an embodiment, in order to determine the total loudness value of two or more audio object signals, for example, an energy value corresponding to the loudness value is determined for each audio object signal, the energy values of all loudness values are summed to obtain the sum of the energy and the value of the loudness value. the loudness value corresponding to the sum of the energies is determined as the total loudness value of the two or more audio object signals. For example, a formula can be used
Ζ, = σ + 101ο<sub>& οβ</sub>"^ ΙΟ<sup>11</sup>'-*<sup>0</sup>
[0092] In some embodiments, loudness values are transmitted for each of the audio object signals, or each of the audio object signals is assigned to one of two or more groups, with each group being transmitted a loudness value.
[0093] However, in some embodiments, no loudness value is transmitted for one or more audio object signals or for one or more groups including audio object signals. Instead, the decoder may assume, for example, that those audio object signals or groups of audio object signals for which no loudness value is transmitted have a predetermined loudness value. The decoder may, e.g., base any further determination on this predetermined loudness value.
[0094] According to an embodiment, the receiver interface 110 is configured to receive a downmix signal including one or more downmix channels as the input audio signal, the one or more downmix channels including the audio object signals and the number of audio object signals being less than the number of one or more downmix channels. The receiving interface 10 is configured to receive downmix information indicative of how the audio object signals are mixed into one or more downmix channels. Moreover, the signal processor 120 is configured to generate one or more output audio channels of the audio output signal from the input audio signal depending on the downmix information, depending on the rendering information and depending on the loudness compensation value. In a particular embodiment, signal processor 120 may, for example, be configured to calculate a loudness compensation value depending on the downmix information.
[0095] For example, the downmix information may be a downmix matrix. In embodiments, the decoder may be a SAOC decoder. In such embodiments, the receiving interface 110 may, for example, be further configured to receive covariance information, e.g., a covariance matrix, as described above.
[0096] With respect to the rendering information indicating whether one or more audio object signals should be enhanced or suppressed, it should be noted that, for example, information that indicates how one or more audio object signals may be enhanced or suppressed. is the render information. For example, the render matrix R, e.g., the SAOC render matrix, is render information.
[0097] Fig. 3 illustrates a system according to an embodiment.
[0098] The system includes an encoder 310 according to one of the above-described embodiments for encoding the plurality of audio object signals to obtain an encoded audio signal comprising the plurality of audio object signals.
[0099] Moreover, the system comprises a decoder 320 according to one of the above-described embodiments for generating an audio output signal including one or more output audio channels. The decoder is configured to receive the encoded audio signal as input audio signal and loudness information. Moreover, the decoder 320 is configured to receive further rendering information. Moreover, the decoder 320 is configured to determine a loudness compensation value depending on the loudness information and depending on the rendering information. Moreover, the decoder 320 is configured to generate one or more output audio channels of the audio output signal from the input audio signal depending on the rendering information and depending on the loudness compensation value.
[0100] Fig. 7 illustrates a conscious loudness estimation according to an embodiment. To the left of transport stream 730, the components of an object-based audio coding encoder are shown. In particular, an object-based coding unit 710 ("object-based audio encoder") and an object-loudness coding unit 720 ("object-loudness estimation") are illustrated.
[0101] The transport stream 730 itself includes the loudness information L, the downmix information D, and the object-based audio coding output 710 B.
[0102] To the right of transport stream 730, signal processor components of an object-based audio encoding decoder are shown. The decoder receiving interface is not illustrated. An output loudness estimator 740 and an object-based audio decoding unit 750 are shown. The output loudness estimator 740 may be configured to derive a loudness compensation value. The object-based audio decoding unit 750 may be configured to determine a modified audio signal from the audio signal input to the decoder by applying the rendering information R. Applying a loudness compensation value on the modified audio signal to compensate for the overall loudness change caused by rendering is not shown in Fig. 7.
[0103] The encoder input signal consists of the input objects S in a minimum. The system estimates the loudness of each object (or some other loudness related information, such as object energies), e.g., with the object loudness coding unit 720, and this L information is transmitted and / or stored. (It is also possible that the loudness of the objects is provided as an input to the system and the system of estimation in the system may be omitted).
[0104] In the embodiment of Fig. 7, the decoder receives at least the object loudness information, e.g., render information R describing the mixing of the objects in the output signal. Based on it, e.g., the output loudness estimator 740 estimates the loudness of the output signal and provides this information at its output.
[0105] Downmix information D may be provided as rendering information, in which case the loudness information provides an estimation of the loudness of the downmix signal. It is also possible to provide the downmix information as an input to the object loudness estimation and transmit and / or record it together with the loudness information. The output loudness estimation may then estimate the loudness of the downmix signal and the rendered output signal simultaneously and provide the two values or the difference thereof as the loudness output information. The difference value (or the reciprocal thereof) describes the required compensation that should be applied on the rendered output signal in order to make its loudness similar to that of the downmix signal. The object loudness information may additionally include coefficient correlation information between different objects, and this correlation information can be used in loudness output estimation to obtain a more accurate estimation.
[0106] In the following, a preferred embodiment for a dialogue enhancement application is described.
[0107] In a dialogue enhancement application as described above, the input audio object signals are grouped and partially downmixed to form two meta objects, FGO and BGO, which can then be simply summed to obtain a final downmix signal.
[0108] After the description of the SAOC [SAOC], the N audio object signals are represented as a matrix S with size N x Nsampies and the downmix information as matrix D with size M xN. The downmix signals can then be obtained as X = DS.
[0109] The downmix information D may now be divided into two parts
D <sup>= +</sup> for meta-objects.
[0110] Since each column of the matrix D corresponds to the original signal of the audio objects, the two component downmix matrices can be obtained by setting the columns that correspond to the different meta object to zero (assuming that no original object can be present in both meta objects). In other words, the columns corresponding to the BGO meta objects are set to zero in Dfgo and vice versa.
[0111] These new downmix matrices describe how both meta objects can be extracted from the input objects, namely:
<sup>=</sup> ^ BGO ~ ^ SGO ^ 'and the actual downmix is simplified to
ZK - iJ<sub>FG0</sub> T ^><sub>SG0</sub> .
[0112] It may also be considered that an object decoder (e.g., SAOC) tries to reconstruct the meta-objects ^ FGO ~ $ FGO <sup>anC</sup>^ ^ BGO ^ BGO <sup>1</sup> and that the DE-specific rendering can be written as a combination of the reconstruction of these two meta objects:
Y ~ Sfgo ^ fgo <sup>+</sup> Sbgo ^ bgo & Sfgc ^ fgo Sbgo ^ bgo
[0113] The object loudness estimation takes the two meta objects Sfgo and Sbgo at its input and estimates the loudness of each of them: where Lfgo is the loudness (total / total) Sfgo and Lbgo is the loudness (total / total) Sbgo. These volume values are transmitted and / or saved.
Alternatively, by using one of the meta objects, e.g., FGO, as reference, it is possible to calculate the loudness difference of the two objects, e.g., <sup>=</sup> L.<sub>bgo</sub> ~ L<sub>FG0</sub> .
[0115] This one value is then transmitted and / or stored.
[0116] Fig. 8 illustrates an encoder according to another embodiment. The encoder in Fig. 8 includes an object downmix module 811 and an object auxiliary information estimator 812. Moreover, the encoder of Fig. 8 further includes an object loudness coding unit 820. Further, the encoder of Fig. 8 includes a meta audio object mixer 805.
[0117] The encoder of Fig. 8 uses intermediate meta audio objects as an input to the object loudness estimation. In embodiments, the encoder of Fig. 8 may be configured to generate two meta audio objects. In other embodiments, the encoder of Fig. 8 may be configured to generate three or more meta audio objects.
[0118] Among other things, the concepts provided provide the new possibility that an encoder can, e.g., estimate the average loudness of all input objects. The objects can, for example, be mixed into one downmix signal that is transmitted. The concepts provided further provide the new feature that the object volume and downmix information can, e.g., be included in the object coding support information that is transmitted.
[0119] The decoder may, e.g., use the coding support information of the objects to (virtual) separate objects and reconnect the objects using the rendering information.
[0120] Furthermore, the concepts provided offer a new possibility, or that the downmix information can be used to estimate the loudness of the default downmix signal, the rendering information and the received object loudness can be used to estimate the average loudness of the output signal and / or the loudness change can be estimated from these two values. Or that the downmix and rendering information can be used to estimate the loudness change from the default downmix is another new feature of the shared concepts.
[0121] Moreover, the disclosed concepts provide the new possibility that the decoder output can be modified to compensate for the loudness change such that the average loudness of the modified signal corresponds to the average loudness of the default downmix.
[0122] A specific embodiment related to SAOC-DE is illustrated in Fig. 9. The system receives input audio object signals, downmix information and object grouping information in meta entities. Based on this, the audio object meta mixer 905 creates two meta objects Sfgo and Sbgo. It is possible that the part of the signal that is processed with SAOC is not the whole signal. For example, in a 5.1 channel configuration, SAOC may be run on a subset of channels, such as the front channel (left, right, center), while the other channels (left surround, right surround, and two low frequency effects channels) are routed around (via bypass) SAOC and delivered unchanged. These channels not processed by SAOC are marked as Xbpass. Potential bypass channels must be provided for the encoder for a more accurate estimation of loudness information.
[0123] The bypass channels may be supported in various manners.
[0124] For example, the bypass channels may, e.g., form an independent meta object. This allows you to define rendering so that all three meta objects are scaled independently.
[0125] Or, for example, the bypass channels may, eg, be connected to one of the other two meta entities. The render settings also control part of the bypass channel. For example, if the dialogue is improved, it may be important to combine the bypass channels with the meta background object: Xbco = Sbgo + 'Kbypass.
[0126] Or, for example, the bypass channels may, eg, be ignored.
[0127] According to an embodiment, an audio coding unit 210 based on the encoder object is configured to receive audio object signals, each of the audio object signals being assigned to exactly one of two groups, each of exactly two groups including one or more. the number of audio object signals. Further, the object-based audio coding unit 210 is configured to downmix the audio object signals contained in exactly two groups to obtain a downmix signal including one or more audio downmix channels as an encoded audio signal, the number of one or more downmix channels is less than the number of audio object signals contained in exactly two groups. The object loudness encoding unit 220 is arranged to receive one or more consecutive audio object bypass signals, each of the one or more consecutive audio object bypass signals being assigned to a third group, each of the one or more consecutive object bypass signals. audio is not included in the first group and is not included in the second group, the object-based audio coding unit 210 is configured not to downmix one or more consecutive audio object bypass signals in the downmix signal.
[0128] In an embodiment, the object loudness coding unit 220 is configured to determine the first loudness value, the second loudness value and the third loudness value of the loudness information, the first loudness value indicating the total loudness of one or more of the first group audio object signals. the second volume value and the loudness value of the loudness information indicates the total loudness of one or more of the second group audio object signals, and the third loudness value value of the loudness information indicates the total loudness of the one or more third group audio object signals. In another embodiment, the object-based audio coding unit 220 is configured to determine the first loudness value and the second loudness value of the loudness information. the first loudness value indicates the total loudness of the one or more audio object signals of the first group and the second loudness value indicates the total loudness of the one or more audio object signals of the second group and the one or more consecutive audio object bypass signals of the third group.
[0129] According to an embodiment, the decoder receiving interface 110 is configured to receive the downmix signal. Moreover, the receiving interface 110 is configured to receive one or more consecutive audio object bypass signals, with the one or more consecutive audio object signals not being mixed into the downmix signal. Moreover, the receiving interface 110 is configured to receive loudness information indicative of loudness information of the audio object signals that are mixed in the downmix signal and indicative of loudness information of one or more consecutive audio object bypass signals that are not mixed in the downmix signal. Moreover, the signal processor 120 is configured to determine a loudness compensation value depending on the loudness information of the audio object signals that are mixed in the downmix signal and depending on the loudness information of one or more consecutive audio object bypass signals that are not mixed in the signal. downmix.
[0130] Fig. 9 illustrates an encoder and decoder according to an embodiment related to SAOC-DE which includes bypass channels. Among others, the encoder of Fig. 9 includes an SAOC encoder 902.
[0131] In the embodiment of Fig. 9, potential coupling of the bypass channels to other meta entities takes place at "bypass connect" blocks 913,914 producing meta entities Xfgo and Xbgo with defined portions of the bypass channels included.
[0132] The perceptual loudness Lbypass, Lfgo, and Lbgo of both of these meta objects are estimated in loudness estimation units 921, 922, 923. This loudness information is then sent to the corresponding coding in the estimator 925 of the meta object loudness information and then transmitted and / or stored.
[0133] The actual SAOC encoder and decoder operate as expected to retrieve the object side information from the objects, form the X downmix signal, and transmit to and / or write the information at the decoder. Potential bypass channels are transmitted to and / or stored together with other information in the decoder.
[0134] The SAOC-DE 945 decoder receives the gain value "Dialog Gain" as a user setting. Based on this setting and the received downmix information, the SAOC decoder 945 determines the render information. The SAOC 945 decoder then produces the rendered output scene as the Y signal. In addition, it produces the gain factor (and delay value) that should be applied to the potential Xbypass signals.
[0135] A "bypass containment" unit 955 receives this information along with the rendered output scene and the bypass signals and produces the full scene output signal. The SAOC 945 decoder also produces a set of meta object gain values, the number of which depends on the grouping of the meta objects and the desired form of the loudness information.
[0136] The gain values are provided to a mixture loudness estimator 960, which also receives meta object loudness information from an encoder.
[0137] The mix loudness estimator 960 may then determine the desired loudness information, which may include, but is not limited to, the downmix loudness, the rendered output scene loudness, and / or the loudness difference between the downmix signal and the rendered output scene.
[0138] In some embodiments, the loudness information alone is sufficient, while in other embodiments, it is desirable to process the full output in dependence on the determined loudness information. This processing can, for example, compensate for any potential loudness difference between the downmix signal and the rendered output scene. Such processing, e.g., with loudness processing unit 970, would make sense in a broadcast situation as it would reduce variations in perceived loudness of the signal independent of user interference (setting the input "dialogue gain).
[0139] The processing related to the loudness in a specific embodiment includes many new functions. Among others, the FGO, BGO and potential bypass channels are pre-mixed to the final channel configuration so that a downmix can be performed simply by adding two pre-mixed signals together (e.g., with downmix matrix 1 factors), which is a new feature. Moreover, as a further new function, the mean volumes of the FGO and BGO are estimated and the difference is calculated. Moreover, the objects are mixed into the downmix signal that is transmitted. Moreover, as a further new function, the loudness difference information is included in the side information that is transmitted. (new) Moreover, the decoder uses the auxiliary information to (virtual) separate the objects and reconnect the objects using the rendering information which is based on the downmix information and the user input modification gain. Moreover, as a further new function, the decoder uses the modification gain and the transmitted loudness information to estimate the change in average loudness at the system output compared to the default downmix.
[0140] A formal description of the embodiments is provided below.
[0141] Assuming that the loudness values of the objects behave similarly to the logarithm of the energy values in summing objects, i.e., the loudness values have to be transformed into the linear domain, added to there, and finally converted back into the linear domain logarithmic . Based on the definition of BS.1770, a loudness measure will now be presented (for the sake of simplicity, the number of channels is set to one, but the same principle can be applied to multi-channel signals with appropriate channel summation).
[0142] The loudness of the ith K-filtered signal z with mean square energy e, is defined as = c + 101og<sub>10</sub> e<sub>p</sub> where c is the displacement constant. For example, c may be -0.691. From this it follows that the signal energy can be determined from the loudness with.
<img file="PL3074971T3_D0001.tif" />
[0143] Sum energy of N uncorrelated signals
N <sup>WITH</sup>CATFISH <sup>=</sup> Σ <sup>WITH</sup>if = l is then ^ = Σ ^ Σ> θ<sup>(Λ</sup>-*<sup>10</sup>. / = 1 / = 1 and the loudness of this sum signal is then<sup>L.</sup>sum = c +! 01og<sub>lo</sub> e<sub>CATFISH</sub> = c + 101og<sub>10</sub>^10<sup>(with</sup>’<sup>c) / 1</sup>°. /=1
[0144] If the signals are not uncorrelated, the correlation coefficients Cj have to be considered when approximating the energy of the sum signal as
NN <sup>e</sup>SUM = ΣΣ<sup>β</sup>Λ7 '/ = 1 7 = 1 where the energy level ej between objects i and j is defined as
<img file="PL3074971T3_D0002.tif" />
where 1 <Ci, j <1 is the correlation coefficient between the two objects ii j. When the two objects are uncorrelated, the correlation coefficient is 0, and when the two objects are identical, the correlation coefficient is 1.
[0145] Moreover, by extending the model to the mix weights gi to be applied to the signals in the mixing operation, i.e.,
AND <sup>WITH</sup>SUM = Σ ^>
? = 1 the energy of the sum signal will be
NN / = 1 y = la the loudness of the mixture signal can be obtained from this, as before, with Aum - <sup>c</sup> +10 log<sub>10</sub> e <sub>CATFISH</sub> ,
[0146] The difference between the loudness of the two signals can be estimated as
[0147] If the loudness definition is now applied as before it can be written as = (c +10 log<sub>10</sub> e ,.) - (c +10 log<sub>10</sub> e<sub>f</sub>)
P = 101og<sub>10</sub> - which, as can be seen, is a function of the energy of the signals. If now it is desired to estimate the loudness difference between two mixtures
<img file="PL3074971T3_D0003.tif" />
with potentially different mix weights of g, and hi, this can be estimated as
<img file="PL3074971T3_D0004.tif" />
[0148] When the objects are uncorrelated (Ci; j = 0, vi ΐ ji Ci, j = 1, vi = j), the estimate of the difference takes the form
<img file="PL3074971T3_D0005.tif" />
[0149] In the following, differential encoding will be considered.
[0150] It is possible to code the loudness values per object as loudness differences with respect to the selected reference object:
Ki <sup>=</sup> Lj - Lref 'where Lref is the loudness of the reference object. This encoding is advantageous if it does not result in the need for absolute loudness values, since now it is only necessary to transmit one less value and an estimate of the loudness difference can be written as
<img file="PL3074971T3_D0006.tif" />
M (45) = 101o<sub>glo NN x</sub>—Iw — i / {^ i ^ i}<sup>1</sup>
ΣΪΑμΑ ™ / = 1 7 = 1 or in the case of objects not correlated as
Σ ^<sup>210</sup>^<sup>0</sup>
AL (A, B) = 10 logs<sub>l0</sub> --------- Σ ^ ιο / = 1
[0151] In the following, we consider situations for improving dialogue.
[0152] Again, consider the situation of applying dialogue enhancement. The freedom to define rendering information at the decoder is limited only to changing the levels of the two meta objects. Furthermore, assume that the two meta objects are uncorrelated, i.e., Cfgo, bgo = 0. If the downmix weights of the meta objects are hFGo and hBGo and they are rendered with fFGo and fBGo gains, the volume of the output downmix relative to the default is z 'FFGO-ήη »<sub>?</sub> (L.<sub>boo</sub>-c) IW ί Λ - 1 ΠIna fpGo} f ________<sup>+</sup> fBGO 1 θ ______
A £ (4.7i) -101og<sub>10</sub> and<sub>FG0</sub>-<sub>c</sub>y<sub>Ui</sub> (lbco-φο hpGO <sup>1</sup> θ + ^ BGO 1 θ / -2, "<sup>£</sup>rao '<sup>10</sup> -i,
And A1 "fpGO ^ <sup>+</sup> fiGO ^ <sup>1υ1</sup>υ £<sub>10</sub> L.<sub>FG0</sub>M>, LjgoIIO ^ FGO 1 θ <sup>+</sup> ^ BGO θ
[0153] This is then also a desirable compensation if the same volume is desired in the output downmix as in the default.
[0154] ∆L (A, B) may be considered as the loudness compensation value that can be transmitted by the decoder signal processor 120. ∆Σ (Α, B) may also be called the loudness change value, thus the actual compensation value may be an inverted value. also may the term "loudness compensation factor" be used for it? Thus, the lcv loudness compensation value mentioned earlier in this document will correspond to gDelta below.
[0155] For example, ^ δ = 10<sup>-Δ £ (A</sup>, B) / 20 1 / Δί (Α, B) can be used as a multiplication factor on each channel of the modified audio signal that results from applying the rendering information on the signal of the audio objects. This equation for gdelt works in the linear domain. In the logarithmic domain, the equation would be different such as 1 / ΔΣ (Α, B and applied accordingly.
[0156] If the downmix operation is simplified such that two meta objects can be mixed with unit weights to obtain a downmix signal, i.e., hFGo and hBGo = 1, now the rendering gains for these two objects are denoted as gFGo and qbgo . This simplifies the equation for changing the volume to
AL (Λ, B) = 101og<sub>in</sub> , <sub>n</sub>(<sup>L.</sup>FGO- ^ o <sub>2</sub> (and<sub>B</sub>ao ~ W
8fgo ^ · ® <sup>+</sup> 8bgo ^ ®,. (J.FGO-) M, ΛΙ-BGO-cW +10 logio
J, LFSO<sup>m</sup> 2 . ^<sup>L.</sup>KGO<sup>m</sup>
Sfgo <sup>1</sup> θ <sup>+</sup> 8bgo ^ θ, "'iFO' BGO<sup>m</sup> +10
[0157] Again, Δί (Α, B) must be considered as a loudness compensation value that is determined by signal processor 120.
[0158] Generally, the gFso can be considered a rendering gain for an FGO (group of foreground objects) and the geso can be considered as a rendering gain for a BGO (group of background objects).
[0159] As stated previously, it is possible to transmit loudness differences instead of absolute loudness. Let us define the reference loudness as the meta loudness of the object FGO Lref = Lfgo, i.e., Kfgo = Lfgo - Lref = 0 and Kbgo = Lbgo - Lref = Lbgo - Lfgo. Now the volume change is at
2 1(\<sup>K.</sup>‘<sup>and</sup>°<sup>n</sup> + 10
[0160] It may also be, as with SAOC-DE, that the two meta objects do not have individual scaling factors, but one of the objects is left unmodified while the other is suppressed in order to obtain the correct mixing ratio between the objects. With this render setting, the output will be louder than the default mix and the change in volume will be
-2 by 2 and / ++ 0+<sup>1</sup>°
Δΐμ.Β) = 101ο<sub>8</sub>,/^<sup>+</sup>^--.
+ 10 where
FGO 1 8 FGO ^ 8 BGO
FGO —8 BGO if 8FGO <8 BGO
8BGO <sup><</sup> 8FGO
BGO - 8 FGO
[0160] This embodiment is already quite simple and is quite indifferent to the loudness measure used. The only actual requirement is that the loudness values are summed in the exponential domain. It is possible to transmit / save the energy values of the signals instead of the loudness values as they are closely related.
[0162] In each of the above formulas, ∆L (A, B) can be considered as a loudness compensation value that can be transmitted by signal processor 120 to a decoder.
[0163] In the following, exemplary cases are considered. The accuracy of the shared concepts is illustrated with two exemplary signals. Both signals include a 5.1 downmix with surround channels and LFE bypassed from SAOC processing.
[0164] Two main approaches have been used: one ("three-member) with three meta FGO objects, BGOs and bypass channels, e.g.,
X ^ FGO <sup>Y</sup> Χ-BGO + ^ SYPASS>
[0165] And a second ("2-membered) with two meta objects, e.g., X - X<sub>FGO</sub> + X<sub>BG0</sub> .
[0166] In the 2-member approach, bypass channels can, for example, be mixed together with the BGO to estimate the loudness of a meta object. The loudness of both (or all three) objects as well as the loudness of the downmix signal are estimated and the values are recorded.
[0167] Rendering instructions are of the form ¥ <sup>=</sup> Sfgo ^ fgo + Sbgo ^ -bgo + Sbgo ^ bypass
Y <sup>_</sup> Sfgo ^ · FGO <sup>+</sup> SbGO ^ BGO for both approaches respectively.
[0168] The gain values are, e.g., determined according to
FGO, dg<sub>FG0</sub> > 1 g<sub>FG0</sub> , otherwise ', otherwise whereby the FGO gFGo gain varies between -24 and +24 dB.
[0169] The output circuit is rendered, the loudness is measured and the attenuation from the loudness of the downmix signal is calculated.
[0170] This result is shown in Fig. 10 and Fig. 11 by a blue line with circular marks. Fig. 10 shows the first illustration and Fig. 11 shows a second illustration of the measured loudness change and the result of applying the provided concepts for loudness change estimation in a purely parametric manner.
[0171] Then, the downmix attenuation is parametrically estimated using the stored loudness values of meta objects and downmix and rendering information. The loudness estimation for the three meta objects is represented by a green line with square markers and the estimate using the loudness of two meta objects is represented by a red line with asterisks.
[0172] It can be seen in the figures that the 2 and 3-membered approaches provide virtually identical results and both approximate the measured values quite well.
[0173] The concepts provided have many advantages. For example, the concepts provided make it possible to estimate the loudness of a mixture signal from the loudness of the component signals that make up the mixture. The advantage of this is that the component signal can be estimated once and the loudness estimation of the mixture signal can be obtained parametrically for any mixture without the need for an actual loudness estimation based on the signal. This offers a significant improvement in the computational performance of the entire system where an estimation of the loudness of different mixtures is needed. For example, when the end user changes the rendering settings, the estimation of the volume of the output signal is available immediately.
[0174] In some applications, such as those following the EBU Recommendation R128, the average loudness of the overall program is important. If the estimation of the loudness at the receiver, e.g., in a broadcast situation, is made based on the received signal, the estimation converges to the average loudness only after receiving the entire program. For this reason, any loudness compensation will either contain errors or exhibit temporal variability. With the proposed loudness estimation of the component objects and the transmission of loudness information, it is possible to estimate the average loudness of the mixture in the receiver without delay.
[0175] If it is desired that the average loudness of the output signal remains (approximately) constant regardless of changes in the rendering information, the concepts provided make it possible to determine a compensation factor for this purpose. The calculations in the decoder for this are negligible from the point of view of computational complexity, and thus it is possible to add functionality to any decoder.
[0176] There are cases where the absolute loudness level of the output is not important, but it is important to determine the change in loudness with respect to the reference scene. In such cases, the absolute levels of the objects are not important, but their relative levels are important. This enables one of the objects to be defined as a reference object and to represent the loudness of other objects in dependence on the loudness of that reference object. This has some advantages in terms of transmitting and / or recording loudness information.
[0177] First, it is not necessary to transmit a reference loudness level. In this case, using two meta objects, it halves the amount of data to be transferred. The second advantage relates to the potential quantization and representation of the loudness value. Since the absolute levels of objects can be almost any, the absolute values of the loudness can also be almost any. However, relative loudness values, on the other hand, are taken to have an average of 0 and have a fairly good distribution around the mean. The difference between the representations makes it possible to define the quantization grid of the related representation in a way that is potentially more precise, with the same number of bits consumed for the quantized representation.
[0178] Fig. 12 illustrates another embodiment for performing loudness compensation. In Fig. 12, loudness compensation may be performed, e.g., to compensate for loudness loss. For this, for example, the values DE_loudness_diff_dialogue (= Kfgo) and DE_loudness_diff_background (= Kbgo) from DE_control_info can be used. In this case, DE_control_info may specify Advanced Clean Audio Dialogue Enhancement (DE) control information.
[0179] Loudness compensation is achieved by applying a gain value "g" on the SAOC-DE output and bypass channels (in the case of a multi-channel signal).
[0180] In the embodiment of Fig. 12, it is implemented as follows:
[0181] A constrained modification gain mG is used to determine the effective gains for the foreground object (FGO, e.g., dialog) and for the background object (BGO, e.g., environment). This is accomplished by "gain mapping" block 1220 which produces the mFGo and mBGo gain values.
[0182] The "output loudness estimator" block 1230 uses the loudness information Kfgo and Kbgo, and the effective gains values mFGo and mBGo to estimate this potential loudness change compared to the default downmix event. The change is then mapped to a "loudness compensation factor" that is used on the output channels to produce the final "output signals".
[0183] For loudness compensation, the following steps are used:
- Receive the limited value of the mG gain from the SAOC-DE decoder (defined in article 12.8 "Modification range control for SAOC-DE [DE]) and determine the applied FGO / BGO gains:
m<sub>FG0</sub>= l,
- Get meta volume information of Kfgo and Kbgo objects.
- Calculate the change in input volume compared to the default downmix
AZ = 101og<sub>10</sub> ^ FGO / K<sub>BCO</sub>/
Kfgo / K<sub>8GO</sub>/ <sup>/10</sup>+10 <sup>/,()</sup>
- Calculate the loudness compensation gain Ρδ = 1 °<sup>-</sup>°,°<sup>5Δ</sup>^ - Calculate the scaling factors
<img file="PL3074971T3_D0007.tif" />
- where if cłiannel i belongs to SAOC-DE output <sup>m</sup>BooS & if channel i is a by-pass channel
- and N is the total number of output channels. In Fig. 12, the gain matching is divided into two steps: the gain of the potential "bypass channels" is matched with mBGo before combining them with the "output SAOC-DE channels" and then a common gA gain is applied to all connected channels. This is only a potential rearrangement of the gain fit operation, and g here combines both fitting steps into a single gain match.
- Apply the g scaling values on Yuull audio channels consisting of Ysaoc "SAOC-DE output channels" and potential Ybyass time-aligned "bypass channels": Yfull = Ysaoc U Ybyfass.
[0184] Applying the scaling value g on the Yfull audio channels is performed by a gain matching unit 1240.
[0185] The ALs calculated above can be considered as the loudness compensation value. Generally, mFGo indicates rendering gain for a foreground object FGO (group of foreground objects) and mBGo indicates rendering gain for a background object BGO (group of background objects).
[0186] While certain aspects have been described in the context of an apparatus, it is evident that these aspects also represent a description of a corresponding method wherein the block or device corresponds to a method step or a feature of a method step. Likewise, aspects described in the context of a method also represent a description of a corresponding block or item or feature of the corresponding device.
[0187] The distributed signal according to the invention can be stored in a digital storage medium or can be transmitted on a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.
[0188] Depending on certain implementation requirements, embodiments of the invention can be implemented in hardware or in software. The implementation may be by means of a non-transient storage medium such as digital storage media, e.g. or are capable of such interaction) with the computer programmed system, such that an appropriate method is performed.
[0189] Some embodiments may include a non-transient data medium containing electronically readable control signals that are operable to interact with a programmable computer system, such that one of the methods described herein is performed.
[0190] Generally, embodiments of the present invention can be implemented as a computer program product with program code, the program code being operable to perform one of the methods of the invention, when the computer program product runs on a computer. For example, the program code may be stored on a machine-readable medium.
[0191] Other embodiments include a computer program for performing one of the methods described herein stored on a machine-readable medium.
In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program product runs on a computer.
[0193] Another embodiment, therefore, is a data medium (or a digital storage medium or a computer readable medium) containing, recorded thereon, the computer program for performing one of the methods described herein.
[0194] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals may e.g. be configured to be transmitted over a data link, e.g.
[0195] Another embodiment includes processing means, for example, a computer or programmable logic device configured or adapted to perform one of the methods described herein.
[0196] Another embodiment includes a computer in which the computer program for performing one of the methods described herein is installed.
[0197] In some embodiments, a programmable logic device (e.g., a user programmable logic table) may be used to perform some or all of the functions of the methods described herein. In some embodiments, a user programmable logic table may interact with a microprocessor to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware device.
Literature
[0198]
[BCC] C. Faller and F. Baumgarte, Binaural Cue Coding - Part II: Schemes and applications, IEEE Trans. On Speech and Audio Proc., Vol. 11, no. 6, Nov. 2003.
[EBU] EBU Recommendation R 128 Loudness normalization and permitted maximum level of audio signals, Geneva, 2011.
[JSC] C. Faller, Parametric Joint-Coding of Audio Sources, 120th AES Convention, Paris, 2006.
[ISS1] M. Parvaix and L. Girin: Informed Source Separation of underdetermined instantaneous Stereo Mixtures using Source Index Embedding, IEEE ICASSP, 2010.
[ISS2] M. Parvaix, L. Girin, JM Brossier: A watermarking-based method for informed source separation of audio signals with a single sensor, IEEE Transactions on Audio, Speech and Language Processing, 2010.
[ISS3] A. Liutkus and J. Pinel and R. Badeau and L. Girin and G. Richard: Informed source separation through spectrogram coding and data embedding, Signal Processing Journal, 2011.
[ISS4] A. Ozerov, A. Liutkus, R. Badeau, G. Richard: Informed source separation: source coding meets source separation, IEEE Workshop on Applications of Signal Processing to Audio and Acoustics, 2011.
[ISS5] S. Zhang and L. Girin: An Informed Source Separation System for Speech Signals, INTERSPEECH, 2011.
[ISS6] L. Girin and J. Pinel: Informed Audio Source Separation from Compressed Linear Stereo Mixtures, AES 42nd International Conference: Semantic Audio, 2011.
[ITU] International Telecommunication Union: Recommendation ITU-R BS.1770-3 Algorithms to measure audio program loudness and true-peak audio level, Geneva, 2012.
[SAOC1] J. Herre, S. Disch, J. Hilpert, O. Hellmuth: From SAC To SAOC - Recent Developments in Parametric Coding of Spatial Audio, 22nd Regional UK AES Conference, Cambridge, UK, April 2007.
[SAOC2] J. EngdegSrd, B. Resch, C. Falch, O. Hellmuth, J. Hilpert, A. Holzer, L. Terentiev, J. Breebaart, J. Koppens, E. Schuijers and W. Oomen: Spatial Audio Object Coding (SAOC) - The Upcoming MPEG Standard on Parametric Object Based Audio Coding, 124th AES Convention, Amsterdam 2008.
[SAOC] ISO / IEC, MPEG audio technologies - Part 2: Spatial Audio Object Coding (SAOC), ISO / IEC JTC1 / SC29 / WG11 (MPEG) International Standard 23003-2.
[EP] EP 2146522 A1: S. Schreiner, W. Fiesel, M. Neusinger, O. Hellmuth, R. Sperschneider, Apparatus and method for generating audio output signals using object based meta data, 2010.
[DE] ISO / IEC, MPEG audio technologies - Part 2: Spatial Audio Object Coding (SAOC) - Amendment 3, Dialogue Enhancement, ISO / IEC 23003-2: 2010 / DAM 3, Dialogue Enhancement.
[BRE] WO 2008/035275 A2.
[SCH] EP 2 146 522 A1.
[ENG] WO 2008/046531 A1.
Fraunhofer-Gesellschaft zur Forderung der angewandten Forschung e. V., Germany
Proxy:
Eliza Stypińska, patent attorney
EP 3 074 971 B1
Z-17157
Contents5
33 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33
74 members in 19 offices
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 13194664 | European Patent Office (EPO) | A | |
| 2014075787 | European Patent Office (EPO) | W |
Members74
| Document | Office | Kind | |
|---|---|---|---|
| EP2879131A1 | European Patent Office (EPO) | A1 | |
| CA2900473A1 | Canada | A1 | |
| CA2931558A1 | Canada | A1 | |
| WO2015078956A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2015078964A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW201525990A | Taiwan Province of China | A | |
| AU2014356475A1 | Australia | A1 | |
| TW201535353A | Taiwan Province of China | A | |
| KR20150123799A | Republic of Korea | A | |
| EP2941771A1 | European Patent Office (EPO) | A1 | |
| US2015348564A1 | United States of America | A1 | |
| CN105144287A | China | A | |
| MX2015013580A | Mexico | A | |
| AR098558A1 | Argentina | A1 | |
| AU2014356467A1 | Australia | A1 | |
| KR20160075756A | Republic of Korea | A | |
| JP2016520865A | Japan | A | |
| AR099360A1 | Argentina | A1 | |
| CN105874532A | China | A | |
| AU2014356475B2 | Australia | B2 | |
| MX2016006880A | Mexico | A | |
| US2016254001A1 | United States of America | A1 | |
| EP3074971A1 | European Patent Office (EPO) | A1 | |
| AU2014356467B2 | Australia | B2 | |
| HK1217245A1 | Hong Kong, China | A1 | |
| JP2017502324A | Japan | A | |
| TWI569259B | Taiwan Province of China | B | |
| TWI569260B | Taiwan Province of China | B | |
| RU2015135181A | Russian Federation | A | |
| EP2941771B1 | European Patent Office (EPO) | B1 | |
| KR101742137B1 | Republic of Korea | B1 | |
| PT2941771T | Portugal | T | |
| BR112015019958A2 | Brazil | A2 | |
| BR112016011988A2 | Brazil | A2 | |
| ES2629527T3 | Spain | T3 | |
| MX350247B | Mexico | B | |
| JP6218928B2 | Japan | B2 | |
| PL2941771T3 | Poland | T3 | |
| ZA201604205B | South Africa | B | |
| RU2016125242A | Russian Federation | A | |
| CA2900473C | Canada | C | |
| EP3074971B1 | European Patent Office (EPO) | B1 | |
| US9947325B2 | United States of America | B2 | |
| RU2651211C2 | Russian Federation | C2 | |
| ES2666127T3 | Spain | T3 | |
| PT3074971T | Portugal | T | |
| KR101852950B1 | Republic of Korea | B1 | |
| JP6346282B2 | Japan | B2 | |
| US2018197554A1 | United States of America | A1 | |
| PL3074971T3This record | Poland | T3 | |
| MX358306B | Mexico | B | |
| RU2672174C2 | Russian Federation | C2 | |
| CA2931558C | Canada | C | |
| US10497376B2 | United States of America | B2 | |
| US2020058313A1 | United States of America | A1 | |
| CN105874532B | China | B | |
| CN111312266A | China | A | |
| US10699722B2 | United States of America | B2 | |
| US2020286496A1 | United States of America | A1 | |
| CN105144287B | China | B | |
| CN112151049A | China | A | |
| US10891963B2 | United States of America | B2 | |
| US2021118454A1 | United States of America | A1 | |
| BR112015019958B1 | Brazil | B1 | |
| MY189823A | Malaysia | A | |
| US11423914B2 | United States of America | B2 | |
| BR112016011988B1 | Brazil | B1 | |
| US2022351736A1 | United States of America | A1 | |
| MY196533A | Malaysia | A | |
| US11688407B2 | United States of America | B2 | |
| US2023306973A1 | United States of America | A1 | |
| CN111312266B | China | B | |
| US11875804B2 | United States of America | B2 | |
| CN112151049B | China | B |
Numbers
- Publication
- 3074971
- Application
- 14805849
Titles2
- English
- DECODER, ENCODER AND METHOD FOR INFORMED LOUDNESS ESTIMATION IN OBJECT-BASED AUDIO CODING SYSTEMS
- Polish
- Dekoder, koder i sposób świadomej estymacji głośności w systemie kodowania audio na bazie obiektów
Classification
- CPC, 5
- G10L19/008
- G10L19/0017
- G10L19/005
- G10L19/265
- H03G3/20
- IPC, 1
- G10L19 008