Decoder, encoder and method for informed loudness estimation employing by-pass audio object signals in object-based audio coding systems
Abstract
This record has no abstract on file.
Term
8.2 yearsto projected expiry
Projected expiry 27 November 2034, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
1 claim: 1 independent, 0 dependent
- 1Zastrzeżenia patentowe 1. Dekoder do generowania wyjściowego sygnału audio zawierającego jeden lub większą liczbę wyjściowych kanałów audio, przy czym dekoder zawiera:interfejs (110) odbiorczy do odbierania wejściowego sygnału audio zawierającego wiele sygnałów obiektów audio, odbierania informacji głośności o sygnałach obiektów audio oraz do odbierania informacji odtwarzania, które wskazują czy jeden lub większa liczba sygnałów obiektów audio powinna zostać wzmocniona lub stłumiona, i procesor (120) sygnału do generowania jednego lub większej liczby wyjściowych kanałów audio wyjściowego sygnału audio, przy czym interfejs (110) odbiorczy jest przystosowany do odbierania downmiksowanego sygnału zawierającego jeden lub większą liczbę downmiksowanych kanałów jako wejściowy sygnał audio, przy czym jeden lub większa liczba downmiksowanych kanałów zawiera sygnały obiektów audio i gdzie liczba jednego lub większej liczby downmiksowanych kanałów jest mniejsza niż liczba sygnałów obiektów audio, przy czym interfejs (110) odbiorczy jest przystosowany do odbierania informacji downmiksu wskazujących, w jaki sposób sygnały obiektów audio są miksowane wewnątrz jednego lub większej liczby downmiksowanych kanałów, przy czym interfejs (110) odbiorczy jest przystosowany do odbierania jednego lub większej liczby kolejnych obejściowych sygnałów obiektów audio, gdzie jeden lub większa liczba kolejnych obejściowych sygnałów obiektów audio nie jest miksowana wewnątrz downmiksowanego sygnału, przy czym interfejs (110) odbiorczy jest przystosowany do odbierania informacji głośności wskazujących informacje o głośności sygnałów obiektów audio, które są miksowane wewnątrz downmiksowanego sygnału i wskazywania informacji o głośności jednego lub większej liczby kolejnych obejściowych sygnałów obiektów audio, które nie są miksowane wewnątrz downmiksowanego sygnału, przy czym procesor (120) sygnału jest przystosowany do określania wartości kompensacji głośności w zależności od informacji o głośności sygnałów obiektów audio, które są miksowane wewnątrz downmiksowanego sygnału i w zależności od informacji głośności jednego lub większej liczby kolejnych obejściowych sygnałów obiektów audio, które nie są miksowane wewnątrz downmiksowanego sygnału, i przy czym procesor (120) sygnału jest przystosowany do generowania jednego lub większej liczby wyjściowych kanałów z audio wyjściowego sygnału audio w zależności od informacji downmiksu, w zależności od informacji odtwarzania i w zależności od wartości kompensacji głośności. 2. Dekoder według zastrz.1, w którym procesor (120) sygnału jest przystosowany do generowania jednego lub większej liczby wyjściowych kanałów audio wyjściowego sygnału audio z wejściowego sygnału audio w zależności od informacji odtwarzania i w zależności od wartości kompensacji głośności w taki sposób, że głośność wyjściowego sygnału audio jest równa głośności wejściowego sygnału audio, lub w taki sposób, że głośność wyjściowego sygnału audio jest bardziej zbliżona do głośności wejściowego sygnału audio niż głośność zmodyfikowanego sygnału audio, który byłby uzyskiwany przez modyfikowanie wejściowego sygnału audio poprzez wzmocnienie lub stłumienie sygnałów obiektów audio wejściowego sygnału audio, zgodnie z informacjami odtwarzania, 3. Dekoder według zastrz. 2, w którym procesor (120) sygnału jest przystosowany do generowania zmodyfikowanego sygnału audio przez modyfikowanie wejściowego sygnału audio poprzez wzmocnienie lub stłumienie sygnałów obiektów audio wejściowego sygnału audio, zgodnie z informacjami odtwarzania, i przy czym procesor (120) sygnału jest przystosowany do generowania wyjściowego sygnału audio przez zastosowanie wartości kompensacji głośności na zmodyfikowanym sygnale audio tak, aby głośność wyjściowego sygnału audio była równa głośności wejściowego sygnału audio, lub tak, aby głośność wyjściowego sygnału audio była bardziej zbliżona do głośności wejściowego sygnału audio niż głośność zmodyfikowanego sygnału audio. 4. Dekoder według jednego z powyższych zastrz., w którym każdy z sygnałów obiektów audio wejściowego sygnału audio jest przypisany do dokładnie jednej grupy spośród dwóch lub większej liczby grup, przy czym każda spośród dwóch lub większej liczby grup zawiera jeden lub większą liczbę sygnałów obiektów audio wejściowego sygnału audio, w którym interfejs (110) odbiorczy jest przystosowany do odbierania wartości głośności każdej grupy spośród dwóch lub większej liczby grup jako informacji głośności, w którym procesor (120) sygnału jest przystosowany do określania wartości kompensacji głośności w zależności od wartości głośności każdej spośród dwóch lub większej liczby grup, i przy czym procesor (120) sygnału jest przystosowany do generowania jednego lub większej liczby wyjściowych kanałów audio wyjściowego sygnału audio z wejściowego sygnału audio, w zależności od wartości kompensacji głośności. 5. Dekoder według dowolnego z powyższych zastrz., w którym co najmniej jedna grupa spośród dwóch lub większej liczby grup zawiera dwa lub większą liczbę sygnałów obiektów audio. 6. Dekoder według dowolnego z powyższych zastrz., w którym każdy z sygnałów obiektów audio wejściowego sygnału audio jest przypisany do dokładnie jednej grupy spośród dokładnie dwóch grup jako dwie lub większą liczba grup, gdzie każdy z sygnałów obiektów audio wejściowego sygnału audio przypisany jest albo do grupy obiektów pierwszoplanowych spośród dokładnie dwóch grup albo do grupy obiektów tła spośród dokładnie dwóch grup, w którym interfejs (110) odbiorczy jest przystosowany do odbierania wartości głośności grupy obiektów pierwszoplanowych, w którym interfejs (110) odbiorczy jest przystosowany do odbierania wartości głośności grupy obiektów tła, w którym procesor (120) sygnału jest przystosowany do określania wartość kompensacji głośności w zależności od wartości głośności grupy obiektów pierwszoplanowych i w zależności od wartości głośności grupy obiektów tła, i przy czym procesor (120) sygnału jest przystosowany do generowania jednego lub większej liczby kanałów wyjściowych wyjściowego sygnału audio z wejściowego sygnału audio, w zależności od wartości kompensacji głośności. 7. Dekoder według zastrz. 6, przy czym procesor (120) sygnału jest przystosowany do określania wartości ΔL zmiany głośności zgodnie ze wzorem gdzie KFGO wskazuje wartość głośności grupy obiektów pierwszoplanowych, gdzie KBGO wskazuje wartość głośności grupy obiektów tła, gdzie mFGO wskazuje wzmocnienie odtwarzania grupy obiektów pierwszoplanowych, i gdzie mBGO wskazuje wzmocnienie odtwarzania grupy obiektów tła. 8. Dekoder według zastrz. 6, w którym procesor (120) sygnału jest przystosowany do określania wartości ΔL kompensacji głośności zgodnie ze wzorem Ł^co/iO + 10 ificoflo gdzie LFGO wskazuje wartość głośności grupy obiektów pierwszoplanowych, gdzie LBGO wskazuje wartość głośności grupy obiektów tła, gdzie gFGO wskazuje wzmocnienie odtwarzania grupy obiektów pierwszoplanowych, i gdzie gBGO wskazuje wzmocnienie odtwarzania grupy obiektów tła. 9. Koder zawierający zespół (210;710) kodowania obiektowego do kodowania wielu sygnałów obiektów audio w celu zakodowania sygnału audio zawierającego wiele sygnałów obiektów audio, i zespół (220;720;820) kodowania głośności obiektów do kodowania informacji głośności w sygnałach obiektów audio, przy czym informacje głośności zawierają jedną lub większą liczbę wartości głośności, gdzie każda spośród jednej lub większej liczby wartości głośności zależy od jednego lub większej liczby sygnałów obiektów audio, przy czym zespół (210;710) kodowania obiektowego jest przystosowany do odbierania sygnałów sygnały obiektów audio, przy czym każdy z sygnałów obiektów audio jest przypisany do dokładnie jednej spośród dwóch lub większej liczby grup, przy czym każda spośród dwóch lub większej liczby grup zawiera jeden lub większą liczbę sygnałów obiektów audio, przy czym zespół (210;710) kodowania obiektowego jest przystosowany do downmiksowania sygnałów obiektów audio, zawartych w dwóch lub większej liczbie grup, w celu uzyskania downmiksowanego sygnału zawierającego jeden lub większą liczbę downmiksowanych kanałów audio jako zakodowany sygnał audio, przy czym liczba jednego lub większej liczby downmiksowanych kanałów jest mniejsza niż liczba sygnałów obiektów audio zawartych w dwóch lub większej liczbie grup, przy czym zespół (220;720;820) kodowania głośności obiektów jest przypisany do odbierania jednego lub większej liczby kolejnych obejściowych sygnałów obiektów audio, przy czym każdy spośród jednego lub większej liczby kolejnych obejściowych sygnałów obiektów audio jest przypisany do trzeciej grupy, przy czym każdy spośród jednego lub większej liczby kolejnych obejściowych sygnałów obiektów audio nie jest zawarty w pierwszej grupie i nie jest zawarty w drugiej grupie, przy czym zespół (210;710) kodowania obiektowego jest przystosowany do nie przeprowadzania downmiksowania jednego lub większej liczby kolejnych obejściowych sygnałów obiektów audio wewnątrz downmiksowanego sygnału, i gdzie zespół (220;720;820) kodowania głośności obiektów jest przystosowany do określania pierwszej wartości głośności, drugiej wartości głośności i trzeciej wartości głośności informacji głośności, przy czym pierwsza wartość głośności wskazuje całkowitą głośność jednego lub większej liczby sygnałów obiektów audio pierwszej grupy, druga wartość głośności wskazuje całkowitą głośność jednego lub większej liczby sygnałów obiektów audio drugiej grupy, a trzecia wartość głośności wskazuje całkowitą głośność jednego lub większej kolejnych obejściowych sygnałów obiektów audio trzeciej grupy, lub jest przystosowany do określania pierwszej wartości głośności i drugiej wartości głośności informacji głośności, przy czym pierwsza wartość głośności wskazuje całkowitą głośność jednego lub większej liczby sygnałów obiektów audio pierwszej grupy, a druga wartość głośności wskazuje całkowitą głośność jednego lub większej liczby sygnałów obiektów audio drugiej grupy i jednego lub większej liczby kolejnych obejściowych sygnałów obiektów audio trzeciej grupy. 10. Koder według zastrz. 9, w którym dwie lub większa liczba grup to dokładnie dwie grupy, w którym każdy z sygnałów obiektów audio jest przypisany do dokładnie jednej grupy spośród dokładnie dwóch grup, przy czym każda z dokładnie dwóch grup zawiera jeden lub większą liczbę sygnałów obiektów audio, w którym zespół (210;710) kodowania obiektowego jest przystosowany do downmiksowania sygnałów obiektów audio zawartych w dokładnie dwóch grupach, w celu uzyskania downmiksowanego sygnału zawierającego jeden lub większą liczbę downmiksowanych kanałów audio jako zakodowany sygnał audio, przy czym liczba jednego lub większej liczby downmiksowanych kanałów jest mniejsza niż liczba sygnałów obiektów audio zawartych w dokładnie dwóch grupach. 11. Układ zawierający: koder (310) określony w zastrz. 9 albo 10 do kodowania wielu sygnałów obiektów audio w celu uzyskania zakodowanego sygnału audio zawierającego wiele sygnałów obiektów audio, i dekoder (320) określony w jednym z zastrz. 1 do 8 do generowania wyjściowego sygnału audio zawierającego jeden lub większą liczbę wyjściowych kanałów audio, przy czym dekoder (320) jest przystosowany do odbierania zakodowanego sygnału audio jako wejściowego sygnału audio oraz do odbierania informacji głośności, przy czym dekoder (320) jest przystosowany do dalszego odbierania informacji odtwarzania, przy czym dekoder (320) jest przystosowany do określania wartości kompensacji głośności w zależności od informacji głośności i w zależności od informacji odtwarzania, i przy czym dekoder (320) jest przystosowany do generowania jednego lub większej liczby wyjściowych kanałów audio wyjściowego sygnału audio z wejściowego sygnału audio w zależności od informacji odtwarzania i w zależności od wartości kompensacji głośności. 12. Sposób generowania wyjściowego sygnału audio zawierającego jeden lub większą liczbę wyjściowych kanałów audio, przy czym sposób obejmuje: odbieranie wejściowego sygnału audio zawierającego wiele sygnałów obiektów audio, odbieranie informacji głośności wskazujących informacje o głośności sygnałów obiektów audio, które są miksowane wewnątrz downmiksowanego sygnału, i wskazywanie informacji o głośności jednego lub większej liczby kolejnych obejściowych sygnałów obiektów audio, które nie są miksowane wewnątrz downmiksowanego sygnału, i odbieranie informacji odtwarzania wskazujących, czy jeden lub większa liczba sygnałów obiektów audio powinna zostać wzmocniona lub stłumiona, odbieranie downmiksowanego sygnału zawierającego jeden lub większą liczbę downmiksowanych kanałów jako wejściowy sygnał audio, przy czym jeden lub większa liczba downmiksowanych kanałów zawiera sygnały obiektów audio, i gdzie liczba jednego lub większej liczby downmiksowanych kanałów jest mniejsza niż liczba sygnałów obiektów audio, odbieranie informacji downmiksu wskazujących, w jaki sposób sygnały obiektów audio są miksowane wewnątrz jednego lub większej liczby downmiksowanych kanałów, odbieranie jednego lub większej liczby kolejnych obejściowych sygnałów obiektów audio, przy czym jeden lub większa liczba kolejnych obejściowych sygnałów obiektów audio nie jest miksowana wewnątrz downmiksowanego sygnału, określanie wartości kompensacji głośności w zależności od informacji o głośności sygnałów obiektów audio, które są miksowane wewnątrz downmiksowanego sygnału, i w zależności od informacji głośności jednego lub większej liczby kolejnych obejściowych sygnałów obiektów audio, które nie są miksowane wewnątrz downmiksowanego sygnału, i generowanie jednego lub większej liczby wyjściowych kanałów audio wyjściowego sygnału audio z wejściowego sygnału audio w zależności od informacji downmiksu, w zależności od informacji odtwarzania i w zależności od wartości kompensacji głośności. 13. Sposób kodowania obejmujący: kodowanie wejściowego sygnału audio zawierającego wiele sygnałów obiektów audio, i kodowanie informacji głośności w sygnałach obiektów audio, przy czym informacje głośności zawierają jedną lub większą liczbę wartości głośności, przy czym każda z jednej lub większej liczby wartości głośności zależy od jednego lub większej liczby sygnałów obiektów audio, przy czym każdy z sygnałów obiektów audio jest przypisany do dokładnie jednej spośród dwóch lub większej liczby grup, przy czym każda spośród dwóch lub większej liczby grup zawiera jeden lub większą liczbę sygnałów obiektów audio, przy czym kodowanie informacji głośności w sygnałach obiektów audio przeprowadzane jest przez downmiksowanie sygnałów obiektów audio zawartych w dwóch lub większej liczbie grup w celu uzyskania downamiksowanego sygnału zawierającego jeden lub większą liczbę downmiksowanych kanałów audio jako zakodowany sygnał audio, przy czym liczba jednego lub większej liczby downmiksowanych kanałów jest mniejsza niż liczba sygnałów obiektów audio zawartych w dwóch lub większej liczbie grup, gdzie każdy spośród jednego lub większej liczby kolejnych obejściowych sygnałów obiektów audio przypisany jest do trzeciej grupy, przy czym każdy spośród jednego lub większej liczby kolejnych obejściowych sygnałów obiektów audio nie jest zawarty w pierwszej grupie i nie jest zawarty w drugiej grupie, gdzie kodowanie informacji głośności w sygnałach obiektów audio jest przeprowadzane bez downmiksowania jednego lub większej liczby kolejnych obejściowych sygnałów obiektów audio wewnątrz downmiksowanego sygnału, i gdzie kodowanie informacji głośności w sygnałach obiektów audio jest przeprowadzane przez określanie pierwszej wartości głośności, drugiej wartości głośności i trzeciej wartości głośności informacji głośności, przy czym pierwsza wartość głośności wskazuje całkowitą głośność jednego lub większej liczby sygnałów obiektów audio pierwszej grupy, druga wartość głośności wskazuje całkowitą głośność jednego lub większej liczby sygnałów obiektów audio drugiej grupy, a trzecia wartość głośności wskazuje całkowitą głośność jednego lub większej kolejnych obejściowych sygnałów obiektów audio trzeciej grupy, lub jest przystosowany do określania pierwszej wartości głośności i drugiej wartości głośności informacji głośności, przy czym pierwsza wartość głośności wskazuje całkowitą głośność jednego lub większej liczby sygnałów obiektów audio pierwszej grupy, a druga wartość głośności wskazuje całkowitą głośność jednego lub większej liczby sygnałów obiektów audio drugiej grupy i jednego lub większej liczby kolejnych obejściowych sygnałów obiektów audio trzeciej grupy. 14. Sposób według zastrz. 13, w którym dwie lub większa liczba grup to dokładnie dwie grupy, w którym każdy z sygnałów obiektów audio jest przypisany do dokładnie jednej grupy spośród dokładnie dwóch grup, przy czym każda z dokładnie dwóch grup zawiera jeden lub większą liczbę sygnałów obiektów audio, w którym kodowanie informacji głośności sygnałów obiektów audio przeprowadzane jest przez downmiksowanie sygnałów obiektów audio zawartych w dokładnie dwóch grupach w celu uzyskania downmiksowanego sygnału zawierającego jeden lub większą liczbę downmiksowanych kanałów audio jako zakodowany sygnału audio, przy czym liczba jednego lub większej liczby downmiksowanych kanałów jest mniejsza niż liczba sygnałów obiektów audio zawartych w dokładnie dwóch grupach. 15. Program komputerowy do przeprowadzenia sposobu określonego w jednym z zastrz. 12 do 14, gdy jest wykonywany na komputerze lub procesorze sygnału. Fraunhofer- Gesellschaft zur Forderung der angewandten Forschung e.V.,Niemcy Pełnomocnik: EP 2 941 771 B1 Z-15848/17 1/12 ο 1_Ι_ <υ ο Ν Ο Λ'Ξ Ο ct σ3 Ν ε £ ο £ μ ο EP 2 941 771 B1 Z-15848/17 3/12 Wyjściowy sygnał audi k CD CO Dekoder a λ O o N O O Ό O 'E? cś cś ’2 cś i£ a cś £ "O o Informacje dotyczące co o LL_ N > s 1 ce dł. N .2 % o ~ c § te tz §3.2 >1 Xł c/j o EP 2 941 771 B1 Z-15848/17 2/12 FIG 2 EP 2 941 771 B1 Z-15848/17 o £ o '5 -CZ) £ "3 c ' 00 X CZ) >~ O N O o o3 .2/2 O 03 σ3 N ε fe o £ m o CD u_ O £ .2 ’o ^ ’a? £ CD £ o 44 τ-, O gCZ) K N CŚ EP 2 941 771 B1 Z-15848/17 5/12 LO o Ll_ EP 2 941 771 B1 Z-15848/17 T Wprowadzanie danych przez użytkownika Rzeczywisty poziom sygnału wyjściowego Czas Estymacja głośności sygnału wyjściowego w oparciu o sygnał Opóźnienie Czas - i 1 i ..... Estymacja głośności sygnału wyjściowego w oparciu o informacje Czas - ------- -------w Czas FIG 6 EP 2 941 771 B1 Z-15848/17 O £ >? O '3 .2 £ '5? O co ω 3 o ΓΪ2 c®' N C i? « £> £ o o Ό Ό σ3 N c3 o S >. > -rt -rt O Ό Ό O OQ CC hO u_ S—< £ c EP 2 941 771 B1 Z-15848/17 8/1 ο c σ3 oo O Ll_ EP 2 941 771 B1 Z-15848/17 9/12 Transportowanie EP 2 941 771 B1 Z-15848/17 10/12 Tłumienie downmiksu (LU) Wzmocnienie FGO (dB) FIG 10 EP 2 941 771 B1 Z-15848/17 11/12 Wzmocnienie FGO (dB) EP 2 941 771 B1 Z-15848/17 12/12 Dane wprowadzane c\j O
325 paragraphs in 10 sections, as filed
[0001] The present invention relates to the coding, processing and decoding of an audio signal, in particular a decoder, an encoder and a method based on loudness estimation information in object-oriented audio coding systems.
[0002] In the field of audio coding, parametric transmission / storage techniques have recently been proposed with the efficient use of bit rate of audio scenes containing multiple signals of audio objects [BCC, JSC, SAOC, SAOC1, SAOC2] and source separation based on information [ISS1, ISS2, ISS3 , ISS4, ISS5, ISS6]. These techniques are intended to reconstruct the desired output audio scene or source audio object based on additional auxiliary information describing the transmitted / stored audio scene and / or source objects in the audio scene. This reconstruction takes place in the decoder using an information-based source separation scheme. Reconstructed objects can be combined to create an output audio scene. Depending on how the objects are combined, the perceptual volume of the output stage may change.
[0003] In television and radio broadcasting, the volume levels of the audio tracks of different programs can be normalized based on various aspects such as peak signal level or volume level. Depending on the dynamic properties of the signals, two signals with the same peak level can have very different levels of perceptual loudness. When switching between programs or channels, the differences in signal volume can be very annoying and are the main cause of complaints from end users in the transmission.
[0004] In the prior art it has been proposed to normalize all programs on all channels to a common reference level based on perceptual signal loudness. In Europe, one such recommendation is the EBU R128 [EBU] recommendation (hereinafter referred to as R128).
[0005] The recommendation states that "program volume", for example average volume over one program (or one advertisement or some other significant program position) should be equal to a certain level (with slight tolerances). If more and more broadcasters follow this recommendation and the required normalization, the differences between the average volume of different programs and channels should be minimized.
[0006] Loudness estimation can be performed in various ways. There are several mathematical models for estimating the perceptual loudness of an audio signal. The EBU R128 recommendation is based on the loudness estimation model presented in ITU-R BS.1770 (hereinafter BS.1770) (see [ITU]).
[0007] As stated above, for example, according to the EBU R128 recommendation, program volume, e.g. average volume over one program should be equal to a certain level with slight tolerance deviations. However, this leads to serious problems during audio playback that have not been solved in the prior art. Performing audio playback in the decoder has a significant effect on the overall / total volume of the received audio signal. However, despite the scene being played back, the total volume of the received audio signal should remain the same.
[0008] There is currently no solution to this problem regarding the decoder.
[0009] EP 2 146 522 A1 ([EP]) relates to the concept of generating audio output signals using object metadata. At least one audio output signal representing the superposition of at least two different audio object signals is generated, but no solution is provided for this problem.
[0010] WO 2008/035275 A2 ([BRE]) describes an audio system comprising an encoder encoding audio objects in a coding unit that generates a downmixed audio signal and parametric data representing a plurality of audio objects. The downmixed audio signal and parametric data are sent to a decoder that includes a decoding assembly that generates approximated copies of audio objects, and a reproduction assembly that generates an output signal from audio objects. The decoder further includes a processor for generating coding modification data that is sent to the encoder. The encoder then modifies the coding of audio objects, and in particular modifies the parametric data in response to the coding modifying data. This approach allows the decoder to control the changed audio objects, but is carried out completely or partly by the encoder. In this way, alteration can be performed on real independent audio objects rather than approximated copies, which provides increased performance.
[0011] EP 2 146 522 A1 ([SCH]) discloses an apparatus for generating at least one output audio signal representing the superposition of at least two different audio objects, comprising a processor for processing the input audio signal to provide an object representation of the input audio signal, wherein this object representation can be generated by parametrically controlled approximation of original objects using an object downmixed signal. The object changer changes individual objects using object-oriented audio metadata about individual audio objects to change the audio objects. The changed audio objects are then mixed using an object mixer to finally obtain an output audio signal having one or more channel signals, depending on the particular shape of the reproduction.
[0012] WO 2008/046531 A1 ([ENG]) describes an audio object encoder for generating an encoded object signal using multiple audio objects and including a downmix information generator for generating downmix information indicating the placement of multiple audio objects in at least two downmixed channels, audio object parameter generator for generating object parameters for audio objects and an output interface for generating the imported audio output signal using downmix information and object parameters. The audio synthesizer uses downmix information to generate output data that can be used to create multiple output channels with a predetermined audio output shape.
[0013] In addition, WO2012 / 125855 relates to the creation, coding, transmission, decoding and reproduction of spatial audio tracks. The coding format is compatible with earlier surround coding formats.
[0014] It is desirable to provide accurate estimation of the average output volume or changes in the average volume without delays, and when the program does not change or the scene is not changing, the average loudness estimation should also remain static.
[0015] The object of the present invention is to provide an improved concept for encoding, processing and decoding an audio signal. The object of the present invention is achieved by a decoder according to claim 1, an encoder according to claim 9, a system according to claim 11, a method according to claim 12, a method according to claim 13 and a computer program according to claim 15.
[0016] An information-based method of estimating output loudness in an object-oriented audio coding system is provided. The concepts provided are based on information about the volume of objects in the audio mix delivered to the decoder. The decoder uses this information together with playback information to estimate the output signal volume. It is then possible, for example, to estimate the volume difference between the default downmix and the reproduced output signal. It is then possible to compensate the difference to obtain approximately constant output volume, regardless of the playback information. The decoder loudness estimation is performed in a completely parametric way and is associated with very low computational complexity and is accurate compared to the concept of signal based loudness estimation.
[0017] Concepts are provided for obtaining loudness information of a given output scene using purely parametric methods, which then enables loudness processing without explicitly estimating loudness based on a signal at the decoder. In addition, specific spatial audio object coding (SAOC) technology is standardized by MPEG (SAOC), but the available concepts can also be used in conjunction with other spatial audio object coding technologies.
[0018] A decoder is provided for generating an output audio signal comprising one or more output audio channels. The decoder includes a receiving interface for receiving an input audio signal comprising a plurality of audio object signals, for receiving loudness information of the audio object signals and for receiving reproduction information indicating whether one or more audio object signals should be amplified or suppressed. In addition, the decoder includes a signal processor for generating one or more audio output channels of the output audio signal. The signal processor is adapted to determine the volume compensation value depending on the volume information and depending on the playback information. In addition, the signal processor is adapted to generate one or more audio output channels of the output audio signal from the input audio signal depending on the playback information and depending on the volume compensation value.
[0019] According to an embodiment, the signal processor may be adapted to generate one or more audio output channels of the output audio signal from the input audio signal depending on the reproduction information and depending on the volume compensation value such that the volume of the output audio signal is equal to the volume of the input audio signal or in such a way that the volume of the output audio signal is closer to the volume of the input audio signal than the volume of the modified audio signal that would be obtained by modifying the input audio signal by amplifying or suppressing the audio objects signals of the input audio signal according to the reproduction information.
[0020] According to another embodiment, each of the audio objects signals of the input audio signal may be assigned to exactly one group among two or more groups, each of the two or more groups may include one or more signals of the input audio objects audio signal. In such an embodiment, the receiving interface may be adapted to receive the loudness value of each group among two or more groups as loudness information, said loudness value indicating the original total loudness of one or more signals of the audio objects of said group. In addition, the receiving interface may be adapted to receive playback information indicating for at least one group among two or more groups whether one or more signals of the audio objects of said group should be amplified or suppressed by indicating the modified total volume of one or more object signals audio of said group. Additionally, in such an embodiment, the signal processor may be adapted to determine a loudness compensation value depending on the modified total loudness of each of said at least one of the two or more groups and depending on the original total loudness of each of the two or more groups. In addition, the signal processor may be adapted to generate one or more audio output channels of the audio output signal from the audio input signal depending on the modified total volume of each of said at least one of the two or more groups and depending on the volume compensation value.
[0021] In specific embodiments, the at least one group among the two or more groups may comprise two or more signals of the audio objects.
[0022] An additional encoder is provided. The encoder includes an object coding assembly for encoding a plurality of audio object signals to obtain an encoded audio signal comprising a plurality of audio object signals. In addition, the encoder includes an object loudness coding assembly for encoding loudness information on the signals of the audio objects. The loudness information includes one or more loudness values, each of the one or more loudness values depending on one or more signals of the audio objects.
[0023] According to an embodiment, each of the audio object signals of the encoded audio signal may be assigned to exactly one group among two or more groups, each of the two or more groups comprising one or more of the audio object signals of the encoded audio signal . The object loudness coding assembly may be adapted to specify one or more loudness information loudness values to determine the loudness value for each group of two or more groups, said loudness value of said group indicating the original total loudness of one or more audio object signals mentioned group.
[0024] An additional system is provided. The system includes an encoder according to one of the above-described embodiments for encoding a plurality of audio object signals to obtain an encoded audio signal comprising a plurality of audio object signals and for encoding loudness information of the audio object signals. In addition, the system includes a decoder according to one of the above-described embodiments for generating an output audio signal comprising one or more output audio channels. The decoder is adapted to receive encoded audio signals as an audio input signal and volume information. In addition, the decoder is adapted to receive signals also of reproduction information. In addition, the decoder is adapted to determine the volume compensation value depending on the volume information and depending on the playback information. In addition, the decoder is adapted to generate one or more audio output channels of the audio output signal from the audio input signal depending on the playback information and depending on the volume compensation value.
[0025] A method for generating an audio output signal comprising one or more audio output channels is additionally provided. The method includes:
- Receiving an input audio signal containing multiple signals of audio objects.
- Receiving volume information of audio object signals.
- Receiving playback information indicating whether one or more signals of audio objects should be amplified or suppressed.
- Determining the volume compensation value depending on the volume information and depending on the playback information. And:
- Generating one or more audio output channels of the audio output signal from the audio input signal depending on the playback information and depending on the volume compensation value.
[0026] An encoding method is further provided. The method includes:
- Encoding of an input audio signal containing multiple signals of audio objects. And:
- Coding of loudness information on audio object signals, the loudness information comprising one or more loudness compensation values, each of one or more loudness values depending on one or more signals of the audio objects.
[0027] A computer program is additionally provided to perform the method described above when it is executed on a computer or signal processor.
[0028] Preferred embodiments are disclosed in the dependent claims.
[0029] In the following, embodiments of the present invention are described in more detail below with reference to the figures in which:
Fig. 1 shows a decoder according to an embodiment for generating an audio output signal comprising one or more audio output channels,
Fig. 2 shows a decoder according to an embodiment,
Fig. 3 shows a system according to an embodiment,
Fig. 4 shows a coding system for spatial audio objects comprising a SAOC encoder and a SAOC decoder,
Fig. 5 shows a SAOC decoder including an auxiliary information decoder, an object separator module and a reproduction module,
Fig. 6 shows the behavior of estimating the output signal volume when changing the volume,
Fig. 7 shows information-based loudness estimation according to an embodiment, shows encoder and decoder components according to an embodiment,
Fig. 8 shows an encoder according to another embodiment,
Fig. 9 shows an encoder and decoder according to an embodiment associated with the SAOC (Dialog Enhancement) dialogue improvement including bypass channels,
Fig. 10 shows a first illustration of the measured loudness change and the result of using the available concepts to estimate the loudness change parametrically,
Fig. 11 shows a second illustration of the measured loudness change and the result of using the available concepts to estimate the loudness change parametrically, and
Fig. 12 shows other embodiments for performing loudness compensation.
[0030] Before describing preferred embodiments in detail, loudness estimation, spatial audio coding (SAOC) dialogues improvement (DE) will be described.
[0031] First, loudness estimation will be described.
[0032] As stated above, the EBU R128 recommendation is based on the loudness estimation model presented in ITU-R BS.1770. This measure will be used as an example, but the concepts described below can also be used for other loudness measures.
[0033] The process of loudness estimation according to BS.1770 is relatively simple and is based on the following main stages [ITU]:
- The xi input signal (or signals in the case of a multi-channel signal) is filtered using a K filter (a combination of shelf filter and high-pass filters) to obtain the yi signal (s),
- the mean square energy z and signal yi is calculated,
- For multi-channel signal, Gi channel weighting is used and weighted signals are added together. The signal volume is then defined as £ = c + i01og<sub>10</sub>] T <7, .z,.
Z with a constant value of c = -0.691. The output is expressed in units of "LKFS" (volume, weighting K, relative to full scale), which scale similarly to the decibel scale.
[0034] In the above formula, Gi can be, for example, equal to 1 for some of the channels, while for some other channels Gi can be, for example, 1.41. For example, considering the left channel, right channel, center channel, left surround channel and right surround channel, the corresponding weights Gi may be, for example, 1 for the left, right and center channel, and may, for example, be equal to 1 , 41 for the left surround channel and right surround channel, see [ITU].
[0035] It can be seen that the loudness value L is closely related to the logarithm of the signal energy.
[0036] Coding of spatial audio objects will be described below.
[0037] Object-oriented audio coding concepts allow much greater flexibility on the side of the chain containing the decoder. An example of the concept of object-oriented audio coding is coding of spatial audio objects (SAOC).
[0038] Fig. 4 is a spatial audio object coding (SAOC) system comprising SAOC 410 encoder and SAOC 420 decoder.
[0039] The SAOC encoder 410 receives Signal Si, SN audio objects as input. In addition, the SAOC 410 encoder also receives "Mixing information D" instructions on how these objects should be combined to obtain a downmixed signal containing M downmixed X1,., And XM channels. The SAOC encoder 410 obtains some auxiliary information from objects and downmixing, and this auxiliary information is sent and / or saved together with the downmixed signals.
[0040] The main feature of the SAOC system is that the downmixed X signal containing the downmixed X1,., And XM channels creates a semantically significant signal. In other words, it is possible to listen to the downmixed signal. If, for example, the receiver does not have the SAOC decoder functionality, the receiver can, however, always provide a downmixed signal at the output.
[0041] Fig. 5 is a SAOC decoder including auxiliary information decoder 510, object separator module 520 and reproduction module 530. The SAOC decoder shown in Fig. 5 receives, for example, from the SAOC encoder, the downmixed signal and auxiliary information. The downmixed signal can be treated as an input audio signal containing audio object signals, because inside the downmixed signal the audio object signals are mixed (the audio object signals are mixed inside one or more downmixed channels of the downmixed signal).
[0042] The SAOC decoder may then, for example, attempt to (virtual) restore the original objects, for example by using the object separator module 520, for example, using decoded auxiliary information. These (virtual) reconstructions of Si, Sn objects, e.g. signals of reconstructed audio objects, are then combined based on the playback information, e.g. the playback matrix R, to form K output channels Y1,., YK audio output signal Y audio.
[0043] In SAOC, the audio object signals are, for example, reconstructed, for example, by using covariance information, e.g. an E signal covariance matrix, which is transmitted from the SAOC encoder to the SAOC decoder.
[0044] For example, the following formula may be used to reconstruct the signals of the audio objects on the decoder side:
S = GX with G ~ ED<sup>H</sup> (DED<sup>H</sup>)<sup>-1</sup> where
N
Nsamples number of audio object signals The number of audio object signal samples included
M
X
D
E
XX<sup>H</sup>
S
Nsamples (·) "
number of downmixed channels downmixed audio signal, size M x Nsamples downmix matrix, size M x N signal covariance matrix, size M x N, defined as E = parametrically reconstructed N signals of audio objects, size N x self-coupled (Hermitian) operator representing coupled transposition (·) [0045] Next, it is possible to use the R playback matrix on the reconstructed S signals of the audio objects to obtain audio output channels of the output Y audio signal, for example according to the formula:
where
K number of output channels Yi, ..., Yk audio output signal Y audio
R reproduction matrix with the dimension K x N
Y audio output signal containing K audio output channels, size M x Nsamples [0046] Fig. 5 shows the process of reconstructing objects, for example carried out by the object separator module 520 to which the term "virtual" or "optional" is connected, because no it must necessarily be carried out, but the desired functionality can be obtained by combining reconstruction and reconstruction steps in the parametric domain (i.e. by combining patterns).
In other words, instead of reconstructing the audio object signals first using the mixing D information and covariance information, and then applying the playback R information on the reconstructed audio object signals to obtain output channels Yi, Yk audio, both steps can be performed in a single stage, thanks to which the output channels Yi, Yk audio are generated directly from downmixed channels.
[0048] For example, it is possible to use the following formula:
Y = RGX with G ~ ED<sup>H</sup> (DED<sup>H</sup>)<sup>-1</sup> [0049] As a general rule, the playback R information may request any combination of original audio object signals. In practice, however, object reconstruction may contain reconstruction errors, and the desired output scene does not have to be achieved. As a general rough rule for many practical cases, the more the desired output scene differs from the downmixed signal, the more audible reconstruction errors occur.
[0050] The following describes the improvement (DE) of the dialogues. SAOC technology can be used, for example, to implement a scenario. It should be noted that although the name 'dialogues improvement' suggests focusing on signals related to dialogues, the same principle can also be used for other types of signals.
[0051] In the DE scenario, the degrees of freedom of the system are limited compared to the general case.
[0052] For example, the Si, ..., Sn = S audio objects are grouped (and optionally mixed) in two metgoobjects of the foreground SFGO object (FGO) and the background SBGO object (BGO).
[0053] In addition, the output stage Y1,., YK = Y output resembles the downmixed signal X1,., XM = X. More specifically, both signals have the same dimensions, i.e. K = M, and the end user can only control relative mixing levels two meta objects. More specifically, the downmix signal received by mixing FGO and BGO using certain scalar weights
^ FGO FGO
BGO ^ BGO 'and the output stage is obtained in a similar way with some scalar weighing of FGO and BGO:
<img file="PL2941771T3_D0001.tif" />
[0054] Depending on the relative values of the mixing weights, the balance between FGO and BGO may vary. For example, if set
<img file="PL2941771T3_D0002.tif" />
[0055] it is possible to increase the relative level of FGO in a mix. If FGO is a dialogue, these settings provide the function of improving the dialogue.
[0056] As an example of a use case, BGO may be stadium noise.
stadium noises) and other background sounds during the sporting event, and FGO is the voice of the commentator. The DE function allows the end user to amplify or suppress the commentator level in relation to background sounds.
[0057] Embodiments are based on the finding that the use of SAOC (or similar) technology in the broadcasting scenario enables the end user to provide an extended signal change function. More functions are provided than just changing the channel and adjusting the playback volume.
[0058] One possibility of using DE technology is briefly described above. If the transmitted signal being the downmixed SAOC signal has a normalized level, for example, according to R128, different programs have similar average volume when no processing (SAOC-) is used (or the playback description is the same as the description of the downmixing). However, when using certain processing (SAOC-), the output signal differs from the default downmixed signal, and the volume of the output signal may differ from the volume of the default downmixed signal. From the end user's point of view, this can lead to a situation where the volume of the output signal between channels or programs may again have undesirable jumps or differences. In other words, the benefits of standardization by the broadcaster are partially lost.
[0059] This problem does not only occur with the SAOC or DE scenario, but can also occur with other audio coding concepts that allow the end user to interact with the content. However, in many cases this does not cause harm when the output signal has a different volume than the default signal.
[0060] As noted above, the total volume of the audio input signal of the program should have a certain level with slight tolerance deviations. However, as indicated above, this can lead to significant problems when performing audio playback because audio playback can have a significant effect on the overall / total volume of the received audio input. However, despite performing the scene playback, the total volume of the received audio input signal should remain the same.
[0061] One approach is to estimate the volume of the signal during playback, and with the right concept of time integration, the estimation may converge to true average loudness after some time has elapsed. However, the time required to achieve convergence is a problem for the end user. If there are changes in loudness estimation, even when no signal changes are used, the loudness compensation should also change and change its course. This leads to an output signal with a time-varying average volume that can be seen as rather annoying.
[0062] Fig. 6 shows the behavior of the output volume estimation when the volume is changed. Here, among others, an estimation of the output signal volume based on the signal is presented, which illustrates the effect of using the above described solution. The estimation is approaching the correct estimation relatively slowly. Instead of estimating the output signal-based loudness, it is preferable to use information-based output loudness estimation that ensures that the correct output loudness is determined immediately.
[0063] In particular, in Fig. 6, the user enters, for example, the level of the dialog object, which changes at time T by increasing its value. The actual level of the input signal and the corresponding volume change at the same time. When the output signal volume estimation is performed based on the output signal with a certain time integration time, the estimation changes gradually and reaches the correct value with a certain delay. During this delay, the estimation values change and cannot be reliably used for further processing of the output signal, example to adjust the volume level.
[0064] As already noted above, it is desirable to provide an accurate estimation of the average output volume or change of the average loudness without delays, and when the program is not changed or the playback scene does not change, the average loudness estimation should also remain static. In other words, if some compensation for the volume change is used, the compensation parameter should change only in the event of a program change or some user interaction.
[0065] The desired operation is shown in the lowest illustration of Fig. 6 (information-based estimation of the output signal volume). The output signal volume estimation should change immediately after the user makes the change.
[0066] Fig. 2 shows an encoder according to an embodiment.
[0067] The encoder includes an object coding assembly 210 for encoding a plurality of audio object signals to encode an audio signal comprising a plurality of audio object signals.
[0068] The encoder further includes an object loudness coding assembly 220 for encoding loudness information on the signals of the audio objects. The loudness information includes one or more loudness values, each of the one or more loudness values depending on one or more signals of the audio objects.
[0069] According to an embodiment, each of the audio object signals of the encoded audio signal is assigned to exactly one group of two or more groups, each of the two or more groups comprising one or more of the audio object signals of the encoded audio signal. The object loudness coding assembly 220 adapted to determine one or more loudness values from the loudness information by determining the loudness value for each group of two or more groups, said loudness value of said group indicating the original total loudness of one or more signals of the said audio objects group.
[0070] Fig. 1 shows a decoder according to embodiments for generating an output audio signal comprising one or more output audio channels.
[0071] The decoder comprises a receiving interface 10 for receiving an input audio signal comprising a plurality of audio object signals, for receiving loudness information of the audio object signals and for receiving reproduction information indicating whether one or more audio object signals should be amplified or suppressed.
[0072] In addition, the decoder includes a signal processor 120 for generating one or more audio output channels of the output audio signal. The signal processor i20 is adapted to determine the value of loudness compensation depending on the loudness information and depending on the reproduction information. In addition, the signal processor i20 is adapted to generate one or more audio output channels of the output audio signal depending on the input audio signal depending on the playback information and depending on the volume compensation value.
[0073] According to an embodiment, the signal processor 110 is adapted to generate one or more audio output channels of the output audio signal from the input audio signal depending on the reproduction information and depending on the volume compensation value such that the volume of the output audio signal is equal to the volume of the input audio signal or in such a way that the volume of the output audio signal is closer to the volume of the input audio signal than the volume of the modified audio signal that would be obtained by modifying the input audio signal by amplifying or suppressing the audio objects signals of the input audio signal according to the reproduction information.
[0074] According to another embodiment, each of the audio objects signals of the input audio signal is assigned to exactly one group among two or more groups, each of the two or more groups includes one or more signals of the audio objects of the input audio signal .
In such an embodiment, the receiving interface ii0 is adapted to receive a loudness value for each group of two or more groups as loudness information, said loudness value indicating the original total loudness of one or more signals of the audio objects of said group. In addition, the receiving interface ii0 is adapted to receive playback information indicating for at least one group among two or more groups whether one or more audio objects signals of said group should be amplified or suppressed by indicating the modified total volume of one or more object signals audio of said group. Additionally, in such an embodiment, the signal processor 120 is adapted to determine a loudness compensation value depending on the modified total loudness of each of said at least one group among two or more groups and depending on the original total loudness of each of two or more groups. In addition, the signal processor 120 is adapted to generate one or more output channels of the output audio signal from the input audio signal depending on the modified total loudness of each of said at least one of the two or more groups and depending on the loudness compensation value.
[0076] In particular embodiments, the at least one group among the two or more groups comprises two or more signals of the audio objects.
[0077] There is a direct relationship between the energy ei of the signal and the audio object and the loudness Li of the signal and the audio object, which is expressed by the formula:
<img file="PL2941771T3_D0003.tif" />
where c is a constant value.
[0078] Embodiments are based on the following findings. Different audio signals of the audio input signal may have different volumes and therefore different energy. If, for example, a user wants to increase the volume of one of the audio object signals, it is possible to change the reproduction information accordingly, and increasing the volume of that audio object signal increases the energy of that audio object. This leads to an increase in the volume of the audio output signal. In order to keep the total volume constant, it is necessary to perform volume compensation. In other words, the modified audio signal obtained by applying reproduction information to the audio input signal must be changed. However, the exact gain effect of one of the audio object signals on the total volume of the modified audio signal depends on the original volume of the amplified audio object signal, e.g., the audio object signal whose volume has been increased. If the original volume of this object corresponds to energy that was relatively low, the overall volume of the input audio signal will be small. However, if the original volume of this object corresponds to energy that was relatively high, the effect on the total volume of the audio input will be significant.
[0079] Two examples can be considered. In both examples, the input audio signal includes two audio object signals, and in both examples, the energy of the first of the audio object signals is increased by 50% by using reproduction information.
[0080] In the first example, the first audio object signal provides 20% and the second audio object signal provides 80% of the total energy of the audio input signal. However, in the second example, the first audio object, and therefore the first audio object signal provides 40%, and the second audio object signal provides 60% of the total energy of the input audio signal. In both examples, these shares can be obtained from the loudness information of the audio objects, because there is a direct relationship between loudness and energy.
[0081] In the first example, increasing the energy of the first audio object by 50% results in the modified audio signal generated by applying reproduction information on the audio input signal having a total energy of 1.5 x 20% + 80% = 110% of the input signal energy audio.
[0082] In the second example, increasing the energy of the first audio object by 50% results in the modified audio signal generated by applying reproduction information on the audio input signal having a total energy of 1.5 x 40% + 80% = 120% of the energy of the input signal audio.
[0083] Thus, after applying the reproduction information to the audio input signal, the total energy of the modified audio signal must be reduced by only 9% (10/110) in the first example to obtain the same energy in both the audio input signal and the audio output signal, while in the second example, the total energy of the modified audio signal must be reduced by 17% (20/120). For this purpose it is possible to calculate the volume compensation value.
[0084] For example, the volume compensation value may be a scalar used for all audio output channels of the audio output signal.
[0085] According to an embodiment, the signal processor is adapted to generate the modified audio signal by modifying the input audio signal by amplifying or suppressing the signals of the audio objects of the input audio signal according to the reproduction information. In addition, the signal processor is adapted to generate an audio output signal by applying a volume compensation value to the audio input signal such that the volume of the audio output signal is equal to the volume of the input audio signal or the volume of the output audio signal is closer to the volume of the input audio signal than the volume modified audio signal.
[0086] For example, in the first example above, it is possible to specify the volume compensation value lcv, for example, as values lcv = 10/11, and a multiplication factor of 10/11 can be applied to all channels obtained by playing the input audio channels according to the information playback.
[0087] Similarly, for example, in the second example above, it is possible to specify the loudness compensation value lcv, for example, as lcv = 10/12 = 5/6, and a multiplication factor of 5/6 can be applied to all channels obtained by playback audio input channels according to the playback information.
[0088] In other embodiments, each of the audio object signals may be assigned to one of a plurality of groups, and a loudness value indicating the value of the total loudness of the audio objects signals of said group may be transmitted for each of the groups. If the reproduction information indicates that the energy of one of the groups is amplified or suppressed, for example 50% amplified, as in the above, it is possible to calculate the total energy increase value and the loudness compensation value can be determined as described above.
[0089] For example, according to an embodiment, each of the audio objects signals of an input audio signal is assigned to exactly one group among exactly two groups constituting two or more groups. Each of the audio object signals of the input audio signal is assigned either to a group of foreground objects from exactly two groups or to a group of background objects from exactly two groups. The receiving interface 110 is adapted to receive signals of the original total loudness of one or more signals of audio objects belonging to the group of foreground objects. In addition, the receiving interface 110 is adapted to receive the signals of the original total loudness of one or more of the audio objects of the background object group. Additionally, the receiving interface 110 is adapted to receive reproduction information signals indicating for at least one group of exactly two groups whether one or more signals of audio objects belonging to each of said at least one group should be amplified or suppressed by indicating the modified total volume one or more signals of the audio objects of said group.
[0090] In such an embodiment, the signal processor 120 is adapted to determine a loudness compensation value depending on the modified total loudness of each of said at least one group, depending on the original total loudness of one or more signals of the audio objects of the foreground group and depending on from the original total volume of one or more signals of the audio objects of the background object group. In addition, the signal processor 120 is adapted to generate one or more audio output channels of the output audio signal from the input audio signal depending on the modified total volume of each of said at least one group and depending on the volume compensation value.
[0091] According to some embodiments, each of the audio object signals is assigned to one of three or more groups, and the receiving interface can be adapted to receive a loudness value regarding each of the three or more groups indicating the total loudness of the audio objects signals of said group.
[0092] According to an embodiment, to determine the total loudness value of two or more audio object signals, for example, an energy value corresponding to the loudness value for each audio object signal is determined, the energy values of all loudness values are summed to obtain the sum of energy, and the volume value corresponding to the sum of energy is defined as the value of the total volume of two or more signals of the audio objects. It is possible, for example, to use the formula:
A = c + 101og<sub>]0</sub> ς, ς. LO =<sup>(L</sup>'“<sup>c) / 1</sup>[0093] In some embodiments, loudness values are transmitted for each of the audio object signals, or each of the audio object signals is assigned to one, two or more groups, with the loudness value being transmitted for each group.
[0094] However, in some embodiments, no loudness value is transmitted for one or more audio object signals or one or more groups containing audio object signals. Instead, the decoder may, for example, assume that those audio object signals or groups of audio object signals for which no volume value is being transmitted have a predetermined volume value. The decoder may, for example, base all further determinations on this predetermined loudness value.
[0095] According to an embodiment, the receiving interface ii0 is adapted to receive downmixed signals comprising one or more downmixed channels as an audio input signal, wherein the one or more downmixed channels comprise audio object signals and wherein the number of audio object signals is smaller than the number of one or more downmixed channels. The receiving interface ii0 is adapted to receive downmix information signals indicating how the signals of audio objects are mixed within one or more downmixed channels. In addition, the signal processor i20 is adapted to generate one or more audio output channels of the output audio signal from the input audio signal depending on the downmix information, depending on the reproduction information and depending on the volume compensation value. In a particular embodiment, the signal processor 120 may be, for example, adapted to calculate a loudness compensation value depending on the downmix information.
[0096] For example, the downmix information may be in the form of a downmix matrix. In embodiments, the decoder may be a SAOC decoder. In such embodiments, the receiving interface may be, for example, also adapted to receive covariance information signals, e.g., a covariance matrix, as described above.
[0097] Regarding reproduction information indicating whether one or more signals of audio objects should be amplified or suppressed, it should be noted that reproduction information may be, for example, information indicating how one or more signals of audio objects should be reinforced or suppressed. For example, the reproduction information is a reproduction matrix R, e.g. a reproduction matrix in SAOC encoding.
[0098] Fig. 3 shows a system according to an embodiment.
[0099] The system includes an encoder 310 according to one of the above-described embodiments for encoding a plurality of audio object signals to obtain an encoded audio signal comprising a plurality of audio object signals.
[0100] In addition, the system includes a decoder 320 according to one of the embodiments described above for generating an output audio signal comprising one or more output audio channels. The decoder is adapted to receive signals of an encoded audio signal constituting an audio input signal and volume information. In addition, the decoder 320 is further adapted to receive reproduction information signals. In addition, the decoder 320 is adapted to determine the volume compensation value depending on the loudness information and reproduction information. In addition, the decoder 320 is adapted to generate one or more audio output channels of the output audio signal from the input audio signal depending on the playback information and depending on the volume compensation value.
[0101] Fig. 7 shows information-based loudness estimation according to an embodiment. To the left of the transport stream 730 are encoder components for object-oriented audio coding. In particular, the object coding assembly 710 ("object audio coder") and the object loudness coding assembly 720 ("object loudness estimation") are shown herein.
[0102] The transport stream 730 itself includes loudness L information, downmix information D and the output of the object audio encoder 710 B.
[0103] On the right side of transport stream 730 are components of the audio object coding decoder signal processor. The decoder receiving interface is not shown here. The module 740 estimating the output volume and the assembly 750 of object audio decoding are presented. The output volume estimation module 740 may be adapted to determine the volume compensation value. The object audio decoding assembly 750 may be adapted to determine a modified audio signal from an audio signal that is input into the decoder, which is performed using the playback R information. Fig. 7 does not show the use of loudness compensation values for a modified audio signal to compensate for the change in total volume caused by playback.
[0104] The encoder input is at least input objects S. The system estimates the volume of each object (or some other loudness- related information , such as the energies of the objects), for example, using the object-oriented loudness coding assembly 720, and this L information is sent and / or saved. (It is also possible to input objects into the system as input, and the estimation step inside the system can be omitted.)
[0105] In the embodiment shown in Fig. 7, the decoder receives at least the loudness information of the objects and, for example, the reproduction R information describing mixing the objects to the output signal. Based on them, for example, output loudness is estimated using output loudness estimation assembly 740, and this information is made available at its output.
[0106] Downmixing information D may be used as reproduction information, in which case the loudness estimation provides an estimation of the downmixed signal loudness. It is also possible to provide information such as inputs for estimating the volume of objects and for sending and / or saving them together with the object volume information. The output loudness estimation can then estimate the simultaneous estimation of the downmix signal volume and the reproduced output and provide these values two values or their difference as output loudness information. The value of the difference (or its inverse) describes the required compensation, which should be applied to the reproduced output signal in order to make its volume similar to the volume of the downmixed signal. Object loudness information may further include information regarding correlation coefficients between different objects, and this correlation information may be used when estimating output loudness to obtain a more accurate estimation.
[0107] A preferred embodiment related to improving dialogs will be described below.
[0108] In dialog enhancement applications as described above, the input signals of audio objects are grouped and partially downmixed to form two metaobjects, FGO and BGO, which can then be easily combined to obtain the final downmixed signal.
[0109] In the following SAOC encoding description [SAOC], N object input signals are represented as a matrix S with the size N x NSamples, and downmixing information as a matrix D with the size M x N. Downmixed signals can then be obtained as X = DS.
[0110] D downmixing information can then be divided into two parts for meta objects:
η - D + D " <sup>J</sup>FGO <sup>r U</sup>BGO [0111] Since each column of matrix D corresponds to the original input signal of an audio object, two-component downmix matrices can be obtained by setting zero column values that correspond to other metaobjects (assuming that no original object can be in two metaobjects). In other words, the value of the columns corresponding to the BGO meta-object is zeroed in DFGO and vice versa.
[0112] These new downmixing matrices describe how two metaobjects can be obtained based on input objects, namely:
and the actual downmixing is simplified to [0113] One may also consider a situation in which the object decoder (e.g. SAOC) attempts to reconstruct the metaobjects:
<img file="PL2941771T3_D0004.tif" />
<img file="PL2941771T3_D0005.tif" />
and DE-specific reproduction can be written as a combination of these two reconstructions of meta-objects:
<img file="PL2941771T3_D0006.tif" />
<img file="PL2941771T3_D0007.tif" />
[0114] For the estimation of object loudness, two metgoobjects SFGO and SBGO are introduced as input and the loudness of each is estimated: LFGO is (total / overall) loudness of SFGO, and LBGO is (total / overall) loudness of SBGO. These volume values are transmitted and / or saved.
[0115] As an alternative, using one of the metaobjects, for example FGO, as a reference, it is possible to calculate the loudness difference between these two objects, for example as ΔΖ /<sub>Μ</sub>·<sub>; ο</sub> = L<sub>H (, (J</sub> - .
[0116] This single value is then transmitted and / or saved.
[0117] Fig. 8 shows an encoder according to another embodiment. The encoder shown in Fig. 8 includes a downmixing module 811 and an module 812 estimating object auxiliary information. The encoder shown in Fig. 8 further includes an assembly 820 encoding the loudness of the object. In addition, the encoder shown in Fig. 8 includes a mixer 805 audio meta objects.
[0118] The encoder shown in Fig. 8 uses intermediate audio meta objects as input for estimating the volume of objects. In the embodiments, the encoder shown in Fig. 8 can be adapted to generate two audio meta objects. In other embodiments, the encoder shown in Fig. 8 may be adapted to generate three or more audio meta objects.
[0119] The concepts provided provide, among other things, a new feature in that the encoder can, for example, estimate the average loudness of all input objects. Objects can, for example, be mixed inside a downmixed signal that is transmitted. The concepts provided additionally provide a new function in that object volume and downmixing information can, for example, be included in object-coded auxiliary information that is transmitted.
[0120] The decoder may, for example, use object-coded ancillary information to (virtual) separating the objects and reconnecting the objects using the reproduction information.
[0121] The available concepts additionally provide a new function in that downmixing information can either be used to estimate the volume of the default downmixed signal, playback information, and received object volumes can be used to estimate the average output volume and / or it is possible to estimate the change volume based on these two values. Or, Downmix and playback information can also be used to estimate the change in volume relative to the default downmix, which is another new feature of the shared concepts.
[0122] In addition, the concepts provided provide a new feature in that the decoder output signal can be modified to compensate for the volume change so that the average volume of the modified signal matches the average volume of the default downmix.
[0123] Fig. 9 shows a specific embodiment associated with SAOCDE. The system receives input signals of audio objects, downmixing information and information about grouping objects in metaobjects. Based on them, the 905 audio meta object mixer creates two SFGO and SBGO meta objects. It is possible that part of the signal processed using SAOC encoding does not constitute the whole signal. For example, in the configuration of 5.1 channels SAOC can be used in a subset of channels, for example in the front channel (left, right and center), while other channels (surround left, surround right and low frequency effects channel) are routed alongside (bypassing) SAOC and delivered in as such. Those channels that are not processed by SAOC are marked as XBYPASS. Possible bypass channels must be provided to the encoder for more accurate estimation of loudness information.
[0124] Bypass channels can be handled in various ways.
[0125] For example, bypass channels may, for example, form an independent meta-object. This allows you to define playback in such a way that all three metaobjects are scaled independently.
[0126] Or, for example, bypass channels may be combined with one of two other meta-objects. The playback settings of this meta object also control the part containing the bypass channel. For example, in the dialog improvement scenario, it may be important to combine bypass channels with the background meta-object: XBGO = SBGO + XBYPASS.
[0127] Or, for example, bypass channels may be omitted.
[0128] According to embodiments, the object coder assembly 210 is adapted to receive audio object signals, each of the audio object signals being assigned to exactly one of exactly two groups, each of exactly two groups having one or more audio object signals. In addition, the object coding assembly 210 is adapted to downmix signals of audio objects contained in exactly two groups to obtain a downmixed signal comprising one or more downmixed audio channels as the encoded audio signal, wherein the number of one or more downmixed channels is less than the number audio object signals contained in exactly two groups. The object loudness coding assembly 220 adapted to receive one or more consecutive bypass audio object signals, each of the one or more successive bypass audio object signals is not included in the first group and is not included in the second group, the coding assembly 210 object adapted not to downmix one or more consecutive bypass signals of audio objects within the downmixed signal.
[0129] In an embodiment, the object loudness coding assembly 220 is adapted to determine the first loudness value, the second loudness value and the third loudness value of the loudness information, wherein the first loudness value indicates the total loudness of one or more signals of the first group audio objects, the second value volume indicates the total volume of one or more of the audio objects of the second group, and the third loudness value indicates the total volume of one or more successive bypass signals of the third group audio objects. In another embodiment, the object loudness coding assembly 220 is adapted to determine the first loudness value and the second loudness value of the loudness information, wherein the first loudness value indicates the total loudness of the one or more audio signals of the first group, and the second loudness value indicates the total volume of one or more second group audio object signals and one or more successive third group audio object bypass signals.
[0130] According to an embodiment, the decoder receiving interface 110 is adapted to receive a downmixed signal. In addition, the decoder receiving interface 110 adapted to receive one or more consecutive bypass audio signals, the one or more consecutive bypass audio signals are not mixed within the downmixed signal. In addition, the decoder receiving interface 110 is adapted to receive loudness information indicating the loudness information of the signals of the audio objects that are mixed inside the downmixed signal and indicating the loudness information of one or more consecutive bypass signals of the audio objects that are not mixed inside the downmixed signal. In addition, the signal processor 120 adapted to determine the loudness compensation value depending on the loudness information of the audio object signals that are mixed inside the downmixed signal and depending on the loudness information of one or more successive bypass signals of the audio object that are not mixed inside the downmixed signal.
[0131] Fig. 9 shows an encoder and a decoder according to SAOC-DE related embodiments which includes bypass channels. The encoder shown in Fig. 9 includes, among others, SAOC 902 encoder.
[0132] In the embodiment shown in Fig. 9, the potential joining of bypass channels with other metaobjects is carried out in two blocks 913, 914 "bypassing" generating XFGO and XBGO metaobjects with specific parts from bypass channels included.
[0133] The perceptual loudness of LBYPASS, LFGO and LBGO of both these metaobjects are estimated in loudness estimation sets 921, 922, 923. The loudness information is then converted to the appropriate encoding in the module estimating the loudness information of the objects 925, after which they are transmitted and / or saved.
[0134] The operation of the actual SAOC encoder and decoder ensures, as expected, obtaining object auxiliary information from objects, generating a downmixed X signal, and transmitting and / or writing information to the decoder / in the decoder. Optional bypass channels are sent along with other information and / or saved to the decoder / in the decoder.
[0135] The SAOC-DE decoder 945 receives the input value of "dialogue enhancement" introduced by the users. Based on this input and downmixing information received, the SAOC decoder 945 determines playback information. The 945 SAOC decoder then generates the reproduced output scene as a Y signal. In addition, it generates a gain factor (and delay value) that should be applied to optional XBYPASS bypass signals.
[0136] The "bypass enable" assembly 955 receives this information along with the reproduced output scene and bypass signals, and generates a full output scene signal. The SAOC-DE 945 decoder also generates a set of metaobject gain values, the amount of which depends on the metaobject grouping and the desired form of loudness information.
[0137] Gain values are provided to the mix volume estimation assembly 960, which also receives metaobject volume information from the encoder.
[0138] Mix volume estimation module 960 may then determine the desired loudness information, which may include, but is not limited to, the volume of the downmixed signal, the volume of the reproduced output stage and / or the volume differences between the downmixed signal and the reproduced output scene.
[0139] In some embodiments, the loudness information alone is sufficient, while in other embodiments, it is desirable to process the full output depending on the specific loudness information. This processing can be, for example, compensating for any possible loudness differences between the downmixed signal and the reproduced output stage. Such processing, for example, carried out by the loudness processing unit 970 makes sense in the transmission scenario because it provides a reduction in changes in the perceptual loudness of the signal regardless of the user interaction (setting of "enhanced dialogues" introduced).
[0140] In this particular embodiment, the loudness-related processing includes many new features, including FGO, BGO and optional bypass channels are pre-mixed into the final channel design so that the downmixing process can be carried out simply by adding two pre-mixed channels to themselves (for example, with downmixing coefficients of 1), which is a new feature. In addition, as another new feature, the average volumes of FGO and BGO are estimated and the difference is calculated. In addition, the objects are mixed into a downmix channel that is transmitted. In addition, as another new feature, loudness difference information is included in the support information that is being transmitted (new). In addition, the decoder uses auxiliary information to (virtual) object separation and reconnection of objects using playback information that is based on downmixing information and user-introduced modification gain. In addition, as another new feature, the decoder uses modifying gain and transmits loudness information to estimate the change in the average output volume of the system compared to the default downmix.
[0141] The following is a formal description of the embodiments.
[0142] It has been assumed that the loudness values of objects behave similar to the logarithm of the energy values during the summation of the objects, i.e. the loudness values must be transformed to the linear domain, summed there, and then transformed back to the logarithmic domain. Justifying this with the definition of BS.1770, the loudness measure will be presented below (for simplicity the number of channels is one here, but the same principle can be applied to multi-channel signals with proper channel summation).
[0143] The loudness of the ith K-filtered signal z and with mean square energy ei is defined as
<img file="PL2941771T3_D0008.tif" />
where c is the shift constant. For example, c may be -0.69i. It follows that the signal energy can be determined on the basis of loudness using _ 1 η (4-4 «0
e.
[0144] Energy of the sum of N uncorrelated signals <sup>FROM</sup>SllM 2 ^ <sup>FROM</sup>ii = t is then
Ν N ,, = Σι «
J = t (= 1 <sup>e</sup>i'UVf (£, -4'ΐο and the volume of this sum signal is then
<img file="PL2941771T3_D0009.tif" />
[0145] If the signals are uncorrelated, the correlation coefficients Ci, j in the form of .VN should be considered when approximating the signal sum energy <sup>e</sup>SUM Σ Σ <sup>e</sup>f, 7 where the energy ei, j cross between the i-th and j-th object is defined as
<img file="PL2941771T3_D0010.tif" />
where -1 <Cj <1 is the correlation coefficient between two objects i and j. When two objects are uncorrelated, the correlation coefficients are equal to 0, and when the two objects are identical, the correlation coefficients are equal to 1.
[0146] Further extending the model to include g and mixed weights for use on signals in the mixing process, i.e.
N 'SUM
Lsp sum signal energy is
SUM / =] y = ł
Isum ~ <sup>C</sup> 1 θ 1θ8ιο
SUM and the volume of the mix signal can be obtained on this basis, as above, as [0147] The difference between the volume of two signals can be estimated as [0148] Now, as above, if a loudness definition can be used that can be written as
<img file="PL2941771T3_D0011.tif" />
which can be considered as a function of signal energy. Now if it is desirable to estimate the volume difference between two mixes
<img file="PL2941771T3_D0012.tif" />
with optionally different weights g and hi mixing, it can be estimated as
Ai (4S) = 10 logs,<sub>0</sub>
14'jB jV ΛΓ
BACKGROUND = 101o<sub>gl</sub> ---------- HO ΛΓ jV f = lj = l
j), pink = 10 logs
ΣΣ> * ΑΛ<sup>10</sup> f = i ./=1
VV and
ZEW>
ioiog / = 1 y = l
1, · + ί.γ-2ί) / 10 (ą + Zj-2r) / 10 g ^ Wio
0, V i Ψ ji Cj = 1, V i =, v jV
YYaac.Kio
Ύ<sup>1</sup> 2ι n, (Α ^ Χ<sup>10</sup>
Σ &<sup>10</sup>
Ai (J, B) = 101og<sub>l0</sub>-i · '; V
Σ ^ ιο and ~ l
Y
Σ ^<sup>1θ</sup> = K> iog
Lj / 10 o JV
Σ ^ ιο i = 1 [0151] It is possible to encode volume values into objects in the form of differences from the selected reference object:
where Lref is the volume of the reference object. This coding is beneficial if no absolute loudness values are required as a result, because it is now necessary to send one less value and the loudness difference estimation can be saved as
<img file="PL2941771T3_D0013.tif" />
or for uncorrelated objects
Ai (Z £) -10log<sub>10</sub>^ [0152] The scenario of improving dialogues was considered below.
[0153] The scenario to improve dialogs was re-considered. The freedom to define playback information in a decoder is limited only to changing the levels of two meta-objects. Let's assume additionally that two metaobjects are uncorrelated, i.e. CFGO, BGO = 0. If the weights of downmixed metaobjects are hFGO and hBGO and they are reproduced with fFGO and fBGO gains, the output signal volume in relation to the default downmix is
<img file="PL2941771T3_D0014.tif" />
[0154] If it is desired to obtain the same output volume as the default downmix then subsequent compensation is required.
[0154] ΔL (A, B) can be considered as a volume compensation value that can be transmitted by the decoder signal processor 120. ΔL (A, B) can also be called a loudness change value, so the actual compensation value can be the inverse. Thus, the lcv value of loudness compensation mentioned earlier in this document corresponds to the following gDelta value.
[0156] For example, it is possible to use gΔ = W-WA B)<sup>2</sup> or gΔ = 1 / ΔL (A, B) as a multiplication factor for each channel of the modified audio signal obtained by using the reproduction information to process the input audio signal. This equation works for gDelt in the linear domain. In the logarithmic domain, the equation has the form 1 / ΔL (A, B) and is used accordingly.
[0157] If the downmixing process is simplified, so that two metaobjects can be mixed with unit weights to obtain a downmixed signal, i.e. hFGO = hBGO = 1, and the reproduction gain of these two objects is designated as gFGO and gBGO. This simplifies the equation for changing the volume to
<img file="PL2941771T3_D0015.tif" />
(t-pGO-c) n (t iFGOllO (Χοοψιο <sub>2</sub> (r-soo- »yio + 10 (Lbco-c) m>
<sup>+</sup> gjSGĆjlO +10 ieao<sup>m</sup> and "GO<sup>/ tC1</sup> [0158] Again, ΔL (A, B) can be treated as a loudness compensation value that can be determined by the signal processor 120.
[0159] In general, gFGO can be regarded as the playback gain associated with the FGO object foreground (group of foreground objects), and gBGO can be treated as the playback gain associated with the background BGO object (group of background objects).
[0160] As mentioned earlier, it is possible to send volume differences instead of absolute volumes. Let us now define the reference volume as the loudness of the FGO LREF = LFGO meta-object, i.e. KFGO = LFGO - LREF - 0 and KBGO = LBGO - LREF = LBGO - LFGO. Now the volume change is expressed as kL
<img file="PL2941771T3_D0016.tif" />
+ 10 <sup>Κ</sup>ΰ (Χ!<sup>Λ <</sup>·<sup>}</sup> [0161] It is also possible that, as in the case of SAOC-DE, two metaobjects do not have individual scaling factors, but one of the objects is left unmodified, while the other is suppressed to obtain the right proportions between the objects during the mix. In this playback design, the output signal has a lower volume than the default mix, and the volume change is
<img file="PL2941771T3_D0017.tif" />
<img file="PL2941771T3_D0018.tif" />
+ 10
ΧΒΟΟ'Ι<sup>0</sup> where
Sfgo
5 Sfgo - Sbgo
Sbgo
Sfgo: <sub><</sub> > <sup>and</sup> Sbgo '<sup>11</sup> Sfgo <sup>v</sup> Sbgo .Sbgo »Sbgo <sup><</sup> Sfgo
Sfgo 1 s ifSbgo -Sfgo [0162] This form is already relatively simple and is also relatively independent of the loudness measure used. The only real requirement is that loudness values should add up in the exponential domain. It is possible to send / save signal energy values instead of volume values because they are closely related to each other.
[0163] In each of the above equations, ΔL (A, B) can be treated as a loudness compensation value that can be transmitted by the decoder signal processor 120.
[0164] Exemplary cases will be discussed below. The accuracy of the available concepts is illustrated using two examples of signals. Both signals contain a 5.1 downmix, and their surround and LFE channels bypass SAOC processing.
[0165] Two main approaches were used: one ("ternary") with three metaobjects, i.e. FGO, BGO and bypass channels, for example:
X "X /? GĆ> <sup>+</sup> ^ BGO <sup>+</sup> ^ BYPASS '[0166] And a second (two-component) with two meta-objects, for example:
X <sup>—</sup> X / W "FX # GO · [0167] In a two-component approach, bypass channels can be, for example, mixed together with BGO when estimating the volume of meta-objects. The volume of both (or all three) objects is estimated, as well as the volume of the downmix signal, and values are saved.
[0168] Playback instructions take the form of two approaches, respectively
BGO <sup>+</sup> SbGO ^ BYPASS and
V = bG0 * FG0 + Sbgo * BGO [0169] Gain values are, for example, determined according to:
<img file="PL2941771T3_D0019.tif" />
where the FGO gFGO gain varies between -24 and +24 dB.
[0170] The output scenario is played, the volume is measured, and the attenuation is calculated based on the volume of the downmix signal.
[0171] The result is marked in Fig. 10 and Fig. 11 with a blue line with round markers. Fig. 10 is a first illustration and Fig. 11 is a second illustration of a measured loudness change and the result of using available concepts for estimating loudness change in a completely parametric manner.
[0172] Then, based on the downmix, the parametric attenuation is estimated using recorded values of the meta-object loudness and downmixing and playback information. Estimation using the volume of three meta-objects is marked with a green line with square markers, and estimation using the volume of two meta-objects is marked with a red line with star-shaped markers.
[0173] In the figures, it can be seen that the binary and ternary approaches give virtually identical results and both provide a relatively good approximation of the measured value.
[0174] The shared concepts have many advantages. For example, the concepts provided allow estimating the volume of a mix signal based on the volume of the component signals constituting the mix. The advantage of this is that component signal loudness can be estimated once, and mix signal estimation can be obtained parametrically for each mix without actually estimating signal-based loudness. This is a significant improvement in the context of the computational performance of the entire system, in which the estimation of loudness of various mixes is required. For example, if the end user changes the playback settings, output volume estimation is immediately available.
[0175] In some applications, for example, when complying with the EBU R128 recommendation, the average loudness of the entire program is important. If the receiver's volume estimation, for example in the transmission scenario, is based on the received signal, the estimation converges to the average volume only after receiving the entire program. For this reason, each volume compensation will contain errors or exhibit time changes. When estimating the volume of component objects in the proposed manner and transmitting loudness information, it is possible to estimate the average volume of the mix in the receiver without delay.
[0176] If it is desired that the average loudness of the output signal remain (approximately) constant regardless of changes in the reproduction information, the available concepts make it possible to determine the compensation factor for this reason. The calculations required for this purpose in the decoder are negligible from the point of view of their computational complexity, and thus it is possible to add such a function to any decoder.
[0177] There are cases where the absolute volume of the output signal is not significant, but it is important to determine the change in volume relative to the reference scene. In these cases, absolute object levels are not relevant, but relative levels. This allows you to define one of the objects as a reference object and to represent the volume of other objects in relation to the volume of that reference object. This has some advantages in the context of transporting and / or storing volume information.
[0178] First of all, it is not necessary to transport the reference volume. If two meta objects are used, the amount of data transferred is reduced by half. The second advantage is related to the possible quantization and representation of loudness values. Because absolute object levels can be almost arbitrary, absolute volume values can also be almost arbitrary. On the other hand, it was assumed that the absolute loudness values have an average of 0 and are relatively well distributed around the average. The difference between the representations allows you to define a relative quantization grid in a way that provides potentially greater accuracy with the same number of bits used in the quantized representation.
[0179] Fig. 12 shows other embodiments for performing loudness compensation. In Fig. 12, loudness compensation can be performed, for example, to compensate for loudness losses. For this purpose, it is possible, for example, to use the DE_loudness_diff_dialogue (= KFGO) and DE_loudness_diff_background (= KBGO) values from DE_control_info. For example, DE_control_info may be the Advanced Clean Audio "Dialogue Enhancement" (DE) control information.
[0180] Loudness compensation is obtained by using the "g" gain value to process the SAOC-DE output signal and bypass channels (in the case of a multi-channel signal).
[0181] In the embodiments shown in Fig. 12, this is carried out as follows:
The limited mg value of dialogue-modifying gain is used to determine effective foreground object enhancements (FGOs, e.g., dialogues) and background objects (BGOs, e.g., ambient sounds). This is done using block 1220 "gain mapping", which generates mFGO and mBGO gain values.
[0182] Block 1230 "output loudness estimation module" uses KFGO and KBGO loudness information and effective mFGO and mBGO gain values to estimate this possible loudness change compared to the default downmix case. The change is then mapped to the "loudness compensation factor" which is applied to the output channels to generate the final "output channels".
[0183] During loudness compensation, the following steps are performed:
- Receiving the limited mg gain value transmitted by the decoder
SAOC-DE (such as defined in clause 12.8 "Modification range control for SAOC-DE" [DE]) and definition of the applied FGO / BGO reinforcements:
BGO
<img file="PL2941771T3_D0020.tif" />
<sup>1</sup> <! ", -! when _<sub>s</sub> when = 1,. , m<sub>about</sub>"= M '0. w_> le inform bgo g etaobif ó and Kbgo.
- Calculation of the output volume change compared to the default downmix using
AL = 10 Log, KfGO / - K<sub>BCO</sub>/ mj,<sub>CM</sub>IN <sup>/0</sup>+<<sub>about</sub>10 <sup>ao</sup><sup>K</sup>rco / ~ Egoo / <sup>/ l0</sup>+10 <sup>/ l0</sup>
Δ = 10<sup>-0,05ńL</sup>.
Calculation of scaling factors
<img file="PL2941771T3_D0021.tif" />
<img file="PL2941771T3_D0022.tif" />
£ δ <sup>m</sup>BGO% b if channel i belongs to the SAOC-DE output if channel with is a bypass channel and N is the total number of output channels.
In Fig. 12, the gain control is divided into two stages: the gain of the optional "bypass channels" is adjusted using mBGO before connecting them to the "SAOC-DE output channels", and then common gΔ gain is used for all connected channels. It is only possible to reproduce gain control operations, while g combines both stages of gain control into one gain control.
- Applying the scaling g value to YFULL audio channels consisting of YSAOC "SAOC-DE output channels" and optional time-aligned Ybypass "bypass channels". Yfull = Ysaoc o Ybypass.
[0184] The use of the scaling g value to process YFULL audio channels is performed by the gain control assembly i240.
[0185] Calculated above ΔL can be regarded as a loudness compensation value. In general, mFGO indicates the gain of reproducing the FGO object foreground (groups of foreground objects), and mBGO indicates the gain of reproducing the background BGO object (groups of background objects).
[0186] Although some aspects have been described with respect to the device, it is clear, however, that these aspects are also a description of the corresponding method in which the block or device corresponds to the method step or function of the method step. Similarly, the aspects described in relation to the method step also represent a description of the corresponding block or element or function of the respective device.
[0187] The striped signal of the invention may be recorded on a digital storage medium or transmitted using a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.
[0188] Depending on certain requirements related to the implementation of the invention, embodiments of the invention may be implemented in hardware or in software. The implementation can also be carried out using a digital data storage medium, e.g. floppy disks, DVDs, Blue-Ray discs, CDs, ROM, PROM, EPROM, EEPROM or FLASH memory, on which control signals that cooperate (or are capable of for such interaction) with a programmable computer system so that a suitable method is implemented.
[0189] Some embodiments of the invention include a durable data carrier having electronically readable control signals that can cooperate with a programmable computer system such that one of the methods described herein is implemented.
[0190] In general, embodiments of the invention may be implemented as a computer program product with a program code, which program code may operate to implement one method of the invention when the computer program product is running on a computer. The program code may, for example, be written on a machine-readable medium.
[0191] Other embodiments include a computer program for performing one of the methods described herein, stored on a machine readable carrier.
In other words, an embodiment of the method of the invention is thus a computer program having a program code for carrying out one of the methods described herein when the computer program product is running on a computer.
[0193] A further embodiment of the methods of the invention is thus a data carrier (or a digital memory carrier, or a computer readable medium) containing a computer program written thereon for carrying out one of the methods described herein.
[0194] A further embodiment of the method of the invention is thus a data stream or a sequence of signals representing a computer program for performing one of the methods described herein. The data stream or signal sequence may, for example, be shaped to be transmitted over a data link, e.g., via the Internet.
[0195] Another embodiment includes processing means, e.g., a computer or programmable logic device, shaped or adapted to implement one of the methods described herein.
[0196] Another embodiment includes a computer having a computer program installed therein for performing one of the methods described herein.
[0197] In some embodiments, the programmable logic device (e.g., a user programmable logic table) may be used to perform some or all of the functions of the methods described herein. In some embodiments, the user programmable logic table may interact with a microprocessor to implement one of the methods described herein. Generally, the methods are preferably carried out by any hardware device.
[0198] The above described embodiments are merely illustrative for the principles of the present invention. It should be understood that modifications and variants of the systems and details described herein are obvious to those skilled in the art. It is therefore intended that the restrictions arise only from the scope of the following claims, and not from the specific details provided for the purpose of describing and explaining the present embodiments.
Bibliography [0199] [BCC] C. Faller and F. Baumgarte, "Binaural Cue Coding - Part II: Schemes and applications", IEEE Trans. on Speech and Audio Proc., volume 11, No. 6, November 2003.
[EBU] EBU Recommendation R 128 "Loudness normalization and permitted maximum level of audio signals", Geneva, 2011.
[JSC] C. Faller, "Parametric Joint-Coding of Audio Sources", 120th AES Convention, Paris, 2006.
[ISS1] M. Parvaix and L. Girin: "Informed Source Separation of underdetermined instantaneous Stereo Mikstures using Source Index Embedding", IEEE ICASSP, 2010.
[ISS2] M. Parvaix, L. Girin, J.-M. Brossier: "A watermarking-based method for informed source separation of audio signals with a single sensor", IEEE Transactions on Audio, Speech and Language Processing, 2010.
[ISS3] A. Liutkus and J. Pinel and R. Badeau and L. Girin and G. Richard: "Informed source separation through spectrogram coding and data embedding", Signal Processing Journal, 2011.
[ISS4] A. Ozerov, A. Liutkus, R. Badeau, G. Richard: "Informed source separation: source coding meets source separation", IEEE Workshop on Applications of Signal Processing to Audio and Acoustics, 2011.
[ISS5] S. Zhang and L. Girin: "An Informed Source Separation System for Speech Signals", INTERSPEECH, 2011.
[ISS6] L. Girin and J. Pinel: "Informed Audio Source Separation from Compressed Linear Stereo Mikstures", AES 42nd International Conference: Semantic Audio, 2011.
[ITU] International Telecommunication Union: "Recommendation ITU-R BS.1770-3 Algorithms to measure audio program loudness and true-peak audio level", Geneva, 2012.
[SAOC1] J. Herre, S. Disch, J. Hilpert, O. Hellmuth: "From SAC To SAOC - Recent
Developments in Parametric Coding of Spatial Audio ”, 22nd Regional UK AES Conference,
Cambridge, United Kingdom, April 2007.
[SAOC2] J. EngdegSrd, B. Resch, C. Falch, O. Hellmuth, J. Hilpert, A. Holzer, L. Terentiev, J. Breebaart, J. Koppens, E. Schuijers and W. Oomen: "Spatial Audio Object Coding (SAOC) The Upcoming MPEG Standard on Parametric Object Based Audio Coding ”, 124th AES Convention, Amsterdam 2008.
[SAOC] ISO / IEC, "MPEG audio technologies - Part 2. Spatial Audio Object Coding (SAOC)", ISO / IEC JTC1 / SC29 / WG11 (MPEG) International Standard 23003-2.
[EP] EP 2i46522 A1. S. Schreiner, W. Fiesel, M. Neusinger, O. Hellmuth, R.
Sperschneider, "Apparatus and method for generating audio output signals using object based metadata", 2010.
[DE] ISO / IEC, "MPEG audio technologies - Part 2. Spatial Audio Object Coding (SAOC) Amendment 3, Dialogue Enhancement", ISO / IEC 23003-2.20i0 / DAM 3, Dialogue
Enhancement.
[BRE] WO 2008/035275 A2.
[SCH] EP 2 i46 522 Ai.
[ENG] WO 2008 / 04653i Ai.
Fraunhofer- Gesellschaft zur Forderung der angewandten Forschung eV, Germany Representative:
EP 2 94i 77i Bi
Z - 15848/17
Contents10
74 members in 19 offices
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 13194664 | European Patent Office (EPO) | A | |
| 2014075801 | European Patent Office (EPO) | W |
Members74
| Document | Office | Kind | |
|---|---|---|---|
| EP2879131A1 | European Patent Office (EPO) | A1 | |
| CA2900473A1 | Canada | A1 | |
| CA2931558A1 | Canada | A1 | |
| WO2015078956A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2015078964A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW201525990A | Taiwan Province of China | A | |
| AU2014356475A1 | Australia | A1 | |
| TW201535353A | Taiwan Province of China | A | |
| KR20150123799A | Republic of Korea | A | |
| EP2941771A1 | European Patent Office (EPO) | A1 | |
| US2015348564A1 | United States of America | A1 | |
| CN105144287A | China | A | |
| MX2015013580A | Mexico | A | |
| AR098558A1 | Argentina | A1 | |
| AU2014356467A1 | Australia | A1 | |
| KR20160075756A | Republic of Korea | A | |
| JP2016520865A | Japan | A | |
| AR099360A1 | Argentina | A1 | |
| CN105874532A | China | A | |
| AU2014356475B2 | Australia | B2 | |
| MX2016006880A | Mexico | A | |
| US2016254001A1 | United States of America | A1 | |
| EP3074971A1 | European Patent Office (EPO) | A1 | |
| AU2014356467B2 | Australia | B2 | |
| HK1217245A1 | Hong Kong, China | A1 | |
| JP2017502324A | Japan | A | |
| TWI569259B | Taiwan Province of China | B | |
| TWI569260B | Taiwan Province of China | B | |
| RU2015135181A | Russian Federation | A | |
| EP2941771B1 | European Patent Office (EPO) | B1 | |
| KR101742137B1 | Republic of Korea | B1 | |
| PT2941771T | Portugal | T | |
| BR112015019958A2 | Brazil | A2 | |
| BR112016011988A2 | Brazil | A2 | |
| ES2629527T3 | Spain | T3 | |
| MX350247B | Mexico | B | |
| JP6218928B2 | Japan | B2 | |
| PL2941771T3This record | Poland | T3 | |
| ZA201604205B | South Africa | B | |
| RU2016125242A | Russian Federation | A | |
| CA2900473C | Canada | C | |
| EP3074971B1 | European Patent Office (EPO) | B1 | |
| US9947325B2 | United States of America | B2 | |
| RU2651211C2 | Russian Federation | C2 | |
| ES2666127T3 | Spain | T3 | |
| PT3074971T | Portugal | T | |
| KR101852950B1 | Republic of Korea | B1 | |
| JP6346282B2 | Japan | B2 | |
| US2018197554A1 | United States of America | A1 | |
| PL3074971T3 | Poland | T3 | |
| MX358306B | Mexico | B | |
| RU2672174C2 | Russian Federation | C2 | |
| CA2931558C | Canada | C | |
| US10497376B2 | United States of America | B2 | |
| US2020058313A1 | United States of America | A1 | |
| CN105874532B | China | B | |
| CN111312266A | China | A | |
| US10699722B2 | United States of America | B2 | |
| US2020286496A1 | United States of America | A1 | |
| CN105144287B | China | B | |
| CN112151049A | China | A | |
| US10891963B2 | United States of America | B2 | |
| US2021118454A1 | United States of America | A1 | |
| BR112015019958B1 | Brazil | B1 | |
| MY189823A | Malaysia | A | |
| US11423914B2 | United States of America | B2 | |
| BR112016011988B1 | Brazil | B1 | |
| US2022351736A1 | United States of America | A1 | |
| MY196533A | Malaysia | A | |
| US11688407B2 | United States of America | B2 | |
| US2023306973A1 | United States of America | A1 | |
| CN111312266B | China | B | |
| US11875804B2 | United States of America | B2 | |
| CN112151049B | China | B |
Numbers
- Application
- 14802914
Titles2
- English
- DECODER, ENCODER AND METHOD FOR INFORMED LOUDNESS ESTIMATION EMPLOYING BY-PASS AUDIO OBJECT SIGNALS IN OBJECT-BASED AUDIO CODING SYSTEMS
- Polish
- Dekoder, koder i sposób opartej o informacje estymacji głośności z wykorzystaniem obejściowych sygnałów obiektów audio w układach obiektowego kodowania audio
Classification
- CPC, 5
- G10L19/008
- G10L19/0017
- G10L19/005
- G10L19/265
- H03G3/20
- IPC, 1
- G10L19 008