Decoder, encoder and method for informed loudness estimation in object-based audio coding systems.
Abstract
A decoder for generating an audio output signal comprising one or more audio output channels is provided. The decoder comprises a receiving interface (110) for receiving an audio input signal comprising a plurality of audio object signals, for receiving loudness information on the audio object signals, and for receiving rendering information indicating whether one or more of the audio object signals shall be amplified or attenuated. Moreover, the decoder comprises a signal processor (120) for generating the one or more audio output channels of the audio output signal. The signal processor (120) is configured to determine a loudness compensation value depending on the loudness information and depending on the rendering information. Furthermore, the signal processor (120) is configured to generate the one or more audio output channels of the audio output signal from the audio input signal depending on the rendering information and depending on the loudness compensation value. Moreover, an encoder is provided.

Term
8.2 yearsleft in the term
Expires 27 November 2034.
- Priority
- Filed
- Granted
- Today
- Expires
20 claims: 11 independent, 9 dependent
- 1REIVINDICACIONES Un decodificador para generar una señal de salida de audio que comprende uno o más canales de salida de audio, caracterizado porque el decodificador comprende:una interfaz de recepción (110) para recibir una señal de entrada de audio que comprende una pluralidad de señales de objetos de audio, para recibir información de sonoridad sobre las señales de objetos de audio, y para recibir información de representación que indica cómo una o más de las señales de objetos de audio debe ser amplificada o atenuada, y un procesador de señal (120) para generar los uno o más canales de salida de audio de la señal de salida de audio, en donde el procesador de señal (120) está configurado para determinar un valor de compensación de sonoridad que depende de la información de sonoridad y que depende de la información de representación, y en donde el procesador de señal (120) está configurado para generar los uno o más canales de salida de audio de la señal de salida de audio a partir de la señal de entrada de audio que depende de la información de representación y que depende del valor de compensación de sonoridad, en donde el procesador de señal (120) está configurado para generar los uno o más canales de salida de audio de la señal de salida de audio a partir de la señal de entrada de audio que depende de la información de representación y que depende del valor de compensación de IMPI nW1 ™T° MEXICANO > DE LA PROPIEDAD INDUSTRIAL sonoridad, de tal manera que una sonoridad de la señal de salida de audio sea igual a una sonoridad de la señal de entrada de audio, o de tal manera que la sonoridad de la señal de salida de audio sea más próxima a la sonoridad de la señal de entrada de audio que una sonoridad de una señal de audio modificada que resultaría de modificar la señal de entrada de audio al amplificar o atenuar las señales de objetos de audio de la señal de entrada de audio de acuerdo con la información de representación.
- 2El decodificador de conformidad con la reivindicación 1, caracterizado además porque el procesador de señal (120) está configurado para generar la señal de audio modificada al modificar la señal de entrada de audio al amplificar o atenuar las señales de objetos de audio de la señal de entrada de audio de acuerdo con la información de representación, y en donde el procesador de señal (120) está configurado para generar la señal de salida de audio al aplicar el valor de compensación de sonoridad en la señal de audio modificada, de tal manera que la sonoridad de la señal de salida de audio sea igual a la sonoridad de la señal de entrada de audio, o de tal manera que la sonoridad de la señal de salida de audio sea más próxima a la sonoridad de la señal de entrada de audio que la sonoridad de la señal de audio modificada. IMPI INSTITUTO MEXICANO DE LA MONEDAD INDUSTRIAL
- 3El decodificador de conformidad con la reivindicación 1 o 2, caracterizado además porque cada una de las señales de objetos de audio de la señal de entrada de audio está asignada a exactamente un grupo de dos o más grupos, en donde cada uno de los dos o más grupos comprende una o más de las señales de objetos de audio de la señal de entrada de audio, en donde la interfaz de recepción (110) está configurada para recibir un valor de sonoridad para cada grupo de los dos o más grupos como la información de sonoridad, en donde el procesador de señal (120) está configurado para determinar el valor de compensación de sonoridad que depende del valor de sonoridad de cada uno de los dos o más grupos, y en donde el procesador de señal (120) está configurado para generar los uno o más canales de salida de audio de la señal de salida de audio a partir de la señal de entrada de audio que depende del valor de compensación de sonoridad.
- 4El decodificador de conformidad con la reivindicación 3, caracterizado además porque al menos un grupo de los dos o más grupos comprende dos o más de las señales de objetos de audio.
- 5El decodificador de conformidad con la reivindicación 1 o 2, caracterizado además porque cada una de las señales de objetos de audio de la señal de entrada de audio está asignada a exactamente un grupo de más de dos ·.* IMPI instttuto mexicano de la propiedad industrial grupos grupos, en donde cada uno de los más de dos comprende una o más de las señales de objetos de audio de la señal de entrada de audio, en donde la interfaz de recepción (110) está configurada para recibir un valor de sonoridad para cada grupo de los más de dos grupos como la información de sonoridad, en donde el procesador de señal (120) está configurado para determinar el valor de compensación de sonoridad que depende del valor de sonoridad de cada uno de los más de dos grupos, y en donde el procesador de señal (120) está configurado para generar los uno o más canales de salida de audio de la señal de salida de audio a partir de la señal de entrada de audio que depende del valor de compensación de sonoridad.
- 6El decodificador de conformidad con la reivindicación 5, caracterizado además porque al menos un grupo de los más de dos grupos comprende dos o más de las señales de objetos de audio.
- 7El decodificador de conformidad con una de las reivindicaciones 3 a 6, caracterizado además porque el procesador de señal (120) está configurado para determinar el valor de compensación de sonoridad de acuerdo con la fórmula AL =101og 10 o de acuerdo con la fórmula IMPI $5^*% INSTITUTO MEXICANO DE LA PROPIEDAD INDUSTRIAL AL- lOlog en sonoridad, en audio de las una sonoridad donde gi obj etos para la es un donde AL es el valor donde i indica una señales de la primer de audio, en i-ava de objetos i-ava señal de compensación señal de objetos de audio, en donde Li de objetos de audio, peso de mezcla para la donde hi es un segundo i-ava señal de objetos de audio, en valor constante, y en donde N es un número. reivindicaciones 3 a 6, procesador de señal (120) el valor de compensación fórmula de de es en i-ava señal de peso de mezcla donde c es un con una de las caracterizado además está configurado para porque el determinar de sonoridad de acuerdo con la .v Σ. g,-10 Ai = 101og„^--- ¿VflO* 1 ;=1 en donde AL es el valor de compensación de sonoridad, en donde i indica una í-ava señal de objetos de audio de las señales de objetos de audio, en donde g Y es un primer peso de mezcla para la ϊ-ava señal de objetos de audio, en donde hi es un segundo peso de mezcla para la iava señal de objetos de audio, en donde N es un número, y en donde se define de acuerdo con K^L^-L^p, en donde Li es una sonoridad de la i-ava señal de objetos de audio, y en donde L REF es la sonoridad de un objeto de referencia.
- 89. El decodificador de conformidad con la reivindicación 3 o 4, caracterizado además porque cada una de las señales de objetos de audio de la señal de entrada de audio está asignada a exactamente un grupo de exactamente dos grupos como los dos o más grupos, en donde cada una de las señales de objetos de audio de la señal de entrada de audio está asignada ya sea a un grupo de objetos en primer plano de los exactamente dos grupos o a un grupo de objetos en segundo plano de los exactamente dos grupos, en donde la interfaz de recepción (110) está configurada para recibir el valor de sonoridad del grupo de objetos en primer plano, en donde la interfaz de recepción (110) está configurada para recibir el valor de sonoridad del grupo de objetos en segundo plano, en donde el procesador de señal (120) está configurado para determinar el valor de compensación de sonoridad que depende del valor de sonoridad del grupo de objetos IMPI INSTITUTO MEXICANO DE LA PROPIEDAD INDUSTRIAL en primer plano, y que depende del valor de sonoridad del grupo de objetos en segundo plano, y en donde el procesador de está configurado para generar los uno o más canales de salida de audio de la señal de salida de audio a partir de la señal de entrada de audio que depende del valor de compensación de sonoridad.
- 910. El decodificador de conformidad con la reivindicación 9, caracterizado además porque el procesador de señal (120) está configurado para determinar el valor de compensación de sonoridad de acuerdo con la fórmula en donde AL es el valor de compensación de sonoridad, en donde Kfgo indica el valor de sonoridad del grupo de objetos en primer plano, en donde Kbgo indica el valor de sonoridad del grupo de objetos en segundo plano, en donde mreo indica una ganancia de representación del grupo de obj etos en primer plano, y en donde iübgo indica una ganancia de representación del grupo de objetos en segundo plano.
- 1011. El decodificador de conformidad con la reivindicación caracterizado además porque el IMPI INSTITUTO MEXICANO DE LA PROPIEDAD industrial procesador de señal (120) está configurado para determinar el valor de compensación de sonoridad de acuerdo con la fórmula en donde AL es el valor de compensación de sonoridad, en donde Lfgo indica el valor de sonoridad del grupo de objetos en primer plano, en donde Lbgo indica el valor de sonoridad del grupo de objetos en segundo plano, en donde greo indica una ganancia de representación del grupo de objetos en primer plano, y en donde geco indica una ganancia de representación del grupo de objetos en segundo plano.
- 1112. El decodificador de conformidad con una de las reivindicaciones anteriores, caracterizado además porque la interfaz de recepción (110) está configurada para recibir una señal de mezcla descendente que comprende uno o más canales de mezcla descendente como la señal de entrada de audio, en donde los uno o más canales de mezcla descendente comprenden las señales de objetos de audio, y en donde el número de los uno o más canales de mezcla descendente es menor que el número de las señales de objetos de audio, en donde la interfaz de recepción (110) está configurada para recibir información de mezcla descendente que indica cómo están mezcladas las señales de IMPI INSTITUTO MEXICANO DE LA MONEDAD INDUSTRIAL objetos de audio dentro de los uno o más canales de mezcla descendente, y en donde el procesador de señal (120) está configurado para generar los uno o más canales de salida de audio de la señal de salida de audio a partir de la señal de entrada de audio que depende de la información de mezcla descendente, que depende de la información de representación y que depende del valor de compensación de sonoridad.
- 1213. El decodificador de conformidad con la reivindicación 12, caracterizado además porque la interfaz de recepción (110) está configurada para recibir una o más señales de obj etos de audio de derivación adicionales en donde las una o más señales de objetos de audio de derivación adicionales no están mezcladas dentro de la señal de mezcla descendente, en donde la interfaz de recepción (110) está configurada para recibir la información de sonoridad que indica información sobre la sonoridad de las señales de objetos de audio que están mezcladas dentro de la señal de mezcla descendente y que indica información sobre la sonoridad de las una o más señales de objetos de audio de derivación adicionales que no están mezcladas dentro de la señal de mezcla descendente, y en donde el procesador de señal (120) está configurado para determinar el valor de compensación de sonoridad que depende de la información sobre la sonoridad IMPI INSTITUTO MEXICANO DE LA PROPIEDAD industrial de las señales de objetos de audio que están mezcladas dentro de la señal de mezcla descendente, y que depende de la información sobre la sonoridad de las una o más señales de objetos de audio de derivación adicionales que no están mezcladas dentro de la señal de mezcla descendente.
- 1314. Un decodificador para generar una señal de salida de audio caracterizado porque comprende uno o más canales de salida de audio, en donde el decodificador comprende:una interfaz de recepción (110) para recibir una señal de entrada de audio que comprende una pluralidad de señales de objetos de audio, para recibir información de sonoridad sobre las señales de objetos de audio, y para recibir información de representación que indica si una o más de las señales de objetos de audio debe ser amplificada o atenuada, y un procesador de señal (120) para generar los uno o más canales de salida de audio de la señal de salida de audio, en donde el procesador de señal (120) está configurado para determinar un valor de compensación de sonoridad que depende de la información de sonoridad y que depende de la información de representación, y en donde el procesador de señal (120) está configurado para generar los uno o más canales de salida de audio de la señal de salida de audio a partir de la señal de entrada de audio que depende de la información de representación y que depende del valor de compensación de sonoridad, en donde la interfaz de recepción (110) está ronfigurada nara recibir INSTITUTO MEXICANO OS LA PROPIEDAD CWZrStülí'í, industrial VwgerarTgjiM 5 una señal de mezcla descendente que comprende uno o más canales de mezcla descendente como la señal de entrada de audio, en donde los uno o más canales de mezcla descendente comprenden las señales de objetos de audio, y en donde el número de los uno o más canales de mezcla descendente es menor que el número de las señales de objetos de audio, en donde la interfaz de recepción (110) está configurada para recibir información de mezcla descendente que indica cómo están mezcladas las señales de objetos de audio dentro de los uno o más canales de mezcla descendente, y en donde el procesador de señal (120) está configurado para generar los uno o más canales de salida de audio de la señal de salida de audio a partir de la señal de entrada de audio que depende de la información de mezcla descendente, que depende de la información de representación y que depende del valor de compensación de sonoridad, en donde la interfaz de recepción (110) está configurada para recibir una o más señales de objetos de audio de derivación adicionales, en donde las una o más señales de objetos de audio de derivación adicionales no están mezcladas dentro de la señal de mezcla descendente, en donde la interfaz de recepción (110) está configurada para recibir la información de sonoridad que indica información sobre la sonoridad de la señales de objetos de audio que están mezcladas dentro de la señal de mezcla rin.^candente y que indica información sobre la sonoridad de las una o más IMPI INSTITUTO MEXICANO DE LA PROPIEDAD INDUSTRIAL señales de objetos de audio de derivación adicionales que no están mezcladas dentro de la señal de mezcla descendente, y en donde el procesador de señal (120) está configurado para determinar el valor de compensación de sonoridad que depende de la información sobre la sonoridad de las señales de objetos de audio que están mezcladas dentro de la señal de mezcla descendente, y que depende de la información sobre la sonoridad de las una o más señales de objetos de audio de derivación adicionales que no están mezcladas dentro de la señal de mezcla descendente.
- 1415. Un codificador, caracterizado porque comprende:una unidad de codificación basada en objetos (210;710) para codificar una pluralidad de señales de objetos de audio para obtener una señal de audio codificada que comprende la pluralidad de señales de objetos de audio, y una unidad de codificación de sonoridad de objetos (220;720;820) para codificar información de sonoridad sobre las señales de objetos de audio, en donde la información de sonoridad comprende uno o más valores de sonoridad, en donde cada uno de los uno o más valores de sonoridad depende de una o más de las señales de objetos de audio, en donde cada una de las señales de objetos de audio de la señal de audio codificada está asignada a exactamente un IMPI INSTITUTO MEXICANO DI LA MONEDAD INDUSTRIAL grupo de dos o más grupos, en donde cada uno de los dos o más grupos comprende una o más de las señales de objetos de audio de la señal de audio codificada, en donde al menos un grupo de los dos o más grupos comprende dos o más de las señales de objetos de audio, en donde la unidad de codificación de sonoridad de objetos (220;720;820) está configurada para determinar los uno o más valores de sonoridad de la información de sonoridad al determinar un valor de sonoridad para cada grupo de los dos o más grupos, en donde dicho valor de sonoridad de dicho grupo indica una sonoridad total de las una o más señales de objetos de audio de dicho grupo.
- 1516. Un codificador, caracterizado porque comprende:una unidad de codificación basada en objetos (210;710) para codificar una pluralidad de señales de objetos de audio para obtener una señal de audio codificada que comprende la pluralidad de señales de objetos de audio, y una unidad de codificación de sonoridad de objetos (220;720;820) para codificar información de sonoridad sobre las señales de objetos de audio, en donde la información de sonoridad comprende uno o más valores de sonoridad, en donde cada uno de los uno o más valores de sonoridad depende de una o más de las señales de objetos de audio, en donde la unidad de codificación basada en objetos (210;710) está configurada para recibir las señales de objetos IMPI INSTITUTO MEXICANO M LA MONEDAD INDUSTRIAL de audio, en donde cada una de las señales de objetos de audio está asignada a exactamente uno de exactamente dos grupos, en donde cada uno de los exactamente dos grupos comprende una o más de las señales de objetos de audio, en donde al menos un grupo de los exactamente dos grupos comprende dos o más de las señales de objetos de audio, en donde la unidad de codificación basada en objetos (210;710) está configurada para efectuar la mezcla descendente de las señales de objetos de audio, siendo comprendidas por los exactamente dos grupos, para obtener una señal de mezcla descendente que comprende uno o más canales de audio de mezcla descendente como la señal de audio codificada, en donde el número de los uno o más canales de mezcla descendente es menor que el número de señales de objetos de audio que están comprendidos por los exactamente dos grupos, en donde la unidad de codificación de sonoridad de objetos (220;720;820) está configurada para recibir una o más señales de objetos de audio de derivación adicionales, en donde cada una de las una o más señales de objetos de audio de derivación adicionales están asignadas a un tercer grupo, en donde cada una de las una o más señales de objetos de audio de derivación adicionales no está comprendida por el primer grupo y no está comprendida por el segundo grupo, en donde la unidad de codificación basada en obj etos (210;710) está configurada para no efectuar la IMPI •Nsmvro MEXICANO DE LA PROPIEDAD INDUSTRIAL mezcla descendente de las una o más señales de objetos de audio de derivación adicionales dentro de la señal de mezcla descendente, y en donde la unidad de codificación de sonoridad de objetos primer determinar un valor de sonoridad, un segundo valor de sonoridad y un tercer valor de sonoridad de la información de sonoridad, el primer valor de sonoridad que indica una sonoridad total de las una o más señales de objetos de audio del primer grupo , el segundo valor de sonoridad que indica una sonoridad total de las una o más el tercer señales de objetos de audio del segundo grupo, y valor de sonoridad que indica una sonoridad total de las una o más señales de objetos de audio de derivación adicionales del tercer grupo, o está configurado para determinar un primer valor de sonoridad y un segundo valor de sonoridad de la información de sonoridad, el primer valor de sonoridad que indica una sonoridad total de las una o más señales de objetos de audio del primer grupo, y el segundo valor de sonoridad que indica una sonoridad total de las una o más señales de objetos de audio del segundo grupo y de las una o más señales de objetos de audio de derivación adicionales del tercer grupo.
- 1617. Un sistema, caracterizado porque comprende:un codificador (310) que comprende: una unidad de IMPI INSTITUTO MEXICANO DE LA PROPIEDAD INDUSTRIAL una pluralidad de señales de objetos de audio para obtener una señal de audio codificada que comprende la pluralidad de señales de objetos de audio, y una unidad de codificación de sonoridad de objetos (220;720;820) para codificar información de sonoridad en las señales de objetos de audio, en donde la información de sonoridad comprende uno o más valores de sonoridad, en donde cada uno de los uno o más valores de sonoridad depende de una o más de las señales de objetos de audio, un decodificador (320) como el que se reclama en una de las reivindicaciones 1 a 14, para generar una señal de salida de audio que comprende uno o más canales de salida de audio, en donde el decodificador (320) está configurado para recibir la señal de audio codificada como una señal de entrada de audio y para recibir la información de sonoridad, en donde el decodificador está configurado para recibir información de representación adicional, en donde el configurado para determinar un valor de compensación de sonoridad que depende de la información de sonoridad y que depende de la información de representación, y en donde el decodificador (320) está configurado para generar los uno o más canales de salida de audio de la señal de salida de audio a partir de la señal de entrada de audio que depende de la información de representación y que depende del valor de compensación de IMPI INSTITUTO MEXICANO DE LA PROPIEDAD INDUSTRIAL sonoridad.
- 1718. Un método para generar una señal de salida de audio caracterizado porque comprende uno o más canales de salida de audio, en donde el método comprende:recibir una señal de entrada de audio que comprende una pluralidad de señales de objetos de audio, recibir información de sonoridad en las señales de objetos de audio, recibir información de representación que indica cómo una o más de las señales de objetos de audio debe ser amplificada o atenuada, determinar un valor de compensación de sonoridad que depende de la información de sonoridad y que depende de la información de representación, y generar los uno o más canales de salida de audio de la señal de salida de audio a partir de la señal de entrada de audio que depende de la información de representación y que depende del valor de compensación de sonoridad, en donde generar los uno o más canales de salida de audio de la señal de salida de audio a partir de la señal de entrada de audio se realiza dependiendo de la información de representación y que depende del valor de compensación de sonoridad, de tal manera que una sonoridad de la señal de salida de audio es igual a una sonoridad de la señal de entrada de audio, o de tal manera que la sonoridad de la señal de salida de audio se aproxima más a la sonoridad de la señal de entrada de audio que una sonoridad de una señal de audio modificada IMPIí^ instituto mexicano de la propiedad CTmJSÍ-Z INDUSTRIAL que puede resultar de modificar la seña·!—4e—&ntr-ada...-da. audio al amplificar o atenuar las señales de objetos de audio de la señal de entrada de audio de acuerdo con la información de representación.
- 1819. Un método para generar una señal de salida de audio, caracterizado porque comprende uno o más canales de salida de audio, en donde el método comprende:recibir una señal de entrada de audio que comprende una pluralidad de señales de objetos de audio, en donde recibir la señal de entrada de audio se realiza al recibir una señal de mezcla descendente que comprende uno o más canales de mezcla descendente como la señal de entrada de audio, en donde los uno o más canales de mezcla descendente comprenden las señales de objetos de audio, y en donde el número de los uno o más canales de mezcla descendente son menores que el número de las señales de objetos de audio, recibir información de representación que indica si una o más de las señales de objetos de audio deben ser amplificadas o atenuadas, recibir información de mezcla descendente que indica cómo las señales de objetos de audio se mezclan dentro de los uno o más canales de mezcla descendente, recibir una o más señales de objetos de audio de derivación adicionales, en donde las una o más señales de objetos de audio de derivación adicionales no se mezclan dentro de la señal de mezcla descendente, recibir información de sonoridad en las señales de objetos IMPI INSTITUTO MEXICANO DE LA PROPIEDAD INDUSTRIAL de audio, en donde la información de sonoridad indica información en la sonoridad de las señales de objetos de audio que se mezclan dentro de la señal de mezcla descendente e indica información en la sonoridad de las una o más señales de objetos de audio de derivación adicionales que no se mezclan dentro de la señal de mezcla descendente, y determinar un valor de compensación de sonoridad que depende de la información de sonoridad y que depende de la información de representación, en donde determinar el valor de compensación de sonoridad se realiza dependiendo de la información en la sonoridad de las señales de objetos de audio que se mezclan dentro de la señal de mezcla descendente, y que depende de la información en la sonoridad de las una o más señales de objetos de audio de derivación adicionales que no se mezclan dentro de la señal de mezcla descendente, y generar los uno o más canales de salida de audio de la señal de salida de audio a partir de la señal de entrada de audio que depende de la información de mezcla descendente, que depende de la información de representación y que depende del valor de compensación de sonoridad.
- 1920. Un método para codificar, caracterizado porque comprende:codificar una pluralidad de señales de objetos de audio para obtener una señal de audio codificada que comprende la pluralidad de señales IMPI INSTITUTO MEXICANO M LA l>ROM£DAD INDUSTRIAL audio, de objetos de y determinar información de sonoridad sobre las señales de objetos de audio, en donde la información de sonoridad comprende uno o más valores de sonoridad, en donde cada uno de los uno o más valores de sonoridad depende de una o más de las señales de objetos de audio, en donde determinar los uno o más valores de sonoridad de la información de sonoridad se realiza al determinar un valor de sonoridad para cada grupo de los dos o más grupos, en donde dicho valor de sonoridad de dicho grupo indica una sonoridad total de las una o más señales de objetos de audio de dicho grupo, codificar la información de sonoridad en las señales de objetos de audio, en donde cada una de las señales de objetos de audio de la señal de audio codificada se asigna a exactamente un grupo de dos o más grupos, en donde cada uno de los dos o más grupos comprende una o más de las señales de obj etos de audio de la señal de audio codificada, en donde al menos un grupo de los dos o más grupos comprende dos o más de las señales de objetos de audio.
- 2021. Un método para codificar, caracterizado porque comprende:recibir las señales de objetos de audio, en donde cada una de las señales de objetos de audio se asignan a exactamente uno de exactamente dos grupos, en donde cada uno de los exactamente dos grupos comprende una o más de las señales de objetos de andio T οη al menos IMPI INSTITUTO MEXICANO DE LA PROPIEDAD INDUSTRIAL un grupo de los exactamente dos grupos comprende dos o más de las señales de objetos de audio, codificar la pluralidad de señales de objetos de audio para obtener una señal de audio codificada que comprende la pluralidad de señales de objetos de audio al efectuar la mezcla descendente de las señales de objetes de audio, siendo comprendido por los exactamente dos grupos, para obtener una señal de mezcla descendente que comprende uno o más canales de audio de mezcla descendente como la señal de audio codificada, en donde el número de los uno o más canales de mezcla descendente es menor que el número de las señales de objetos de audio siendo comprendidas por los exactamente dos grupos, determinar información de sonoridad en las señales de objetos de audio, en donde la información de sonoridad comprende uno o más valores de sonoridad, en donde cada uno de los uno o más valores de sonoridad depende de una o más de las señales de objetos de audio, al determinar un primer valor de sonoridad, un segundo valor de sonoridad y un tercer valor de sonoridad de información de sonoridad, el primer valor de sonoridad que indica una sonoridad total de las una o más señales de objetos de audio del primer grupo, el segundo valor de sonoridad que indica una sonoridad total de las una o más señales de objetos de audio del segundo grupo, y el tercer valor de 100 IMPI mexicano °E LA PROPIEDAD industrial sonoridad que indica una sonoridad total de las una o más señales de objetos de audio de derivación adicionales del tercer grupo, o al determinar un primer valor de sonoridad y un segundo valor de sonoridad de la información de sonoridad, el primer valor de sonoridad que indica una sonoridad total de las una o más señales de objetos de audio del primer grupo, y el segundo valor de sonoridad que indica una sonoridad total de las una o más señales de obj etos de audio del segundo grupo y de las una o más señales de objetos de audio de derivación adicionales del tercer grupo, codifica r la información de sonoridad en las señales de objetos de audio, recibir una o más señales de objetos de audio de derivación adicionales, en donde cada una de las una o más señales de objetos de audio de derivación adicionales se asignan a un tercer grupo, en donde cada una de las una o más señales de objetos de audio de derivación adicionales no se comprenden por el primer grupo y no se comprenden por el segundo grupo, y no efectuar la mezcla descendente de las una o más señales de objetos de audio de derivación adicionales dentro de la señal de mezcla descendente. Un medio legible por computadora caracterizado porque comprende el método como el que se reclama en una de las reivindicaciones 18 a 21. 101 IMPI ΙΝΠΠυΤΟ Mexicano OELA PROPIEDAD INDUSTRIAL
Independent claims20
695 paragraphs in 109 sections, as filed
(54) Title: DECODER, ENCODER AND METHOD FOR THE INFORMED ESTIMATION OF LOUDNESS IN OBJECT-BASED ENCODING SYSTEMS.
(54) Title: DECODER, ENCODER AND METHOD FOR INFORMED LOUDNESS ESTIMATION IN OBJECT-BASED AUDIO CODING SYSTEMS.
(57) Summary
A decoder is provided to generate an audio output signal that comprises one or more audio output channels. The decoder comprises a receiving interface (110) for receiving an audio input signal comprising a plurality of audio object signals, for receiving loudness information on the audio object signals, and for receiving display information indicating whether or not one or more of the audio object signals should be amplified or attenuated. Furthermore, the decoder comprises a signal processor (120) to generate one or more audio output channels from the audio output signal. The signal processor 120 is configured to determine a loudness compensation value depending on the loudness information and depending on the display information. Furthermore, the signal processor 120 is configured to generate one or more audio output channels of the audio output signal from the audio input signal depending on the display information and depending on the offset value of sonority. Furthermore, an encoder is provided.
(57) Abstract
A decoder for generating an audio output signal comprising one or more audio output channels is provided. The decoder comprises a receiving interface (110) for receiving an audio input signal comprising a plurality of audio object signáis, for receiving loudness information on the audio object signáis, and for receiving rendering information indicating whether one or more of the audio object signáis shall be amplified or attenuated. Moreover, the decoder comprises a signal processor (120) for generating the one or more audio output channels of the audio output signal. The signal processor (120) is configured to determine a loudness compensation valué depending on the loudness information and depending on the rendering information. Furthermore, the signal processor (120) is configured to generate the one or more audio output channels of the audio output signal from the audio input signal depending on the rendering reported and depending on the loudness compensated value. Moreover, an encoder is provided.
I Μ ΡI f
<img file="MX358306B_D0001.tif" />
PATENT TITLE No. 358306
<td>Headlines):</td><td>FRAUNHOFER-GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV</td>
<td>Home:</td><td>HansastraBe 27c, 80686, München, GERMANY</td>
Denomination:
Classification:
Inventor (s):
DECODER, ENCODER AND METHOD FOR THE INFORMED ESTIMATION OF LOUDNESS IN OBJECT-BASED AUDIO CODING SYSTEMS. ; .
CIP: G10L19 / 008
CPC: G10L19 / 008; G10L1S / 2W? Nft§Q3 / 20
JOUNI PAULUS; SASCHA H * CHS; BERNHARD GRILL; OLIVER
LfELLMUTH; ADRIAM ^ JRWANFA ^ K «) 'R | pb ^ RJBUSCH; LION TERENTIV
JOUNI PAULUS; SASCHA WCH!
HELLMUTH; ADRIAM ^ RWANFA ^ K €) RU ^ R ^ U · '' 's J<img file="MX358306B_D0002.tif" />International:
! 'L ·' '·· <' X ^ ft ^ i0Ür ^ b ^ 2O14
·. · - '2 / by nbMeáí ^ .á & ^ p3
Veheatient Date ^ 27 φ November-2034
Expedition date:
The patent of referendKserattfrga with melt it in the
Pursuant to article 28 of the given Law ProprietaryAJrteétqriJa prnea ^ e patéate tefe upé valid for twenty years (Sq.lmptorrogables, counted from the date of the present the international request and this payment with payment rights.
Whoever subscribes to the present title made it cpp funjfSfnento in the provisions of teaafticul ^ s 6<sup>or</sup> fraccidna »III and 7 ° 4is Z dpiatíy of Industrial Property (Official Gazette of the Federation <0.0 F.)<sup>-</sup>27/0 ^ 1991, reform on | 2'®8Í1994, Λ / 1ΙΜΜ96 12/26/1997, 05/17/1999. 01/26/2004, 06/16/2005, 01/25/2006, 06/05/2009, 06/01/2010, 1 (1 ^ / ^ 0.4, 08/06 /; jM (M27 / W2012ÍOWCfe / 2012íar | j £ gJI «JR °, 3? Jraádsm ^ r ^ 0a), 4 'and 12' fractions I and III of the Regulations of the Mexican Institute of te- £ f (jp.ie <ad ínSetriab (aLO.F. 14liai999, uifofmadó * 'el Ο1 | Ρ7 / 2β02'> $ í8t / 2004, 07/28/2004 and 09/07/2007); articles 1 », 3º, 4º, 5º fraction V subsection a), 1¿fF» ecjorra4 1 * 111 and 30'WNh4 «tea» Organij |> rafO> stitu * S'M «* ef> o de la Industrial Property (DO F. 12/27/1999, amended on 10/10/2002, 29 / OTOXH, 64 / βάδ®βνΤ ^ Μβ07), ¿«TtJaFAca prdo that delegates powers to the Directors
Deputy Generals, Coordinator, Directors'WrMqteéfMyTjiírarjte * | ejl les, Divisional Deputy Directors, Coordinators
Departmental and other subordinates of the MexidMoM Institute larrPr ^ e ^ l ^^^^^ 05/12/1999, reformed 02/04/2000, 07/29/2004,
04/08/2004 and 09/13/2007) '
Number:
MX / a / 2016/006880
Country:
EP
Validity: Twenty years' K
<img file="MX358306B_D0003.tif" />
Number:
131 ^ 4664.2 'Λ'
or. ^ aemOh ^^ fr ^ ji ^, y 5Hde la uy de la P In ^ tearf; Ha MflMme peteete tj ^ fe wfe validity di '*
I and beSradJwe * pay for Chiefrtfepgfg. pgjratener en lo ®¡
<img file="MX358306B_D0004.tif" />
<img file="MX358306B_D0005.tif" />
Industrial.
, »,. ™, ._________ J Í2Ü¡) 8Í1994, ¿/ 1Qtí9 $ §,
I ^ es / ^ I0.4} 06/06 / ^ 01 (X'27 / W2012 H, 0 $ / d »/ 2012» ar¡jpp $ si le £ f $> te <ad (rwtótriaL (IXP F 12 / 14 (1999, vjfofmadó-fel Ofgl7M02j
i), ItfFae aorraa 1 * 111 and 3O'WMlteteieWrganig & ^ tí.JnstitiAf Mejfcño dr 0TO0 (í, ό4 / βέδ®βνΐ «ΜΜ0Ο7), ¿íjt
Me'xidbfoeS I
<img file="MX358306B_D0006.tif" />
This letter is signed with an advanced electronic signature (FIEL), based on articles 7 BIS 2 of the Industrial Property Law; 3 of its Regulations, and 1 fraction III, 2 fraction V, 26 BIS and 26 TER of the Agreement establishing the guidelines for the use of the Electronic Payment and Services Portal (PASE) of the Mexican Institute of Industrial Property, in the procedures indicated.
THE DIVISIONAL DIRECTOR OF PATENTS
NAHANNY CANAL REYES
<img file="MX358306B_D0007.tif" />
Original string:
NAHANNY MARISOL CANAL REYES | 00001090000403252793 | Administration Service
Tax | 1695 || MX / 2018/68215 | MX / a / 2016/006880 | PCT patent title | 1220 | RRGO | Page (s) | awDTPzvjewFs7z / PNZvmCZbt / 3g =
Digital stamp:
pZn6 + 6zxludl91 / 8TvHpihOTvOVEYpSYxXTqlEwEgfr2tZjG1xHdvdGQWVAUdDnksOBU1hDOu6xD2AVqahNGYxcOBX SY3s1V2Sc8vnwhon2O / m8aX6EXu2mvUVSUjnprcGI6EVT2WIVJJjiuJ9E7ht7 + dSw2tlW5WLdJGbd8Qc1tcaaüyL5a Rzs274VtaciHuGStLlw0muKc¡Aixabf6S5WW8BW7uZ3B40zrXCx6LfpDo + kxSBI19XKO / IMJpYxQuRMeCikRR5WT6 / == nhZrvoJ4lzagfwhBiz4CD¡BmaFLoaS1iQ7YjlNTnaUpgBOdNMCq6YXg47tGa9WPCKI46LWTw
Areu-ií Nú VO. Γ · η; ο i. F <inu, Í1i, ii l epepart, Xochimilco. 16020. l (', 1 1- r ir “11 O. · I iir ·, 4 i / g.> R mx / impi
<img file="MX358306B_D0008.tif" />
<img file="MX358306B_D0009.tif" />
3583 bé
<img file="MX358306B_D0010.tif" />
IMPI
MEXICAN INSTITUTE ps.LAjaonEPAP.
DECODER, ENCODER AND METHOD FOR LAÍi
<img file="MX358306B_D0011.tif" />
INFORMED OF SOUNDING IN SYSTEMS OF
OBJECT BASED
FIELD OF THE INVENTION
The present invention relates to the encoding, processing and decoding of audio signals, and in particular to a decoder, an encoder, and a method for informed loudness estimation in object-based audio encoding systems.
BACKGROUND OF THE INVENTION
Recently, parametric techniques have been proposed for efficient bit rate transmission / storage of audio scenes comprising multiple object signals in the field of audio encoding [BCC, JSC, SAOC, SAOC1, SAOC2] and reported separation from sources [ISS1, ISS2, ISS3, ISS4, ISS5, ISS6]. These techniques are intended to reconstruct an output audio scene or a desired audio source object based on additional secondary information describing the transmitted / stored audio scene and / or the source objects in the audio scene. This reconstruction takes place in the decoder using an informed source separation scheme. Reconstructed objects can be combined
IMPI
MEXICAN INSTITUTE • OF PROPERTY for producí'P<sup>UST</sup>t '^ · é
<img file="MX358306B_D0012.tif" />
audio output. Depending on the lsr —memeiie -— in · which- ·· -o »the objects combine, the perceptual loudness of the output scene may vary.
In TV and radio transmission, the volume levels of the audio tracks of various programs can be normalized based on various aspects, such as the peak signal level or loudness level. Depending on the dynamic properties of the signals, two signals with the same peak level can have a perceived loudness level with large differences. Now switching the differences between programs or channels is very irritating in the loudness signal and has been a major source of complaints from end users in the broadcast.
In the prior art, it has been proposed to standardize all programs on all channels in a similar way at a common reference level using a measure based on the perceptual loudness of the signal. One such recommendation in Europe is EBU Recommendation R128 [EBU] (hereinafter referred to as R128).
The recommendation says that the loudness of the program, for example, the average loudness with respect to a program (or a commercial programming entity or
<img file="MX358306B_D0013.tif" />
<img file="MX358306B_D0014.tif" />
MEXICAN INSTITUTE OF PROPERTY (some other significant) must be equal<sup>STR</sup>to<sup>L </sup>specified (with small deviations ^ 'P> él'llll Llddij)'. 'CujhiIu' more and more stations comply with this recommendation and the required standardization, differences in average loudness between programs and channels should be minimized.
Loudness estimation can be done in several ways. There are several mathematical models to estimate the perceptual loudness of an audio signal. The EBU R128 recommendation is based on the model presented in ITU-R BS.1770 (hereinafter referred to as BS.1770) (see [ITU]) for loudness estimation.
As stated above, for example, in accordance with EBU Recommendation R128, the loudness of programs, for example, the average loudness over a program must be equal to a specified level with small deviations allowed. However, this leads to considerable problems in carrying out audio rendering, hitherto without solution in the prior art.
Performing the audio representation on the decoder side has a significant effect on the overall / overall loudness of the received audio input signal. However, even though scene rendering is performed, the total loudness of the received audio signal must be kept constant.
IMPI
<img file="MX358306B_D0015.tif" />
Nowadays,
MEXICAN INSTITUTE OF PROPERTY does not exist nincfCñY? L<sup>RIAL</sup>only specifies decoder side for piublomct tutu
<td>EP 2</td><td> 146 522</td><td>To the</td><td>([EP]),</td><td>it's related</td><td>with</td>
<td>concepts for</td><td>generate</td><td colspan="2">signals of</td><td>output of</td><td>Audio</td>
<td colspan="4">using object based metadata</td><td>. It is generated by</td><td>less</td>
<td>a sign of</td><td>departure</td><td>of</td><td colspan="2">audio representing</td><td>a</td>
<td>overlapping</td><td>at least</td><td>two</td><td>signs</td><td>of objects of</td><td>Audio</td>
<td>different but</td><td>Does not offer</td><td>a</td><td>solution</td><td colspan="2">for this problem.</td>
WO 2008/035275 A2 ([BRE]) describes an audio system comprising an encoder encoding audio objects in an encoding unit that generates a downmix audio signal and parametric data representing the plurality of audio objects. The downmixed audio signal and parametric data are transmitted to a decoder comprising a decoding unit that generates approximate replicas of the audio objects and a display unit that generates an output signal from the audio objects. The decoder further contains a processor to generate encoding modification data that is sent to the encoder. The encoder then modifies the encoding of the audio objects, and in particular modifies the parametric data in response to the encoding modification data. The procedure allows the manipulation of the audio objects that must be controlled by the
IMPI
MEXICAN INSTITUTE
<img file="MX358306B_D0016.tif" />
MEXICAN INSTITUTE
FROM PROPERTY decoder but which is done total or paer? B = r-po: r the encoder. In this way, the BU 11 ρτιτι clu is performed on real independent audio objects instead of on rough replicas, thus leading to improved performance.
EP 2 146 522 Al ([SCH]) describes an apparatus for generating at least one audio output signal representing an overlay of at least two different audio objects comprising a processor to process an audio input signal to provide a object representation of the audio input signal, where this object representation can be generated by a parametrically guided approximation of the original objects using an object downmix signal. An object manipulator individually manipulates objects using audio object-based metadata referring to individual audio objects to get manipulated audio objects. The manipulated audio objects are mixed using an object mixer to ultimately obtain an audio output signal with one or more channel signals, depending on the specific configuration for the rendering.
WO 2008/046531 Al ([ENG]) describes an audio object encoder for generating an object signal
<img file="MX358306B_D0017.tif" />
I
MEXICAN INSTITUTE OF PROPERTY, INDUSTRIAL coded using a plurality of objects
<img file="MX358306B_D0018.tif" />
audio including a downlink generator for generating mix information indicating a distribution of the plurality of downmixed audio objects on at least two downmix channels, an audio object parameter generator for generating object parameters for audio objects, and an output interface to generate the imported audio output signal using downmix information and object parameters. An audio synthesizer uses the downmix information to generate output data that can be used to create a plurality of output channels from the predetermined audio output configuration.
It would be desirable to have an accurate estimate of the average output loudness or change in the average loudness without delay, and when the program does not change or the rendering scene does not change, the average loudness estimate should also be kept static.
The object of the present invention is to provide improved concepts on encoding, processing and decoding of audio signals. The object of the present invention is solved by a decoder according to claim 1, by an encoder according to claim 15,
<img file="MX358306B_D0019.tif" />
by using by using
IMPI
<td>a</td><td>system</td><td>of</td><td>agreement</td><td>MEXICAN INSTITUTE A DE LA MONEDAD ÓúÜtí <17 INDUSTRIAL with claim 18,</td>
<td>a</td><td>method</td><td>of</td><td>agreement</td><td>with claim 19,</td>
<td>a</td><td>method</td><td>of</td><td>agreement</td><td>with claim 20 and</td>
<td>a</td><td colspan="2">Program</td><td colspan="2">computer according to the</td>
claim 21.
An informed way to estimate the loudness of the output is provided in an object-based audio coding system. The concepts provided are based on information about the loudness of the objects in the audio mix that will be provided to the decoder. The decoder uses this information together with the display information to estimate the loudness of the output signal. This then allows, for example, to estimate the loudness difference between the default downmix and the displayed output. Therefore, it is possible to compensate the difference to obtain an approximately constant loudness at the output regardless of the display information. The loudness estimation in the decoder takes place in a totally parametric manner and is very light, in computer terms, and accurate compared to the concepts of signal-based loudness estimation.
Concepts are provided to obtain information about the loudness of the specific output scene using purely parametric concepts, which
IMPI MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY la sonorida
<img file="MX358306B_D0020.tif" />
which then results in loudness estimation processing based on the explicit signal 'CTt the decoder. In addition, the specific spatial audio object encoding (SAOC) technology standardized by MPEG [SAOC] is described, although the concepts provided can also be used in combination with other audio object encoding technologies.
A decoder is provided to generate an audio output signal that comprises one or more audio output channels. The decoder comprises a receiving interface for receiving an audio input signal comprising a plurality of audio object signals, for receiving loudness information on the audio object signals, and for receiving display information indicating whether one or more of the audio object signals must be amplified or attenuated. Furthermore, the decoder comprises a signal processor to generate one or more audio output channels from the audio output signal. The signal processor is configured to determine a loudness compensation value depending on the loudness information and depending on the display information. Furthermore, the signal processor is configured to generate one or more audio output channels of the audio output signal from the
<img file="MX358306B_D0021.tif" />
IMPI
MEXICAN INSTITUTE OF PROPERTY
INDUSTRIAL __ audio input signal depending -on the representation information and depending on the loudness compensation value.
According to one embodiment, the signal processor can be configured to generate one or more audio output channels of the audio output signal from the audio input signal depending on the rendering information and depending on the offset value. loudness, so that the loudness of the audio output signal is equal to the loudness of the audio input signal, or in such a way that the loudness of the audio output signal is closer to the loudness of the audio input signal than the loudness of a modified audio signal that would be produced as a result of the modification of the audio input signal. audio by amplifying or attenuating the audio object signals of the audio input signal according to the rendering information.
According to another embodiment, each of the audio object signals of the audio input signal can be assigned exactly to a group of two or more groups, where each of the two or more groups can comprise one or more of the Audio object signals from the audio input signal. In that mode, the receiving interface can be configured to receive a value of
<img file="MX358306B_D0022.tif" />
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<img file="MX358306B_D0023.tif" />
loudness for each group of the two or more groups as loudness information, where such value of sonoriTacl<sup>1 </sup>indicates an original total loudness of one or more audio object signals in such a group. Furthermore, the receiving interface can be configured to receive the rendering information indicating, with respect to at least one group of the two or more groups, whether one or more audio object signals of such group should be amplified or attenuated by the indication of a modified total loudness of one or more audio object signals of such group. Furthermore, in that mode, the signal processor can be configured to determine the loudness compensation value depending on the modified total loudness of each of at least one group of two or more groups and depending on the original total loudness of each of two or more groups. Furthermore, the signal processor can be configured to generate one or more audio output channels of the audio output signal from the audio input signal depending on the total loudness modified from each of at least a group of two. or more groups and depending on the loudness compensation value.
In particular embodiments, at least one group of two or more groups may comprise two or more of the audio object signals.
In addition an encoder is provided.
The 'Ύϊν
MEXICAN " <sup>THE</sup>, í ''<sup>Or, EDA</sup>D INDUSTRIAL encoder comprises an object-based encoding unit for encoding a plurality of audio object signals to obtain an encoded audio signal comprising the plurality of audio object signals. Furthermore, the encoder comprises an object loudness encoding unit for encoding loudness information on the audio object signals. Loudness information comprises one or more loudness values, where each of one or more loudness values depend on one or more of the audio object signals.
According to one embodiment, each of the audio object signals of the encoded audio signal can be assigned exactly to a group of two or more groups, where each of the two or more groups comprises one or more of the signals audio object of the encoded audio signal. The loudness encoding unit of objects can be configured to determine one or more loudness values of the loudness information by determining a loudness value for each group of the two or more groups, wherein such loudness value of such group indicates an original total loudness of one or more audio object signals in such a group.
In addition, a system is provided. The system comprises an encoder according to one of the previously described modalities for encoding a plurality
<img file="MX358306B_D0024.tif" />
from audio object signals to encoded audio comprising of audio objects and to encode loudness information about the audio object signals. Furthermore, the system comprises a decoder according to one of the previously described modalities for generating an audio output signal that comprises one or more audio output channels. The decoder is configured to receive the encoded audio signal as the audio input signal and the loudness information. Furthermore, the decoder is configured to additionally receive rendering information. Furthermore, the decoder is configured to determine a loudness compensation value depending on the loudness information and depending on the display information. Furthermore, the decoder is configured to generate one or more audio output channels of the audio output signal from the audio input signal depending on the rendering information and depending on the loudness compensation value.
In addition, a method is provided for generating an audio output signal comprising one or more audio output channels. The method comprises:
Receive an audio input signal that comprises a plurality of audio object signals.
<img file="MX358306B_D0025.tif" />
Receive loudness information about audio object signals.
Receive rendering information indicating whether one or more of the audio object signals should be amplified or attenuated.
Determine a loudness compensation value depending on the loudness information and depending on the rendering information. AND:
Generate one or more audio output channels of the audio output signal from the audio input signal depending on the rendering information and depending on the loudness compensation value.
Furthermore, a coding method is provided. The method comprises:
Encoding an audio input signal that comprises a plurality of audio object signals. AND:
Encoding loudness information on audio object signals, where the loudness information comprises one or more loudness values, where each of one or more loudness values depends on one or more of the audio object signals.
Furthermore, a computer program is provided to implement the above-described method when running on a computer or signal processor.
IMPI
MEXICAN INSTITUTE *% · £ (? J
OWNERSHIP OR «*» ¿DÍá J? Í
INDUSTRIAL ^ a »^ SB-Tag-ji»
Preferred embodiments are provided in the dependent claims.
BRIEF DESCRIPTION OF THE FIGURES
The embodiments of the present invention are described in more detail below with reference to the figures, in which:
Figure 1 illustrates a decoder for generating an audio output signal that comprises one or more audio output channels according to one embodiment,
<td>The</td><td>Figure</td><td> 2</td><td>illustrates</td><td>a</td><td colspan="2">encoder</td><td>agree</td><td>with</td>
<td colspan="2">a modality,</td><td></td><td></td><td></td><td></td><td></td><td></td><td></td>
<td>The</td><td>Figure</td><td> 3</td><td>illustrates</td><td>a</td><td>system</td><td>of</td><td>agree with</td><td>a</td>
<td>modality,</td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td>
<td>The</td><td>Figure</td><td> 4</td><td>illustrates</td><td>a</td><td>system</td><td>of</td><td>Coding</td><td>of</td>
Space Audio Objects comprising a SAOC encoder and a SAOC decoder,
Figure 5 illustrates a SAOC decoder comprising a secondary information decoder, an object separator and a presenter,
Figure 6 illustrates the behavior of the output signal loudness estimates when faced with a change in loudness,
Figure 7 represents the reported loudness estimate according to one embodiment, illustrating the components of an encoder and a decoder according
IMPI
ICνΤΟ MEXICANO DE LA MOHEDAL INDUSTRIAL with a modality,
Figure 8 illustrates an encoder according to another embodiment,
Figure 9 illustrates an encoder and a decoder according to an embodiment relating to SAOC-dialog enhancement, comprising bypassed channels,
Figure 10 represents a first illustration of a measured loudness change and the result of using the concepts provided to estimate the loudness change parametrically,
<img file="MX358306B_D0026.tif" />
Figure 11 represents a second illustration
<td>of a measured loudness change and</td><td>the result</td><td>of the</td><td>use</td><td>of</td>
<td>the concepts provided for</td><td>estimate the</td><td colspan="2">change</td><td>of</td>
<td>loudness in a parametric way, and</td><td></td><td></td><td></td><td></td>
<td>Figure 12 illustrates</td><td colspan="2">another modality</td><td>of</td><td>the</td>
loudness compensation execution.
DETAILED DESCRIPTION OF THE INVENTION
Before describing the preferred modalities in detail, we describe loudness estimation, Spatial Audio Object Coding (SAOC), and dialog enhancement (DE).
First, the loudness estimation is described.
As already mentioned above, EBU R128 recommendation is based on the model presented
MEXICAN INSTITUTE 1
OF THE PROPERTY Cteígfefjfe
INDUSTRIAL in ITU-R BS.1770 for loudness estimation. This measurement will be used as an example, although the concepts described below can also be applied to other loudness measurements.
The loudness estimation operation according to BS.1770 is relatively simple and based on the following main steps [ITU]:
Input signal x is filtered<sub>{</sub> (or signals, in the case of the multi-channel signal) with a K filter (a combination of shelving and high-pass filters) to obtain the signal (or signals) and ¡.
The mean quadratic energy zj of the signal yj is calculated.
For multichannel signal, G-channel weighting is applied<sub>i</sub> and weighted signals are added. Next, the signal loudness is defined as follows: = c + 101og<sub>10</sub>^ G, .z,., I
with the constant value c = -0.691. Then the output is expressed in the LKFS units (Loudness, weighted with K, relative to Full Scale) that is scaled similarly to the decibel scale.
In the preceding formula, Gi can be, for example, equal to 1 in some of the channels, while
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
Gi may be, for example, 1.41 for some other channels. For example, if a left channel, a right channel, a center channel, a left surround channel and a right surround channel are contemplated, the respective Gi weights can be, for example, 1 in the case of the left, right and center channel and it can be, for example, 1.41 in the case of the left surround channel and the right surround channel, see [ITU].
<img file="MX358306B_D0027.tif" />
It can be seen that the loudness value L is
<td>closely related to</td><td>the logarithm</td><td>energy</td><td>of</td>
<td>the signal.</td><td></td><td></td><td></td>
<td>Below is</td><td>describes the</td><td>Coding</td><td>of</td>
<td>Space Audio Objects.</td><td></td><td></td><td></td>
<td>Coding</td><td colspan="2">object-based audio</td><td>gives</td>
<td>place to great flexibility</td><td>on the side of the</td><td>decoder</td><td>of</td>
chain. An example of an object-based audio coding concept is Spatial Audio Object Coding (SAOC).
Figure 4 illustrates a Spatial Audio Object Coding (SAOC) system comprising a 410 SAOC encoder and a 420 SAOC decoder.
SAOC Encoder 410 receives N audio object signals Yes, S<sub>N</sub> as input. In addition, the 410 SAOC encoder also receives Mix D Information instructions on how these should be combined.
IMPI Mexican institute OF INDUSTRIAL PROPERTY objects to obtain a downmix signal that
<img file="MX358306B_D0028.tif" />
comprises M downmix channels Xi, X<sub>M</sub>. SAOC encoder 410 extracts some secondary information from the objects and the downmix process, and this secondary information is transmitted and / or stored along with the downmix signals.
A fundamental property of the SAOC system is that the downmix signal X comprising the downmix channels Xi, Xm forms a semantically significant signal. In other words, it is possible to hear the downmix signal. If, for example, the receiver does not have the SAOC decoder functionality, the receiver can still output the downmix signal as output.
Figure 5 illustrates a SAOC decoder comprising a secondary information decoder 510, an object separator 520 and a presenter 530. The SAOC decoder illustrated in Figure 5 receives, for example, from a SAOC encoder, the downmix signal and secondary information. The downmix signal can be considered an audio input signal that comprises the audio object signals, since the audio object signals are mixed within the downmix signal (the audio object signals are mixed within one or more mixing channels
The SAOC decoder can try, for
IMPI MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<img file="MX358306B_D0029.tif" />
example, rebuild (virtually) the original objects, for example, using object separator 520, for example, using the
In these example reconstructions, the signals are then combined with object information based representation, eg, a decoded secondary.
of of in objects, by reconstructed audio, the representation matrix information
R, to produce K output channels of an Y audio output signal.
In information signals covariance audio encoder Vi, ·· '/ ík of SAOC, audio objects are often reconstructed, for example, using de
SAOC covariance, eg a signal E, which is transmitted to the SAOC decoder.
matrix from him
For example, the following formula can be used to reconstruct the decoder-side audio object signals:
where
Our samples
S = GX number number signal of audio objects of audio object signals, of samples considered of a
Μ number of channels
<img file="MX358306B_D0030.tif" />
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL FRORITY of
<img file="MX358306B_D0031.tif" />
downstream, downmixed audio signal, size MX Ours, downmix matrix, size M signal covariance matrix, size N x N defined by
E = XX «parametrically reconstructed audio object signals, size
NX ^ Self-attached samples (Hermitian) representing the conjugate transpose of
Then, a representation matrix R can be applied to the reconstructed audio object signals S to obtain the audio output channels of the audio output signal Y, for example, according to the formula:
<td></td><td>in</td><td colspan="2">where</td>
<td></td><td>K</td><td>number</td><td>of</td>
<td>Yi audio,</td><td> ... <sub>r</sub></td><td>signal ϊκ</td><td>of</td>
<td></td><td>R</td><td>matrix</td><td>of</td>
<td>x N</td><td></td><td></td><td></td>
<td></td><td>AND</td><td colspan="2">exit sign</td>
audio output channels,
Y = RS the Y audio output channels.
representation of an audio K size comprising the K
<img file="MX358306B_D0032.tif" />
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY size KX ^ Samples
In Figure 5, reference is made to the object reconstruction process, for example, carried out the object separator 520, with the virtual or optional notion, since it may not be essential that it necessarily occur, but rather that the Desired functionality by combining the reconstruction and representation steps in the parametric domain (i.e. combining the equations).
In other words, instead of reconstructing the audio object signals using the mix information D and the covariance information E first, and then applying the representation information R to the reconstructed audio object signals to obtain the channels Yi, ..., Yk audio output channels, both steps can be performed in one step, so that the Yi, Yk audio output channels are generated directly from the downmix channels.
For example, the following formula can be used:
can
<td>Y = RGX</td><td>with G «ED<sup>H</sup></td><td>(D</td><td>ED<sup>H</sup>) -<sup>1</sup> .</td>
<td>In principle, the</td><td>information</td><td>of</td><td>representation R</td>
<td>sue any</td><td>combination</td><td>of</td><td>the signs of</td>
original audio objects. In practice, however, object reconstructions can involve errors
IΜ ΡI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY reconstruction and the requested exit scene may not necessarily be achieved. As a rough rule of thumb that covers many case studies, the more the requested output scene differs from the downmix signal, the more audible reconstruction errors there will be.
Dialog enhancement (DE) is described below. For example, SAOC technology can be used to achieve the scenario. It should be noted that although the name Dialog Enhancement suggests a primary interest in dialogue-oriented cues, the same principle can also be used with other types of cues.
In the DE scenario, the degrees of freedom of the system are limited with respect to the general case.
For example, audio object signals = S are grouped together (and possibly mixed together) forming two meta objects of a foreground object (FGO) S<sub>/ w</sub> and a background object (BGO) S<sub>BC0</sub>.
Also, the exit scene y¡, ..., Y<sub>K</sub> = Y resembles the downmix signal X<sub>x</sub>, ..., X<sub>M</sub> = X. More specifically, both signals have the same dimensionalities, i.e. KM, and the end user can only control the relative mix levels of the two FGO and BGO meta-objects. To be more precise, the downmix signal is obtained by mixing the FGO and BGO with some scale weights
<img file="MX358306B_D0033.tif" />
γ _ z. and<sup>Λ</sup> 'fcofgo
IMPI
MEXICAN INSTITUTE OF LA MOHEDA!), Z. C INDUSTRIAL <sup>n</sup>BCO<sup>;and</sup>'BGO'
<img file="MX358306B_D0034.tif" />
and the exit scene is obtained in a similar way with some scale weighting of the FGO and the BGO:
+ Sbg <^
Depending on the relative values of the mix weights, the balance between the FGO and the BGO may change. For example, with the configuration
Sfgo <sup>></sup> h-pco .Sbgo “^ bgo it is possible to increase the relative level of FGO in the mix. If FGO is dialog, this setting provides dialog enhancement functionality.
As a case study example, the BGO may consist of stadium noise and other background noise during a sporting event and the FGO is the voice of the commentator. The DE functionality allows the end user to amplify or attenuate the commentator level relative to the background.
The modalities are based on the finding that the use of SAOC technology (or similar) in a transmission strategy allows the end user to be offered a
IMPIOS
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY increased signal manipulation functionality. More functionality is offered than just changing the channel and adjusting the playback volume.
In the above, a possibility of using DE technology has been briefly described. If the transmit signal level, which is the downmix signal for SAOC, is normalized, for example, according to R128, the different programs have a similar average loudness when processing (SAOC) is not applied (or the description representation equals the description of the downmix). However, when a certain degree of processing (SAOC) is applied, the output signal differs from the default downmix signal and the loudness of the output signal may be different from the loudness of the default downmix signal. From the end user point of view, this can lead to a situation where the loudness of the output signal between channels or programs may, once again, have undesirable jumps or differences. In other words, the benefits of the standardization applied by the issuer are partially lost.
This issue is not specific to SAOC or the DE scenario only, but may also appear with other audio encoding concepts that allow the end user to interact with the content. Nevertheless, *
<img file="MX358306B_D0035.tif" />
<img file="MX358306B_D0036.tif" />
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY in many cases does not cause any harm if the output signal has a different loudness from the default downmix.
As stated above, the total loudness of an audio input signal program must be equal to a specified level with small deviations allowed. However, as already mentioned, this leads to significant problems in performing audio rendering, since rendering can have a significant effect on the overall / overall loudness of the received audio input signal. However, despite the scene representation being performed, the total loudness of the received audio signal must remain the same.
One strategy would be to estimate the loudness of a signal while it is playing, and with an appropriate concept of temporal integration, the estimate may converge to the true average loudness over time. However, the time required for convergence is problematic from the end user's point of view. When the loudness estimate changes even if no changes are applied to the signal, the loudness change compensation should also react and change its behavior. This would lead to an output signal with a temporarily variable average loudness, which can be perceived as somewhat irritating.
<img file="MX358306B_D0037.tif" />
Figure 6 illustrates the behavior of loudness estimates of the. exit signal before a change in loudness. Among other things, an estimate of the loudness of the output signal is represented based on the signal, illustrating the effect of a solution like the one just described. The estimate approaches the correct estimate rather slowly. Instead of a signal-based estimate of the output signal loudness, an informed estimate of the output signal loudness would be preferable, which immediately correctly determines the loudness of the output signal.
In particular, in Figure 6, the user input, for example, the level of the dialog object changes at time T by the increment of value. The true level of the output signal and therefore the loudness will change in the same instant of time. When the loudness estimate of the output signal is made from the output signal with a certain time integration time, the estimate gradually changes and reaches the correct value after a certain delay. During this delay, the estimate values are changing and cannot be used reliably for further processing of the output signal, for example for correction of loudness level.
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<img file="MX358306B_D0038.tif" />
As already mentioned average loudness of output it would be desirable to have one or the average loudness change without delay and when the program does not change or change the scene of the representation, the loudness estimate should remain static as well. In other words, when applying some loudness change compensation, the compensation parameter should change only when the program changes or when there is some interaction with the user.
The desired behavior is illustrated in the lower diagram of Figure 6 (reported estimation of the loudness of the output signal). The loudness estimate of the output signal should change immediately when the user input changes.
Figure 2 illustrates an encoder according to one embodiment.
The encoder comprises an object-based encoding unit 21Ü for encoding a plurality of audio object signals to obtain an encoded audio signal comprising the plurality of audio object signals.
Furthermore, the encoder comprises an object loudness encoding unit 220 for encoding loudness information on the audio object signals. Loudness information comprises one or more loudness values, where loudness values depend on
A IVAJT 1
MEXICAN INSTITUTE OF LA MONEDAD
INDUSTRIAL
<img file="MX358306B_D0039.tif" />
plus each of one one or more of the audio object signals.
According to one embodiment, each of the audio object signals of the encoded audio signal is exactly assigned to a group of two or more groups, where each of the two or more groups comprises one or more of the signals audio object of the encoded audio signal. The object loudness encoding unit 220 is configured to determine one or more loudness values of the loudness information by determining a loudness value for each group of the two or more groups, wherein such loudness value of such group indicates an original total loudness of one or more audio object signals from that group.
Figure 1 illustrates a decoder for generating an audio output signal that comprises one or more audio output channels according to one embodiment.
The decoder comprises a receiving interface 110 for receiving an audio input signal comprising a plurality of audio object signals, for receiving loudness information on the audio object signals, and for receiving display information indicating whether a or more of the audio object signals must be amplified or attenuated.
IMPI
MEXICAN INSTITUTE OF THE «INDUSTRY OPIITY!
In addition, the decoder comprises a processor
<img file="MX358306B_D0040.tif" />
/ 120 signal to generate one or more audio output channels of the audio output signal. Signal processor 120 is configured to determine a representation loudness compensation value.
configure for loudness depending on the information depending
Also, generating one or more of the processor channels information
120 of the audio output signal of the audio output signal audio input depending representation and depending loudness.
According to one of the signals it is configured for from the mode value, generating one of the signal information compensation processor of from
110 more channels of audio output from the audio output signal from that of the input signal representation and loudness, of such audio output being audio input, or signal of the audio signal of audio depending on the information depending on the compensation value in the same way of such audio output is of modified input modification of which the loudness of the signal to the loudness so that it approximates more than that of the Signal Loudness of Audio Loudness The loudness of a signal that would be produced as a result of audio input signal by amplifying or attenuating the signals of audio objects from the signal audio input
IMPI MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY according
<img file="MX358306B_D0041.tif" />
with representation information.
According to another embodiment, each of the audio object signals of the audio input signal is assigned exactly to a group of two or more groups, where each of the two or more groups comprises one or more of the Audio object signals from the audio input signal.
In that embodiment, the receiving interface 110 is configured to receive a loudness value for each group of the two or more groups as loudness information, wherein such loudness value indicates an original total loudness of one or more signal objects. audio of such a group. Furthermore, the receiving interface 110 is configured to receive the rendering information indicating, for at least one group of the two or more groups, whether one or more audio object signals of such group should be amplified or attenuated by indication of a modified total loudness of one or more audio object signals of such group. Furthermore, in that mode, the signal processor 120 is configured to determine the loudness compensation value depending on the modified total loudness of each of at least one group of two or more groups and depending on the original total loudness of each of two or more groups. Furthermore, the
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<img file="MX358306B_D0042.tif" />
Signal processor 120 plus output channels is configured to generate one or more audio from the audio output signal from the audio input signal depending on the total loudness modified from each of at least a group of two or more groups and depending on the loudness compensation value.
In particular embodiments, at least one group of two or more groups comprises two or more of the audio object signals.
There is a direct relationship between the energy ei of an audio object signal i and the loudness Li of the audio object signal i according to the formulas:
L¡ = c +10 log<sub>)0</sub> and<sub>¡</sub>, ς = 10<sup>(TO</sup><sup>c) / 1</sup>° where c is a constant value.
The modes are based on the following findings: Different audio object signals from the audio input signal may have different loudness, and therefore different energy. If, for example, a user wishes to increase the loudness of one of the audio object signals, it can be adjusted correspondingly to the rendering information, and increasing the loudness of this audio object signal increases the energy of that audio object. This would result in increased loudness of the audio output signal. To keep the total loudness constant, it is necessary to
MEXICAN INSTITUTE Dt THE INDUSTRIAL PROPERTY that carry out a loudness compensation.
<img file="MX358306B_D0043.tif" />
In other words, the modified audio signal that it would generate as a result of the application of the rendering information on the audio input signal should be adjusted. However, the exact effect of the amplification of one of the audio object signals on the total loudness of the audio signal depends on the original loudness of the amplified audio object signal, for example, the obj signal audio etos, the loudness of which is increased. If the original loudness of this object corresponded to a rather low energy, the effect on the total loudness of the audio input signal will be less. If, on the other hand, the original loudness of this object corresponds to a rather high energy, the effect on the total loudness of the audio input signal will be significant.
Two examples can be considered. In both examples, an audio input signal comprises two audio object signals, and in both the representation information, is powered from the first of the example signals, applying increases the audio objects 50%.
In the first example, the first audio object signal contributes 20% and the second audio object signal contributes 80% to the total power of the audio input signal. However, in the second example, the
IMPI
INSTITUTO MEXICANO of the industrial promise first audio object, the first audio object signal contributes 40% and the second audio object signal contributes 60% to the total energy of the audio input signal. In both examples these contributions are derived from the loudness information on the audio object signals, since there is a direct relationship between loudness and energy.
In the first example, a 50% increase in the energy of the first audio object results in a modified audio signal that is generated by applying the rendering information on the audio input signal having a total energy of 1.5 x 20% + 80% = 110% of the audio input signal energy.
In the second example, a 50% increase in the energy of the first audio object results in the modified audio signal being generated by the application of the rendering information on the audio input signal having a total energy of 1.5 x 40% + 60% = 120% of the audio input signal energy.
Therefore, after applying the rendering information on the audio input signal, in the first example, you only have to reduce the total energy of the modified audio signal 9% (10/110) to obtain the same energy both in the audio input signal as well as the audio output signal, while
<img file="MX358306B_D0044.tif" />
β
<img file="MX358306B_D0045.tif" />
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<img file="MX358306B_D0046.tif" />
the second example the total energy of the modified audio signal has to be reduced 17% (20/120). For this purpose, a loudness compensation value can be calculated.
For example, the loudness compensation value may be a scalar value that is applied to all audio output channels of the audio output signal.
According to one embodiment, the signal processor is configured to generate the modified audio signal by modifying the audio input signal by amplifying or attenuating the audio object signals of the audio input signal accordingly. with representation information. Furthermore, the signal processor is configured to generate the audio output signal by applying the loudness compensation value to the modified audio signal so that the loudness of the audio output signal is equal to the loudness of the audio input signal, or in such a way that the loudness of the audio output signal is closer to the loudness of the audio input signal than the loudness of the modified audio signal.
<td></td><td>For example,</td><td colspan="3">In the first example above, you can</td>
<td>set</td><td>the value of</td><td>compensation</td><td>of</td><td>loudness lev, for</td>
<td>example,</td><td>in a value</td><td>lev = 11/10,</td><td>and</td><td>a</td>
IMPI
INSTITUTO MEXICANO DE LA MONEDAD INDUSTRIAL multiplication factor of 10/11 to all channels resulting from the representation of the audio input channels according to the representation information.
Therefore, for example, in the second example above, the loudness compensation value lev can be set, for example, to a value lev = 10/12 = 5/6, and a multiplication factor of 5/6 can be applied to all channels resulting from the representation of the audio input channels according to the representation information.
In other embodiments, each of the audio object signals may be assigned to one of a plurality of groups, and a loudness value may be transmitted for each of the groups indicating a total loudness value of the audio object signals. of such a group. If the representation information specifies that the energy of one of the groups is attenuated or amplified, for example, 50% amplified as in the previous case, an increase in the total energy can be calculated and a loudness compensation value can be determined as described in the above.
<img file="MX358306B_D0047.tif" />
<td></td><td colspan="2">For example,</td><td colspan="2">according to a</td><td colspan="3">modality, each</td>
<td>one of</td><td>the</td><td>signals of</td><td>objets of</td><td>Audio</td><td>of the</td><td>signal</td><td>of</td>
<td>entry</td><td>of</td><td>audio se</td><td colspan="2">assign exactly</td><td>yet</td><td>group</td><td>of</td>
<td colspan="2">exactly</td><td>two groups</td><td>like the two</td><td>or more</td><td>groups.</td><td>Every</td><td>a</td>
IMPI
The INSTITUTO MEXICANO DE LA PROPIEDAD INDUSTRIAL of the audio object signals of the audio signal is assigned to a group of objects in the foreground of exactly the two groups or to a group of objects in the background of exactly the two groups. The receiving interface 110 is configured to receive the original full loudness of one or more audio object signals from the foreground object group. In addition, the receiving interface 110 is configured to receive the original full loudness of one or more audio object signals from the background object group. Furthermore, the receiving interface 110 is configured to receive the rendering information indicating, for at least one group of exactly two groups, whether one or more audio object signals from each of the at least one group should be amplified or attenuated by indicating a modified total loudness of one or more audio object signals of such a group.
In that mode, the signal processor 120 is configured to determine the loudness compensation value depending on the modified total loudness of each of at least one group, depending on the original total loudness of one or more audio object signals of the group of objects in the foreground, and depending on the original loudness of one or more signals of audio objects in the group of objects in the background. Furthermore, the
<img file="MX358306B_D0048.tif" />
IMPI
INSTITUTO MEXICANO DE LA PROPIEDAD INDUSTRIAL signal processor 120 is configured to generate one or more channels of audio output from the audio output signal from the audio input signal depending on the total modified loudness of each of at least a group and depending on the loudness compensation value.
According to some modes, each of the audio object signals is assigned to one of three or more groups, and the receiving interface can be configured to receive a loudness value for each of the three or more groups, indicating the total loudness of the audio object signals of such a group.
According to one embodiment, to determine the total loudness value of two or more of the audio object signals, for example, the energy value for the loudness value for each audio object signal is determined, the energy values of all loudness values to obtain an energy sum and the loudness value corresponding to the sum of energies is determined as the total loudness value of the two or more of the audio object signals. For example, the following formulas can be used
L<sub>t</sub>= c + 101og<sub>10</sub> and<sub>(</sub>,
<img file="MX358306B_D0049.tif" />
<img file="MX358306B_D0050.tif" />
In some modes, the loudness values corresponding to each of the audio object signals, or each of the
IMPI
MEXICAN INSTITUTE OF PROPERTY
INDUSTRIAL
<img file="MX358306B_D0051.tif" />
Audio object signals are assigned to one or two or more groups, where for each of the groups, a loudness value is transmitted.
However, in some embodiments, for one or more audio object signals or for one or more of the groups comprising audio object signals, no loudness value is transmitted. Rather, the decoder, for example, may assume that these audio object signals or groups of audio object signals, for which no loudness value is transmitted, have a predefined loudness value.
The decoder can base all subsequent determinations, for example, on this predefined loudness value.
According to one modality, the interface
110 Receive signal is configured to receive a downmix signal comprising one or more downmix channels as the audio input signal, wherein one or more downmix channels comprise the audio object signals, and wherein the number of Audio object signals is less than the number of the one or more downmix channels. The receive interface 110 is configured to receive downmix information indicating how the audio object signals are mixed within one or more downmix channels.
IMPI 'NSTnVTOMSXKUNO OF PROPERTY. INDUSTRIAL
Also, the signal processor 120 is configured to
<img file="MX358306B_D0052.tif" />
generating one or more audio output channels of the audio output signal from the audio input signal depending on the downmix information, depending on the rendering information and depending on the loudness compensation value. In a particular embodiment, processor 120 is configured, for example, to calculate loudness compensation depending on the downmix information value.
For example, the downmix information may be a downmix matrix. In embodiments, the decoder may be a SAOC decoder. In such embodiments, the receiving interface 110 may further be configured, for example, to receive covariance information, for example, a covariance matrix as described above.
With respect to the rendering information indicating whether one or more of the audio object signals should be amplified or attenuated, it should be noted that, for example, the information indicating how one or more of the amplified signals or grayed out, is For example, an array of audio objects must be rendering information, R rendering, eg a SAOC rendering array, is rendering information.
<img file="MX358306B_D0053.tif" />
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<img file="MX358306B_D0054.tif" />
Figure 3 illustrates a system according to one embodiment.
The system comprises an encoder 310 according to one of the above-described embodiments for encoding a plurality of audio object signals to obtain an encoded audio signal comprising the plurality of audio object signals.
Furthermore, the system comprises a decoder 320 according to one of the previously described modalities for generating an audio output signal comprising one or more audio output channels. The decoder is configured to receive the encoded audio signal as the audio input signal and the loudness information. Furthermore, decoder 320 is configured to also receive rendering information. Furthermore, decoder 320 is configured to determine a loudness compensation value depending on the loudness information and depending on the display information. Furthermore, decoder 320 is configured to generate one or more audio output channels of the audio output signal from the audio input signal depending on the rendering information and depending on the loudness compensation value.
Figure 7 illustrates the loudness estimate
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY informed in accordance with a modality. To the left of transport stream 730, the components of an encoder for object-based audio encoding are illustrated. In particular, an object-based encoding unit 710 (object-based audio encoder) and an object loudness encoding unit 720 (object loudness estimation) are illustrated.
Transport stream 730 itself comprises loudness information L, downmix information D and the output of object-based audio encoder 710 B.
To the right of transport stream 730, the components of a signal processor of a decoder for object-based audio encoding are illustrated. Decoder receiving interface is not illustrated. An output loudness estimator 740 and an object-based audio decoding unit 750 are depicted. The output loudness estimator 740 can be configured to determine the loudness compensation value. The object-based audio decoding unit 750 can be configured to determine a modified audio signal from an audio signal, which is input to the decoder, by applying the rendering information R. The illustration is not illustrated.
<img file="MX358306B_D0055.tif" />
<img file="MX358306B_D0056.tif" />
IMPI
INSTITUTO MEXICANO DE LA PROPIEDAD INDUSTRIAL applying the loudness compensation value to the modified audio signal to compensate for a total loudness change caused by the representation in Figure 7.
The input to the encoder consists of the minimum S input objects. The system estimates the loudness of each object (or some other loudness-related information, such as the energies of the objects), for example, by the loudness encoding unit 720 of objects, and this information L is transmitted and / or stored. (It is also possible that the loudness of the objects is provided as an input to the system, and the estimation step within the system may be allowed.)
In the embodiment of Figure 7, the decoder receives at least the loudness information of the objects and, for example, the rendering information R that describes the mixing of the objects in the output signal. Based on these, for example, the output loudness estimator 740 estimates the loudness of the output signal and outputs this information as its output.
The downmix information D may be provided as display information, in which case the loudness estimate provides an estimate of the loudness of the downmix signal. It is also possible to provide descendant as input to objects to transmit the estimation information of the
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL MONEDAD
<img file="MX358306B_D0057.tif" />
Mix information loudness estimate and / or store it together with the loudness of the objects. After the output loudness you can simultaneously estimate the loudness of the downmix signal and the displayed output and provide these two values or their difference as output loudness information. The difference value (or its inverse) describes the necessary compensation that should be applied to the displayed output signal to make its loudness similar to the loudness of the downmix signal. The loudness information of the objects may further include information regarding the correlation coefficients between various objects and this correlation can be used in the output loudness estimation for a more accurate estimate.
Next, a preferred modality for dialogue enhancement is described.
In the dialog enhancement application, as described above, input audio object signals are grouped and partial downmixing is performed to form two meta-objects, FGO and BGO, which can then be easily added together to determine the mix signal descending end.
After the description of
IMPI
MEXICAN INSTRUMENT OF INDUSTRIAL PROPERTY
<img file="MX358306B_D0058.tif" />
SAOC [SAOC], N input object signals are represented in the form of matrix S of size N x NMamples, and the downmix information as matrix D of size Μ x N. Then the downmix signals can be obtained as X = DS.
The downmix information D can now be divided into two parts for the meta-objects.
Since each column of matrix D corresponds to an original audio object signal, the two component downmix matrices can be obtained by setting the columns, which correspond to the other meta-object to zero (assuming that no original object can be present in both meta-objects).
In other words, the columns corresponding to the BGO meta-object are set to zero in D<sub>convict</sub>, and vice versa.
These new downmix matrices describe how two meta objects can be obtained from the input objects, i.e .:
and
ΙΜΡΙ mexican institute
OF INDUSTRIAL PROPERTY
<img file="MX358306B_D0059.tif" />
and the actual downmix is simpli f ió3 “3 ·
The object decoder (eg SAOC) can also be seen as trying to reconstruct the meta-objects:
& o ~ Q <sub>v</sub> & o _ e <sup>J</sup>FGO <sup>J</sup> FGO Y <sup>J</sup> BGO ~ <sup>J</sup> BGO · and the specific representation of DE can be expressed as a combination of these two meta-object reconstructions:
<img file="MX358306B_D0060.tif" />
The loudness estimation of objects receives the two meta-objects S<sub>TO0</sub> and S<sub>BG0</sub> as input and estimates the loudness of each of them: which is the loudness (total / general) of S<sub>TO0</sub>, and L<sub>BG0</sub> which is the loudness (total / general) of S<sub>RC0</sub> · These loudness values are transmitted and / or stored.
As an alternative, using one of the meta-objects, for example the FGO, as a reference, it is possible to calculate the difference in loudness of these two objects, for example, as
<img file="MX358306B_D0061.tif" />
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<img file="MX358306B_D0062.tif" />
^ FGO L<sub>BC</sub>or L<sub>fco</sub> .
This unique value is transmitted and / or stored.
Figure 8 illustrates an encoder according to another embodiment. The encoder of Figure 8 comprises a downstream object mixer 811 and an estimator 812 of secondary object information. Furthermore, the encoder of Figure 8 comprises an object loudness encoding unit 820. Furthermore, the encoder of Figure 8 comprises an audio meta-object mixer 805.
The encoder in Figure 8 uses intermediate audio meta-objects as input to the loudness estimation of objects. In modalities, the encoder of Figure 8 can be configured to generate two audio meta-objects. In other embodiments, the encoder of Figure 8 can be configured to generate three or more audio meta-objects.
Among other things, the concepts provided provide a new feature, and that is that the encoder can, for example, estimate the average loudness of all input objects. Objects can be mixed, for example, to obtain a downmix signal that
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY is transmitted. The concepts provided further provide the new feature that for example object loudness and downmix information can be included in the transmitted object encoding secondary information.
<img file="MX358306B_D0063.tif" />
<td>The</td><td colspan="2">decoder</td><td>you can use,</td><td>for example,</td><td>the</td>
<td>information</td><td>high school</td><td>of</td><td>coding</td><td>objects for</td><td>the</td>
<td>separation</td><td>(virtual)</td><td>of</td><td>the objects and</td><td>recombine</td><td>the</td>
objects using representation information.
In addition, the concepts provided provide the new feature that downmix information can be used to estimate the loudness of the default downmix signal, rendering information, and received object loudness to estimate the average loudness of the downlink signal. output, and / or the change in loudness of these two values can be estimated. 0, downmix and proxy information can be used to estimate the change in loudness of the default downmix, another new feature of the concepts provided.
In addition, the concepts provided provide the new feature that the decoder output can be modified to compensate for the change in loudness so that the average loudness of the modified signal matches the average loudness of the downmix.
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<img file="MX358306B_D0064.tif" />
default.
A particular modality related to SAOC-DE is illustrated in Figure 9. The system receives the input audio object signals, the downmix information, and the information about grouping the objects into meta objects. Based on these, the audio meta-object mixer 905 forms the two meta-objects S<sub>convict</sub> and S<sub>BGO</sub>. The portion of the signal that is processed with SAOC may not constitute the entire signal. For example, in a 5.1 channel configuration, it can be implemented
SAOC in a sub-series of channels, as in the front right channel and while the other channels (left surround, right surround and drop effects are directed to circumscribe (passing by presented as such. These channels not processed by
SAOC are indicated as X<sub>BYPASS</sub>.
It is necessary to provide possible bypassed channels to the encoder for a more accurate estimate of the loudness information.
Bypassed channels can be treated in various ways.
For example, the canceled channels can form, for example, a separate meta-object. This allows the representation to be defined so that the three
IMPI, INSTITUTO MEXICANO DE LA MONEDAD INDUSTRIAL meta-objects are scaled independently.
<img file="MX358306B_D0065.tif" />
Or, for example, channels can be combined
<td>canceled,</td><td>for example with</td><td>one</td><td>of the</td><td>others</td><td>two</td>
<td>metaobjects</td><td>. The configurations</td><td>of</td><td>It represents</td><td>tion of</td><td>that</td>
<td>metaobj eto</td><td>they also control</td><td>the</td><td>portion</td><td colspan="2">of channels</td>
<td>canceled.</td><td>For example, in the</td><td colspan="2">scenario of</td><td>improvement</td><td>of</td>
Dialogues, it may be advantageous to combine the bypassed channels with the background metaobject: X<sub>BG0</sub> - S<sub>BG0</sub><sup>+</sup> 'X-bvpass ·
0, for example, bypassed channels, for example can be ignored
In accordance with some embodiments, the encoder object-based encoding unit 210 is configured to receive the audio object signals, where each of the audio object signals is assigned to exactly one of exactly two groups, where each of the exactly two groups comprises one or more of the audio object signals. Furthermore, the object-based encoding unit 210 is configured to downmix the audio object signals, which are made up of exactly the two groups, to obtain a downmix signal comprising one or more audio channels of downmix as an encoded audio signal, where the number of the one or more downmix channels is less than the number of audio object signals that are made up of
<img file="MX358306B_D0066.tif" />
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL MONEDAD exactly the two groups. The object loudness encoding unit 220 is indicated to receive one or more additional bypass audio object signals, wherein each of one or more additional bypass audio object signals is assigned to a third group, in where each of one or more additional bypass audio object signals is not comprised in the first group and is not comprised in the second group, wherein the object-based encoding unit 210 is configured not to downmix one or more additional bypass audio object signals within the downmix signal.
In one embodiment, the object loudness encoding unit 220 is configured to determine a first loudness value, a second loudness value, and a third loudness value from the loudness information, the first loudness value indicating total loudness. of one or more audio object signals from the first group, the second loudness value indicating a total loudness of one or more audio object signals from the second group and the third loudness value indicating a total loudness of one or more additional bypass audio object signals from the third group. In another embodiment, the object loudness encoding unit 220 is configured to determine a first loudness value and a second value of loudness.
IMPI msTmrro Mexican of LA PROPCDAD INDUSTRIAL sonoridad de la in
<img file="MX358306B_D0067.tif" />
loudness, where the first loudness value indicates a total loudness of one or more audio object signals from the first group, and where the second loudness value indicates a total loudness of one or more audio object signals from the second group group and one or more additional bypass audio object signals from the third group.
In accordance with one embodiment, the decoder receiving interface 110 is configured to receive the downmix signal. Furthermore, the receive interface 110 is configured to receive one or more additional bypass audio object signals, where one or more additional bypass audio object signals are not mixed within the downmix signal. Furthermore, the receiving interface 110 is configured to receive loudness information indicating information about the loudness of the audio object signals being mixed into the downmix signal and indicating information about the loudness of one or more signal signals. Additional bypass audio objects that are not mixed in the downmix signal.
Furthermore, the signal processor 120 is configured to determine the loudness compensation value depending on the information about the loudness of the audio object signals that are mixed into the signal depending on the information about
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY descending mixes the sound of
<img file="MX358306B_D0068.tif" />
plus additional bypass audio object signals that are not mixed in the downmix signal.
Figure 9 illustrates an encoder and a decoder according to an SAOC-DE related embodiment, comprising bypassed channels. Among other things, the encoder of Figure 9 comprises a 902 SAOC encoder.
In the embodiment of Figure 9, the possible combination of the canceled channels with the other meta-objects takes place in the two cancellation inclusion blocks 913, 914, which produce the meta-objects X<sub>FGO</sub> and
X<sub>gG0</sub> where the defined parts of the voided channels are included.
The perceptual loudness L<sub>BYPAS</sub>s' ^ fgo> Y ^ bgo of these two meta-objects is estimated in loudness estimation units 921, 922, 923. This loudness information is then transformed into appropriate encoding in a meta-loudness loudness information estimator 925 and then transmitted and / or stored.
The actual SAOC encoder and decoder operate properly by extracting the secondary object information from the objects, creating the mix signal
ΙΜΡΙ @ Β ^
MEXICAN INSTITUTE OF PROPERTY descending X, and storing and / or tran ^^ ieMT ^ u information to the decoder. The possible T Cáftá'léS 'á'Ul'dUóü are transmitted and / or stored together with the other information in the decoder.
The 945 SAOC-DE decoder receives a Dialog gain gain value as user input. Based on this input and the received downmix information, the decoder 945 SAOC determines the rendering information. The 945 SAOC decoder then produces the output scene presented as the Y signal. In addition to this, it produces a gain factor (and a delay value) that would be applied to the possible null signals Y ^ BYPASS '
The override inclusion unit 955 receives this information along with the displayed exit scene and the bypassed signals and creates the total exit scene signal. The 945 SAOC decoder also produces a series of meta-object gain values, the number of which depends on the grouping of meta-objects and the form of loudness information desired.
The gain values are provided to the loudness estimator 960 of the mix which also receives loudness information from the meta-objects of the encoder.
Then the loudness estimator 960 of the mix is able to determine the desired loudness information, which may include, but is not limited to.
IMPI Mexican institute OF INDUSTRIAL MONETY
<img file="MX358306B_D0069.tif" />
of limitation, the loudness of the downmix signal, the loudness of the displayed output scene, and / or the difference in loudness between the downmix signal and the displayed output scene.
In some modes, the loudness information itself is sufficient, while in other modes it is desirable to process the entire output depending on the loudness information determined. This processing may, for example, be the compensation of any possible difference in loudness between the downmix signal and the displayed output scene. Such processing, for example, by a loudness processing unit 970, would make sense in a transmission situation, as it would reduce changes in the perceived loudness of the signal regardless of user interaction (setting the gain of input dialog).
Loudness-related processing in this specific embodiment comprises a plurality of new features. Among other things, FGO, BGO, and potential bypassed channels are premixed to form the final channel configuration so that downmixing can be performed by simply adding the two premixed channels together (for example, downmix matrix coefficients of 1 ), what
IMPI Mexican iNrrmrro
DI INDUSTRIAL PROPERTY constitutes
<img file="MX358306B_D0070.tif" />
a new feature.
In addition, as a new additional feature, the average loudness of the FGO and BGO is estimated and the difference is calculated. In addition, the objects are mixed to obtain a downmix signal that is transmitted. In addition, as a new additional feature, the loudness difference information is included in the secondary information that is transmitted, (new) In addition, the decoder uses the secondary information for the (virtual) separation of the objects and recombines the objects using the information representation that is based on downmix information and user-entered gain mod. In addition, as a new additional feature, the decoder uses the modifying gain and transmitted loudness information to estimate the change in the average loudness of the system output compared to the default downmix.
The following is a formal description of the modalities.
Assuming that the loudness values of the objects behave similarly to the logarithm of the energy values when adding the objects, that is, that the loudness values must be transformed to the linear domain, added there and finally transformed again to the logarithmic domain .
Presents itself
IMPI
ICβΠΤυΤΟ MEXICAN. FROM PROPERTY now the fWWa
<img file="MX358306B_D0071.tif" />
this by means of the definition of the 'meiílida' of<sup>1</sup> yCJiiCJilllJLl BS.1770 (for simplicity, the number of channels is set to one, although the same principle can be applied to multi-channel signals with the appropriate sum over the channels).
The sound of the 1<sup>ava</sup> K filtered signal z, where the mean quadratic energy e, is defined as
L¡ = c + 101og<sub>10</sub> e., where c is a constant of displacement.
For example, c can be
-0.691. From this it appears that the energy of the signal can be determined from the loudness with
The sum of
No signs is not then
Í = 1 / = 1 and the loudness of this sum signal is then
IMPI
JV MEXICAN INSTITUTE OF PROPERTY
L<sub>sutl</sub> = e + 101og<sub>the</sub>and<sub>yes</sub>,<sub>w</sub> = <sub>C</sub> + 101og „/ = 1
<img file="MX358306B_D0072.tif" />
If the signals are not correlated, the correlation coefficients C¡ should be taken into account. by approximating the summed signal energy as follows<sup>and</sup>SUM
NN = ΣΣ ^
Í = 1 y = i where the cross energy e. . between objects i<sup>ava</sup> and J<sup>ava</sup> is defined as follows
<img file="MX358306B_D0073.tif" />
= c. .7io<sup>(to</sup>-<sup>c) / 1st</sup>io<sup>(£</sup>^<sup>c) / 10</sup><sup>l</sup>'J f <sub>= C</sub> J<sub>10</sub>(A + L-2<sub>C</sub>) / I0 where —1 <C<sub>;</sub>. <1 is the correlation coefficient between the two objects i and j. When two objects are uncorrelated, the correlation coefficient is equal to 0, and when the two objects are identical, the correlation coefficient is equal to 1.
Extending the model further with the g mix weights<sub>i</sub> to be applied to the signals
Λ 'in the mixing process, i.e. z<sub>sl¡M</sub> = Xg, z,, energy, = i of the summed signal is <sup>and</sup>suM ΣΣ SiS¡<sup>and</sup>i, /,, = l / = 1
IMPI
MEXICAN INSTITUTE
OF THE PROPERTY
INDUSTRIAL
<img file="MX358306B_D0074.tif" />
and from this the loudness of the mixing signal can be obtained, as before, with
Lsum c +10 log<sub>10</sub> .
The difference between the loudness of the two signals can be estimated as follows
If the aforementioned definition of loudness is now used, this can be expressed as (i,;) =!, - £.
= (c + 10ÍDg<sub>ia</sub>eJ ~ (c4401og<sub>]0</sub>and<sub>;</sub>), = 101og<sub>10</sub> - which can be seen depending on the signal energies. If now you want to estimate the difference in loudness between two mixes, VA '^ = Σ &<sup>Ζ</sup><· And <sup>ζ</sup>5 = Σ<sup>/ ζ</sup>α · / = 1 f = I possibly with mix weights
IMPI different g<sub>i</sub> and Z<sub>;</sub>., this can be estimated
Mexican Institute of Industrial Property
<img file="MX358306B_D0075.tif" />
ΔΖ (Λ, 2?) = 101og<sub>]0</sub>- e
B
NN
ΣΣ & ^ α, <sup>= 1</sup>01ogio4 - ^ ----- ΣΣ<sup>Α</sup>Λ% ·
Z = 1 j = \ = 10 log
-AI (wW<sup>111</sup>
ΣΣ ^ αΛ =<sup>10 the</sup>8i »dd ------- 7—, rΣΣΜΑΛίο / = iy = i
ΣΣ ^ ΑΑ '= i> i ________________________ <sup>1(1</sup> Λ X /
1 = 1 J = 1
In case the objects are not correlated (C¡. = 0, Vz Ψ jy
Cj j = l, Vz - j), the estimate of the difference becomes
Λ '
Σ-<sup>(:</sup>
A¿ (d, B) = 101og<sub>10</sub>^ --- A / iO / = 1
Σ ^<sup>10</sup>*
-lfllnn / = 1___ ___
Coding is considered below
IMPI
MEXICAN INSTITUTE
OF THE PROPERTY
INDUSTRIAL
<img file="MX358306B_D0076.tif" />
differential.
Loudness values per object can be coded as differences from the loudness of a selected reference object:
<img file="MX358306B_D0077.tif" />
where
Lref is the loudness of the reference object. This coding is advantageous if absolute loudness values are not needed as a result, since it is now necessary to transmit a less value, and the loudness difference estimate can be expressed as
<img file="MX358306B_D0078.tif" />
or in the case of uncorrelated objects
Next, a dialogue improvement scenario is considered.
Consider, once again, the dialog enhancement application scenario. The freedom to define the rendering information in the decoder is limited only to the level changes of the two meta-objects.
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<img file="MX358306B_D0079.tif" />
Let us further suppose that the two meta-objects are uncorrelated, that is, C<sub>FGOBGO</sub>-0. If the down-mix weights of the meta-objects are h<sub>FG0</sub> yh<sub>BG0</sub>, and are presented with earnings f<sub>FGO</sub> and f<sub>BCO</sub>, the loudness of the output with respect to the default downmix is
<img file="MX358306B_D0080.tif" />
<td></td><td>This</td><td>it is then also</td><td>the</td><td colspan="2">compensation</td>
<td>necessary</td><td>whether</td><td colspan="2">you want to have the same loudness</td><td>in</td><td>the exit</td>
<td>that in the</td><td>mixture</td><td>descending by default.</td><td></td><td></td><td></td>
<td></td><td>TO,</td><td>B) It can be considered</td><td>how</td><td>a</td><td>value of</td>
loudness compensation, which can be transmitted by the decoder signal processor 120. AL (A, B) can also be designated as the loudness change value and thus the actual offset value can be an inverse value. Or is it correct to use the name loudness compensation factor for it too? In this way, the loudness compensation value lev mentioned earlier in this document would correspond to the following value gDeita.
-ΛΛ (, Ι.β).
For example, you can apply g<sub>TO</sub>=10 <sup>20</sup> 1 /
IMPI
MEXICAN INSTITUTE DK LA PROPIEDAD
AL (A, B) as a multiplication factor for each<sup>IN</sup>'channel
<img file="MX358306B_D0081.tif" />
e a modified audio signal that appears cofflo 'YéáüTfádb — cTe — Ta application of the representation information to the audio input signal. This linear goeite equation. In the logarithmic, different domain, such as 1 / AL (A, B) and the domain is acted upon, the serious equation would apply accordingly.
If the downmix process is simplified in such a way that the two meta-objects can be mixed with unit weights to obtain the downmix signal, i.e. h<sub>FG0</sub> = h<sub>BGO</sub> = 1, and now the representation gains for these two objects are
<td>indicated</td><td>with g<sub>FCO</sub> yg<sub>BCO</sub> This simplifies the equation of</td>
<td>Change of</td><td>loudness to get 2 Ul-ai -'-) 2. > ΔΛ (Α5) = 101og<sub>10</sub> ---- 10 +10 2 .λ '· ™<sup>10</sup> 2 = 101oe <sup>+</sup>Sbgo ^ , Re, O ™ i-BGom 10 +10 Again, AL (A<sub>Z</sub> B) how</td>
<td>value of</td><td>loudness compensation determined by the</td>
signal processor 120.
In general, greo can be considered a representation gain for the foreground object FGO (group of foreground objects), and gsGo can be considered a representation gain for the background object BGO (group of objects in the foreground). second
As mentioned above, it is possible to transmit loudness differences instead of absolute loudness. Let's define the flat reference loudness).
IMPI
MEXICAN INSTITUTE
OF PROPERTY, INDUSTMAL
<img file="MX358306B_D0082.tif" />
as loudness of the FGO meta-object
Lref
<img file="MX358306B_D0083.tif" />
f ie
K-FGO ~ L<sub>FG0</sub> L<sub>RE</sub>f - θ and K-bgo L<sub>bco</sub> L<sub>ref</sub> L<sub>bgo</sub> L<sub>fgo</sub> . Now the change in loudness is
2 . «^ ΛΠΟ '<sup>10</sup>
Δ £ (4δ) = 10log „ <sup>gTO + g</sup>X „.-- 1 + 10
It can also occur, as in the case of
SAOC-DE, that two meta-objects do not have individual scale factors, but that one of the objects remains unchanged, while the other is attenuated to obtain the correct mixing ratio between the objects. In this rendering setting, the output will be lower in loudness than the default mix, and the loudness change is '2 ·'?
Ai (4 β) = io; og, „+ 10 with
Sfgo <
Sfgo ^ Sbgo
Sfgo - Sbgo
Sfgo <sup><</sup> Sbgo
Sbgo 'Sbgo <sup><</sup> Sfgo 'if Sbgo - Sfgo
This form is already quite simple, and quite independent with respect to the loudness measurement used. The only real requirement is that the values of
IMPI
MEXICAN INSTITUTE OF PROPERTY
INDUSTRIAL
<img file="MX358306B_D0084.tif" />
loudness must be added in the exponential domain. It is possible to transmit / store the values of the signal energies instead of the loudness values, since there is a close connection between the two.
In each of the formulas presented above, Δί (Α, B) can be considered a loudness compensation value, which can be transmitted by the decoder signal processor 120.
Practical examples are considered below.
The precision of the concepts provided is illustrated by two exemplary signals. Both signals have 5.1 downmix with the surround and LFE channels bypassed for SAOC processing.
Two main procedures are used: one (of 3 terms) with three meta-objects: FGO, BGO, and overridden channels, for example,
X <sup>=</sup> XFGO <sup>+</sup> BGO <sup>+</sup> X-BYPASS '
And another (of 2 terms) with two meta-objects, for example emplo:
In the 2-term strategy, the canceled channels can be mixed, for example, together with the BGO for the estimation of the loudness of the meta-objects. The loudness of both (or all three) objects is estimated, as well as the loudness of the downmix signal, and the values are saved.
IMPI
MEXICAN INSTITUTE
OF INDUSTRIAL PROPERTY
<img file="MX358306B_D0085.tif" />
Representation instructions are on the form
<img file="MX358306B_D0086.tif" />
<img file="MX358306B_D0087.tif" />
for the two procedures, respectively.
Gain values are determined, for example, according to:
<img file="MX358306B_D0088.tif" />
, or .Sfgo, or where the gain of the FGO Sfgo varies between -24 to +24 dB.
The output scenario is presented, the loudness is measured, and the loudness attenuation of the downmix signal is calculated.
This result is represented in Figure 10 and Figure 11 with the blue line with circular markers. Figure 10 represents a first illustration and the Figure represents a second illustration of a change in loudness measured and the result of using the concepts provided to estimate the change in a purely parametric way.
IMPI
MEXICAN INSTITUTE
OF THE PROPERTY
INDUSTRIAL
<img file="MX358306B_D0089.tif" />
loudness of
The downmix attenuation is then estimated parametrically using the loudness values of the stored meta-objects and the downmix and display information.
The estimation using the loudness of three meta-objects with the green line with square markers is illustrated and the estimation using the loudness of two meta-objects with the red line with star-shaped markers is illustrated.
It can be seen from the figures that the 2 and 3 term procedures provide practically identical results, and that both approximate the measured value quite effectively.
The concepts provided show a plurality of advantages. For example, the concepts provided allow us to estimate the loudness of a mix signal from the loudness of the component signals that make up the mix. The benefit of this is that the loudness of the component signals can be estimated once and that the estimate of the loudness of the mix signal can be obtained parametrically with respect to any mix without the need for loudness estimation based on the actual signal. This constitutes a considerable improvement in the computing efficiency of the system.
IMPI «βπτυτο muKano DE LA MONEDAD INDUSTRIAL ensemble in which the loudness estimation of various mixtures is necessary. For example, when the end user changes the rendering settings, the loudness estimation of the output is immediately available.
adapt loudness
In some applications, such as the EBU R128 recommendation, average of the total scenario estimate received, average in the case is an important program.
loudness in transmission, only the program.
temporary loudness.
components of the estimate after
Due will contain
According to the loudness, it is from the receiver, for example, it is carried out based on the convergence of this, errors estimating what has been proposed towards the received
<img file="MX358306B_D0090.tif" />
Loudness signal all compensation of the o will show loudness variations of the objects and transmit the information possible to estimate the average loudness of mix in the receiver without delay.
If it is convenient for the output signal to be independently constant representation information, keep from the average loudness changes in the provided concepts allow to determine a compensation factor for this reason.
The calculations necessary for this in the decoder are negligible from the
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<img file="MX358306B_D0091.tif" />
point of its computing complexity and therefore it is possible to add the functionality to any decoder.
There are cases in which the absolute loudness level of the output is not important, but the importance lies in determining the change in loudness with respect to a reference scene. In these cases, the absolute levels of the objects are not important, although their relative levels are. This allows one of the objects to be defined as a reference object and represents the loudness of the other objects in relation to the loudness of this reference object. This offers some benefits, taking into account the transport and / or storage of loudness information.
First of all, it is not necessary to transport the reference loudness level. In the case of application of the two meta-objects, this cuts in half the amount of data to be transmitted. The second benefit is related to the possible quantification and representation of loudness values. Since the absolute levels of the objects can be almost any, the absolute loudness values can also be almost any. Relative loudness values, on the other hand, are assumed to have a mean of 0 and a well-formed distribution around the mean. The difference between the representations allows you to define the quantization grid of the relative representation so that it has potentially higher precision with the same number of bits
IMPI
My MEXICAN IITUTE OF THE INDUSTRIAL rKOriEDAB
<img file="MX358306B_D0092.tif" />
used for quantized representation.
Figure 12 illustrates another embodiment of loudness compensation execution. In Figure 12, loudness compensation can be carried out, for example, to compensate for loudness loss. For this purpose, for example, the values can be used
DE_loudness_diff_dialogue (= Kfgo) and DE_loudness_diff_background (= DE_control_info's Kbgo'i. Here, DE_control_info can specify Advanced Clean Audio dialog enhancement control (DE) information
Loudness compensation is obtained by applying a gain value to the SAOC-DE output signal and
<td>to</td><td>the channels canceled</td><td>(in</td><td>the case</td><td>of a signal</td>
<td colspan="2">multichannel).</td><td></td><td></td><td></td>
<td></td><td>In the modality</td><td>of the</td><td>Figure 12,</td><td>this is made of</td>
<td>the</td><td>Following way:</td><td></td><td></td><td></td>
<td></td><td>A</td><td>value</td><td>1imited</td><td>of profit by</td>
dialog modification rr¡<sub>G</sub> to determine the effective gains for the foreground object (FGO, for example, dialog) and the background object (BGO, for example, environment).
This is done using block 1220
IMPI
Mexican DEντΟ OF INDUSTRIAL PROPERTY
<img file="MX358306B_D0093.tif" />
Gain mapping that produces the gain values ^ FGO 7 ^ BGO '
Block 1230 Output Loudness Estimator uses the loudness information y, and the effective gain values Ηθο and / 77<sub>β</sub>θ<sub>0</sub> to estimate this possible change in loudness compared to the default downmix case. The change is then mapped with the Loudness Compensation Factor that is applied to the output channels to produce the final Output Signals.
The following steps apply for loudness compensation:
Receive the limited gain value m<sub>Q</sub> the SAOC-DE decoder (as defined in clause 12.8 Modification margin control for SAOC-DE [DE]), and determine the applied FGO / BGO gains:
<sup>m</sup>FG0 = <sup>m</sup>G 'and <sup>m</sup>BG<sub>OR</sub> =<sup>1</sup> Yes <I <
<sup>m</sup>FGO = 1 »Y <sup>m</sup>BGO = <sup>m</sup>G <sup>YES m</sup>G> 1
Get the loudness information of metabjects K<sub>FG0</sub> and ^ BGO ·
Calculate the change in output loudness compared to the default downmix with
IMPI
MEXICAN OE INDUSTRIAL PROPERTY
<img file="MX358306B_D0094.tif" />
A¿ = 101og<sub>10</sub><sub>3</sub> Kfgo / ~ Krgo / m<sup>2</sup>FCOW <sup>/,0</sup>+ m<sup>2</sup>GOlQ 4¡L.
K / 'Go / ^ BG () / <sup>/10</sup>+10
Calculate the compensation profit of. ,<sub>n</sub> _ 1 n-0.05AL loudness and <sub>Δ</sub> - lu
Calculate the scale factors where
If = g<sub>TO</sub> if channel i belongs to the SAOC-DE output if channel i is a bypassed channel, and <sup>m</sup>BGoS \ total output channels. In the figure
12, g<sub>N</sub> is the number the gain adjustment is divided into two steps: the gain of the possible channels bypassed is adjusted with m<sub>BGO</sub> before combining them with the SAOC-DE output channels, and then a common gain is applied to all the channels combined. This is only possible by rearranging the gain adjustment operations, while g here combines
<td>both steps</td><td>adjustment</td><td>of</td><td>gain in</td><td>a solo</td><td>adjustment</td><td>of</td>
<td>gain.</td><td></td><td></td><td></td><td></td><td></td><td></td>
<td> -</td><td>Apply</td><td>the</td><td>values of</td><td>scale</td><td>g <sup>to</sup></td><td>the</td>
<td>Channels of</td><td>audio Y<sub>FULL</sub></td><td colspan="2">what consist</td><td colspan="2">in the channels</td><td>of</td>
<td>output of</td><td colspan="2">SAOC- DE</td><td>^ SAOC Y <sup>the</sup></td><td>possible</td><td colspan="2">channels</td>
canceled aligned in time Y<sub>BYPASS</sub> : AND<sub>FULL</sub><sup>=</sup>Ysaoc '' ^ svpass
The application of values ** · **** ·>
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY of scale g
<img file="MX358306B_D0095.tif" />
to audio channels AND<sub>FULL</sub> it is executed by the gain adjustment unit 1240.
AL previously calculated can be considered as loudness compensation value. In general, iufgo indicates a representation gain for the foreground object FGO (group of foreground objects) and iübgo indicates a representation gain for the background object BGO (group of foreground objects).
Although some aspects have been described in the context of an apparatus, it is obvious that these aspects also represent a description of the corresponding method, in which a block or device corresponds to a method step or a characteristic of a method step. Similarly, the aspects described in the context of a method step also represent a description of a corresponding block or item or of a characteristic of a corresponding apparatus.
The inventive decomposed signal can be stored in a digital medium or can be transmitted by a transmission medium such as a wireless transmission medium or a cable transmission medium such as the Internet.
Depending on certain implementation requirements, the embodiments of the invention can be implemented in hardware or in software. The implementation
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY can be carried out using a digital storage medium, for example a floppy disk, a DVD, a CD, a ROM, a PROM, an EPROM, an EEPROM or a FLASH memory, which has stored in the same signals readable controllers, which cooperate (or have the ability to cooperate) with a programmable computer system such that the respective method is executed.
Some embodiments in accordance with the invention comprise a non-transient data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is executed.
In general, the embodiments of the present invention can be implemented as a computer program product with a program code, the program code can be operated to execute one of the methods when it is run when the computer program product runs on a computer . The program code can be stored, for example, on a machine-readable carrier.
Other modalities include the computer program to execute one of the methods described herein, stored in a machine-readable carrier.
In other words, a modality of the
<img file="MX358306B_D0096.tif" />
<img file="MX358306B_D0097.tif" />
IMPI
The MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY therefore consists of a program of —l ll.
computer that has a program code to perform one of the methods described here, when the computer program runs on a computer.
Another embodiment of the methods of the invention therefore consists of a data carrier (or digital storage medium, or computer readable medium) comprising, recorded therein, the computer program for executing one of the methods. described herein.
Another embodiment of the method of the invention is, therefore, a data stream or signal sequence representing the computer program for executing one of the methods described herein. The data stream or signal sequence may be configured, for example, to be transferred over a data communication connection, for example over the Internet.
Another embodiment comprises a processing means, for example a computer, or a programmable logic device configured or adapted to execute one of the methods described herein.
Another modality comprises a computer in which the computer program is installed therein to execute one of the methods described herein.
In some modalities, a
IMPI
INSTITUTO MEXICANO DE LA PROPIEDAD INDUSTRIAL programmable logic device (for example an array of programmable gates in the field) to execute some or all the functionalities of the methods described herein. In some embodiments, an array of field programmable gates can cooperate with a microprocessor to execute one of the methods described herein. In general, the methods are preferably executed by any hardware device.
The embodiments described above are only illustrative of the principles of the present invention. It is understood that modifications and variations of the arrangements and details described herein will be apparent to persons skilled in the art. Therefore, it is only intended to be limited to the scope of the following patent claims and not to the specific details presented by way of description and explanation of the modalities presented herein.
<img file="MX358306B_D0098.tif" />
References
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<img file="MX358306B_D0099.tif" />
[BCC]
C. Faller and F. Baumgarte, Binaural Cue
Coding - Part II: Schemes and applications, IEEE Trans. on
Speech and Audio Proc., Vol. 11, no. 6, Nov. 2003.
[EBU] EBU Recommendation R 128 Loudness normalization and permitted maximum level of audio signáis, Geneva, 2011.
[JSC] C. Faller, Parametric Joint-Coding of
Audio Sources, 120th AES Convention, Paris, 2006.
[ISS1] M. Parvaix and L. Girin: Informed Source
Separation of underdetermined instantaneous Stereo Mixtures using Source Index Embedding, IEEE ICASSP, 2010.
[ISS2] M. Parvaix, L. Girin, J.-M. Brossier:
<td>TO</td><td colspan="3">watermarking-based method for informed</td><td>source</td><td>separation</td>
<td>of</td><td>Audio</td><td colspan="2">you sign with a single sensor,</td><td colspan="2">IEEE Transactions</td>
<td>on</td><td>Audio,</td><td>Speech</td><td>and Language Processing,</td><td> 2010.</td><td></td>
<td></td><td></td><td>[ISS3]</td><td colspan="2">A. Liutkus and J. Pinel and R.</td><td>Badeau and</td>
<td>L.</td><td>Girin</td><td>and</td><td>G. Richard: Informed</td><td>source</td><td>separation</td>
through spectrogram coding and data embedding, Signal Processing Journal, 2011.
[ISS4] A. Ozerov, A. Liutkus, R. Badeau, G. Richard: Informed source separation: source coding meets source separation, IEEE Workshop on Applications of Signal Processing to Audio and Acoustics, 2011.
[ISS5]
S. Zhang and L. Girin:
An Informed
Source
Separation
IMPI
MEXICAN INSTITUTE OF PROPERTY „ <sub>n</sub> INDUSTRIAL
System for Speech Signáis, INTERSP
<img file="MX358306B_D0100.tif" />
F
2011.
[ISS6]
L.
Girin and J.
Pinel: Informed Audio
Source
Separation from
Compressed
Linear Stereo Mixtures,
AES 42nd International
Conference:
Semantic Audio, 2011.
<td>[AND YOU]</td><td colspan="2">International</td><td colspan="2">Telecommunication Union:</td>
<td colspan="2">Recommendation ITU-R</td><td>BS.1770-3</td><td colspan="2">- Algorithms to measure</td>
<td>audio program</td><td colspan="2">loudness and</td><td colspan="2">true-peak audio level,</td>
<td>Geneva, 2012.</td><td></td><td></td><td></td><td></td>
<td>[SAOC1]</td><td>J.</td><td>Herre, S.</td><td>Disch, J. Hilpert,</td><td>OR.</td>
<td>Hellmuth: From</td><td>SAO</td><td>To SAOC</td><td>Recent Developments</td><td>in</td>
<td>Parametric Coding</td><td>of</td><td colspan="2">Spatial Audio, 22nd Regional UK</td><td>AES</td>
<td colspan="2">Conference, Cambridge,</td><td>UK, April</td><td> 2007 .</td><td></td>
<td>[SAOC2]</td><td>J.</td><td>Engdegárd,</td><td>B. Resch, C. Falch,</td><td> 0.</td>
<td colspan="4">Hellmuth, J. Hilpert, A. Holzer, L. Terentiev,</td><td>J.</td>
Breebaart, J. Koppens, E. Schuijers and W. Oomen: Spatial Audio Object Coding (SAOC) - The Upcoming MPEG Standard on Parametric Object Based Audio Coding, 124th AES Convention, Amsterdam 2008.
[SAOC] ISO / IEC, MPEG audio technologies Part 2: Spatial Audio Object Coding (SAOC), ISO / IEC JTC1 / SC29 / WG11 (MPEG) International Standard 23003-2.
[EP]
EP 2146522 Al: S. Schreiner, W. Fiesel,
M. Neusinger, O. Hellmuth, R. Sperschneider, Apparatus and method for generating audio output signáis using object
IMPI
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<img file="MX358306B_D0101.tif" />
based metadata, 2010.
<td>[OF]</td><td>ISO / IEC, MPEG audio technologies -</td><td>Part 2:</td>
<td>Spatial Audio</td><td>Object Coding (SAOC) - Amendment 3,</td><td>Dialogue</td>
<td>Enhancement, </td><td>ISO / IEC 23003-2: 2010 / DAM 3,</td><td>Dialogue</td>
<td>Enhancement.</td><td></td><td></td>
<td>[BRE]</td><td>WO 2008/035275 A2.</td><td></td>
<td>[SCH]</td><td>EP 2 146 522 Al.</td><td></td>
<td>[ENG]</td><td>WO 2008/046531 Al.</td><td></td>
<img file="MX358306B_D0102.tif" />
Contents109
114 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41 Sheet 42 Sheet 43 Sheet 44 Sheet 45 Sheet 46 Sheet 47 Sheet 48 Sheet 49 Sheet 50 Sheet 51 Sheet 52 Sheet 53 Sheet 54 Sheet 55 Sheet 56 Sheet 57 Sheet 58 Sheet 59 Sheet 60 Sheet 61 Sheet 62 Sheet 63 Sheet 64 Sheet 65 Sheet 66 Sheet 67 Sheet 68 Sheet 69 Sheet 70 Sheet 71 Sheet 72 Sheet 73 Sheet 74 Sheet 75 Sheet 76 Sheet 77 Sheet 78 Sheet 79 Sheet 80 Sheet 81 Sheet 82 Sheet 83 Sheet 84 Sheet 85 Sheet 86 Sheet 87 Sheet 88 Sheet 89 Sheet 90 Sheet 91 Sheet 92 Sheet 93 Sheet 94 Sheet 95 Sheet 96 Sheet 97 Sheet 98 Sheet 99 Sheet 100 Sheet 101 Sheet 102 Sheet 103 Sheet 104 Sheet 105 Sheet 106 Sheet 107 Sheet 108 Sheet 109 Sheet 110 Sheet 111 Sheet 112 Sheet 113 Sheet 114
74 members in 19 offices
Priority claims3
| Document | Office | Kind | Date |
|---|---|---|---|
| 131946642 | European Patent Office (EPO) | – | |
| 13194664 | European Patent Office (EPO) | A | |
| 2014075787 | European Patent Office (EPO) | W |
Members74
| Document | Office | Kind | |
|---|---|---|---|
| EP2879131A1 | European Patent Office (EPO) | A1 | |
| CA2900473A1 | Canada | A1 | |
| CA2931558A1 | Canada | A1 | |
| WO2015078956A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2015078964A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW201525990A | Taiwan Province of China | A | |
| AU2014356475A1 | Australia | A1 | |
| TW201535353A | Taiwan Province of China | A | |
| KR20150123799A | Republic of Korea | A | |
| EP2941771A1 | European Patent Office (EPO) | A1 | |
| US2015348564A1 | United States of America | A1 | |
| CN105144287A | China | A | |
| MX2015013580A | Mexico | A | |
| AR098558A1 | Argentina | A1 | |
| AU2014356467A1 | Australia | A1 | |
| KR20160075756A | Republic of Korea | A | |
| JP2016520865A | Japan | A | |
| AR099360A1 | Argentina | A1 | |
| CN105874532A | China | A | |
| AU2014356475B2 | Australia | B2 | |
| MX2016006880A | Mexico | A | |
| US2016254001A1 | United States of America | A1 | |
| EP3074971A1 | European Patent Office (EPO) | A1 | |
| AU2014356467B2 | Australia | B2 | |
| HK1217245A1 | Hong Kong, China | A1 | |
| JP2017502324A | Japan | A | |
| TWI569259B | Taiwan Province of China | B | |
| TWI569260B | Taiwan Province of China | B | |
| RU2015135181A | Russian Federation | A | |
| EP2941771B1 | European Patent Office (EPO) | B1 | |
| KR101742137B1 | Republic of Korea | B1 | |
| PT2941771T | Portugal | T | |
| BR112015019958A2 | Brazil | A2 | |
| BR112016011988A2 | Brazil | A2 | |
| ES2629527T3 | Spain | T3 | |
| MX350247B | Mexico | B | |
| JP6218928B2 | Japan | B2 | |
| PL2941771T3 | Poland | T3 | |
| ZA201604205B | South Africa | B | |
| RU2016125242A | Russian Federation | A | |
| CA2900473C | Canada | C | |
| EP3074971B1 | European Patent Office (EPO) | B1 | |
| US9947325B2 | United States of America | B2 | |
| RU2651211C2 | Russian Federation | C2 | |
| ES2666127T3 | Spain | T3 | |
| PT3074971T | Portugal | T | |
| KR101852950B1 | Republic of Korea | B1 | |
| JP6346282B2 | Japan | B2 | |
| US2018197554A1 | United States of America | A1 | |
| PL3074971T3 | Poland | T3 | |
| MX358306BThis record | Mexico | B | |
| RU2672174C2 | Russian Federation | C2 | |
| CA2931558C | Canada | C | |
| US10497376B2 | United States of America | B2 | |
| US2020058313A1 | United States of America | A1 | |
| CN105874532B | China | B | |
| CN111312266A | China | A | |
| US10699722B2 | United States of America | B2 | |
| US2020286496A1 | United States of America | A1 | |
| CN105144287B | China | B | |
| CN112151049A | China | A | |
| US10891963B2 | United States of America | B2 | |
| US2021118454A1 | United States of America | A1 | |
| BR112015019958B1 | Brazil | B1 | |
| MY189823A | Malaysia | A | |
| US11423914B2 | United States of America | B2 | |
| BR112016011988B1 | Brazil | B1 | |
| US2022351736A1 | United States of America | A1 | |
| MY196533A | Malaysia | A | |
| US11688407B2 | United States of America | B2 | |
| US2023306973A1 | United States of America | A1 | |
| CN111312266B | China | B | |
| US11875804B2 | United States of America | B2 | |
| CN112151049B | China | B |
1 legal event, as the office reported them to INPADOC
Events
| Event | Code | |
|---|---|---|
| Grant or registrationFG | FG |
Numbers
- Publication
- 358306
- Application
- 6880
Titles2
- Spanish
- DECODIFICADOR, CODIFICADOR Y METODO PARA LA ESTIMACION INFORMADA DE SONORIDAD EN SISTEMAS DE CODIFICACION BASADA EN OBJETOS.
- English
- DECODER, ENCODER AND METHOD FOR INFORMED LOUDNESS ESTIMATION IN OBJECT-BASED AUDIO CODING SYSTEMS.
Classification
- CPC, 5
- G10L19/008
- G10L19/0017
- G10L19/005
- G10L19/265
- H03G3/20
- IPC, 1
- G10L19 008