Concept for audio encoding and decoding for audio channels and audio objects.
Abstract
Audio encoder for encoding audio input data (101) to obtain audio output data (501) comprises an input interface (100) for receiving a plurality of audio channels, a plurality of audio objects and metadata related to one or more of the plurality of audio objects; a mixer (200) for mixing the plurality of objects and the plurality of channels to obtain a plurality of pre-mixed channels, each pre-mixed channel comprising audio data of a channel and audio data of at least one object; a core encoder (300) for core encoding core encoder input data; and a metadata compressor (400) for compressing the metadata related to the one or more of the plurality of audio objects, wherein the audio encoder is configured to operate in at least one mode of the group of two modes comprising a first mode, in which the core encoder is configured to encode the plurality of audio channels and the plurality of audio objects received by the input interface as core encoder input data, and a second mode, in which the core encoder (300) is configured for receiving, as the core encoder input data, the plurality of pre-mixed channels generated by the mixer (200).

Term
7.8 yearsleft in the term
Expires 16 July 2034.
- Priority
- Filed
- Granted
- Today
- Expires
23 claims: 14 independent, 9 dependent
- 1REIVINDICACIONES 1. Un codificador de audio para codificar datos de entrada de audio para obtener datos de salida de audio caracterizado porque comprende:una interfaz de entrada configurada para recibir una pluralidad de canales de audio, una pluralidad de objetos de audio y meta-datos relacionados con uno o más de la pluralidad de objetos de audio;un mezclador configurado para mezclar la pluralidad de objetos y la pluralidad de canales para obtener una pluralidad de canales mezclados previamente, cada canal mezclado previamente que comprende datos de audio de un canal y datos de audio de por lo menos un objeto;un codificador central configurado para codificar en forma central datos de entrada del codificador central;y un compresor de meta-datos configurado para comprimir los meta-datos relacionados con uno o más de la pluralidad de objetos de audio, en donde el codificador de audio se configura para operar en ambos modos de un grupo de por lo menos dos modos que comprende un primer modo, en el cual el codificador central se configura para codificar la pluralidad de canales de audio y la pluralidad de objetos de audio recibidos por la interfaz de entrada como datos de entrada del codificador central, y un segundo modo, en el cual el codificador central se configura para recibir, como los datos de entrada del codificador central, la pluralidad de canales mezclados previamente generados por el mezclador y para codificar la pluralidad de canales previamente mezclados.
- 2El codificador de audio de conformidad con la reivindicación 1, caracterizado porque además comprende:un codificador de objeto de audio espacial para generar uno o más canales de transporte y datos paramétricos a partir de los datos de entrada del codificador de objetos de audio espacial, en donde el codificador de audio se configura para operar en forma adicional en un tercer modo, en el cual el codificador central codifica uno o más canales de transporte derivados de los datos de entrada del codificador de objeto de audio espacial, los datos de entrada del codificador de objeto de audio espacial que comprende la pluralidad de objetos de audio dos o más de la pluralidad de canales de audio.
- 3El codificador de audio de conformidad con cualquiera de las reivindicaciones 1 o 2, caracterizado porque además comprende:un codificador de objeto de audio espacial para generar uno o más canales de transporte y datos paramétricos a partir de los datos de entrada del codificador de objetos de audio espacial, en donde el codificador de audio se configura para operar en forma adicional en un cuarto modo, en el cual el codificador central codifica canales de transporte derivado del codificador de objeto de audio espacial a partir de los canales mezclados previamente como los datos de entrada del codificador de objeto de audio espacial.
- 4El codificador de audio de cualquiera de las reivindicaciones que anteceden, caracterizado porque además comprende un conector para conectar una salida de la interfaz de entrada a una entrada del codificador central en el primer modo y para conectar la emisión de la interfaz de entrada a una entrada del mezclador y a conectar una salida del mezclador a la entrada del codificador central en el segundo modo, y un controlador de modos para controlar el conector de acuerdo con una indicación de modo recibida de una interfaz de usuario o extraída de los datos de entrada de audio.
- 5El codificador de audio de conformidad con cualquiera de las reivindicaciones anteriores, caracterizado porque además comprende:una interfaz de salida para proporcionar una señal de salida como los datos de salida de audio, la señal de salida que comprende, en el primer modo, una salida del codificador central y metadatos comprimidos, y que comprende, en el segundo modo, una salida del codificador central sin ningún meta-dato, y que comprende, en el tercer modo, una salida del codificador central, información lateral SAOC y los meta-datos comprimidos y que comprende, en el cuarto modo, una salida del codificador central e información lateral SAOC.
- 6El codificador de audio de conformidad con cualquiera de las reivindicaciones anteriores, caracterizado porque el mezclador se configura para pre- renderizar la pluralidad de objetos de audio con el uso de los meta-datos y una indicación de la posición de cada canal en una configuración de reproducción, con la cual se asocia la pluralidad de canales, en donde el mezclador se configura para mezclar un objeto de audio con por lo menos dos canales de audio y con esto luego el total de cantidad de canales de audio, cuando el objeto de audio deberá colocarse entre por lo menos dos canales de audio en la configuración de reproducción, según lo determinan los meta-datos.
- 7El codificador de audio de conformidad con cualquiera de las reivindicaciones anteriores, caracterizado porque además comprende un descompresor de meta-datos para descomprimir meta-datos comprimidos emitidos por el compresor de meta-datos, y en donde el mezclador se configura para mezclar la pluralidad de objetos de acuerdo con meta-datos descomprimidos, en donde una operación de compresión realizada por el compresor de meta-datos es una operación de compresión con pérdida que comprende un paso de cuantificación.
- 8Un descodificador de audio para descodificar datos de audio codificados, caracterizado porque comprende:una interfaz de salida configurada para recibir los datos de audio codificados, los datos de audio codificados que comprende una pluralidad de pluralidad de canales codificados o una objetos codificados o comprimir meta-datos relacionados con la pluralidad de objetos;un descodificador central configurado para descodificar la pluralidad de canales codificados y la pluralidad de objetos codificados;un descompresor de meta-datos configurado para descomprimir los meta-datos comprimidos, un procesador de objetos configurado para procesar la pluralidad de objetos descodificados con el uso de meta datos comprimidos para obtener una cantidad de canales de salida que comprende datos de audio a partir de los objetos y los canales descodificados;y un post-procesador configurado para convertir la cantidad de canales de salida en un formato de salida, en donde el descodificador de audio se configura para desviar el procesador de objetos y para alimentar una pluralidad de canales descodificados en el post-procesador, cuando los datos de audio codificados no contienen ningún objeto de audio y para alimentar la pluralidad de objetos descodificados y la pluralidad de canales descodificados en el procesador de objetos, cuando los datos de audio codificados comprenden canales codificados y objetos codificados.
- 9El descodificador de audio de conformidad con la reivindicación 8, caracterizado porque el postprocesador se configura para convertir la cantidad de canales de salida en una renderización binaural o en un formato de reproducción que tiene una menor cantidad de canales que la cantidad de canales de salida, en donde el descodificador de audio se configura para controlar el post-procesador de acuerdo con entrada de control derivada de la interfaz del usuario o extraída de la señal de audio codificada.
- 10El descodificador de audio de conformidad con la reivindicación 8 ó 9, caracterizado porque el procesador de objetos comprende:un renderizador de objetos para renderizar objetos descodificados con el uso de meta-datos descomprimidos;y un mezclador para mezclar objetos renderizados y canales descodificados para obtener la cantidad de canales de salida.
- 11El descodificador de audio de conformidad con cualquiera de las reivindicaciones 8 a 10, caracterizado porque el procesador de objetos comprende:una codificación de un objeto de descodificador de audio espacial para descodificar uno o más canales de transporte e información lateral paramétrica asociada que renderiza objetos codificados de audio, en donde la codificación de un objeto de descodificador de audio espacial se configura para renderizar los objetos descodificados de audio de acuerdo con la información de renderización relacionada con la colocación de los objetos de audio y para controlar el procesador de objetos para mezclar los objetos de audio renderizados y los canales de audio descodificados para obtener la cantidad de canales de salida.
- 12El descodificador de audio de conformidad con cualquiera de las reivindicaciones 8 a 10, caracterizado porque el procesador de objetos comprende una codificación de un objeto de descodificador de audio espacial para descodificar uno o más canales de transporte e información lateral paramétrica asociada que renderiza objetos codificados de audio y canales de audio codificados, en donde la codificación de un objeto de descodificador de audio espacial se configura para descodificar los objetos codificados de audio y los canales de audio codificados con el uso de uno o más canales de transporte y la información lateral paramétrica y en donde el procesador de objetos se configura para renderizar la pluralidad de objetos de audio con el uso de meta-datos comprimidos y para descodificar los canales y se mezcla con los objetos renderizados para obtener la cantidad de canales de salida.
- 13El descodificador de audio de conformidad con cualquiera de las reivindicaciones 8 a 10, caracterizado porque el procesador de objetos comprende una codificación de un objeto de descodificador de audio espacial para descodificar uno o más canales de transporte e información lateral paramétrica asociada que renderiza objetos codificados de audio o canales de audio codificados, en donde la codificación de un objeto de descodificador de audio espacial se configura para transcodificar la información paramétrica asociada y los meta-datos descomprimidos en información lateral paramétrica transcodificada susceptible de usarse para procesar directamente el formato de salida, y en donde el postprocesador se configura para calcular canales de audio del formato de salida con el uso de los canales de transporte codificados y la información lateral paramétrica transcodificada, o en donde la codificación de un objeto de descodificador de audio espacial se configura para directamente realizar una mezcla ascendente y procesar señales de canales para el formato de salida con el uso de los canales de transporte codificados y la información lateral paramétrica.
- 14El descodificador de audio de conformidad con cualquiera de las reivindicaciones anteriores, caracterizado porque el procesador de objetos comprende una codificación de un objeto de descodificador de audio espacial para descodificar uno o más canales de transporte emitidos por el descodificador central y datos paramétricos asociados y meta-datos descomprimidos para obtener una pluralidad de objetos de audio procesados, en donde el procesador de objetos se configura en forma adicional para procesar objetos descodificados emitidos por el descodificador central; en donde el procesador de objetos se configura en forma adicional para mezclar objetos descodificados renderizados con canales descodificados, en donde el descodificador de audio que comprende, además, una interfaz de salida para emitir una salida del mezclador a los alta voz, en donde el post-procesador en forma adicional comprende:un renderizador binaural para renderizar los canales de salida en dos canales binaurales con el uso de funciones de transferencia relacionadas con el cabezal o respuesta de impulso binaural, y un conversor de formato para convertir los canales de salida en un formato de salida que tiene una cantidad menor de canales que los canales de salida del mezclador con el uso de información sobre una disposición de reproducción.
- 15El descodificador de audio de conformidad con cualquiera de las reivindicaciones 8 a 14, caracterizado porque la pluralidad de elementos de canal codificados o la pluralidad de objetos codificados de audio se codifican como elementos del par de canales, elementos de canal simple, elementos de baja frecuencia o elementos del canal quad, en donde un elemento de canal quad comprende cuatro canales originales u objetos, y en donde el descodificador central se configura para descodificar el elementos del par de canales, elementos de canal simple, elementos de baja frecuencia o elementos del canal quad de acuerdo con información lateral incluida en los datos de audio codificados lo que indica un elemento de par de canales, un elemento de canal simple, un elemento de baja frecuencia o un elemento de canal quad.
- 16El descodificador de audio de conformidad con cualquiera de las reivindicaciones 8 a 15, caracterizado porque el descodificador central se configura para aplicar la operación de descodificación de banda completa con el uso de una operación de llenado de ruido sin una operación de replicación de banda espectral.
- 17El descodificador de audio de conformidad con la reivindicación 14, caracterizado porque los elementos que comprenden el renderizador binaural, el conversor de formato, el mezclador, el descodificador de SAOC y el descodificador central y el procesador de objetos operan en un dominio de banco de filtros en espejo con cuadratura (QMF) y en donde los datos del dominio de filtro en espejo con cuadratura se transmite de uno de los elementos a otro de los elementos sin ningún banco de filtros de síntesis y el procesamiento posterior del banco de filtro de análisis.
- 18El descodificador de audio de conformidad con cualquiera de las reivindicaciones 8 a 17, caracterizado porque el post-procesador se configura con canales para mezcla descendente emitidos por el procesador de objetos a un formato que tiene tres o más canales y que tienen menos canales que la cantidad de canales de salida del procesador de objetos para obtener una mezcla descendente intermedio, y para renderizar binauralmente los canales de la mezcla descendente intermedio en una señal de salida binaural de dos canales.
- 19El descodificador de audio de conformidad con cualquiera de las reivindicaciones 8 a 15, caracterizado porque el post-procesador comprende:un dispositivo de control para mezcla descendente para aplicar una matriz de mezcla descendente;y un controlador para determinar una matriz específica de mezcla descendente con el uso de información sobre una configuración de canal de una salida del procesador de objetos e información sobre una disposición de reproducción pretendida.
- 20El descodificador de audio de conformidad con cualquiera de las reivindicaciones 8 a 19, caracterizado porque el descodificador central o el procesador de objetos son susceptibles de controlarse, y en el cual el postprocesador se configura para controlar el descodificador central o el procesador de objetos de acuerdo con información sobre el formato de salida de modo tal que una renderización que incurre en procesamiento de eliminación de correlación de objetos o canales que no ocurren como canales separados en el formato de salida se reduce o elimina, o de modo tal que para los objetos o canales no ocurren como los canales separados en el formato de salida, las operaciones de mezcla ascendente o descodificación se realizan como si los objetos o canales ocurrieran como canales separados en el formato de salida, excepto que cualquier procesamiento de eliminación de correlación para los objetos o los canales que no ocurren como los canales separados en el formato de salida se desactivan.
- 21El descodificador de audio de conformidad con cualquiera de las reivindicaciones 8 a 20, caracterizado porque el descodificador central se configura para realizar descodificación por transformación y una replicación de banda espectral que descodifica un elemento de canal simple, y para realizar descodificación por transformación, descodificación estéreo paramétrica y descodificación con reproducción de banda espectral para elementos del par de canales y elementos del canal quad.
- 22Un método para codificación de datos de entrada de audio para obtener datos de salida de audio caracterizado porque comprende:recibir una pluralidad de canales de audio, una pluralidad de objetos de audio y meta-datos relacionados con uno o más de la pluralidad de objetos de audio;mezclar la pluralidad de objetos y la pluralidad de canales para obtener una pluralidad de canales mezclados previamente, cada canal mezclado previamente comprende datos de audio de un canal y datos de audio de por lo menos un objeto;codificación central de datos de entrada de codificación central;y comprimir los meta-datos relacionados con uno o más de la pluralidad de objetos de audio, en donde el método de codificación de audio opera en dos modos de un grupo de dos o más modos que comprende un primer modo, en el cual la codificación central codifica la pluralidad de canales de audio y la pluralidad de objetos de audio recibidos como datos de entrada de codificación central, y un segundo modo, en el cual la codificación central recibe, como datos de entrada de codificación central, la pluralidad de canales mezclados previamente generados por el mezclado y la codificación central de la pluralidad de canales mezclados previamente.
- 23Un método de descodificación de datos de audio codificados, caracterizado porque comprende:recibir los datos de audio codificados, los datos de audio codificados que comprende una pluralidad de canales codificados o una pluralidad de objetos codificados o metadatos comprimidos relacionados con la pluralidad de objetos;descodificar centralmente la pluralidad de canales codificados y la pluralidad de objetos codificados;descomprimir los meta-datos comprimidos, procesar la pluralidad de objetos descodificados con el uso de metadatos comprimidos para obtener una cantidad de canales de salida que comprende datos de audio a partir de los objetos y los canales descodificados;y convertir la cantidad de canales de salida en un formato de salida, en donde, en el método de descodificación de audio, el procesamiento la pluralidad de objetos descodificados se traspasa y una pluralidad de canales descodificados se alimenta en el post-procesamiento, cuando los datos de audio codificados no contienen ningún objeto de audio y la pluralidad de objetos descodificados y la pluralidad de canales descodificados se colocan en el procesamiento la pluralidad de objetos descodificados, cuando los datos de audio codificados comprenden canales codificados y obj etos codificados. 24 . Un programa para computadoras para realizar, cuando se ejecuta en una computadora o un procesador, el método de conformidad con la reivindicación 22 o la reivindicación 23.
Independent claims23
153 paragraphs in 8 sections, as filed
(54) Title: CONCEPT FOR AUDIO CODING AND DECODING FOR AUDIO CHANNELS AND AUDIO OBJECTS.
(54) Title: CONCEPT FOR AUDIO ENCODING AND DECODING FOR AUDIO CHANNELS AND AUDIO OBJECTS.
(57) Summary
Audio encoder to encode audio input data (101) to obtain audio output data (501) comprises an output interface (100) to receive a plurality of audio channels, a plurality of audio objects and meta-data related to one or more of the plurality of audio objects; a mixer (200) for mixing the plurality of objects and the plurality of channels to obtain a plurality of pre-mixed channels, each pre-mixed channel comprising audio data from one channel and audio data from at least one object; a central encoder (300) for centrally encoding input data from the central encoder; and a metadata compressor (400) for compressing the metadata related to one or more of the plurality of audio objects, wherein the audio encoder is configured to operate in at least one mode from the group of two modes. comprising a first mode, in which the central encoder is configured to encode the plurality of audio channels and the plurality of audio objects received by the input interface as input data from the central encoder, and a second mode, wherein the core encoder (300) is configured to receive, as the core encoder input data, the plurality of previously mixed channels generated by the mixer (200).
(57) Abstract
Audio encoder for encoding audio input data (101) to obtain audio output data (501) comprises an input interface (100) for receiving a plurality of audio channels, a plurality of audio objects and metadata related to one or more of the plurality of audio objects; a mixer (200) for mixing the plurality of objects and the plurality of channels to obtain a plurality of premixed channels, each pre-mixed channel comprising audio data of a channel and audio data of at least one object; a core encoder (300) for core encoding core encoder input data; and a metadata compressor (400) for compressing the metadata related to the one or more of the plurality of audio objects, where the audio encoder is configured to operate in at least one mode of the group of two modes comprising a first mode, in which the core encoder is configured to encode the plurality of audio channels and the plurality of audio objects received by the input interface as core encoder input data, and a second mode, in which the core encoder (300) is configured for receiving, as the core encoder input data, the plurality of pre-mixed channels generated by the mixer (200).
CONCEPT FOR AUDIO CODING AND DECODING FOR
AUDIO CHANNELS AND AUDIO OBJECTS
FIELD OF THE INVENTION
The present invention relates to audio encoding / decoding and, in particular, spatial audio encoding and encoding of a spatial audio object.
BACKGROUND OF THE INVENTION
Spatial audio encoding tools are well known in the art and are, for example, standardized on the MPEG surround standard. Spatial audio encoding begins with original input channels such as five or seven channels that are identified by their placement in a playback configuration, i.e. a left channel, a center channel, a right channel, a left surround channel, a right surround channel and a low frequency power channel. A spatial audio encoder normally derives one or more downmix channels from the original channels, and additionally derives parametric data related to spatial signals such as inter-channel level differences in channel coherence values, phase differences between channels, time differences between channels, etc. One or more downmix channels are transmitted along with the parametric side information indicating spatial signals to a spatial audio decoder that decodes the downmix channel and associated parametric data in order to finally obtain output channels that are a rough version of the original input channels. The placement of the channels in the output configuration is normally fixed and is, for example, a 5.1 format, a 7.1 format, etc.
Additionally, the spatial audio object encoding tools are well known in the art and standardized on the MPEG SAOC standard (SAOC = encoding a spatial audio object). In contrast to spatial audio encoding that starts on original channels, encoding a spatial audio object begins with audio objects that are not automatically dedicated for a particular render rendering setting. Instead, the placement of the audio objects in the playback scene is flexible and can be determined by the user by entering certain rendering information into an encoding of a spatial audio decoder object. Alternatively or additionally, the rendering information, that is, the information at whose position in the playback configuration a certain audio object should normally be placed over time can be transmitted as additional side information or meta-data. In order to obtain a certain data compression, a number of audio objects are encoded by means of a SAOC encoder that calculates, from the input objects, one or more transport channels by performing downmixing of objects according to certain information from the downmix process. Additionally, the SAOC encoder calculates parametric lateral information representing signals between objects such as object level differences (OLD), object coherence values, etc. As it happens in SAC
Spatial audio encoding), the parametric data between objects is calculated for individual time / frequency mosaics, i.e. for a given frame of the audio signal comprising, for example, 1024 or 2048 samples, 24, 32, or 64, etc., frequency bands are considered in such a way that, in the end, there are parametric data for each frame and each frequency band. As an example, when an audio piece has frames and when each frame is sub-divided into 32 frequency bands, then the amount of time / frequency tiles is 640.
Until now, there is no flexible technology that combines channel coding on the one hand and object coding on the other hand, so that acceptable audio qualities are obtained at low bit transfers.
It is an objective of the present invention to provide an improved concept for audio encoding and audio decoding.
This objective is achieved by an audio decoder of claim 1, an audio decoder of claim 8, an audio encoding method of claim 22, an audio decoding method of claim 23 or a computer program for claim 24.
BRIEF DESCRIPTION OF THE INVENTION
The present invention is based on the finding that, for an optimal system that is flexible on the one hand and provides good compression efficiency with good audio quality on the other hand is achieved by the combination of spatial audio encoding, that is, channel-based audio encoding with encoding of a spatial audio object, ie, based on object encoding. In particular, providing a mixer to mix objects and channels already on the encoder side provides good flexibility, particularly for low bit-transfer applications, since any object transmission may then be unnecessary or the amount of objects to be transmitted can be reduced. On the other hand, flexibility is required so that the audio encoder can be controlled in two different ways, i.e. in the way in which objects are mixed with channels before being encoded, while in the other mode the data of objects on the one hand and channel data on the other hand are encoded directly to the core without any mixing between them.
This ensures that the user can either separate the processed objects and channels on the encoder side so that full flexibility is available on the decoder side but at the price of enhanced bit transfer. On the other hand, when the bit transfer requirements are more stringent, then the present invention already allows to perform a pre-presentation / mixing on the encoder side, i.e. that some or all of the audio objects are already mixed with the channels in such a way that the central encoder only encodes channel data and any bits required to transmit object audio data either in the form of a downmix or in the form Data between parametric objects are not required.
On the decoder side, the user again has high flexibility due to the fact that the same audio decoder allows operation in two different modes, i.e. the first mode where encoding of individual or separate objects and channels takes place and the decoder It has complete flexibility to process objects and mix with channel data. On the other hand, when a preset / mix has already been developed on the encoder side, the decoder is configured to perform post-processing without processing any intermediate objects. On the other hand, post-processing can also be applied to the data in the other mode, that is, when object processing / mixing takes place on the decoder side. In this way, the present invention allows a framework of processing tasks which allows a great reuse of resources not only on the encoder side but also on the decoder side. Post-processing may refer to downmixing and binarization, or any other processing to obtain a final channel scenario such as an intended playback arrangement.
Additionally, in the case of very low bit transfer requirements, the present invention provides the user with sufficient flexibility to react to low bit transfer requirements, i.e. by pre-display on the encoder side in such a way that, for the price of some flexibility, however very good audio quality is obtained from the decoder side is obtained due to the fact that the bits that have been saved by no longer providing any encoder object data to the decoder can be used to better encode channel data such as by Finer quantization of channel data or by other means to improve quality or to reduce encoding loss when sufficient bits are available.
In a preferred embodiment of the present invention, the encoder additionally comprises an SAOC encoder and additionally allows not only to encode the input of objects in the encoder but also to encode channel data by SAOC in order to obtain good quality audio even at a lower bit rate. Furthermore, the modalities of the present invention allow the functionality of a post processing that includes a binaural renderer and / or a format converter. Additionally, it is preferred that full decoder-side processing already take place for a certain high amount of high voice such as a 22 or 32 channel high voice configuration. However, then the format converter, for example, determines that only at a 5.1 output, that is, an output is required for a playback arrangement that has less than the maximum number of channels, then it is preferred that the format converter controls both the USAC decoder and the SAOC decoder or both devices to restrict the central decoding operation and the SAOC decoding operation such that any channel that is nevertheless ultimately mixed in a format conversion does not result in decoding. Typically, upmixing channel generation requires correlation removal processing, and each correlation removal processing introduces a certain level of artifacts. Therefore, by controlling the central decoder and / or SAOC decoder by the finally required output format, a lot of additional correlation removal processing is saved when compared to a situation when this interaction does not exist, which it not only produces improved audio quality but also produces reduced decoder complexity and ultimately low power consumption, which is particularly useful for mobile devices that encompass the encoder of the invention or the decoder of the invention. The encoders / decoders of the invention, however, can not only be inserted into mobile devices such as mobile phones, smart phones, notebook computers or navigation devices, but can also be used in desktop computers or other non-mobile devices .
The previous implementation, that is, not generating some channels, may not be optimal, since some information may be lost (such as the level difference between the channels that will be downmixed). This level difference information may not be critical, but it may produce a different downmix output signal, if the downmix applies different downmix increases to channels that are upmixed. An improved solution only turns off correlation removal in the upmix, but still generates all the upmix channels with correct level differences (as signaled by the parametric SAC). The second solution produces better audio quality, but the first solution produces more complexity reduction.
BRIEF DESCRIPTION OF THE FIGURES
The preferred modalities are discussed below with respect to the attached figures, in which:
FIG. 1 illustrates a first embodiment of a Mode 1 encoder: channel / individual object encoding; and Mode 2: mixing channels and rendered objects.
<td></td><td>FIG.</td><td>2 illustrates</td><td>a first</td><td>modality</td><td>of</td><td>a</td>
<td colspan="2">decoder</td><td>Mode 1: the</td><td>processor</td><td>of objects</td><td>not</td><td>I know</td>
<td>deviation</td><td>(channels</td><td>and objects</td><td>received)</td><td>; and Mode</td><td> 2:</td><td>the</td>
object processor is bypassed (received (pre-rendered) channels only).
<td>The</td><td>FIG. 3</td><td>illustrates</td><td>A second</td><td>modality of</td><td>a</td>
<td>encoder.</td><td></td><td></td><td></td><td></td><td></td>
<td>The</td><td>FIG. 4</td><td>illustrates</td><td>A second</td><td>modality of</td><td>a</td>
<td colspan="2">decoder; in</td><td>where I know</td><td>exhibits a</td><td>processing</td><td>of</td>
<td>QMF domain</td><td>direct</td><td colspan="3">in binaural renderer, converter</td><td>of</td>
Format; STOC decoder, USAC mode decoder
SBR.
<td>The</td><td>FIG.</td><td> 5</td><td>illustrates</td><td>a</td><td>third</td><td>modality</td><td>of</td><td>a</td>
<td>encoder.</td><td></td><td></td><td></td><td></td><td></td><td></td><td></td><td></td>
<td>The</td><td>FIG.</td><td> 6</td><td>illustrates</td><td>a</td><td>third</td><td>modality</td><td>of</td><td>a</td>
<td colspan="2">decoder.</td><td></td><td></td><td></td><td></td><td></td><td></td><td></td>
<td>The</td><td>FIG.</td><td> 7</td><td>illustrates</td><td>a</td><td>map it</td><td>which indicates</td><td colspan="2">modes</td>
singles in which encoders / decoders in accordance with the embodiments of the present invention can be operated.
FIG. 8 illustrates a specific implementation of the format converter.
FIG. 9 illustrates a specific implementation of the binaural converter.
FIG. 10 illustrates a specific implementation of the central decoder.
FIG. 11 illustrates a specific implementation of an encoder for processing a quad channel element (QCE) and the corresponding QCE decoder.
DETAILED DESCRIPTION OF THE INVENTION
Figure 1 illustrates an encoder in accordance with an embodiment of the present invention. The encoder is configured to encode audio input data 101 to obtain audio output data 501. The encoder comprises an output interface for receiving a plurality of audio channels indicated by CH and a plurality of audio objects indicated by OBJ. Additionally, as illustrated in Figure 1, input interface 100 additionally receives metadata related to one or more of the plurality of OBJ audio objects. Additionally, the encoder comprises a mixer 200 to mix the plurality of objects and the plurality of channels to obtain a plurality of pre-mixed channels, wherein each pre-mixed channel comprises audio data of one channel and audio data of at least minus one object.
Additionally, the encoder comprises a core encoder 300 for centrally encoding input data from the core encoder, a metadata compressor 400 for compressing the metadata related to one or more of the plurality of audio objects. Additionally, the encoder may comprise a mode controller 600 to control the mixer, the core encoder, and / or an output interface 500 in one of many modes of operation, where in the first mode, the core encoder is configured to encoding the plurality of audio channels and the plurality of audio objects received by the input interface 100 without any interaction by the mixer, that is, without any mixing done by the mixer 200. In a second mode, however, in which mixer 200 was active, the central encoder encodes the plurality of mixed channels, that is, the output generated by block 200. In the latter case, it is preferred to no longer encode any data of objects. Instead, the metadata indicates audio object positions already used by mixer 200 to process the objects on the channels as indicated by the metadata. In other words, mixer 200 uses metadata related to the plurality of audio objects to pre-process the audio objects and then the pre-processed audio objects are mixed with the channels to obtain mixed channels at the output of the mixer . In this mode, any object may not necessarily be transmitted and this also applies to compressed meta-data as output for block 400. However, if not all objects entering interface 100 are mixed but only a certain number of objects are mixed, then not only the previously unmixed objects and associated metadata are nonetheless transmitted to the central encoder 300 or the compressor. of meta-data 400, respectively.
Figure 3 illustrates a further embodiment of an encoder that additionally comprises an SAOC 800 encoder. The SAOC 800 encoder is configured to generate one or more transport channels and parametric data from the encoder input data. of spatial audio objects. As illustrated in Figure 3, the input data from the spatial audio object encoder are objects that have not been processed by the pre-processor / mixer. Alternatively, whenever the preprocessor / mixer has been passed on as in mode one where an individual channel / object encoding is active, all object inputs on input interface 100 are encoded by means of the SAOC 800 encoder.
Additionally, as illustrated in Figure
3, the core encoder 300 is preferably implemented as a USAC encoder, that is, as an encoder as defined and standardized in the MPEG-USAC standard (USAC = Unified Speech and Audio Coding). The broadcast of the complete encoder illustrated in Figure 3 is an MPEG 4 data stream having container-like structures for individual data types. Additionally, the metadata is indicated as OAM data and the metadata compressor 400 in Figure 1 corresponds to the OAM 400 encoder to obtain compressed OAM data that is entered into the USAC 300 encoder which, as can be seen in the Figure 3 additionally comprises the output interface to obtain the MP4 replay data stream which not only has encoded object / channel data but also has the compressed OAM data.
Figure 5 illustrates a further embodiment of the encoder, where in contrast to Figure 3, the SAOC encoder can be configured interchangeably to encode, with the SAOC encoding algorithm, the channels provided in the pre-render / mixer 200 that are not It is active in this mode or, alternatively, to SAOC encode channels without pre-rendered objects. Thus, in Figure 5, the SAOC 800 encoder can operate on three different kinds of input data, that is, channels without any pre-processed objects, pre-rendered channels and objects, or objects only. Additionally, it is preferred to provide an additional OAM decoder 420 in Figure 5 such that the SAOC 800 encoder uses the same decoder-side data, i.e. data obtained by compression, for processing. loss instead of the original OAM data.
The encoder in Figure 5 can operate in several individual modes.
In addition to the first and second modes as discussed in the context of Figure 1, the encoder of Figure 5 may additionally operate in a third mode in which the central encoder generates one or more transport channels from the objects. individual when the pre-processor / mixer 200 was not active. Alternatively or additionally, in this third mode the SAOC 800 encoder can generate one or more alternative or additional transport channels from the original channels, i.e. again when the preprocessor / mixer 200 corresponding to mixer 200 in Figure 1 does not was active.
Finally, the SAOC 800 encoder can encode, when the encoder is configured in the fourth mode, the channels plus objects previously rendered as generated by the pre-processor / mixer. Thus, in the fourth mode, the applications with lower bit transfer will provide good quality due to the fact that the channels and objects have been completely transformed into individual SAOC transport channels and the associated lateral information as indicated in Figs. 3 and 5 as SAOC-SI and, additionally, any uncompressed meta-data does not have to be transmitted in this fourth mode.
Figure 2 illustrates a decoder according
<td>with a</td><td colspan="2">modality of the present</td><td colspan="2">invention.</td><td>The</td>
<td colspan="2">decoder receives, as input,</td><td>the</td><td>data</td><td colspan="2">audio</td>
<td>coded,</td><td>that is, the 501 data from</td><td>the</td><td>Figure</td><td> 1.</td><td></td>
<td>The</td><td>decoder comprises</td><td>a</td><td colspan="2">decompressor</td><td>of</td>
<td>metadata</td><td>1400, a decoder</td><td colspan="2">central</td><td> 1300,</td><td>a</td>
1200 object renderer, 1600 mode controller and 1700 post processor.
Specifically, the audio encoder is configured to decode encoded audio data and the input interface is configured to receive the encoded audio data, the encoded audio data comprising a plurality of encoded channels, and the plurality of encoded and meta-objects. compressed data related to the plurality of objects in a certain mode.
Additionally, the central decoder 1300 is configured to decode the plurality of encrypted channels and the plurality of encoded objects, and additionally, the metadata decompressor is configured to decompress the compressed metadata.
Additionally, the object renderer 1200 is configured to process the plurality of objects
<td>decoded according to</td><td>are generated by the decoder</td>
<td>central 1300 with the</td><td>use of compressed meta-data to</td>
<td>get a quantity</td><td>default output channels</td>
comprising object data and decoded channels. These output channels as indicated in 1205 are then input into a 1700 post-processor. The 1700 post-processor is configured to convert the number of 1205 output channels into a certain input format which may be a binaural playback format or a high-voice playback format such as a 5.1, 7.1 playback format, etc.
Preferably, the decoder comprises a mode controller 1600 which is configured to analyze the encoded data to detect a mode indication. Therefore, the mode controller 1600 connects to the input interface 1100 in Figure 2. However, alternatively, the mode controller does not necessarily have to be there. Instead, the flexible decoder can be pre-configured by any other kind of control data such as user input or any other control. The audio decoder in Figure 2, and preferably controlled by the 1600 mode controller, is configured either to bypass the object renderer and to feed the plurality of decoded channels in the post processor 1700. This is operation in The mode
2, that is, in which only the previously rendered channels are received, that is, when mode 2 has been applied in the encoder of Figure 1. Alternatively, when mode 1 has been applied to the encoder, i.e. when the encoder has performed individual channel / object encoding, then the object renderer 1200 is not deviated, but the plurality of decoded channels and the plurality of decoded objects they are placed in the object renderer 1200 along with unzipped metadata generated by the 1400 metadata decompressor.
Preferably, the indication of whether to apply mode 1 or mode 2 includes the encoded audio data, and then the mode controller 1600 analyzes the encoded data to detect a mode indication. Mode 1 is used when the mode indication indicates that the encoded audio data comprises encoded channels and encoded objects and mode 2 is applied when the mode indication indicates that the encoded audio data does not contain any audio object, i.e. , only contain pre-processed channels obtained by mode 2 of the encoder of Figure 1.
Figure 4 illustrates a preferred embodiment compared to that of the decoder of Figure 2 and the embodiment of Figure 4 corresponds to the encoder of Figure 3. In addition to the implementation of the decoder of Figure 2, the decoder in Figure 4 comprises a SAOC 1800 decoder. Additionally, the object renderer 1200 of Figure 2 is implemented as a separate object renderer 1210 and mixer 1220 while, depending on the mode, the functionality of the object renderer 1210 may also be implemented by the SAOC 1800 decoder. .
Additionally, the post processor 1700 can be implemented as a 1710 binaural renderer or a 1720 format converter. Alternatively, a direct data stream 1205 of Figure 2 can also be implemented as illustrated by 1730. Therefore, it is preferred to perform processing in the decoder on the highest number of channels such as 22.2 or 32 in order to have flexibility and then post-process if a smaller format is required. However, when it becomes clear from the very beginning that only a small format such as a 5.1 format is required, then it is preferred, as indicated by Figure 2 or 6 by the simplified method 1727, that a certain control over the decoder SAOC and / or USAC decoder can be applied in order to avoid unnecessary upmix operations and subsequent downmix operations.
In a preferred embodiment of the present invention, the object renderer 1200 comprises the SAOC decoder 1800 and the SAOC decoder is configured to decode one or more transport channels emitted by the central decoder and associated parametric data and with the use of meta -compressed data to obtain the plurality of processed audio objects. Up to this point, the OAM output connects to mailbox 1800.
Additionally, the object renderer 1200 is configured to process decoded objects emitted by the central decoder that are not encoded on the SAOC transport channels but are individually encoded on normally elements on individual channels as indicated by the object renderer 1210. Additionally, the decoder comprises an output interface that corresponds to output 1730 to output a mixer output to the loud voices.
In a further embodiment, the object renderer 1200 comprises encoding a spatial audio decoder object 1800 to decode one or more transport channels and associated parametric side information that renders encoded audio objects or encoded audio channels, wherein the encoding of a spatial audio decoder object is configured to transcode the associated parametric information and the decompressed meta-data into transcoded parametric side information capable of being used to directly process the output format, as defined for example in a version previous SAOC. Post processor 1700 is configured to compute audio channels from the output format using the encoded transport channels and the transcoded parametric side information. The processing performed by the post processor may be similar to MPEG Envelope processing or it may be any other processing such as BCC processing and so on.
In a further embodiment, the object renderer 1200 comprises an encoding of a spatial audio decoder object 1800 configured to directly mix rendered channel signals and for the output format with the use of the decoded transport channels (by the central decoder ) and the parametric lateral information.
Additionally, and very importantly, the object renderer 1200 of Figure 2 additionally comprises mixer 1220 that receives, as input, data generated by the USAC 1300 decoder directly when there are previously rendered objects mixed with channels, that is, when mixer 200 in Figure 1 was active. Additionally, mixer 1220 receives data from the object renderer that renders objects without SAOC decoding. Additionally, the mixer receives output data from the SAOC decoder, that is, objects rendered by SAOC.
The 1220 mixer connects to the 1730 output interface, the 1710 binaural renderer, and the 1720 format converter. The 1710 binaural renderer is configured to render the output channels to two binaural channels using the head-related transfer functions or responses to binaural room impulses (BRIR). The 1720 format converter is configured to convert the output channels to an output format that has fewer channels than the mixer's 1205 output channels, and the 1720 format converter requires output layout information such as loud voice. 5.1 and others.
The decoder in Figure 6 is different from the decoder in Figure 4 in that SAOC decoder can not only render rendered objects but also rendered channels and this is the case where the encoder of Figure 5 has been used and
<td>the connection</td><td>900 among</td><td colspan="2">channels / objects previously</td>
<td>rendered</td><td>and the interface of</td><td>encoder input</td><td>of</td>
<td>SAOC 800 is</td><td>active.</td><td></td><td></td>
<td>In</td><td>additional form,</td><td>a panning stage</td><td>of</td>
vector base amplitude (VPAP) 1810 is configured to receive, from the SAOC decoder, the information on the output layout and to output a render matrix to the SAOC decoder such that the SAOC decoder can, in the end, provide rendered channels without any additional mixer operation in the 1205 high channel format, i.e. 32 high voices.
The VBAP block preferably receives the decoded OAM data to derive the render (render) matrices. More generally, with
<td>preference</td><td>requires geometric information not only from the</td>
<td>provision</td><td>output but also from the positions where</td>
<td>the signs</td><td>input should be rendered (processed) in</td>
the output arrangement. This geometric input data may be OAM data for channel position information or objects for channels that have been transmitted with the use of SAOCs.
However, if only a specific exit interface is required then the VBAP 1810 state can already provide the required rendering matrix for the exit, for example 5.1. The SAOC 1800 decoder then performs a direct render of the SAOC transport channels, the associated parametric data, and decompressed meta-data, a direct render in the required output format without any interaction from the 1220 mixer. However, when a certain mix is applied between modes, i.e. where multiple channels are encoded with SAOC but not all channels are encoded with SAOC or where multiple objects are encoded with
SAOC but not all the objects are encoded with SAOC or when only a certain amount of objects previously rendered with channels are decoded by SAOC and the remaining channels are not processed with SAOC then the mixer will unify the data of the individual input portions, i.e. , directly from the 1300 core decoder, 1210 object renderer, and SAOC 1800 decoder.
Subsequently, Figure 7 is discussed for what it indicates certain encoder / decoder modes that can be applied by the highly flexible and high quality audio encoder / decoder concept of the invention.
In accordance with the first encoding mode, mixer 200 in the encoder of Figure 1 is carried forward, and therefore the object processor in the decoder of Figure 2 is not carried over.
In the second mode, mixer 200 in Figure 1 is active and the object processor in Figure 2 is bypassed.
So, in the third encoding mode, the SAOC encoder in Figure 3 is active but only SAOC encodes the objects instead of channels or channels as output by the mixer. Therefore, mode 3 requires that, on the decoder side illustrated in Figure 4, the SAOC decoder is only active for the objects and generates rendered objects.
In a fourth encoding mode as illustrated in Figure 5, the SAOC encoder is configured for SAOC encoding of pre-displayed channels, ie the mixer is active as in the second mode. On the decoder side, SAOC decoding is performed for previously rendered objects in such a way that the object processor is passed through as in the second encoding mode.
Additionally, there is a fifth encoding mode that can be mixed by any one of modes 1 through 4. In particular, a mixing encoding mode will exist when mixer 1220 in Figure 6 receives channels directly from the USAC decoder and, in the form Additionally, it receives channels with previously rendered objects from the USAC decoder. Additionally, in this mixed encoding mode, objects are encoded directly with the use of, preferably, a single channel element of the USAC decoder. In this context, the object renderer 1210 will then render these objects decoded and send them to mixer 1220. Additionally, multiple objects are additionally encoded by a SAOC encoder such that the SAOC decoder will generate rendered objects to the mixer and / or rendered channels when there are multiple channels encoded by SAOC technology.
Each input portion of mixer 1220 can then, in exemplary form, have at least one potential to receive the number of channels such as 32 as indicated at 1205. In this way, basically, the mixer could receive 32 channels from the USAC decoder and additionally 32 mixed / pre-presented channels from the USAC decoder and additionally 32 channels from the object renderer and additionally 32 SAOC decoder channels, where each channel between blocks 1210 and 1218 on the one hand and block 1220 on the other hand has a contribution of the corresponding objects on a corresponding loud voice channel and then mixer 1220 mixes, i.e. adds individual contributions for each channel of loud voice.
In a preferred embodiment of the present invention, the encoding / decoding system relies on a USAC MPEG-D codee to encode the channel and the object signals. To increase the efficiency of encoding large numbers of objects, SAOC technology from MPEG has been adapted. Three types of renderers perform the task of rendering objects to channels, rendering channels to headphones, or rendering channels to different high-voice settings. When object signals are explicitly transmitted or parametrically encoded with the use of SAOCs, the corresponding object meta-data information is compressed and multiplexed in the encoded output data.
In one embodiment, mixer / pre-presenter 200 is used to convert an object plus channel input scene to a channel scene before encoding. Functionally, this is identical to the decoder-side mixer / object processor combination as illustrated in Figure 4 or Figure 6 and as indicated by the object processor 1200 in Figure 2. Pre-display of objects ensures a determining signal entropy at the encoder input that is basically independent of the amount of simultaneously active object signals. With object pre-presentation, no meta-data transmission of objects is required. The individual object signals are processed at the channel layout such that the encoder is configured to use. The weight of the objects for each channel is derived from the associated OAM object metadata as indicated in row 402.
As the encoder / decoder / hub for high-voice channel signals, individual object signals, object downmix signals, and pre-displayed signals, a USAC technology is preferred. It handles the encoding of large numbers of signals by creating channel and object mapping information (the semantic and geometric information of the input channel and object mapping). This mapping information describes how input channels and objects are mapped to elements of the USAC channel as illustrated in
Figure
10, ie channel pair elements (CPEs), quad channel element elements (QCEs) and the corresponding information is transmitted to the central decoder of the central encoder. All additional loads such as SAOC data or object meta-data have been passed through extension elements and considered in the encoder speed control.
Object encoding is possible in different modes, depending on the ratio / distortion requirements and the interactivity requirements for the renderer.
The following encoding of object variants is possible:
Pre-rendered Objects: Object signals are pre-rendered and mixed with channel 22.2 signals before encoding. The subsequent encoding chain sees signals from channel 22.2.
• Individual object waveforms: Objects are supplied as monophonic waveforms to the encoder. The encoder uses elements of single channel SCEs to transmit the objects in addition to the channel signals. Decoded objects are processed and mixed on the receiver side. Compressed object meta-data information is transmitted to the receiver / presenter throughout its journey.
• Parametric object waveforms: The properties of the objects and their relationship to each other are described using the SAOC parameters. The downmix of the object signals is encoded with USAC.
Parametric information is transmitted in its entire length. The number of channels for downmixing is chosen depending on the number of objects and the overall data rate. The compressed object metadata information is transmitted to the SAOC renderer.
The SAOC encoder and decoder for the object signals are based on SAOC MPEG technology. The system is capable of recreating, modifying and rendering a number of audio objects based on a smaller number of transmitted channels and additional parametric data (OLDs, IOCs (Coherence Between Objects), DMGs (Down Mix Increases)). The additional parametric data exhibits a significantly lower data transfer rate than that required to transmit all objects individually, making encoding very efficient.
The SAOC encoder takes the object / channel signals as monophonic waveforms as input and produces the parametric information (which is packaged in the Audio 3D bit rate) and the SAOC transport channels (which are encoded with the use of simple and transmitted channel elements).
The SAOC decoder reconstructs the object / channel signals of the decoded SAOC transport channels and the parametric information, and generates the output audio stream based on the output layout, the information from the decompressed object metadata. and optionally in interaction with user information.
For each object, the associated metadata that specify the geometric position and volume of the object in 3D space is efficiently encoded by quantifying the object's properties in time and space. The compressed metadata of cOAM objects is transmitted to the receiver as lateral information. The volume of the object may comprise information about a spatial degree and / or the signal level of the audio signal of this audio object.
The object renderer uses compressed object metadata to generate object waveforms according to the given playback format. Each object is rendered to certain output channels according to its metadata. The emission of this block is the result of the sum of the partial results.
If both channel-based content as well as individual / parametric objects are decoded, the channel-based waveforms and the rendered object's waveforms are mixed before outputting the resulting waveforms (or before placing them in a post module -processor like binaural renderer or loud voice renderer module).
The binaural renderer module produces a binaural downmix of the multichannel audio material, such that each input channel is rendered by a virtual sound source. Processing is conducted in the form of tables in the QMF (Quadrature Mirror Filterbank) domain.
Binarization is based on measured responses to binaural room impulses
Figure 8 illustrates a preferred embodiment of the 1720 format converter. The high-voice renderer or format converter converts between the transmitter channel configuration and the desired playback format. This format converter performs conversions up to a smaller number of output channels, that is, it creates downstream mixes. Up to this point, a downmix device 1722 that preferably operates in the QMF domain receives output signals from mixer 1205 and output high voice signals. Preferably, a controller 1724 is provided to configure downmix 1722 that receives, as a control input, a mixer output arrangement, i.e., the arrangement for which data 1205 is determined and a desired playback arrangement is normally put in the 1720 format conversion block illustrated in Figure 6. Based on this information, controller 1724 preferably automatically generates downmix matrices optimized for the given combination of input and output formats and applies these matrices to the downmix device block 1722 in the downmix process. The format converter allows standard loud voice settings as well as random settings with non-standard loud voice positions.
As illustrated in the context of Figure 6, the SAOC decoder is designed to make the channel layout predefined such as 22.2 with a subsequent format conversion to the desired playback layout. Alternatively, however, the SAOC decoder is implemented to support low power mode where the SAOC decoder is configured to decode to the output arrangement directly without subsequent format conversion. In this implementation, the SAOC 1800 decoder directly produces the loud voice signal such as 5.1 loud voice signals, and the SAOC 1800 decoder requires the output layout information and the rendering matrix such that Pan Vector Base Amplitude or any other kind of processor to generate downmix information can operate.
Figure 9 illustrates a further embodiment of the binaural renderer 1710 of Figure 6. Specifically, for mobile devices binaural rendering is required for headphones attached to such mobile devices or for loud voices directly attached to normally small mobile devices.
For such mobile devices, there may be limitations to limit the complexity of the decoder rendering.
In addition to omitting correlation removal in such processing scenarios, downmixing with the use of downmixing device 1712 is preferred in the first instance to an intermediate downmix, i.e., to a smaller number of output channels which then produces a less input channel for the 1714 binaural converter. As an example, the channel material 22.2 is downmixed by downmix device 1712 to an intermediate downmix 5.1 or, alternatively, the downmix intermediate is calculated directly by the SAOC 1800 decoder of Figure 6 in a simplified method mode class. So binaural rendering only needs to apply ten HRTFs (Head Related Transfer Functions) or BRIR functions to render the five individual channels in different positions in contrast to applying 44 HRTF for BRIR functions if the 22.2 input channels have already been rendered directly.
Specifically, the convolution operations required for binaural rendering require a large amount of processing power, and therefore reducing this processing power while still obtaining acceptable audio quality is particularly useful for mobile devices.
Preferably, the simplified method as illustrated by means of control line 1727 comprises controlling decoder 1300 to decode to a smaller number of channels, i.e. skip the entire OTT processing block in the decoder or a format that is converted to a smaller number of channels and, as illustrated in Figure 9, a binaural rendering is performed for the smaller number of channels. The same processing can be applied not only for binaural processing but also for a format conversion as illustrated by line 1727 in Figure 6.
In a further embodiment, efficient interface generation between processing blocks is required. In particular in Figure 6, the audio channel path between the different processing blocks is rendered. The 1710 binaural renderer, 1720 format converter, SAOC 1800 decoder, and USAC 1300 decoder, in case SBR (Spectral Band Replication) is applied, all operate in a QMF or hybrid QMF domain. According to one embodiment, all of these processing blocks provide a hybrid QMF or QMF interface to allow the passage of audio signals to each other in the QMF domain in an efficient way. Additionally, it is preferred to implement the mixer module and the object processor module to work in the QMF or hybrid QMF domain as well. As a consequence, separate QMF or hybrid QMF synthesis and analysis stages can be avoided which produces considerable complexity savings and then only one final QMF synthesis stage is required to generate the high voices indicated in 1730 or to generate the binaural data in the broadcast of block 1710 or to generate the arrangement of outgoing high-voice signals in the broadcast of block 1720.
Subsequently, reference is made to Figure 11 in order to explain the Quad Channel Elements (QCE). In contrast to a channel pair element as defined in the US AC-MPEG standard, a quad channel element requires four input channels 90 and produces an encoded QCE element 91. In one modality, a hierarchy of two 2-1-2 Mode MPEG Surround boxes or two TTO boxes (Two to One, two to one) and additional joint stereo encoding tools (eg MS-Stereo) as defined in MPEG USAC or MPEG envelope are provided and the QCE element not only comprises two stereo-coded downmix channels together and optionally two stereo-coded residual channels together and, additionally, parametric data derived from, for example, two TTO boxes. On the decoder side, a structure is applied where the joint stereo decoding of the two channels for downmixing and optionally two channels
<td>residual</td><td>I know</td><td>apply</td><td>and in one</td><td>second</td><td>stage</td><td>with two OTTs</td>
<td>lockers</td><td>the</td><td>mixture</td><td colspan="2">descending and</td><td>channels</td><td>residual</td>
<td>optional</td><td>I know</td><td>subdue</td><td>to mix</td><td colspan="2">ascending to</td><td>the four of them</td>
output channels. However, alternative processing operations for a QCE encoder can be applied in place of the hierarchical operation. Thus, in addition to incorporating a two-channel group co-channel encoding, the core encoder / decoder additionally uses a four-channel group co-channel encoding.
Additionally, it is preferred to perform an enhanced noise fill procedure to allow uncommitted full band (18 kHz) encoding at 1200 kbps.
The encoder has been operated in a 'bit-shifted constant transfer' mode, using a maximum of 6144 bits per channel as the ratio buffer for dynamic data.
All additional payloads such as SAOC data or object meta-data have been passed through the extension elements and considered in the encoder ratio control.
<td>With the</td><td>end</td><td>to take advantage</td><td>of</td><td>the</td>
<td>functionalities of</td><td>SAOC</td><td>also for the content</td><td>of</td><td>Audio</td>
<td colspan="2">3D, have been implemented</td><td colspan="2">the following extensions</td><td>for</td>
MPEG SAOC:
• Downward mix to arbitrary number of SAOC transport channels.
• Enhanced presentation for high-volume high-voice output configurations (up to 22.2).
The binaural renderer module produces a binaural downmix of the multichannel audio material, such that each input channel (not including the LFEs) is rendered by a virtual sound source. Processing is conducted in the form of tables in the QMF domain.
Binarization is based on measured responses to binaural room impulses. Direct sound and early reflections are printed on the audio material by means of a convolutional approach in a pseudo FFT domain with the use of a fast convolution above the QMF domain.
Although some aspects have been described in the context of an apparatus, it is clear that these aspects furthermore represent a corresponding method description, where a block or device corresponds to a method stage or a characteristic of a method stage. Similarly, the aspects described in the context of a method step further represent a description of a corresponding block or element or feature of a corresponding apparatus. Some or all of the steps in the method may be performed by (or with the use of) a hardware device, such as a microprocessor, a programmable computer, or an electronic circuit. In some modalities, some of the most important steps of the method can be executed by said apparatus.
Depending on certain implementation requirements, the embodiments of the invention can be implemented in hardware or software. The implementation can be done with the use of a non-transient storage medium such as a digital storage medium, for example a floppy disk, a DVD, a Blu-Ray, a CD, a ROM, a PROM, and EPROM, an EEPROM or a flash memory, which has electronically readable control signals stored therein, which cooperate (or are capable of cooperating) with a programmable computer system such that the respective method is performed.
Therefore, the digital storage medium can be read by computer.
Some embodiments in accordance with the invention comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.
In general, the embodiments of the present invention can be implemented as a computer program product with a program code, the program code is operative to perform one of the methods when the computer program product is run on a computer. The program code can, for example, be stored in a machine-readable carrier.
Other modalities include the computer program to perform one of the methods described herein, stored in a machine-readable carrier.
In other words, one embodiment of the method of the invention is therefore a computer program that has program code to perform one of the methods described herein when the computer program is run on a computer.
A further embodiment of the method of the invention is therefore a data carrier (or a digital storage medium, or a computer-readable medium) comprising, recorded there, the computer program for performing one of the methods described herein. The data bearer, the digital storage medium or the recording medium are usually tangible and / or non-transient.
A further embodiment of the method of the invention is, therefore, a data stream or a sequence of signals that render the computer program to perform one of the methods described herein.
The data stream or signal sequence can, for example, be configured to be transferred via a data communication connection, for example, over the Internet.
A further embodiment comprises a processing means, eg, a computer or a programmable logic device, configured to, or adapted to, perform one of the methods described herein.
An additional embodiment comprises a computer that has the computer program installed to perform one of the methods described herein.
A further embodiment in accordance with the invention comprises an apparatus or system configured to transfer (for example, electronically or optically) a computer program to perform one of the methods described herein to a receiver. The receiver may, for example, be a computer, a mobile device, a memory device, or the like. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver.
In some embodiments, a programmable logic device (eg, field programmable access ordering) can be used to perform all or some of the functionality of the methods described herein. In some embodiments, a field programmable access ordering may cooperate with a microprocessor in order to perform one of the methods described herein. In general, the preferred methods are performed by any hardware appliance.
The embodiments described above are merely illustrative for the principles of the present invention. It is understood that modifications and variations of the provisions and details described herein will be obvious to others with experience in the art. It is the intention, therefore, to be limited only by the scope of the patent pending claims and not by the specific details represented by way of description and explanation of the modalities of this document.
Contents8
10 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10
190 members in 21 offices
Priority claims3
| Document | Office | Kind | Date |
|---|---|---|---|
| 131773780 | European Patent Office (EPO) | – | |
| 13177378 | European Patent Office (EPO) | A | |
| 2014065289 | European Patent Office (EPO) | W |
Members190
| Document | Office | Kind | |
|---|---|---|---|
| EP2830045A1 | European Patent Office (EPO) | A1 | |
| EP2830047A1 | European Patent Office (EPO) | A1 | |
| EP2830048A1 | European Patent Office (EPO) | A1 | |
| EP2830049A1 | European Patent Office (EPO) | A1 | |
| EP2830050A1 | European Patent Office (EPO) | A1 | |
| CA2918148A1 | Canada | A1 | |
| CA2918166A1 | Canada | A1 | |
| CA2918529A1 | Canada | A1 | |
| CA2918860A1 | Canada | A1 | |
| CA2918869A1 | Canada | A1 | |
| WO2015010996A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2015010998A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2015010999A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2015011000A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2015011024A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW201519216A | Taiwan Province of China | A | |
| TW201519217A | Taiwan Province of China | A | |
| TW201523591A | Taiwan Province of China | A | |
| TW201528251A | Taiwan Province of China | A | |
| TW201528252A | Taiwan Province of China | A | |
| AR096997A1 | Argentina | A1 | |
| AR096998A1 | Argentina | A1 | |
| AR096999A1 | Argentina | A1 | |
| AR097000A1 | Argentina | A1 | |
| AR097003A1 | Argentina | A1 | |
| AU2014295267A1 | Australia | A1 | |
| SG11201600396QA | Singapore | A | |
| SG11201600460UA | Singapore | A | |
| SG11201600469TA | Singapore | A | |
| SG11201600471YA | Singapore | A | |
| SG11201600476RA | Singapore | A | |
| AU2014295216A1 | Australia | A1 | |
| AU2014295269A1 | Australia | A1 | |
| AU2014295270A1 | Australia | A1 | |
| AU2014295271A1 | Australia | A1 | |
| KR20160033769A | Republic of Korea | A | |
| KR20160033775A | Republic of Korea | A | |
| KR20160036585A | Republic of Korea | A | |
| CN105474309A | China | A | |
| CN105474310A | China | A | |
| KR20160041941A | Republic of Korea | A | |
| MX2016000851A | Mexico | A | |
| MX2016000907A | Mexico | A | |
| MX2016000908A | Mexico | A | |
| MX2016000910AThis record | Mexico | A | |
| MX2016000914A | Mexico | A | |
| US2016133263A1 | United States of America | A1 | |
| US2016133267A1 | United States of America | A1 | |
| KR20160053910A | Republic of Korea | A | |
| CN105593929A | China | A | |
| CN105593930A | China | A | |
| US2016142846A1 | United States of America | A1 | |
| US2016142847A1 | United States of America | A1 | |
| US2016142850A1 | United States of America | A1 | |
| CN105612577A | China | A | |
| EP3025329A1 | European Patent Office (EPO) | A1 | |
| EP3025330A1 | European Patent Office (EPO) | A1 | |
| EP3025332A1 | European Patent Office (EPO) | A1 | |
| EP3025333A1 | European Patent Office (EPO) | A1 | |
| EP3025335A1 | European Patent Office (EPO) | A1 | |
| JP2016525714A | Japan | A | |
| JP2016525715A | Japan | A | |
| JP2016527558A | Japan | A | |
| JP2016528541A | Japan | A | |
| JP2016528542A | Japan | A | |
| AU2014295270B2 | Australia | B2 | |
| TWI560699B | Taiwan Province of China | B | |
| TWI560700B | Taiwan Province of China | B | |
| TWI560701B | Taiwan Province of China | B | |
| TWI560703B | Taiwan Province of China | B | |
| TWI566235B | Taiwan Province of China | B | |
| US9578435B2 | United States of America | B2 | |
| AU2014295269B2 | Australia | B2 | |
| US9699584B2 | United States of America | B2 | |
| BR112016001139A2 | Brazil | A2 | |
| BR112016001140A2 | Brazil | A2 | |
| BR112016001143A2 | Brazil | A2 | |
| BR112016001243A2 | Brazil | A2 | |
| BR112016001244A2 | Brazil | A2 | |
| US9743210B2 | United States of America | B2 | |
| RU2016105469A | Russian Federation | A | |
| RU2016105518A | Russian Federation | A | |
| RU2016105472A | Russian Federation | A | |
| RU2016105682A | Russian Federation | A | |
| RU2016105691A | Russian Federation | A | |
| ZA201601044B | South Africa | B | |
| ZA201601076B | South Africa | B | |
| KR101774796B1 | Republic of Korea | B1 | |
| HK1225505A | Hong Kong, China | A | |
| HK1225505A1 | Hong Kong, China | A1 | |
| US2017272883A1 | United States of America | A1 | |
| AU2014295267B2 | Australia | B2 | |
| US9788136B2 | United States of America | B2 | |
| AU2014295271B2 | Australia | B2 | |
| AU2014295216B2 | Australia | B2 | |
| US2017311106A1 | United States of America | A1 | |
| JP6239109B2 | Japan | B2 | |
| JP6239110B2 | Japan | B2 | |
| ZA201601045B | South Africa | B | |
| US2017366911A1 | United States of America | A1 |
1 legal event, as the office reported them to INPADOC
Events
| Event | Code | |
|---|---|---|
| Grant or registrationFG | FG |
Numbers
- Publication
- 2016000910
- Application
- 910
Titles2
- Spanish
- CONCEPTO PARA CODIFICACION Y DESCODIFICACION DE AUDIO PARA CANALES DE AUDIO Y OBJETOS DE AUDIO.
- English
- CONCEPT FOR AUDIO ENCODING AND DECODING FOR AUDIO CHANNELS AND AUDIO OBJECTS.
Classification
- CPC, 8
- G10L19/008
- G10L19/20
- H04S3/008
- G10L19/18
- G10L19/22
- G10L19/028
- H04S2400/03
- H04S2400/11
- IPC, 1
- G10L19 008