Audio decoder and decoding method using efficient downmixing
Abstract
A method that operates an audio decoder (200) to decode audio data that includes encoded blocks of Nn audio data channels to form decoded audio data including Mm decoded audio channels, M³, n being the number of channels of low frequency effects in the encoded audio data, and m being the number of channels of low frequency effects in the decoded audio data, the method comprising: - accept audio data that includes blocks of Nn channels of encoded audio data that have been encoded by an encoding method, including the coding method transform Nn channels of digital audio data and form and package exponent and mantissa data frequency domain; and - decode the accepted audio data, including decoding: - unpack and decode (403) the exponent data and frequency domain mantissa, - determine transform coefficients (605) from the exponent data and frequency domain mantissa unpacked and decoded, - submit a transformation Reverse (607) frequency domain data and apply additional processing to determine sampled audio data, and - mixing downwardly in the time domain (613) at least some blocks of the sampled audio data determined according to downstream mixing data for the case where M <N; wherein the downstream mixing in the time domain includes (1100) checking whether the downstream mixing data has changed over time with respect to previously used downstream mixing data, and, if they have changed, applying cross-damping to determine data of downstream mixing down and cross-mixing in the time domain according to cross-attenuated downward mixing data, and, if they have not changed, directly perform a downward mixing in the time domain according to the downstream mixing data.

Term
4.4 yearsto projected expiry
Projected expiry 17 February 2031, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
12 claims: 6 independent, 6 dependent
- 1REIVINDICACIONES 1.- Un método que hace funcionar un descodificador de audio (200) para descodificar datos de audio que incluyen bloques codificados de N.n canales de datos de audio para formar datos de audio descodificados que incluyen M.m 5 canales de audio descodificado, M≥1, siendo n el número de canales de efectos de baja frecuencia en los datos de audio codificados, y siendo m el número de canales de efectos de baja frecuencia en los datos de audio descodificados, comprendiendo el método:• aceptar los datos de audio que incluyen bloques de N.n canales de datos de audio codificados que han sido codificados mediante un método de codificación, incluyendo el método de codificación transformar N.n canales de datos de audio digital y formar y empaquetar datos de exponente y de mantisa de dominio de frecuencia;y • descodificar los datos de audio aceptados, incluyendo la descodificación: 15 - desempaquetar y descodificar (403) los datos de exponente y de mantisa de dominio de frecuencia, - determinar coeficientes de transformada (605) a partir de los datos de exponente y de mantisa de dominio de frecuencia desempaquetados y descodificados, - someter a una transformación inversa (607) los datos de dominio de frecuencia y aplicar un procesamiento adicional para determinar datos de audio muestreados, y - mezclar de manera descendente en el dominio de tiempo (613) al menos algunos bloques de los datos de audio muestreados determinados según datos de mezclado descendente para el caso en que M N;25 en el que el mezclado descendente en el dominio de tiempo incluye (1100) comprobar si los datos de mezclado descendente han cambiado en el tiempo con respecto a datos de mezclado descendente usados anteriormente, y , si han cambiado, aplicar una atenuación cruzada para determinar datos de mezclado descendente atenuados de manera cruzada y un mezclado descendente en el domino de tiempo según los datos de mezclado descendente atenuados de manera cruzada, y, si no han cambiado, realizar directamente un mezclado descendente en el dominio de tiempo según los datos de mezclado descendente.
- 2- El método según la reivindicación 1, donde el método incluye identificar (835) uno o más canales no contribuyentes de los N.n canales de entrada, siendo un canal no contribuyente un canal que no contribuye en los 35 M.m canales, y donde el método no lleva a cabo una transformación inversa de los datos de dominio de frecuencia ni aplica un procesamiento adicional en el uno o más canales no contribuyentes identificados.
- 3- El método según cualquiera de las reivindicaciones anteriores, en el que la transformación en el método de codificación usa una transformada solapada, y donde el procesamiento adicional incluye aplicar operaciones de división en ventanas y de solapamiento y suma (609) para determinar datos de audio muestreados.
- 4- El método según cualquiera de las reivindicaciones anteriores, en el que el método de codificación incluye formar y empaquetar metadatos relacionados con los datos de exponente y de mantisa de dominio de frecuencia, donde los metadatos incluyen opcionalmente metadatos relacionados con un procesamiento de pre-ruido transitorio y el 45 mezclado descendente.
- 5- El método según cualquiera de las reivindicaciones anteriores, en el que el descodificador (200) usa al menos un procesador x86 cuyo conjunto de instrucciones incluye difundir en flujo continuo extensiones (SSE) de una sola instrucción y múltiples datos que comprenden instrucciones vectoriales, y en el que el mezclado descendente en el dominio de tiempo incluye ejecutar instrucciones vectoriales en al menos un procesador del uno o más procesadores x86.
- 6- El método según la reivindicación 2, en el que n=1 y m=0, de manera que la transformación inversa y la aplicación de un procesamiento adicional no se llevan cabo en el canal de efecto de baja frecuencia. 55 7.- El método según la reivindicación 2, en el que los datos de audio que incluyen bloques codificados incluyen información que define el mezclado descendente, y en el que la identificación de uno o más canales no contribuyentes usa la información que define el mezclado descendente.
- 8- El método según la reivindicación 7, en el que la información que define el mezclado descendente incluye parámetros de nivel de mezclado que tienen valores predefinidos que indican que uno o más canales son canales no contribuyentes.
- 9- El método según la reivindicación 2, en el que la identificación de uno o más canales no contribuyentes incluye 65 además identificar si uno o más canales tienen una cantidad insignificante de contenido con respecto a uno o más otros canales, en el que la identificación de si uno o más canales tienen una cantidad insignificante de contenido con respecto a uno o más otros canales incluye comparar la diferencia de una medida de cantidad de contenido entre pares de canales con un umbral ajustable y/o en el que un canal tiene una cantidad insignificante de contenido con respecto a otro canal si su energía o nivel absoluto es al menos 15 dB inferior a los del otro canal o si su energía o nivel absoluto es al menos 18 dB inferior a los del otro canal o si su energía o nivel absoluto es al menos 25 dB 5 inferior a los del otro canal.
- 10- El método según cualquier reivindicación anterior, en el que los datos de audio aceptados están en forma de un flujo de bits de tramas de datos codificados, y en el que la descodificación se divide en un conjunto de operaciones de descodificación de sección de entrada (201) y en un conjunto de operaciones de descodificación de sección de 10 procesamiento (203), incluyendo las operaciones de descodificación de sección de entrada el desempaquetado y la descodificación de los datos de exponente y de mantisa de dominio de frecuencia de una trama del flujo de bits en datos de exponente y de mantisa de dominio de frecuencia desempaquetados y descodificados para la trama, y los metadatos incluidos en la trama, incluyendo las operaciones de descodificación de sección de procesamiento la determinación de los coeficientes de transformada, la transformación inversa y la aplicación de un procesamiento 15 adicional, la aplicación de cualquier procesamiento de pre-ruido transitorio requerido de descodificación y el mezclado descendente en el caso en que M N.
- 11- El método según la reivindicación 10, en el que las operaciones de descodificación de sección de entrada se llevan a cabo en una primera pasada seguida de una segunda pasada, comprendiendo la primera pasada 20 desempaquetar metadatos bloque a bloque y guardar punteros que apuntan a la ubicación en la que están almacenados los datos de exponente y de mantisa empaquetados, y comprendiendo la segunda pasada usar los punteros guardados que apuntan a los exponentes y mantisas empaquetados, y desempaquetar y descodificar datos de exponente y de mantisa canal a canal. 25 12.- El método según cualquier reivindicación anterior, en el que los datos de audio codificados se codifican según una norma del conjunto de normas que consiste en la norma AC-3, la norma E-AC-3 y la norma HE-AAC.
- 13- Un medio de almacenamiento legible por ordenador que almacena instrucciones de descodificación que cuando son ejecutadas por uno o más procesadores de un sistema de procesamiento hacen que el sistema de 30 procesamiento lleve a cabo el método de cualquiera de las reivindicaciones anteriores.
- 14- Un aparato (1200) que procesa datos de audio para descodificar los datos de audio que incluyen bloques codificados de N.n canales de datos de audio para formar datos de audio descodificados que incluyen M.m canales de audio descodificado, M≥1, siendo n el número de canales de efectos de baja frecuencia en los datos de audio 35 codificados, y siendo m el número de canales de efectos de baja frecuencia en los datos de audio descodificados, comprendiendo el aparato medios para llevar a cabo el método de cualquiera de las reivindicaciones 1 a 12.
Independent claims12
451 paragraphs in 1 section, as filed
Audio decoding using efficient downstream mixing
5 Field of the Invention
The present disclosure relates generally to the processing of audio signals.
Background
Compression of digital audio data has become an important technique in the audio industry. New formats have been introduced that allow high quality audio reproduction without the need for large data bandwidth, which is necessary in traditional techniques. The Advanced Television Systems Committee (ATSC) has adopted the AC-3 encoding technology and, more recently, the enhanced AC-3 (E-AC-3)
fifteen as the standard of audio service for high definition television (HDTV) in the United States. The EAC-3 standard is also used in consumer multimedia systems (digital video disc) and in direct satellite broadcasting. The E-AC-3 standard is an example of perceptual coding and allows the coding of multiple digital audio channels in a bit stream of encoded audio and metadata.
There is a need to efficiently decode a bit stream of encoded audio. For example, the battery life of portable devices is largely limited by the power consumption of its main processing unit. The energy consumption of a processing unit is closely related to the computational complexity of its tasks. Therefore, reducing the average computational complexity of a portable audio processing system will increase the battery life of such a system.
25 The term x86 is commonly used by those skilled in the art to designate a family of processor instruction set architectures whose origins date back to the Intel 8086 processor. As a result of the ubiquity of the x86 instruction set architecture, there is also a need to efficiently decode a stream of encoded audio bits in a processor or processing system that has an x86 instruction set architecture. Many decoder implementations have a generic character, while others are designed specifically for integrated processors. New processors such as the AMD Geode processor and the new Intel Atom are examples of 32-bit and 64-bit designs that use the x86 instruction set and are being used in small portable devices.
35 Summary
One aspect of the invention relates to a method that operates an audio decoder for decoding audio data according to claim 1. Additional aspects are defined in the manner set forth in dependent claims 2 to 12.
A further aspect of the invention relates to a computer readable storage medium that stores decoding instructions that when executed by one or more processors of a processing system causes the processing system to carry out the method described above.
Four. Five A further aspect of the invention relates to an apparatus that processes audio data to decode audio data that includes encoded blocks of Nn audio data channels to form decoded audio data that includes Mm channels of decoded audio, M≥1 , where n is the number of low frequency effect channels in the encoded audio data, and m being the number of low frequency effect channels in the decoded audio data, the apparatus comprising means for carrying out the above method.
Brief description of the drawings
Figure 1 shows a pseudocode 100 of instructions that, when executed, carry out a typical AC-3 decoding process.
55 Figures 2A to 2D show, in the form of a simplified block diagram, some different decoder configurations that can advantageously use one or more common modules.
Figure 3 shows a pseudocode and a simplified block diagram of an embodiment of an input section decoding module.
Figure 4 shows a simplified data flow diagram of the operation of an embodiment of an input section decoding module.
65 Figure 5A shows a pseudocode and a simplified block diagram of an embodiment of a processing section decoding module.
Figure 5B shows a pseudocode and a simplified block diagram of another embodiment of a processing section decoding module.
5 Figure 6 shows a simplified data flow diagram of the operation of an embodiment of a processing section decoding module.
Figure 7 shows a simplified data flow diagram of the operation of another embodiment of a processing section decoding module.
10 Figure 8 shows a flow chart of an embodiment of a processing of a processing section decoding module such as that shown in Figure 7.
Figure 9 shows an example of five block processing that includes a downward mixing of 5.1 to
fifteen 2.0 using an embodiment of the present invention in the case of a non-overlapping transform that includes a downstream mixing of 5.1 to 2.0.
Figure 10 shows another example of five-block processing that includes a downward mixing of 5.1 to 2.0 using an embodiment of the present invention for the case of an overlapping transform.
twenty Figure 11 shows a simplified pseudocode of an embodiment of downstream mixing in the time domain.
Figure 12 shows a simplified block diagram of an embodiment of a processing system that
25 It includes at least one processor and can perform decodes, which includes one or more features of the present invention.
Description of example embodiments
30 General features
Embodiments of the present invention include a method, an apparatus and logic encoded in one or more computer-readable tangible means for carrying out actions.
35 Particular embodiments include a method that operates an audio decoder to decode audio data that includes encoded blocks of Nn audio data channels to form decoded audio data that include Mm decoded audio channels, M≥1, where n is the number of low frequency effect channels in the encoded audio data, and m being the number of low frequency effect channels in the decoded audio data. The method comprises accepting audio data that includes blocks of Nn
40 encoded audio data channels that have been encoded by an encoding method that includes transforming Nn digital audio data channels and forming and packaging frequency domain mantissa and exponent data; and decode the accepted audio data. Decoding includes: unpacking and decoding the exponent data and frequency domain mantissa; determine transform coefficients from the exponent and mantissa frequency domain data unpacked and decoded;
Four. Five subject the frequency domain data to an inverse transformation and apply additional processing to determine sampled audio data; and mixing down in the time domain at least some blocks of the sampled audio data determined according to downstream mixing data for the case where M <N. At least one of A1, B1 and C1 is met:
fifty being A1 that the decoding includes determining block by block if a downward mixing in the frequency domain or a downward mixing in the time domain must be applied, and if for a particular block it is determined that a downward mixing in the domain must be applied frequency, apply a downward mixing in the frequency domain for the particular block,
55 B1 being that the downstream mixing in the time domain includes checking whether the downstream mixing data has changed with respect to previously used downstream mixing data, and, if they have changed, applying cross-attenuation to determine attenuated downstream mixing data cross and down mixing in the time domain based on cross-down attenuated mixing data, and, if they have not changed, directly perform downward mixing in the time domain
60 according to the downstream mixing data, and
C1 being that the method includes identifying one or more non-contributing channels of the Nn input channels, a non-contributing channel being a channel that does not contribute in the Mm channels, and that the method does not carry out an inverse transformation of the data of frequency domain or apply additional processing on the one or
65 more non-contributing channels identified.
Particular embodiments of the invention include a computer readable storage medium that stores decoding instructions that when executed by one or more processors of a processing system causes the processing system to perform decoding of audio data that includes blocks Nn encoded audio data channels to form decoded audio data that include Mm decoded audio channels, M≥1, n being the number of low frequency effect channels in the encoded audio data, and m being the number of low frequency effect channels in the decoded audio data. Decoding instructions include: instructions that when executed cause audio data that includes blocks of Nn encoded audio data channels that have been encoded by an encoding method to be accepted, including the encoding method transform Nn data channels digital audio and form and package exponent data and frequency domain mantissa; and instructions that when executed cause the accepted audio data to be decoded. The instructions that when executed cause decoding, include: instructions that when executed cause the unpacking and decoding of the frequency domain mantissa and exponent data; instructions that when executed cause transform coefficients to be determined from the exponent data and
fifteen of frequency domain mantissa unpacked and decoded; instructions that when executed cause an inverse transformation of the frequency domain data and the application of additional processing to determine sampled audio data; and instructions that, when executed, determine whether M <N; and instructions that when executed cause at least some blocks of the sampled audio data determined according to downstream mixing data to be mixed down in the time domain if M <N. At least one of A2, B2 and C2 is met:
being A2 that the instructions that when executed cause decoding include instructions that when executed cause block-to-block determination if a downward mixing in the frequency domain or a downstream mixing in the time domain must be applied, and instructions that when they run
25 cause a downward mixing in the frequency domain to be applied if for a particular block it is determined that a downward mixing in the frequency domain must be applied,
B2 being that the downstream mixing in the time domain includes checking whether the downstream mixing data has changed with respect to previously used downstream mixing data, and, if they have changed, applying a cross-damping to determine attenuated downstream mixing data cross and down mixing in the time domain based on cross-down attenuated mixing data, and, if they have not changed, directly perform a downstream mixing in the time domain according to the downstream mixing data, and
35 being C2 that the instructions that when executed cause decoding include identifying one or more non-contributing channels of the Nn input channels, a non-contributing channel being a channel that does not contribute to the Mm channels, and that the method does not carry out an inverse transformation of the frequency domain data or the application of additional processing in the one or more identified non-contributing channels.
Particular embodiments include an apparatus that processes audio data to decode audio data that includes encoded blocks of Nn audio data channels to form decoded audio data that includes Mm decoded audio channels, M≥1, where n is the number of channels of low frequency effects in the encoded audio data, and m being the number of channels of low frequency effects in the decoded audio data. The apparatus comprises: means for accepting audio data that includes blocks of Nn encoded audio data channels that have been encoded by an encoding method, including the encoding method transforming Nn digital audio data channels and forming and packaging data of exponent and mantissa of frequency domain; and means to decode the accepted audio data. The decoding means include: means for unpacking and decoding the frequency domain mantissa and exponent data; means for determining transform coefficients from the exponent data and frequency domain mantissa unpacked and decoded; means for subjecting frequency domain data to an inverse transformation and for applying additional processing to determine sampled audio data; and means for mixing down in the time domain at least some blocks of the sampled audio data determined according to downstream mixing data for the case where M <N. To the
55 least one of A3, B3 and C3 is met:
A3 being that the decoding means include means for determining block by block if a downward mixing in the frequency domain or a downward mixing in the time domain is to be applied, and means for applying a downward mixing in the frequency domain, where the means of applying a downstream mix in the frequency domain apply a downstream mix in the frequency domain for the particular block if it is determined that a downstream mix in the frequency domain must be applied,
B3 being that the downstream mixing media in the time domain checks whether the mixing data
65 downstream have changed with respect to previously used downstream mixing data, and, if they have changed, they apply a cross-damping to determine cross-attenuated downstream mixing data and a downstream mixing in the time domain based on the attenuated downstream mixing data of cross-way, and, if they have not changed, directly apply a downstream mix in the time domain based on the downstream mix data, and
5 C3 being that the apparatus includes means for identifying one or more non-contributing channels of the Nn input channels, a non-contributing channel being a channel that does not contribute in the Mm channels, and that the apparatus does not carry out an inverse transformation of the Frequency domain data or apply additional processing on the one or more identified non-contributing channels.
Particular embodiments include an audio data processing apparatus that include Nn channels of encoded audio data to form decoded audio data that include Mm decoded audio channels, M≥1, where n = 0 or 1 is the number of effect channels of low frequency in the encoded audio data, and m = 0 or 1 being the number of channels of low frequency effects in the decoded audio data. The apparatus comprises: means for accepting audio data that includes Nn encoded audio data channels 15 that have been encoded by an encoding method, the encoding method comprising transforming Nn digital audio data channels so that a reverse transformation and further processing can retrieve time domain samples without overlapping errors, form and package frequency domain mantissa exponent and data and form and package metadata related to frequency domain exponent and mantissa data, where the metadata optionally includes metadata related to transient pre-noise processing; and means to decode the accepted audio data. The decoding means comprises: one or more input section decoding means and one or more processing section decoding means. The input section decoding means includes means for unpacking the metadata, for unpacking and for decoding the frequency domain mantissa and exponent data. The means of decoding
25 Processing section includes means for determining transform coefficients from the exponent data and frequency domain mantissa unpacked and decoded; to subject the frequency domain data to an inverse transformation; to apply window splitting and overlapping and adding operations to determine sampled audio data; to apply any transient pre-routing processing required for decoding according to the metadata related to the transient pre-noise processing; and to perform a downstream mixing in the time domain according to downstream mixing data, the downstream mixing being configured to mix downwardly in the time domain at least some data blocks according to downstream mixing data in the case where M < N. At least one of A4, B4 and C4 is met:
35 A4 being that the processing section decoding means includes means for determining block by block if a downward mixing in the frequency domain or a downward mixing in the time domain is to be applied, and means for applying a downward mixing in the domain of frequency, where the downstream mixing application means in the frequency domain applies a downstream mixing in the frequency domain for the particular block if it is determined for a particular block that a downstream mixing must be applied in the frequency domain,
B4 being that the downstream mixing means in the time domain checks whether the downstream mixing data has changed with respect to previously used downstream mixing data, and, if they have changed, they apply cross-attenuation to determine attenuated downstream mixing data of way
Four. Five cross and down mixing in the time domain according to cross-attenuated down mixing data, and, if they have not changed, directly apply down mixing in the time domain according to the down mixing data, and
C4 being that the apparatus includes means for identifying one or more non-contributing channels of the Nn input channels, a non-contributing channel being a channel that does not contribute in the Mm channels, and that the processing section decoding means does not lead to perform a reverse transformation of the frequency domain data or apply additional processing in the one or more identified non-contributing channels.
Particular embodiments include an audio data decoding system that includes Nn channels of
55 encoded audio data to form decoded audio data including Mm decoded audio channels, M≥1, where n is the number of low frequency effect channels in the encoded audio data, and m being the number of effect channels of Low frequency in decoded audio data. The system comprises: one or more processors; and a storage subsystem coupled to one or more processors. The system accepts audio data that includes blocks of Nn channels of encoded audio data that have been encoded by an encoding method, including the coding method of transforming Nn channels of digital audio data, and forming and packaging exponent data and mantissa frequency domain; it also decodes the accepted audio data, which includes: unpacking and decoding the exponent data and frequency domain mantissa data; determine transform coefficients from the exponent data and frequency domino mantissa unpacked and decoded; undergo an inverse transformation the
65 frequency domain data and apply additional processing to determine sampled audio data; and mixing down in the time domain at least some blocks of the sampled audio data determined according to downstream mixing data for the case where M <N. At least one of A5, B5 and C5 is met:
A5 being that decoding includes determining block by block if a downstream mixing must be applied in
5 the frequency domain or a downward mixing in the time domain, and if for a particular block it is determined that a downward mixing in the frequency domain must be applied, a downward mixing in the frequency domain for the particular block is applied,
B5 being that downstream mixing in the time domain includes checking whether the mixing data
10 downstream have changed with respect to previously used downstream mixing data, and, if they have changed, apply cross-attenuation to determine cross-attenuated downstream mixing and downstream mixing in the time domain based on the attenuated downstream mixing data of cross-way, and, if they have not changed, directly perform a downstream mix in the time domain based on the downstream mix data, and
fifteen C5 being that the method includes identifying one or more non-contributing channels of the Nn input channels, a non-contributing channel being a channel that does not contribute in the Mm channels, and that the method does not carry out an inverse transformation of the data of frequency domain or apply additional processing in the one or more identified non-contributing channels.
twenty In some versions of the system embodiment, the accepted audio data is in the form of a bit stream of encoded data frames, and the storage subsystem is configured with instructions that when executed by one or more of the system processors Processing causes decoding of accepted audio data.
25 Some versions of the system embodiment include one or more subsystems that are interconnected through a network link, each subsystem including at least one processor.
In some embodiments in which A1, A2, A3, A4 or A5 is met, determine whether mixing should be applied
30 descending in the frequency domain or descending mixing in the time domain includes determining if there is a transient pre-noise processing and determining if any of the N channels has a different type of block, so that the descending mixing in the domain Frequency is only applied in a block that has the same type of block in the N channels, if there is no transient pre-noise processing and if M <N.
35 In some embodiments where A1, A2, A3, A4 or A5 are met and in which the transformation in the coding method uses an overlapping transform and the additional processing includes applying window splitting and overlapping and addition operations to determine sampled audio data, (i) the application of a downstream mixing in the frequency domain for the particular block includes determining whether the downstream mixing for the previous block was performed by a downstream mixing in the time domain, and whether mixing
40 The descending of the previous block was done by means of a descending mixing in the time domain, applying a descending mixing in the time domain (or a descending mixing in a pseudo-time domain) in the data of the previous block that will overlap with the decoded data of the particular block, and (ii) the application of a downstream mixing in the time domain for a particular block includes determining whether the downstream mixing of the previous block was performed by a downstream mixing in the domain of
Four. Five frequency, and if the downstream mixing of the previous block was performed by a downstream mixing in the frequency domain, process the particular block differently so that if the downstream mixing of the previous block had not been performed by a downstream mixing in the domain of frequency.
In some embodiments in which B1, B2, B3, B4 or B5 is met, at least one x86 processor is used, whose
fifty Instruction set includes continuous stream diffusion (SSE) extensions of a single instruction and multiple data comprising vector instructions, and downstream mixing in the time domain includes executing vector instructions on at least one processor of the one or more x86 processors.
In some embodiments in which C1, C2, C3, C4 or C5 is met, n = 1 and m = 0, so that the transformation
55 Reverse and the application of additional processing are not carried out in the low frequency effect channel. In addition, in some embodiments where C is met, the audio data that includes encoded blocks includes information defining the downstream mixing, and where the identification of one or more non-contributing channels uses the information defining the downstream mixing. In addition, in some embodiments where C is met, the identification of one or more non-contributing channels also includes identifying if one or more
60 channels have an insignificant amount of content with respect to one or more other channels, where one channel has an insignificant amount of content with respect to another channel if its energy or absolute level is at least 15 dB lower than those of the other channel. In some cases, a channel has an insignificant amount of content with respect to another channel if its energy or absolute level is at least 18 dB lower than those of the other channel, while in other applications a channel has an insignificant amount of content with respect to to another channel if your energy or
65 Absolute level is at least 25 dB lower than those of the other channel.
In some embodiments, the encoded audio data is encoded according to a standard in the set of standards consisting of the AC-3 standard, the E-AC-3 standard, a standard compatible with earlier versions of the E-AC-3 standard, the MPEG-2 AAC standard and the HE-AAC standard.
5 In some embodiments of the invention, the transformation in the coding method uses an overlapping transform, and further processing includes applying window splitting and overlapping and summing operations to determine sampled audio data.
In some embodiments of the invention, the coding method includes forming and packaging metadata.
10 related to the exponent and mantissa data in the frequency domain, where the metadata optionally includes metadata related to transient pre-noise processing and downstream mixing.
Particular embodiments may provide some, all or none of these aspects, features or advantages. Particular embodiments may provide one or more other aspects, features or advantages,
fifteen where one or more of which can be readily apparent to one skilled in the art from the figures, descriptions and claims of this document.
Decoding an encoded stream
twenty Embodiments of the present invention are described for decoding audio that has been encoded according to the extended AC-3 (E-AC-3) standard in an encoded bit stream. The E-AC-3 standard and the previous version, AC-3, are described in detail in the document “Digital Audio Compression Standard (AC-3, E-AC-3)”, Revision B, Document A / 52B, of the Advanced Television Systems Committee, Inc., (ATSC), June 14, 2005, obtained on December 1, 2009 on the website www ^ dot ^ atsc ^ dot ^ org / standards / a_52b ^ dot ^ pdf, (where ^ dot ^ denotes the period (".") in the
25 real web address). However, the invention is not limited to decoding a bit stream encoded in E-AC-3, and can be applied to a decoder that decodes a bit stream encoded according to another encoding method, and to methods of such decoding, decoding apparatus , systems that perform such decoding, to software that when executed causes one or more processors to perform such decoding and / or tangible storage media in which such software is stored. For example, embodiments of the present
30 The invention can also be applied to decode audio that has been encoded according to MPEG-2 AAC standards (ISO / IEC 13818-7) and MPEG-4 Audio (ISO / IEC 14496-3). The MPEG-4 Audio standard includes high-efficiency AAC encoding (version 1) (HE-AAC v1) and high-efficiency AAC encoding (version 2) (HE-AAC v2), collectively referred to as HE-AAC in this document .
35 AC-3 and E-AC-3 standards are also known as DOLBY ® DIGITAL and DOLBY ® DIGITAL PLUS. A HE-AAC version that incorporates some additional, compatible enhancements is also known as DOLBY ® PULSE. They are registered trademarks of Dolby Laboratories Licensing Corporation, the assignee of the present invention, and may be registered in one or more jurisdictions. The E-AC-3 standard is compatible with the AC-3 standard and includes additional functionality.
40 X86 architecture
The term x86 is commonly used by those skilled in the art to designate a family of processor instruction set architectures whose origins date back to the Intel 8086 processor. The architecture has been implemented in processors of companies such as Intel, Cyrix, AMD, VIA and many others. It is generally understood that the term implies binary compatibility with the 32-bit instruction set of the Intel 80386 processor. At present (early 2010), the x86 architecture is widely used in desktop computers and agenda-sized computers, as well as in a growing number of servers and workstations. A large amount of software supports the platform, including operating systems such as MS-DOS,
fifty Windows, Linux, BSD, Solaris and Mac OS X.
As used herein, the term "x86" refers to an x86 processor instruction set architecture that also supports a single instruction and multiple data instruction set extension (SSE) extension (SIMD). SSE is an instruction set extension of a single instruction and multiple data
55 (SIMD) of the original x86 architecture introduced in 1999 in Intel Pentium III series processors and now common in x86 architectures manufactured by many vendors.
Bit streams AC-3 and E-AC-3
60 An AC-3 bit stream of a multichannel audio signal is made up of frames representing a constant time interval of 1536 pulse code modulated (PCM) samples of the audio signal across all encoded channels. A maximum of five main channels are provided and, optionally, a low frequency effects (LFE) channel denoted as ".1", that is, a maximum of 5.1 audio channels is provided. Each frame has a fixed size that depends only on the sampling frequency and the data rate
65 coded
Simply put, AC-3 encoding includes using an overlapping transform (the discrete modified cosine transform (MDCT) with a window derived from Kaiser Bessel (KBD) with a 50% overlap) to convert time data into frequency data . Frequency data is encoded perceptually to compress the data to form a compressed bit stream of frames that each include audio data.
5 encoded and metadata. Each AC-3 frame is an independent entity that does not share data with previous frames other than the intrinsically overlapped transform in the MDCT used to convert time data into frequency data.
At the beginning of each AC-3 frame are the SI (synchronization information) and BSI (bit flow information) fields. The SI and BSI fields describe the configuration of the bit stream, including the sampling frequency, the data rate, the number of encoded channels and various other elements at the system level. There are also two words of CRC (cyclic redundancy code) per frame, one at the beginning and one at the end, which provide a means of error detection.
fifteen Within each frame there are six audio blocks, each representing 256 PCM samples per encoded channel of audio data. The audio block contains block switching flags, coupling coordinates, exponents, bit allocation parameters and mantissa. The data of a frame can be shared, so that the information present in Block 0 can be reused in subsequent blocks.
An optional auxiliary data field is located at the end of the frame. This field allows system designers to enter private control or status information into the AC-3 bit stream for transmission through the system.
The E-AC-3 standard retains the six-coefficient AC-3 frame structure of 256 coefficients, allowing the
25 same time smaller frames formed by one, two and three blocks of 256 coefficient transforms. This allows audio transport at data rates greater than 640 kbps. Each E-AC-3 frame includes metadata and audio data.
The E-AC-3 standard allows a much larger number of channels than the AC-3 5.1; in particular, the E-AC-3 standard can use the 6.1 and 7.1 audio channels of today and at least 13.1 channels to support, for example, future multichannel audio sound tracks. Additional channels to the 5.1 channels are obtained by associating the main audio program bit stream with a maximum of eight additional dependent subflows, which are multiplexed into an E-AC-3 bit stream. This allows the main audio program to use the 5.1-channel format of AC-3, while the additional channel capacity comes from dependent bit streams. This
35 It means that a 5.1-channel version and the various conventional downstream mixes are always available and that encoding artifacts induced by matrix subtractions are removed by using a channel substitution process.
The ability to support multiple programs is also due to the ability to provide seven more independent audio streams, each with possible associated dependent subflows, to increase the number of channels supported by each program beyond 5.1 channels.
The AC-3 standard uses a relatively short transform and simple scalar quantification to significantly encode the audio material. Although the E-AC-3 standard is compatible with the AC-3, it provides greater
Four. Five spectral resolution, improved quantification and improved coding. With the E-AC-3 standard, the coding efficiency has increased with respect to that of the AC-3 standard to allow the beneficial use of lower data rates. This is achieved by using an improved filter bank to convert time data into frequency domain data, improved quantization, improved channel coupling, spectral extension and a technique called transient pre-noise processing (TPNP).
In addition to the overlapped MDCT transform to convert time data to frequency data, the E-AC3 standard uses an adaptive hybrid transform (AHT) for stationary audio signals. The AHT includes the MDCT with the window derived from Kaiser Bessel (KBD) followed by, for stationary signals, a secondary block transform in the form of discrete cosine transform (DCT) of type II not divided into windows and not overlapped.
55 Therefore, the AHT adds a second phase DCT after the existing AC-3 MDCT / KBD filter bank when there is audio with stationary characteristics to convert the six 256-coefficient transform blocks into a single 1536 coefficient hybrid transform block with Higher frequency resolution. This higher frequency resolution is combined with a 6-dimensional vector quantification (VQ) and adaptive gain quantification (GAQ) to improve coding efficiency for some signals, for example, "hard to encode" signals. The VQ is used to efficiently encode frequency bands that need less precision, while the GAQ provides greater efficiency when quantification is needed with greater precision.
Higher coding efficiency is also obtained through the use of a channel coupling with
65 phase conservation This method extends the channel coupling method of the AC-3 standard, which uses a high frequency composite mono channel that reconstitutes the high frequency portion of each channel in decoding.
The addition of phase information and encoder-controlled processing of the spectral amplitude information sent in the bit stream improves the fidelity of this process, so that the composite mono channel can be extended at frequencies lower than previously possible. . This reduces the encoded effective bandwidth and therefore increases the coding efficiency.
5 The E-AC-3 standard also includes spectral extension. The spectral extension includes replacing higher frequency transform coefficients with lower frequency spectral segments converted into frequency in ascending order. The spectral characteristics of the converted segments are matched to the original ones by spectral modulation of the transform coefficients and also by mixing components
10 of noise conformed with the spectral segments of lower frequency converted.
The E-AC-3 standard includes a low frequency effects (LFE) channel. It is a single optional channel with a limited bandwidth (<120 Hz), intended to be played at a level of +10 dB with respect to the total bandwidth channels. The optional LFE channel allows to provide high levels of sound pressure for low sounds
fifteen frequency. Other coding standards, for example, AC-3 and HE-AAC, also include an optional LFE channel.
An additional technique for improving audio quality at low data rates is to use the transient pre-noise processing, described later in detail.
twenty AC-3 decoding
In typical AC-3 decoder implementations, to keep the memory and decoder latency requirements as low as possible, each AC-3 frame is decoded in a series of nested loops.
25 A first stage establishes the frame alignment. This implies finding the AC-3 synchronization word and then confirming that the CRC error detection words do not indicate any errors. Once frame synchronization has been obtained, the BSI data is unpacked to determine important frame information, such as the number of encoded channels. One of the channels may be an LFE channel. The number of encoded channels is denoted as Nn in this document, where n is the number of LFE channels and N is the
30 number of main channels. In the coding standards currently used, n = 0 or 1. In the future there may be cases in which n> 1.
The next stage in decoding is unpacking each of the six audio blocks. To minimize the memory requirements of the output data buffer modulated by pulse code (PCM),
35 The audio blocks are unpacked one at a time. At the end of each block period, the results of PCM, in many implementations, are copied into output buffers, which for real-time operation in a hardware decoder are normally duplicated or stored in buffers circularly for access Direct interruption using a digital to analog converter (DAC).
40 The AC-3 decoder audio block processing can be divided into two different stages, referred to herein as input processing and output processing. Input processing includes unpacking of all bit streams and manipulation of encrypted channels. The output processing mainly refers to the stages of splitting into windows and overlapping and adding the inverse MDCT transform.
Four. Five This distinction is made because the number of main output channels, denoted in this document as M≥1, generated by an AC-3 decoder does not necessarily coincide with the number of main input channels, denoted in this document as N, N≥ 1, encoded in the bit stream, where normally, although not necessarily, N≥M. Using downstream mixing, a decoder can accept a bit stream with
fifty any number N of encoded channels and produce an arbitrary number M, M≥1, of output channels. It should be noted that, in general, the number of output channels is denoted as Mm in this document, where M is the number of main channels and m is the number of LFE output channels. In current applications, m = 0 or 1. In the future it is possible that m is greater than 1.
55 It should be noted that in the downstream mixing not all encoded channels are included in the output channels. For example, in a 5.1 to stereo down mix, the LFE channel information is normally discarded. Therefore, in some downstream mixes, n = 1 and m = 0, that is, there is no output LFE channel.
60 Figure 1 shows a pseudocode 100 of instructions that, when executed, carry out a typical AC-3 decoding process.
The input processing in AC-3 decoding normally begins when the decoder unpacks the fixed audio block data, which is a collection of parameters and flags located at the beginning of the audio block. Fixed data includes elements such as block switching flags,
coupling information, exponents and bit allocation parameters. The term "fixed data" refers to the fact that word sizes for these bit stream elements are known a priori and, therefore, a variable length decoding process is not necessary to retrieve such elements.
5 The exponents form the largest field in the fixed data region, since they include all the exponents of each encoded channel. Depending on the coding mode, in AC-3 there can only be one exponent per mantissa, up to a maximum of 253 mantissa per channel. Instead of unpacking all these exponents in local memory, many decoder implementations store pointers that point to the exponent fields and unpack them as needed, one channel at a time.
Once the fixed data has been unpacked, many known AC-3 decoders begin processing each encoded channel. First, the exponents for the given channel are unpacked from the input frame. A bit allocation calculation is then carried out, which takes the exponents and the bit allocation parameters and calculates the word sizes for each packaged mantissa. Then the
fifteen Mantisas are normally unpacked from the input frame. Mantissans are scaled to provide appropriate dynamic margin control and, if necessary, to undo the coupling operation, and then denormalized by exponents. Finally, an inverse transform is calculated to determine pre-overlap and sum data, data in what is called a "window domain," and the results are mixed down in the downstream mixing buffers appropriate for outgoing processing. subsequent.
In some implementations, the exponents for the individual channel are unpacked in a buffer with a length of 256 samples, called "MDCT buffer". Then, these exponents are grouped into a maximum of 50 bands for bit allocation purposes. The number of exponents in
25 each band increases towards higher audio frequencies, following approximately a logarithmic division that models psychoacoustic critical bands.
For each of these bit allocation bands, the exponents and bit allocation parameters are combined to generate a mantissa word size for each mantissa of that band. These word sizes are stored in a buffer of bands with a length of 24 samples, where the broadest bit allocation band is formed by 24 frequency containers. Once the word sizes have been calculated, the corresponding mantissa is unpacked from the input frame and stored again in its place in the band buffer. These mantises are scaled and denormalized by the corresponding exponent and are written, for example they are rewritten in their place, in memory
35 MDCT intermediate. After all bands have been processed and all mantras have been unpacked, the remaining locations of the MDCT buffer are usually filled with zeros.
A reverse transformation is carried out, for example carried out in the MDCT buffer. Then, the output of this processing, the window domain data, can be mixed downwardly in the appropriate downstream mixing buffers according to downstream mixing parameters determined according to the metadata, for example extracted from predefined data according to the metadata.
Once the input processing is complete and the downstream mixing buffers have been fully generated with data mixed down in the domain of
Four. Five window, the decoder can carry out the output processing. For each output channel, a downstream mixing buffer and its corresponding half block delay buffer with a length of 128 samples are divided into windows and combined to produce 256 PCM output samples. In a hardware sound system that includes a decoder and one or more DACs, these samples are rounded to the DAC word width and copied to the output buffer. Once this has been done, half of the downstream mixing buffer is copied into its corresponding intermediate delay memory , providing 50% overlapping information necessary for a correct reconstruction of the next audio block.
E-AC-3 decoding
55 Particular embodiments of the present invention include a method that operates an audio decoder for decoding audio data that includes a number, denoted as Nn, of encoded audio data channels, for example, an E-AC-3 audio decoder. to decode E-AC-3 encoded audio data to form decoded audio data that includes Mm decoded audio channels, n = 0 or 1, m = 0 or 1 and M≥1. n = 1 indicates an input LFE channel, m = 1 indicates an output LFE channel. M <N indicates a down mix, M> N indicates an up mix.
The method includes accepting audio data that includes Nn channels of encoded audio data that have been encoded by the encoding method, for example by an encoding method that includes transforming N channels of digital audio data using an overlapping transform, forming and package exponent data and frequency domain mantissa, and form and package metadata related to the data of
exponent and frequency domain mantissa, where the metadata optionally includes metadata related to transient pre-noise processing, for example by an E-AC-3 coding method.
Some embodiments described in this document are designed to accept encoded audio data that has been encoded according to the E-AC-3 standard or according to a standard compatible with earlier versions of the EAC-3 standard, and may include more than 5 main channels coded
As will be described later in greater detail, the method includes decoding the accepted audio data, where decoding includes: unpacking the metadata and unpacking and decoding the frequency domain mantissa and exponent data; determine transform coefficients from the exponent and mantissa frequency domain data unpacked and decoded; subject frequency domain data to an inverse transformation; apply window division and overlap and sum operations to determine sampled audio data; apply any transient pre-noise processing required for decoding according to the metadata related to pre-noise processing
fifteen transient; and, in case M <N, perform a downstream mixing according to downstream mixing data. The downstream mixing includes checking whether the downstream mixing data has changed with respect to previously used downstream mixing data, and, if they have changed, applying cross-damping to determine cross-attenuated downstream mixing data and a downstream mixing according to the data of downstream mixes attenuated crosswise, and, if they have not changed, directly perform a downstream mixing according to the downstream mixing data.
In some embodiments of the present invention, the decoder uses at least one x86 processor that executes a continuous flow of single instruction and multiple data extension (SSE) instructions, including vector instructions. In such embodiments, downstream mixing includes executing instructions.
25 vector in at least one processor of the one or more x86 processors.
In some embodiments of the present invention, the E-AC-3 audio decoding method, which can be AC-3 audio, is divided into operation modules that can be applied more than once, that is, instantiated more than once. in different decoder implementations. In the case of a method that includes decoding, the decoding is divided into a set of input section decoding operations (FED) and a set of processing section decoding operations (BED). As will be explained later in detail, the input section decoding operations include unpacking and decoding exponent data and frequency domain mantissa of a frame of an AC-3 or E-AC-3 bit stream in exponent data and unpacked frequency domain mantissa and
35 decoded for the frame and metadata included in the frame. The processing section decoding operations include determining the transform coefficients, subjecting the determined transformation coefficients to an inverse transformation, applying window splitting and overlapping and summing operations, applying any transient pre-noise processing required for decoding and apply downstream mixing in case there are fewer output channels than channels encoded in the bit stream.
Some embodiments of the present invention include a computer readable storage medium that stores instructions that when executed by one or more processors of a processing system causes the processing system to perform decoding of audio data that includes Nn channels of encoded audio data to form decoded audio data including Mm audio channels 45 decoded, M≥1. In current standards, n = 0 or 1 and m = 0 or 1, but the invention is not limited to this. The instructions include instructions that when executed cause the audio data included to be accepted
Nn channels of encoded audio data that have been encoded by an encoding method, for example AC-3 or E-AC-3. The instructions also include instructions that, when executed, cause the accepted audio data to be decoded.
In some such embodiments, the accepted audio data is in the form of an AC-3 or E-AC-3 bit stream of encoded data frames. The instructions that when executed cause the accepted audio data to be decoded are divided into a set of reusable instruction modules, which include an input section decoding module (FED) and a processing section decoding module (BED) ). He
55 Input section decoding module includes instructions that, when executed, cause the frequency domain exponent and mantissa data of a bit stream frame to be unpacked and frequency domain mantissa data unpacked and decoded. and decoded for the frame and metadata included in the frame. The processing section decoding module includes instructions that when executed cause the transform coefficients to be determined, an inverse transformation is performed, window splitting and overlapping and summing operations are applied, any pre-noise processing is applied transient decoding required and downstream mixing is applied in case there are fewer output channels than encoded input channels.
Figures 2A to 2D show in simplified block diagrams some configurations of
65 different decoders that can advantageously use one or more common modules. Figure 2A shows a simplified block diagram of an example of an E-AC-3 200 decoder for 5.1-encoded audio AC-3 or E-AC-3. Obviously, the use of the term "block" when referring to the blocks of a block diagram is different from when it refers to a block of audio data, where the latter refers to a quantity of audio data. The decoder 200 includes an input section decoding module (EDF) 201 that accepts AC-3 or E-AC-3 frames and performs, frame by frame, unpacking the metadata
5 of frame and decoding of frame audio data in exponent data and frequency domain mantissa. The decoder 200 further includes a processing section decoding module (BED) 203 that accepts the exponent and frequency domain mantissa data from the input section decoding module 201 and decodes them for a maximum of 5.1 data channels. PCM audio
10 The decomposition of the decoder into an input section decoding module and a processing section decoding module is a design choice, not a necessary division. Such division provides the benefit of having common modules in several alternative configurations. The EDF module may be common to such alternative configurations, and many configurations have in common the unpacking of frame metadata and decoding of frame audio data in exponent and display data.
fifteen Frequency domain mantissa carried out by an EDF module.
As an example of an alternative configuration, Figure 2B shows a simplified block diagram of an E-AC-3 210 decoder / converter for E-AC-3 encoded 5.1 audio that decodes AC-3 encoded 5.1 audio
or E-AC-3 and that also converts an E-AC-3 encoded frame of up to 5.1 audio channels into one frame
twenty AC-3 encoded up to 5.1 channels. Decoder / converter 210 includes an input section decoding module (EDF) 201 that accepts AC-3 or E-AC-3 frames and that performs, frame by frame, unpacking the frame metadata and decoding of frame audio data in exponent data and frequency domain mantissa. The decoder / converter 210 further includes a processing section decoding module (BED) 203 which is the same or similar to the BED module 203 of the
25 decoder 200 and that accepts the frequency domain mantissa exponent and data of the input section decoding module 201 and decodes it for a maximum of 5.1 channels of PCM audio data. The decoder / converter 210 further includes a metadata converter module 205 for metadata conversion and a processing section coding module 207 that accepts the frequency domain exponent and mantissa data from the input section decoding module 201 and encode the data
30 as an AC-3 frame of up to 5.1 channels of audio data without exceeding the maximum data rate of 640 kbps that is possible with AC-3.
As an example of an alternative configuration, Figure 2C shows a simplified block diagram of an E-AC-3 decoder that decodes an AC-3 frame of up to 5.1 channels of encoded audio and also decodes an E-AC encoded frame. -3 of up to 7.1 audio channels. The decoder 220 includes a frame information analysis module 221 that unpacks the BSI data and identifies the frames and frame types and provides the frames to appropriate input section decoder elements. In a typical implementation that includes one or more processors and a memory that stores instructions that when executed cause the functionality of the modules to be carried out, multiple instances of a decoding module with an input section and multiple instances of a module Decoding section processing may be working. In some embodiments of an E-AC-3 decoder, the BSI unpacking functionality is not present in the input section decoding module that queries the BSI data. This allows common modules to be used in several alternative implementations. Figure 2C shows a simplified block diagram of a decoder with an architecture suitable for a maximum of 7.1 data channels
Four. Five audio Figure 2D shows a simplified block diagram of a 5.1 240 decoder with such architecture. The decoder 240 includes a frame information analysis module 241, an input section decoding module 243 and a processing section decoding module 245. These FED and BED modules may have a structure similar to that of the FED modules and BED used in the architecture of Figure 2C.
fifty Returning to Figure 2C, the frame information analysis module 221 provides the data of an independent AC-3 / E-AC-3 encoded frame up to 5.1 channels to an input section decoding module 223 that accepts the frames AC-3 or E-AC-3 and which performs, frame by frame, unpacking the frame metadata and decoding the frame audio data in exponent and mantissa domain data
55 of frequency. The frequency domain mantissa exponent and data are accepted by a decoding module of processing section 225 that is identical or similar to the BED module 203 of the decoder 200 and which accepts the frequency domain exponent and mantissa data of the decoding module input section 223 and decoding the data up to 5.1 channels of PCM audio data. Any frame encoded AC-3 / E-AC-3 dependent on additional channel data is provided to another module of
60 decoding of input section 227 that is similar to the other EDF module and, therefore, unpacks the frame metadata and decodes the audio data of the frame into exponent data and frequency domain mantissa data. A processing section decoding module 229 accepts the data from the FED module 227 and decodes the data into PCM audio data of any additional channel. A PCM 231 channel mapping module is used to combine the decoded data of the respective BED modules to provide
65 Up to 7.1 PCM data channels.
If there are more than 5 coded main channels, that is, in the case where N> 5, for example, there are 7.1 coded channels, the coded bitstream includes an independent frame of up to 5.1 coded channels and at least one data dependent frame coded In software embodiments for this case, for example embodiments comprising a computer-readable medium that stores instructions for execution, the instructions 5 are arranged as a plurality of 5.1-channel decoding modules, each 5.1-channel decoding module including a respective instance of an input section decoding module and a respective instance of a processing section decoding module. The plurality of 5.1-channel decoding modules includes a first 5.1-channel decoding module which, when executed, causes decoding of the independent frame, and one or more other channel decoding modules for each respective dependent frame. In some of these embodiments, the instructions include a frame information analysis instruction module that when executed causes the bit flow information (BSI) field of each frame to be unpacked to identify frames and frame types and the identified frames are provided to the appropriate input section decoder module instance, and a channel correlation instruction module that when executed, and in case N> 5, they make
fifteen combine the decoded data from respective processing section decoding modules to form the N main channels of decoded data.
A method to operate an AC-3 / E-AC-3 dual decoder / converter
An embodiment of the invention takes the form of a dual decoder / converter (DDC) that decodes two input bit streams AC-3 / E-AC-3, designated as "main" and "associated", each with a maximum 5.1-channel, in PCM audio, and in case of conversion, converts the main audio bit stream from E-AC-3 to AC-3, and in case of decoding, decodes the main bit stream and, if any , the associated bit stream. The dual decoder / converter optionally mixes the two PCM outputs using mixing metadata
25 extracted from the associated audio bit stream.
An embodiment of the dual decoder / converter performs a method that operates a decoder to carry out the processes included in the decoding and / or conversion of the two input bit streams AC-3 / E-AC-3. Another embodiment takes the form of a tangible storage medium that includes instructions, for example software instructions therein, which when executed by one or more processors of a processing system, causes the processing system to carry out the processes included in the decoding and / or conversion of the two input bit streams AC-3 / E-AC-3.
An embodiment of the AC-3 / E-AC-3 dual decoder / converter has six subcomponents, some of which include common subcomponents. The modules are:
Decoder-converter: The decoder-converter is configured, when executed, to decode an input bit stream AC-3 / E-AC-3 (up to 5.1 channels) into PCM audio, and / or to convert the bit stream input from E-AC-3 to AC-3. The decoder-converter has three main subcomponents and can implement the embodiment 210 shown in the previous figure 2B. The main subcomponents are:
Input section decoder: The EDF module is configured, when executed, to decode a frame of an AC-3 / E-AC-3 bit stream in raw audio data of frequency domain and its associated metadata.
Four. Five Processing section decoder: The BED module is configured, when executed, to complete the rest of the decoding process initiated by the FED module. In particular, the BED module decodes audio data (in mantissa and exponent format) into PCM audio data.
Processing section encoder: The processing section coding module is configured, when executed, to encode an AC-3 frame using six blocks of audio data from the EDF. The processing section coding module is also configured, when executed, to synchronize, resolve and convert E-AC-3 metadata into Dolby Digital metadata using a metadata converter module.
55 5.1 Decoder: The 5.1 decoding module is configured, when executed, to decode an AC-3 / E-AC-3 input bit stream (up to 5.1 channels) in PCM audio. The decoder 5.1 also optionally provides mixing metadata for use by an external application to mix two bit streams AC-3 / E-AC-3. The decoding module includes two main subcomponents: an EDF module as described earlier in this document and a BED module as described earlier in this document. A block diagram of an example decoder 5.1 is shown in Figure 2D.
Frame information: The frame information module is configured, when executed, to analyze an AC-3 / E-AC-3 frame and unpack its bit stream information. A CRC check is carried out in the frame as part of the unpacking process.
65 Buffer descriptors: Buffer descriptor module contains descriptions of AC-3, E-AC-3 and PCM buffers, and functions for buffers.
Sample Rate Converter: The sample rate converter module is optional and is configured, when executed, to sample PCM audio up to a factor of two.
5 External mixer: The external mixer module is optional and is configured, when executed, to mix a main audio program and an associated audio program to form a single output audio program using mixing metadata provided in the associated audio program .
Design of an input section decoding module
The input section decoding module decodes data according to the methods of standard AC-3 and according to additional E-AC-3 decoding aspects, including decoding of AHT data for stationary signals, the enhanced channel coupling of standard E -AC-3 and the spectral extension.
fifteen In the case of an embodiment in the form of a tangible storage medium, the input section decoding module comprises software instructions stored in a tangible storage medium which, when executed by one or more processors of a processing system, carry out the actions described in the details provided in this document for the operation of the input section decoding module. In a hardware implementation, the input section decoding module includes elements that are configured in operation to perform the actions described in the details provided in this document for the operation of the input section decoding module.
25 In AC-3 decoding, block-to-block decoding is possible. With E-AC-3, the first audio block, the audio block 0 of a frame, includes the AHT mantissa of the 6 blocks. Therefore, block-to-block decoding is not normally used, but several blocks are processed at the same time. However, the processing of the actual data is evidently carried out in each block.
In one embodiment, to use a uniform decoding / architecture method of a decoder regardless of whether AHT is used, the EDF module performs, channel by channel, two passes. A first pass includes unpacking block-by-block metadata and saving pointers that point to the location in which the exponent and mantissa data is stored packed, and a second pass includes using the saved pointers that point to the exponents and packaged mantissa, and unpack and decode
35 exponent and mantissa data channel by channel.
Figure 3 shows a simplified block diagram of an embodiment of an input section decoding module, for example implemented as a set of instructions stored in a memory that when executed causes FED processing to be carried out. Figure 3 also shows an instruction pseudocode for a first pass of a two-pass input section decoding module 300, as well as an instruction pseudocode for the second pass of the two-pass input section decoding module. The EDF module includes the following modules, each including instructions, where some of these instructions are definition instructions since they define structures and parameters:
Four. Five Channel: The channel module defines structures to represent an audio channel in memory and provides instructions to unpack and decode an audio channel from an AC-3 or E-AC-3 bit stream.
Bit allocation: The bit allocation module provides instructions for calculating the masking curve and calculating the bit allocation for encoded data.
Bit stream operations: The bit stream operations module provides instructions for unpacking data from an AC-3 or E-AC-3 bit stream.
55 Exponents: The exponent module defines structures to represent exponents in memory and provides configured instructions, when executed, to unpack and decode exponents of an AC-3 or E-AC-3 bit stream.
Exponents and mantissa: The exponent and mantissa module defines structures to represent exponents and mantissa in memory and provides configured instructions, when executed, to unpack and decode exponents and mantissa of a bit stream AC-3 or E-AC-3.
Matrix: The matrix module provides configured instructions, when executed, to support the dematrization of matrix channels.
65 Auxiliary data: The auxiliary data module defines auxiliary data structures used in the EDF module for
carry out EDF processing.
Mantissa: The mantissa module defines structures to represent mantissa in memory and provides configured instructions, when executed, to unpack and decode mantissa of a bit stream AC5 3 or E-AC-3.
Adaptive Hybrid Transform: The AHT module provides configured instructions, when executed, to unpack and decode adaptive hybrid transform data from an E-AC-3 bit stream.
Audio frame: The audio frame module defines structures to represent an audio frame in memory and provides configured instructions, when executed, to unpack and decode an audio frame from an AC-3 or E-AC- bit stream. 3.
Enhanced coupling: The enhanced coupling module defines structures to represent a channel of
fifteen Enhanced coupling in memory and provides configured instructions, when executed, to unpack and decode an enhanced coupling channel of an AC-3 or E-AC-3 bit stream. Enhanced coupling extends traditional coupling in an E-AC-3 bit stream by providing phase and chaos information.
Audio block: The audio block module defines structures to represent an audio block in memory and provides configured instructions, when executed, to unpack and decode an audio block of an AC-3 or E-AC- bit stream. 3.
Spectral extension: The spectral extension module provides support for decoding of spectral extension in an E-AC-3 bit stream.
Coupling: The coupling module defines structures to represent a coupling channel in memory and provides configured instructions, when executed, to unpack and decode a coupling channel of an AC-3 or E-AC-3 bit stream.
Figure 4 shows a simplified data flow diagram of the operation of an embodiment of the input section decoding module 300 of Figure 3, which describes how the pseudocode and submodule elements shown in Figure 3 act together to lead to Perform the functions of an input section decoding module. Functional element means an element that performs a processing function. Each element of this type can be a hardware element, or a processing system and a storage medium that includes instructions that, when executed, carry out the function. A bit stream unpacking functional element 403 accepts an AC-3 / E-AC-3 frame and generates bit allocation parameters for a standard bit allocation element and / or AHT 405 that produces additional unpacking data. bit streams to ultimately generate exponent and mantissa data for a standard / improved decoupling functional element 407. Functional element 407 generates exponent and mantissa data so that an included rematrization functional element 409 performs any necessary rematrization. Functional element 409 generates exponent and mantissa data so that a functional spectral extension decoding element included 411 performs any necessary spectral extension. Functional elements 407 to 411 use data obtained through the
Four. Five unpacking operation of functional element 403. The result of the decoding of the input section is exponent and mantissa data, as well as unpacked audio frame parameters and additional audio block parameters.
Referring in greater detail to the first pass and second pass pseudocode shown in Figure 3, the first pass instructions are configured, when executed, to unpack metadata from an AC-3 / E-AC-3 frame. In particular, the first pass includes unpacking the BSI information and unpacking the audio frame information. For each block, starting from block 0 to block 5 (6 blocks per frame), the fixed data is unpacked, and for each channel a pointer pointing to the exponents packed in the bit stream is stored, the exponents are unpacked and the flow position of
55 bits in which the packed mantissa resides. The bit allocation is calculated and, depending on the bit allocation, the mantissa can be omitted.
The instructions of the second pass are configured, when executed, to decode the audio data of a frame to form mantissa and exponent data. For each block, starting with block 0, unpacking includes loading the saved pointer that points to packaged exponents and unpacking the exponents pointed in this way, calculating the bit allocation, loading the saved pointer that points to the packaged mantissa and unpacking the Mantisas pointed this way. Decoding includes carrying out a standard and improved decoupling and generating the spectral extension band (s) and, to be independent of other modules, transfer the resulting data to a memory, for example an external memory at 65. internal memory of the pass, so that other modules, for example the BED module, can access the resulting data. For convenience, this memory is called “external” memory, although, as will be apparent to a
skilled in the art, can be part of a single memory structure used for all modules.
In some embodiments, for unpacking the exponents, the exponents unpacked during the first pass are not saved to minimize memory transfers. If AHT is used for a channel, the
5 exponents are unpacked from block 0 and copied to the other five blocks, numbered 1 to 5. If AHT is not used for a channel, pointers pointing to packaged exponents are saved. If the channel exponents strategy is to reuse the exponents, the exponents are unpacked again using the saved pointers.
In some embodiments, for unpacking coupling mantras, if AHT is used for the coupling channel, the six blocks of the AHT coupling channel mantras are unpacked in block 0 and regenerated by interpolation for each channel, i.e. a coupled channel to produce an uncorrelated interpolation. If AHT is not used for the coupling channel, pointers pointing to the coupling mantras are saved. These saved pointers are used to unpack the blankets of
fifteen coupling for each channel that is a channel coupled in a given block.
Design of a processing section decoding module
The processing section decoding module (BED) can be operated to take the exponent and mantissa frequency domain data and decode it into PCM audio data. PCM audio data is represented based on the modes selected by the user, dynamic margin compression and downmixing modes.
In some embodiments, in which the input section decoding module stores data from
25 exponent and mantissa in a memory, the so-called external memory, different from the working memory of the input section module, the BED module uses block-to-block frame processing to minimize the requirements of downstream mixing and delay buffer and, to be compatible with the output of the input section module, use external memory transfers to access exponent and mantissa data to be processed.
In the case of an embodiment in the form of a tangible storage medium, the processing section decoding module comprises software instructions stored in a tangible storage medium which, when executed by one or more processors of a processing system, carry carry out the actions described in the details provided in this document for the operation of the module
35 decoding of processing section. In a hardware implementation, the processing section decoding module includes elements that are configured in operation to perform the actions described in the details provided in this document for the operation of the processing section decoding module.
Figure 5A shows a simplified block diagram of an embodiment of a processing section decoding module 500 implemented as a set of instructions stored in a memory that, when executed, causes BED processing to be carried out. Figure 5A also shows an instruction pseudocode for the processing section decoding module 500. The BED 500 module includes the following modules, each including instructions, where some of these instructions are
Four. Five instructions that define:
Dynamic margin control: The dynamic margin control module provides instructions that, when executed, perform functions to control the dynamic range of the decoded signal, including applying a gain adjustment and applying a dynamic margin control.
Transformed: The transform module provides instructions that, when executed, carry out the reverse transforms, including performing a discrete reverse modified cosine transform (IMDCT), which includes carrying out a previous rotation used to calculate the DCT transform Inverse, perform a subsequent rotation used to calculate the inverse DCT transform and determine the fast Fourier transform
55 inverse (IFFT).
Transient pre-noise processing: The transient pre-noise processing module provides instructions that, when executed, carry out the transient pre-noise processing.
Window & overlap and sum: The window and overlap and sum module with delay buffer provides instructions that, when executed, perform the operation of splitting into windows and overlapping / summing to reconstruct output samples from transformed samples inverse
Downstream mixing in the time domain (TD): The TD downstream mixing module provides
65 instructions that, when executed, perform a downstream mixing in the time domain, as necessary, to a smaller number of channels.
Figure 6 shows a simplified data flow diagram of the operation of an embodiment of the processing section decoding module 500 of Figure 5A describing how the code and submodule elements shown in Figure 5A act together to perform The functions of a 5 decoding module processing section. A gain control functional element 603 accepts exponent and mantissa data from the input section decoding module 300 and applies any required dynamic range control, a dialog normalization and a gain adjustment according to metadata. The data resulting from exponent and mantissa are accepted by a functional element of denormalization of mantissa by exponents 605 that generates the transform coefficients for the inverse transformation. An inverse transform functional element 607 applies the IMDCT to the transform coefficients to generate time samples that are previously subjected to window splitting and overlapping and summing operations. Such time domain samples previously submitted to overlapping and summing operations are referred to herein as "pseudo-time domain" samples, and these samples are in what is referred to herein as the pseudo-time domain. These are accepted by a functional element of division in windows and of
fifteen overlap and sum 609 generated by PCM samples applying window division and overlap operations and added to the pseudo-time domain samples. A functional element of transient preruid processing 611 applies any transient pre-noise processing according to metadata. If, for example, it is indicated in the metadata or otherwise, the resulting transient pre-noise postprocessing PCM samples are mixed down to the Mm number of PCM sample output channels by a downstream mixing functional element 613.
Referring again to Figure 5A, the pseudocode for processing the BED module includes, for each data block, transferring mantissa and exponent data for blocks of a channel from the external memory and, for each channel: applying any dynamic margin control required, a dialog normalization and
25 a gain adjustment according to metadata; denormalize mantissa using exponents to generate the transform coefficients for the inverse transformation; calculate an IMDCT for the transform coefficients to generate pseudo-time domain samples; apply window splitting and overlapping operations and addition to pseudo-time domain samples; apply any transient pre-noise processing according to metadata; and, if necessary, perform a downstream mixing in the time domain towards the number Mm of output channels of PCM samples.
Decoding embodiments shown in Figure 5A include making such gain adjustments by applying dialog normalization offsets according to the metadata and applying dynamic range control gain factors according to the metadata. Carry out such gain adjustments in the phase in which the data is
35 Provided in the form of mantissa and exponent in the frequency domain is advantageous. Gain changes may vary over time, and such gain changes made in the frequency domain result in gradual cross attenuations once the inverse transform and window splitting and overlapping and summing operations have been performed. .
Transient Pre-Noise Processing
E-AC-3 encoding and decoding are designed to allow and provide better audio quality at lower data rates than in AC-3. At lower data rates, the audio quality of the encoded audio can be adversely affected, especially for a relatively difficult transient material.
Four. Five encode. This impact on audio quality is mainly due to the limited number of data bits available to precisely encode these types of signals. Transient coding artifacts manifest as a reduction in the definition of the transient signal as well as in the “transient pre-noise” artifact, which propagates audible noise throughout the coding window due to coding quantification errors.
As described above and in Figures 5 and 6, the BED provides a transitory pre-noise processing. The E-AC-3 encoding includes transient pre-noise coding processing to reduce transient pre-noise artifacts that can be introduced when audio containing transients is encoded by replacing the appropriate audio segment with audio that is synthesized using the located audio before the transitory pre-noise. The audio is processed using time scaling synthesis, so its duration increases so that
55 have an appropriate length to replace the audio that contains the transient pre-noise. The audio synthesis buffer is analyzed using an audio scene analysis and maximum similarity processing and then scaled over time so that its duration increases sufficiently to replace the audio that contains the transient pre-noise . Synthesized audio of increased length is used to replace the transient pre-noise and is cross-attenuated in the existing transient pre-noise just before the location of the transient to ensure a gradual transition from the synthesized audio to originally encoded audio data . Using transient pre-noise processing, the length of transient pre-noise can be drastically reduced or suppressed, even when block switching is disabled.
In an embodiment of an E-AC-3 encoder, the time scaling synthesis analysis and processing for the
65 Transient pre-noise processing tool are carried out in time domain data to determine metadata information, for example time scaling parameters. Metadata information is accepted by the decoder together with the encoded bit stream. Transient pre-noise metadata transmitted is used to perform time domain processing in decoded audio to reduce or eliminate the transient pre-noise introduced by low bit rate audio coding at low data rates.
5 The E-AC-3 encoder performs a time scaling synthesis analysis and determines time scaling parameters, based on the audio content, for each transient detected. The time scaling parameters are transmitted as additional metadata along with the encoded audio data.
In an E-AC-3 decoder, the optimal time-scaling parameters provided in the EAC-3 metadata are accepted as part of the accepted E-AC-3 metadata for use in transient pre-noise processing. The decoder performs an audio buffer split and cross-attenuation using the transmitted time scaling parameters obtained from the E-AC-3 metadata.
fifteen Using the optimal time scaling information and applying it with the appropriate cross-attenuation processing, the transient pre-noise introduced by the low bit rate audio coding can be drastically reduced or suppressed in decoding.
Therefore, the transient pre-noise processing overwrites the pre-noise with an audio segment that is very similar to the original content. Transient pre-noise processing instructions, when executed, maintain a four-block delay buffer for use in copying. Transient pre-noise processing instructions, when executed, in the event of an overwriting, cross-attenuation takes place inside and outside the overwritten pre-noise.
25 Mixed down
Denote Nn the number of channels encoded in the E-AC-3 bit stream, where N is the number of main channels and n = 0 or 1 is the number of LFE channels. Frequently it is desired to mix down the N main channels to a lower number, denoted as M, of main output channels. Embodiments of the present invention support downstream mixing of N to M channels, M <N. Up mixing is also possible, in which case M> N.
Therefore, in the more general implementation, audio decoder embodiments are operated to decode audio data that includes Nn audio data channels encoded in audio data.
35 decoded that include Mm decoded audio channels, and M≥1, where n, m indicate the number of LFE channels at the input, output respectively. Down mixing is the case where M <N and according to a set of down mixing coefficients is included in the case where M <N.
Mixing down frequency domain vs. time domain
The downstream mixing can be done completely in the frequency domain, before the inverse transform, in the time domain after the inverse transform, but in the case of overlapping and summing block processing before the window splitting operations and overlap and addition, or in the time domain after the operation of splitting into windows and overlapping and adding.
Four. Five Mixing down in the frequency domain (FD) is much more efficient than mixing down in the time domain. Its effectiveness is due, for example, to the fact that any processing stage after the downstream mixing stage is only carried out in the remaining number of channels, which is generally lower after the downstream mixing. Therefore, the computational complexity of all processing stages after the downstream mixing stage is reduced by at least the ratio of input channels to the output channels.
As an example, consider a downward mixing of 5.0 channels to stereo. In this case, the computational complexity of any subsequent processing stage will be reduced approximately by a factor of 5/2 =
55 2,5.
The downstream mixing in the time domain (TD) is used in typical E-AC-3 decoders and in the embodiments described above and illustrated in Figures 5A and 6. There are three main reasons why the E-AC-3 decoders Typical use the downstream mixing in the time domain:
Channels with different types of blocks
Depending on the audio content to be encoded, an E-AC-3 encoder can choose between two different block types, short block and long block, to segment the audio data. Harmonic audio data that
65 They change slowly and are normally segmented and encoded using long blocks, while the transient signals are segmented and coded into short blocks. As a result, the representation of short blocks and long blocks in the frequency domain is intrinsically different and cannot be combined in a downward mixing operation in the frequency domain.
Only after the decoder has undone the specific coding stages of type of
5 block, the channels can be mixed together. Therefore, in the case of block-switched transforms, a partial inverse transform process is used and the results of the two different transforms cannot be combined directly until just before the window phase.
However, methods for first converting short length transform data into longer frequency domain data are known, in which case the downstream mixing can be carried out in the frequency domain. However, in most known decoder implementations, downstream mixing is carried out after inverse transformation according to downstream mixing coefficients.
Mixed up
fifteen If the number of main output channels is greater than the number of main input channels, M> N, a mixing approach in the time domain is beneficial, since this takes the upward mixing stage at the end of the processing, reducing the number of channels in processing.
TPNP
Blocks that are subject to transient pre-noise processing (TPNP) cannot be mixed down in the frequency domain, since the TPNP operates in the time domain. The TPNP requires a history of up to four blocks of PCM data (1024 samples), which must be present for the channel in which
25 TPNP applies. Therefore, moving to downstream mixing in the time domain is necessary to fill in the PCM data history and carry out the pre-noise replacement.
Hybrid downstream mixing that uses downstream mixing in both the frequency domain and the time domain
The inventors recognize that the channels in most encoded audio signals use the same type of block more than 90% of the time. That means that the most efficient frequency domain downlink will be used for more than 90% of the data in typical encoded audio, assuming there is no TPNP. The remaining 10%, or less, will require downward mixing in the time domain, as in E-AC-3 decoders.
35 typical of the prior art.
Embodiments of the present invention include a downstream mixing method selection logic to determine block by block which downstream mixing method to apply, and both a downstream mixing logic in the time domain and a downstream mixing logic in the frequency domain. to apply the particular downstream mixing method as appropriate. Therefore, one method embodiment includes determining block by block if a frequency domain downlink or a time domain downmix is to be applied. The downstream mixing method selection logic works to determine whether a downstream mixing in the frequency domain or a downstream mixing in the time domain must be applied, and includes determining if there is any transient pre-noise processing and
Four. Five determine if any of the N channels has a different type of block. The selection logic determines that the downstream mixing in the frequency domain is only applied for a block that has the same type of block in the N channels, that there is no transient pre-noise processing and that M <N.
Figure 5B shows a simplified block diagram of an embodiment of a processing section decoding module 520 implemented as a set of instructions stored in a memory that, when executed, causes BED processing to be carried out. Figure 5B also shows an instruction pseudocode for the processing section decoding module 520. The BED module 520 includes the modules shown in Figure 5A that only use downstream mixing in the time domain, and subsequent additional modules, each including instructions, where some such instructions are
55 instructions that define:
Down-mix method selection module that checks (i) block type changes, (ii) if there is no true down-mix (M <N), but up-mix, and (iii) if the block is subject to TPNP, and if none of these cases are met, select the frequency domain downstream mix. This module determines block by block if a downstream mix in the frequency domain or a downstream mix in the time domain must be applied.
Down-mixing module in the frequency domain that performs, after the denormalization of the mantras by means of the exponents, a down-mixing in the frequency domain. It should be noted that the downstream mixing module in the frequency domain further includes a time domain to frequency domain transition logic module that checks whether the previous block used downstream mixing in the frequency domain.
time domain, in which case the block is treated differently than described later in greater detail. In addition, the transition logic module also handles processing steps associated with certain events that do not occur regularly, for example program changes such as attenuated output channels.
5 FD to TD downstream mixing transition logic module that checks if the previous block used downstream mixing in the frequency domain, in which case the block is treated differently than described later in greater detail. In addition, the transition logic module also handles processing steps associated with certain events that do not occur regularly, for example program changes such as attenuated output channels.
In addition, the modules that are in Figure 5A may have a different behavior in embodiments that include a hybrid downstream mixing, that is, a both FD and TD downstream mixing, depending on one or more conditions for the current block.
fifteen Referring to the pseudocode of Figure 5B, some embodiments of the processing section decoding method include, after transferring the data of a block frame from the external memory, determining whether an FD downstream mix or a TD downstream mix must be applied . For downstream mixing FD, for each channel, the method includes (i) applying a dynamic margin control and dialog normalization but, as described below, disabling the gain adjustment; (ii) denormalize mantissa by exponents; (iii) carry out a downstream mixing FD; and (iv) determine if there are attenuated output channels or if the previous block was mixed downwards by means of a downward mixing in the time domain, in which case the processing is carried out differently than described later in greater detail. . In the case of TD downstream mixing, and also for mixed data of
25 descending in the frequency domain, the process includes for each channel: (i) processing blocks that are going to be mixed down in the time domain in a different way in case the previous block has been mixed down in the frequency domain, in addition to treating any program change; (ii) determine the inverse transform; (iii) carry out window vision and overlap and sum operations; and in case of TD downstream mixing, (iv) carry out any TPNP and downstream mixing in the appropriate output channel.
Figure 7 shows a simple data flow diagram. Block 701 corresponds to the downstream mixing method selection logic that checks the three conditions: block type change, TPNP or upmixing, and if any condition is met, it directs the data flow to a mixing branch
35 TD 721 downstream which includes in 723 a FD downstream mixing transition logic to process differently a block that occurs immediately after a block processed by a downstream FFD mixing, program change processing, and in 725 denatures the mantras by the exponents After block 721, the data flow is processed by a common processing block 731. If the checks of the downstream mixing method selection logic block 701 determine that the block is for downstream mixing FD, the data flow is directed to a downstream mixing processing FD 711 that includes a downstream mixing process of domain frequency 713 that disables the gain adjustment and, for each channel, denatures the mantras by means of the exponents and performs a downward mixing FD, and a TD 715 downstream mixing transition logic block to determine if the previous block was processed by a TD downstream mixing, to process such block differently
Four. Five and also to detect and treat any changes in the program, such as attenuated output channels. After the downstream mixing transition block TD 715, the data flow is directed to the common processing block 731.
The common processing block 731 includes a reverse transformation and any additional processing in the time domain. Additional processing in the time domain includes undoing the gain adjustment and a window split and overlap and sum processing. If the block is derived from the TD 721 downstream mixing block, the additional processing in the time domain further includes any TPNP processing and a downstream mixing in the time domain.
55 Figure 8 shows a flow chart of a processing embodiment for a decoding module of processing section such as that shown in Figure 7. The flow chart is divided as follows, with the same reference numbers used in Figure 7 for similar respective functional data flow blocks: a section of downstream mixing method selection logic 701 in which a logical flag FD_dmx that when 1 is valid indicates that a frequency domain downstream mix is used for the block; a TD 721 downstream mixing logic section that includes a program change logic and downstream mixing transition logic section FD 723 for differently processing a block that occurs immediately after a block processed by a downstream mixing FD and to carry out a program change processing, and a section to denormalize the mantissa using the exponents for each input channel. After block 721, the flow of
65 Data is processed using a common processing section 731. If the downstream mixing method selection logic block 701 determines that the block is for a downstream mixing FD, the data flow is directed to a downstream mixing processing section FD 711 that includes a downstream mixing process of domain frequency that disables the gain adjustment, and for each channel, denatures the mantras by means of the exponents and performs a downward mixing FD, and a TD 715 downstream mixing transition logic section to determine for each channel of the previous block if
5 produces a channel attenuation or if the previous block was processed by TD downstream mixing, and to process such block in a different way. Following the downstream mixing transition section TD 715, the data flow is directed to the common processing logic section 731. The common processing logic section 731 includes for each channel an inverse transformation and any additional processing in the domain of weather. Additional processing in the time domain includes undoing the gain adjustment and a window split and overlap and sum processing. If FD_dmx is 0, which indicates a TD downstream mix, the additional time domain processing in 731 also includes any TPNP processing and a downstream mix in the time domain.
It should be noted that after downstream mixing FD, in the mixing transition logic section
fifteen descending TD 715, at 817, the number of input channels N is set to be the same as the number of output channels M, so that the rest of the processing, for example the processing in the processing logic section common 713, is carried out only in the mixed data in descending fashion. This reduces the amount of calculation. Obviously, the time domain downstream mixing of the data of the previous block when there is a transition from a block that was mixed downwardly in the time domain, such as the TD downstream mixing shown as 819 in section 715, is carried conducted on all N input channels involved in downstream mixing.
Transition Treatment
25 In decoding it is necessary to have gradual transitions between audio blocks. The E-AC-3 standard and many other coding methods use an overlapping transform, for example an MDCT with a 50% overlap. Therefore, when a current block is processed there is a 50% overlap with the previous block and, in addition, there will be a 50% overlap with the next block in the time domain. Some embodiments of the present invention use overlapping and addition logic that includes an overlapping and summing buffer. When a current block is processed, the overlap and sum buffer contains data from the previous audio block. Since it is necessary to have gradual transitions between audio blocks, the logic is included to treat transitions from a TD down mix to a FD down mix, and from a FD down mix to a TD down mix.
35 Figure 9 shows an example of five block processing, denoted as block k, k + 1, ..., k + 4 of five audio channels that include, as usual: a left channel, a central channel, a channel right, a left surround channel and a right surround channel, denoted as L, C, R, LS and RS, respectively, and a downward mix to a stereo mix, using the formula:
The left output is denoted as L '= aC + bL + cLS and the right output is denoted as R' = aC + bR + cRS.
Figure 9 assumes that a non-overlapping transform is used. Each rectangle represents the audio contents of a block. The horizontal axis from left to right represents the blocks k,…, k + 4 and the vertical axis from top to bottom represents the decoding progress of the data. Suppose block k is processed by a
Four. Five TD downstream mixing, that blocks k + 1 and k + 2 are processed by FD downstream mixing and that blocks k + 3 and k + 4 are processed by TD downstream mixing. As can be seen, for each of the TD downstream mixing blocks, the downstream mixing does not occur until after the downstream mixing of time domain towards the bottom, after which the contents are the mixed L 'and R' channels of way down, while for the block mixed down way FD, the left and right channels in the frequency domain have already been mixed down after the downstream mixing in the frequency domain, and the data of the C, LS and RS channels are ignored. Since there is no overlap between blocks, it is not necessary to deal with any spatial case when switching from a TD down mix to a FD down mix or from a FD down mix to a TD down mix.
55 Figure 10 describes the case of 50% overlapping transforms. Suppose that the overlap and addition are carried out by decoding of overlapping and summing using an overlapping and summing buffer. In this diagram, when the data block is shown as two triangles, the lower left triangle is data in the overlapping and summing buffer of the previous block, while the upper right triangle shows the data of the current block.
Treatment of a transition from a TD down mix to a FD down mix
Consider that block k + 1 is a downward mixing block FD that immediately follows a block
65 mixing down TD. After downstream mixing TD, the overlapping and summing buffer contains the data L, C, R, LS and RS of the last block to be included in the current block. I also know
It includes the contribution of the current block k + 1, already mixed down in the frequency domain. To correctly determine the mixed PCM data downwards for the output, it is necessary to include the data of both the current block and the previous block. To do this, the data in the previous block must be obtained and, since they have not yet been mixed downwards, mixed downwards in the time domain. The two contributions must be added to determine the mixed PCM data in descending order for the output. This processing is included in the downstream mixing transition logic TD 715 of Figures 7 and 8, and by the code in the downstream mixing transition logic TD included in the downstream mixing module FD shown in Figure 5B. The processing carried out therein is summarized in the downstream mixing transition logic section TD 715 of Figure 8. In greater
10 In detail, the treatment of a transition from a TD down mix to a FD down mix includes:
• Empty the overlapping buffers by introducing zeros in the overlapping and summing logic and carrying out window splitting and overlapping and summing operations. Copy the output obtained from the
fifteen logic of overlap and sum. This is the PCM data of the previous block of the particular channel before the downstream mixing. The buffer overlay now includes zeros.
• Mix down in the time domain the PCM data of the buffers of
overlap to generate PCM data of the TD downstream mixing of the previous block. twenty
• Mix down the new data of the current block in the frequency domain. Carry out the inverse transform and enter new data after the FD down mix and the inverse transform in the overlap and sum logic. Carry out window splitting and overlapping and summing operations, and so on, with the new data to generate PCM data of the downstream mixing FD of the current block.
• Add the PCM data of the TD downstream mix and the FD downstream mix to generate a PCM output.
It should be noted that in an alternative embodiment, assuming that TPNP has not been applied in the previous block,
30 the data of the overlapping and summing buffers are mixed downwardly and then an overlapping and summing operation is carried out on the mixed output channels in descending order. This avoids the need to carry out an overlapping and summing operation on each channel of the previous block. In addition, as described above for AC-3 decoding, when a downstream mixing buffer and its corresponding half block delay buffer with a length of
35 128 samples are used, divided into windows and combined to produce 256 PCM output samples, the downstream mixing operation is simpler since the delay buffer only has 128 samples instead of 256. This aspect reduces maximum complexity Computational intrinsic to transition processing. Therefore, in some embodiments, for a particular block mixed downwardly in the frequency domain that follows a block whose data was mixed downwardly in the domain of
40 time, the transition processing includes applying a downward mixing in the pseudo-time domain in the data of the previous block that will overlap with the decoded data of the particular block.
Treatment of a transition from an FD down mix to a TD down mix
Four. Five Consider that the block k + 3 is a downward mixing block TD that immediately follows a downward mixing block FD k + 2. Since the previous block was an FD downstream mixing block, the overlapping and summing buffer in the previous phases, for example before the TD downstream mixing, contains the data mixed downwards in the left and right channels and no data in The other channels. The contributions of the current block are not mixed down until after the
fifty mixed down TD. To correctly determine the mixed PCM data downwards for the output, it is necessary to include the data of both the current block and the previous block. To do this, the data from the previous block must be obtained. The data of the current block must be mixed downward in the time domain and added to the inverse transformed data that was obtained to determine the mixed PCM data downward for the output. This processing is included in the transition logic of
55 downstream mixing FD 723 of Figures 7 and 8, and using the code in the downstream mixing transition logic module FD shown in Figure 5B. The processing carried out therein is summarized in the downstream mixing transition logic section FD 723 of Figure 8. In greater detail, assuming that there are PCM output buffers for each output channel, the treatment of a transition from a FD downstream mix to a TD downstream mix includes:
• Empty the overlapping buffers by introducing zeros in the overlapping and summing logic and carrying out window splitting and overlapping and summing operations. Copy the output to the output PCM buffer. The data obtained are the PCM data of the FD downstream mixing of the previous block. The buffer overlay now includes zeros.
65 • Carry out an inverse transformation of the new data of the current block to generate premixed data descending from the current block. Enter this new time domain data (after the transformation) in the overlap and sum logic.
5 • Carry out window splitting and overlapping and summing operations, TPNP if any, and a TD downstream mix with the new data from the current block to generate PCM data from the TD downstream mix of the current block.
• Add the PCM data of the TD downstream mix and the FD downstream mix to generate a 10 PCM output.
In addition to transitions from time domain down mix to frequency domain down mix, program changes are addressed in the time domain down mix transition logic and in a program change handler. New emerging channels are included
fifteen automatically in the downstream mixing and, therefore, do not need any special treatment. Channels that are no longer present in the new program should be dimmed. This is carried out, as shown in section 715 of Figure 8 for the case of the FD downstream mixing, by emptying the overlapping buffers of the attenuated channels. The emptying is carried out by introducing zeros in the logic of overlapping and addition and carrying out operations of division in windows and of overlapping and addition.
twenty It should be noted that in the flow chart shown and in some embodiments, the frequency domain downstream mixing section 711 includes disabling the optional gain adjustment feature for all channels that are part of the frequency domain downstream mixing. Channels can have different gain adjustment parameters that will induce a different scaling of the spectral coefficients of
25 a channel, thus preventing a downward mixing.
In an alternative implementation, the downstream mixing logic section FD 711 is modified so that the minimum of all gains is used to carry out a gain adjustment for a downstream mixed channel (in the frequency domain).
30 Mixing down in the time domain with varying downward mixing coefficients and need for explicit cross attenuation
Downstream mixing can generate several problems. Different downward mixing equations are
35 used in different circumstances, so it may be necessary to dynamically modify the downmix coefficients depending on the signal conditions. There are metadata parameters available that allow you to customize the downward mixing coefficients for optimal results.
Therefore, the downward mixing coefficients may change over time. When a change occurs
40 From a first set of downstream mixing coefficients to a second set of downstream mixing coefficients, the data should be cross-attenuated from the first set to the second set.
When a mixing down is carried out in the frequency domain, and also in many
Four. Five Decoder implementations, for example in an AC-3 decoder of the prior art, such as that shown in Figure 1, the downstream mixing is carried out before the window splitting and overlapping and summing operations. The advantage of performing downstream mixing in the frequency domain, or in the time domain before window splitting and overlapping and summing operations, is that there is intrinsic cross attenuation as a result of overlapping and sum. Therefore, in many
fifty AC-3 decoders and known decoding methods in which downstream mixing takes place in the window domain after inverse transformation, or in the frequency domain in hybrid downstream implementations, there is no cross-attenuation operation explicit
In the case of downward mixing in the time domain and of a transitory pre-noise processing
55 (TPNP), there will be a delay of a block in transient decoding pre-noise processing caused by program change issues, for example in a decoder 7.1 Therefore, in embodiments of the present invention, when mixing is carried out descending in the time domain and TPNP is used, the downstream mixing in the time domain is carried out after window splitting and overlapping and summing operations. The order of processing in case you use downstream mixing in the domain of
60 time is: carry out the inverse transform, for example MDCT, carry out window splitting and overlapping and summing operations, carry out any transient pre-noise decoding processing (without delay) and then downstream mixing in The time domain.
In such a case, downstream mixing in the time domain requires cross-damping of mixing data.
65 Previous and current downstream, for example downstream mixing coefficients or downstream mixing tables, to ensure that any change in the downstream mixing coefficients is gradual.
One option is to carry out the cross-attenuation operation to calculate the resulting coefficient. Denote c [i] the mixing coefficient to be used, where i denotes the time index of 256 samples in the time domain, so that the interval is i = 0, ..., 255. Denote w2 [i] a positive window function so that w2 [i] + w2 [255 i] = 1 for i = 0,…, 255. Denote the old mixing coefficient updated above and the updated mixing coefficient again. The cross-damping operation to apply is:
10 After each pass of the cross-coefficient attenuation operation, the old coefficients are updated with the new ones, which is expressed as old ← new.
In the next pass, if the coefficients are not updated,
In other words, the influence of the set of old coefficients has completely disappeared.
The inventors have observed that in many audio streams and downstream mixing situations, the
twenty mixing coefficients do not vary frequently. To improve the performance of the down-mixing process in the time domain, embodiments of the down-mixing module in the time domain include determining if the down-mixing coefficients have changed from their previous value, and, if they have not changed, perform a downstream mixing; otherwise, if they have changed, they perform a cross attenuation of the downward mixing coefficients according to a positive window function
25 preselected In one embodiment, the window function is the same window function as that used in the window splitting and overlapping and summing operations. In another embodiment a different window function is used.
Figure 11 shows a simplified pseudocode for a downstream mixing embodiment. He
30 Decoder for such an embodiment uses at least one x86 processor that executes SSE vector instructions. The downstream mixing includes determining whether the new downstream mixing data has not changed with respect to the old downstream mixing data. If so, the downstream mixing includes making adjustments for the execution of SSE vector instructions on at least one processor of the one or more x86 processors and the downstream mixing using the unmodified downstream mixing data, which
35 It includes executing at least one SSE vector instruction. Otherwise, if the new downstream mixing data has changed from the old downstream mixing data, the method includes determining cross-attenuated downstream mixing data by means of a cross-damping operation.
Deletion of unnecessary data processing
40 In some situations of downstream mixing there is at least one channel that does not contribute to the mixed output in a downstream manner. For example, in many cases of downstream mixing of 5.1 audio to stereo, the LFE channel is not included, so that the downstream mixing is 5.1 to 2.0. The exclusion of the LFE channel in the downstream mixing may be intrinsic to the coding format, as is the case of AC-3, or it may
Four. Five controlled by metadata, as is the case with E-AC-3. In E-AC-3, the lfemixlevcode parameter determines whether or not the LFE channel is included in the downstream mix. When the parameter lfemixlevcode is 0, the LFE channel is not included in the downstream mixing.
It should be remembered that downstream mixing can take place in the frequency domain, in the domain of
fifty pseudo-time after the inverse transformation but before the window and overlap and sum division operations, or in the time domain after the inverse transformation and after the windows and overlap and sum division operations. A downstream mixing in the pure time domain is carried out in many known E-AC-3 decoders, and in some embodiments of the present invention, and is advantageous, for example, due to the presence of TPNP; a downward mixing in the pseudo-time domain
55 it is carried out in many AC-3 decoders and in some embodiments of the present invention, and is advantageous because the overlapping and summing operation provides an intrinsic cross attenuation that is advantageous when the downmix coefficients change; and downward mixing in the frequency domain is carried out in some embodiments of the present invention when conditions permit.
60 As described in this document, downstream mixing in the frequency domain is the most efficient downstream mixing method, since it minimizes the number of inverse transforms and window splitting and overlapping and summing operations necessary to produce a Two channel output from a 5.1 channel input. In some embodiments of the present invention, when a downstream mixing FD is carried out, for example, in Figure 8, in the downstream mixing loop section FD 711 in the loop starting at element 813, it ends at 814 and Increases by 815 to the next channel, these channels not included not included in the downstream mixing are excluded in the processing.
5 The downstream mixing in the pseudo-time domain after inverse transformation but before the window and overlapping and summing operations, or in the time domain after the inverse transformation and the window and window division operations overlap and addition is less computationally effective than in the frequency domain. In many current decoders, such as the current AC-3 decoders, downstream mixing takes place in the pseudo-time domain. The reverse transform operation is carried out independently of the downstream mixing operation, for example, in different modules. Reverse transformation into such decoders is carried out on all input channels. This is relatively ineffective from a computational point of view because, if the LFE channel is not included, the reverse transform is still carried out for this channel. East
fifteen unnecessary processing is significant because, although the LFE channel has limited bandwidth, applying the reverse transform to the LFE channel requires an equivalent calculation to apply the reverse transform to any full bandwidth channel. The inventors have recognized this poor efficiency. Some embodiments of the present invention include identifying one or more non-contributing channels of the Nn input channels, a non-contributing channel being a channel that does not contribute to the Mm encoded audio output channels. In some embodiments, the identification uses information, for example, metadata, that defines the downstream mixing. In the example of mixing down from 5.1 to 2.0, the LFE channel is identified as a non-contributing channel. Some embodiments of the invention include carrying out a frequency transformation in time on each channel that contributes to the Mm output channels, and does not carry out any frequency transformation in time on each identified channel that does not contribute to the Mm signal. channels In the example from 5.1 to 2.0 in the
25 that the LFE channel does not contribute to the downstream mixing, the inverse transform, for example an IMCDT, is only carried out in the five full bandwidth channels, so that the inverse transform part is carried out with a reduction Approximately 16% of the computational resources required for the 5.1 channels. Since the IMDCT is an important source of computational complexity in the decoding method, this reduction can be significant.
In many current decoders, such as the current E-AC-3 decoders, downstream mixing takes place in the time domain. The reverse transform operation and the overlapping and summing operations are carried out before any TPNP and before the downstream mixing, regardless of the downstream mixing operation, for example in different modules. The inverse transform and the operations of division in windows and of overlapping and summing in such decoders are carried out in all the input channels. This is relatively inefficient from a computational point of view because, if the LFE channel is not included, the inverse transform and the window and overlapping and summing operations continue to be carried out for this channel. This unnecessary processing is significant because, although the LFE channel has a limited bandwidth, applying the inverse transform and overlapping operations and adding to the LFE channel requires an equivalent calculation to apply the inverse transform and the window and window division operations. overlap and add to any full bandwidth channel. In some embodiments of the present invention, the downstream mixing is carried out in the time domain, and in other embodiments the downstream mixing can be performed in the time domain depending on the result of applying the mixing method selection logic. falling. Some embodiments of the present invention in which TD downstream mixing is used include identifying one or more non-contributing channels of the Nn input channels. In some embodiments, the identification uses information, for example metadata, that defines the downstream mixing. In the example of mixing down from 5.1 to 2.0, the LFE channel is identified as a non-contributing channel. Some embodiments of the invention include performing an inverse transform, that is, a frequency-to-time transformation in each channel that contributes to the Mm output channels, and does not perform any frequency-in-time transformation or other domain processing. of time on each identified channel that does not contribute to the signal of Mm channels. In the example from 5.1 to 2.0 in which the LFE channel does not contribute to the downstream mixing, the inverse transformation, for example an IMCDT, the overlapping and adding operations and the TPNP are only carried out in the five channels of complete band, so that the inverse transform and the operations of division in windows and of overlapping and summation are carried out
55 with an approximate reduction of 16% of the computational resources required for the 5.1 channels. In the flowchart of Figure 8, in the common processing logic section 731, a feature of some embodiments includes that the processing in the block starting with element 833, continues with 834 and that includes the increment up to Next element of channel 835, is carried out for all channels except non-contributing channels. This happens intrinsically in a block that has been mixed down in the frequency domain.
Although in some embodiments the LFE channel is a non-contributing channel, that is, it is not included in the downstream mixing output channels, as is common in AC-3 and E-AC-3, in other embodiments a channel other than the LFE It is also, or instead, a non-contributing channel and is not included in the mixed output in a descending manner. Some embodiments of the invention include checking such conditions to identify which one or more channels, if any, are non-contributing so that such channels are not included in the mixing.
descending and, in the case of descending mixing in the time domain, they do not carry out the processing through the inverse transform or overlapping and overlapping and adding operations for any identified non-contributing channel.
5 For example, in AC-3 and E-AC-3, there are certain conditions under which the surround channels and / or the central channel are not included in the mixed output channels in descending order. These conditions are defined by metadata included in the encoded bit stream that take predefined values. Metadata, for example, may include information that defines downstream mixing, including mixing level parameters.
Some examples of such mixing level parameters for the case of E-AC-3 are described below for illustrative purposes. Two types of downstream mixing are provided in stereo downlink mixing in E-AC-3: downward mixing to a coded stereo pair of matrix LtRt and downstream mixing to a conventional stereo signal, LoRo. The stereo signal mixed downwards
fifteen (LoRo or LtRt) can also be mixed to obtain a mono signal. A 3-bit LtRt surround mixing level code denoted as ltrtsurmixlev, and a 3-bit LoRo envelope mixing level code denoted as lorosurmixlev indicate the nominal downstream mixing level of the envelope channels with respect to the left and right channels in a downstream mixing LtRt or LoRo, respectively. The binary value '111' indicates a mixing level of 0, that is, -∞dB. 3-bit LtRt and LoRo central mixing level codes denoted as ltrtcmixlev, lorocmixlev, indicate the nominal downstream mixing level of the central channel with respect to the left and right channels in a LtRt and LoRo downstream mixing, respectively. The binary value '111' indicates a mixing level of 0, that is, -∞dB.
There are conditions in which the surround channels are not included in the output channels mixed so
25 falling. In E-AC-3, these conditions are identified by metadata. These conditions include cases in which surmixlev = '10 '(only in AC-3), ltrtsurmixlev =' 111 'and lorosurmixlev =' 111 '. For these conditions, in some embodiments, a decoder includes using the mixing level metadata to identify that such metadata indicates that the envelope channels are not included in the downstream mixing, and does not process the envelope channels through the reverse transform phase. nor of the phase of division into windows / overlap and sum. In addition, there are conditions in which the central channel is not included in the downstream mixed output channels, identified by ltrtcmixlevel = '111' and lorocmixlev = '111'. For these conditions, in some embodiments, a decoder includes using the mixing level metadata to identify that such metadata indicates that the central channel is not included in the downstream mixing, and does not process the central channel through the reverse transform phase. neither of the phase of division in windows /
35 overlap and sum.
In some embodiments, the identification of one or more non-contributing channels depends on the content. As an example, identification includes identifying if one or more channels have an insignificant amount of content with respect to one or more other channels. A measure of the amount of content is used. In one embodiment, the measure of quantity of content is energy, while in another embodiment the measure of quantity of content is the absolute level. The identification includes comparing the difference in the measure of the amount of content between pairs of channels with an adjustable threshold. As an example, in one embodiment, identifying one or more non-contributing channels includes determining whether the amount of envelope content of a block is less than each amount of forward channel content at least one adjustable threshold to determine whether the envelope channel is
Four. Five a non-contributing channel
Ideally, the threshold is selected to be as low as possible without introducing notable artifacts in the mixed downward version of the signal to maximize the identification of non-contributing channels to reduce the amount of calculation required, while minimizing the loss of quality In some embodiments, different thresholds are provided for different decoding applications, with the option of a threshold for a particular decoding application representing an acceptable balance between downstream mixing quality (higher thresholds) and reduced computational complexity (lower thresholds). for the specific application.
55 In some embodiments of the present invention, one channel is considered insignificant with respect to another channel if its energy or absolute level is at least 15 dB lower than those of the other channel. Ideally, one channel is insignificant with respect to another channel if its energy or absolute level is at least 25 dB lower than those of the other channel.
Using a threshold for the difference between two channels, denoted as A and B, which is equivalent to 25dB is almost the same as saying that the level of the sum of the absolute values of the two channels does not exceed 0.5 dB of the level of the dominant channel. That is, if channel A is at -6 dBFS (dB with respect to the total scale) and channel B is at -31 dBFS, the sum of the absolute values of channels A and B will be -5.5 dBFS approximately, or almost 0.5 dB higher than the level of channel A.
65 If the audio has a relatively low quality, and in low cost applications, it may be acceptable to sacrifice quality to reduce complexity; The threshold may be less than 25 dB. In one example, a threshold of 18 dB is used.
In this case, the sum of the two channels will not exceed approximately 1 dB of the channel level with the upper level. This may be audible in certain cases, but it should not be too excessive. In another embodiment a threshold of 15 dB is used, in which case the sum of the two channels does not exceed 1.5 dB of the level of the dominant channel.
5 In some embodiments, several thresholds are used, for example 15 dB, 18 dB and 25 dB.
It should be noted that although the identification of non-contributing channels has been described above for AC-3 and E-AC-3, the identification feature of non-contributing channels of the invention is not limited to such formats. Other formats, for example, also provide information, for example, metadata related to downstream mixing that can be used for the identification of one or more non-contributing channels. Both MPEG-2 AAC (ISO / IEC 13818-7) and MPEG-4 Audio (ISO / IEC 14496-3) can transmit what the standard calls “matrix downward mixing coefficient”. Some embodiments of the invention that decode such formats use this coefficient to construct a stereo or mono signal from a 3/2 signal, that is, a left, center, right, surround left and surround right signal. Mixing coefficient
fifteen Matrix descending determines how the surround channels blend with the front channels to build the stereo or mono output. Four possible values of the matrix downward mixing coefficient are possible according to each of these standards, one of which is 0. A value of 0 results in the envelope channels not being included in the downstream mixing. Some embodiments of MPEG-2 AAC decoder or MPEG-4 Audio decoder of the invention include generating a stereo or mono downstream mixing from a 3/2 signal using the downstream mixing coefficients signaled in the bit stream, and also include identify a non-contributing channel using a matrix downward mixing coefficient of 0, in which case the inverse transformation and the processing of window division and overlapping and addition are not carried out.
25 Figure 12 shows a simplified block diagram of an embodiment of a processing system 1200 that includes at least one processor 1203. This example shows an x86 processor whose instruction set includes SSE vector instructions. A bus subsystem 1205 is also shown as a simplified block through which the various components of the processing system are coupled. The processing system includes a storage subsystem 1211 coupled to the processor (s), for example through the bus subsystem 1205, the storage subsystem 1211 presenting one or more storage devices, including at least one memory and , in some embodiments, one or more other storage devices, such as magnetic and / or optical storage components. Some embodiments further include at least one network interface 1207 and an audio input / output subsystem 1209 that can accept PCM data and that includes one or more DACs to convert PCM data into waveforms.
35 electrical to activate a set of speakers or headphones. The processing system may also include other elements, as will be apparent to those skilled in the art, not shown in Figure 12 for simplicity.
The storage subsystem 1211 includes instructions 1213 which when executed in the processing system causes the processing system to perform decoding of audio data that includes Nn channels of encoded audio data, for example E-AC-3 data , to form decoded audio data including Mm decoded audio channels, M≥1, and in the case of downstream mixing, M <N. In the coding formats known at present, n = 0 or 1 and m = 0 or 1, but the invention is not limited to this. In some embodiments, instructions 1211 are divided into modules. Other instructions (other 45 software) 1215 are also normally included in the storage subsystem. The embodiment shown includes the following modules in instructions 1211: two decoder modules: an independent 5.1-channel decoder module 1223 that includes an input section decoding module 1231 and a processing section decoding module 1233, a dependent frame decoder module 1225 that includes a section decoding module input 1235 and a decoding module of processing section 1237, a 1221 frame information analysis instruction module that when executed causes unbundled bit stream information field (BSI) data of each frame to identify frames and frame types and to provide frames identified at instances appropriate input section decoding module 1231 or 1235, and a 1227 channel mapping instruction module that when executed and in case N> 5 cause the data to be combined
55 decoded of respective processing section decoding modules to form the Nn decoded data channels.
Alternative embodiments of the processing system may include one or more processors coupled via at least one network link, that is, distributed. That is, one or more of the modules can be in other processing systems coupled to a main processing system via a network link. Such alternative embodiments will be apparent to one skilled in the art. Therefore, in some embodiments, the system comprises one or more subsystems that are interconnected through a network link, each subsystem including at least one processor.
65 Thus, the processing system of Figure 12 forms an embodiment of an audio data processing apparatus that includes Nn channels of encoded audio data to form decoded audio data including Mm decoded audio channels, M≥1, in the case of down mixing, M <N, and for up mixing, M> N. Although in current standards, n = 0 or 1 and m = 0 or 1, other embodiments are possible. The apparatus includes several functional elements expressed functionally as a means to carry out a function. Functional element means an element that performs a processing function. Every
5 Such an element may be a hardware element, for example special purpose hardware, or a processing system that includes a storage medium that includes instructions that when executed perform the function. The apparatus of Figure 12 includes means for accepting audio data that includes N channels of encoded audio data that have been encoded by an encoding method, for example an E-AC-3 encoding method, and more generally , an encoding method comprising transforming, using an overlapping transform, N channels of digital audio data, forming and packaging exponent data and frequency domain mantissa, and form and package metadata related to exponent data and frequency domain mantissa, where metadata optionally includes metadata related to transient pre-noise processing.
fifteen The device includes means to decode the accepted audio data.
In some embodiments, the decoding means includes means for unpacking the metadata and means for unpacking and decoding the frequency domain mantissa and exponent data, means for determining transform coefficients from the exponent and domain mantissa data. frequency unpacked and decoded; means for subjecting frequency domain data to an inverse transformation; means for applying window splitting and overlapping and adding operations to determine sampled audio data; means for applying any transient pre-noise processing required for decoding according to the metadata related to the transient pre-noise processing; and means for performing TD downstream mixing according to downstream mixing data. The means of
25 TD downstream mixing, in the case where M <N, mixes downwardly according to downstream mixing data, including in some embodiment checking whether the downstream mixing data has changed with respect to previously used downstream mixing data, and, if changed, they apply cross-attenuation to determine cross-attenuated downward mixing data and downward mixing according to cross-attenuated downward mixing data, and, if they have not changed, they perform a downstream mixing directly according to the downstream mixing data.
Some embodiments include means for determining whether a TD downstream mixing or FD downstream mixing is used in a block, and FD downstream mixing means that are activated if the means determining whether a TD downstream mixing or downstream mixing is used in a block FD determine a mixed
35 downstream FD, including means for downstream mixing transition processing from TD to FD. Such embodiments further include means for a downstream mixing transition processing from FD to TD. The operation of these elements is as described in this document.
In some embodiments, the apparatus includes means for identifying one or more non-contributing channels of the Nn input channels, a non-contributing channel being a channel that does not contribute in the Mm channels. The apparatus does not perform the inverse transformation of the frequency domain data nor does it apply the additional processing, such as TPNP or overlap and sum, in the one or more identified non-contributing channels.
In some embodiments, the apparatus includes at least one x86 processor whose instruction set includes
Four. Five disseminate in continuous flow extensions (SSE) of a single instruction and multiple data comprising vector instructions The downstream mixing means, in operation, execute vector instructions in at least one processor of the one or more x86 processors.
Alternative devices to those shown in Figure 12 are also possible. For example, one or more of the elements can be implemented by hardware devices, while others can be implemented by operating an x86 processor. Such variations will be apparent to those skilled in the art.
In some embodiments of the apparatus, the decoding means includes one or more input section decoding means and one or more processing section decoding means. The input section decoding means includes the unpacking means of metadata and the unpacking and decoding means of the exponent data and frequency domain mantissa. The processing section decoding means includes the means that determine whether in a block TD downstream mixing or FD downstream mixing is used, the downstream mixing means FD including the means for downstream mixing transition processing from TD to FD, the means for the downstream mixing transition processing from FD to TD, means for determining transform coefficients from the exponent and mantissa frequency domain data unpacked and decoded; the means that undergo inverse transformation the frequency domain data; the means that apply any transient decoding pre-noise processing according to the metadata related to the transient pre-noise processing; and the 65-time domain downstream mixing media according to downstream mixing data. The downstream mixing in the time domain, in the case where M <N, performs a downstream mixing according to downstream mixing data, including, in some
embodiments, check if the downstream mixing data has changed with respect to previously used downstream mixing data, and, if they have changed, apply cross attenuation to determine cross-attenuated downstream mixing data and downstream mixing according to the mixing data descending cross-attenuated, and, if they have not changed, perform descending mixing according to the
5 mixing data down.
To process E-AC-3 data from more than 5.1 channels of encoded data, the decoding means includes multiple instances of the input section decoding means and the processing section decoding means, including a first decoding means input section and a first processing section decoding means to decode the independent frame of up to 5.1 channels, a second input section decoding means and a second processing section decoding means for decoding one or more dependent data frames. The apparatus further includes means for unpacking bitstream information field data to identify frames and frame types and to provide the identified frames to appropriate input section decoding means, and means for
fifteen combine the decoded data from respective processing section decoding means to form the N decoded data channels.
It should be noted that although the E-AC-3 standard and other coding methods use an overlapping and addition transform, and that the inverse transformation includes window splitting and coupling and summing operations, it is known that other forms of transform are possible. , which work in such a way that reverse transformation and additional processing can recover time domain samples without overlapping errors. Therefore, the invention is not limited to overlapping and summing transforms, and each time mentioning the inverse transformation of frequency domain data and the execution of window splitting and overlapping and summing operations to determine samples of time domain, the
25 Those skilled in the art will understand that, in general, these operations may be referred to as "reverse transformation of frequency domain data and additional processing application to determine sampled audio data."
Although the terms exponent and mantissa are used throughout the description since they are the terms used in AC-3 and E-AC-3, other coding formats may use other terms, for example, scale factors and spectral coefficients in the case of HE-AAC, and the use of the terms exponent and mantissa does not limit the scope of the invention to formats that use the terms exponent and mantissa.
Unless specifically stated otherwise, as is evident from the following description, you must
35 it should be noted that throughout the specification, the analyzes that use terms such as "processing", "computation", "calculation", "determination", "generation", or the like, refer to the action and / or processes of a hardware element, for example a computer or a computer system, a processing system or similar electronic computing device, which manipulates and / or transforms physically represented data, such as electronic data or amounts in other data similarly represented as physical quantities.
Similarly, the term "processor" may refer to any device or part of a device that processes electronic data, for example, from records and / or a memory, to transform that electronic data into other electronic data that, for example, can be stored in records and / or memory. A "processing system" or "computer" or a "computer machine" or a "computer platform" may include one or more
Four. Five processors
It should be noted that when describing a method that includes several elements, for example several stages, there is no implicit order of such elements, for example stages, unless specifically indicated.
In some embodiments, a computer readable storage medium is configured with, for example, encoded with, for example, stores instructions that when executed by one or more processors of a processing system, such as a digital signal processing device or a subsystem that includes at least one processing element and a storage subsystem, causes a method such as that described in this document to be carried out. It should be noted that in the description above, when
55 It mentions that the instructions are configured, when executed, to carry out a process, it should be understood that this means that the instructions, when executed, make one or more processors work so that a hardware device, for example the system processing, carry out the process.
In some embodiments, the methodologies described in this document may be carried out by one or more processors that accept logic, instructions encoded in one or more computer-readable media. When executed by one or more of the processors, the instructions cause at least one of the methods described in this document to be carried out. It includes any processor that can execute a set of instructions (sequential or otherwise) that specify the actions to be performed. Therefore, an example 65 is a typical processing system that includes one or more processors. Each processor may include one or more of a CPU or similar element, a graphics processing unit (GPU) and / or a programmable DSP unit. The processing system further includes a storage subsystem with at least one storage medium, which may include a memory integrated in a semiconductor device, or a separate memory subsystem that includes a main RAM and / or a static RAM and / or a ROM as well as cache memory. The storage subsystem may further include one or more other storage devices, such as magnetic and / or optical and / or solid-state storage devices. A bus subsystem can be included for communication between the components. The processing system can also be a distributed processing system with processors coupled via a network, for example through network interface devices or wireless network interface devices. If the processing system requires a display device, such a display device may include, for example a liquid crystal display (LCD), an organic light emission screen (OLED) or a cathode ray tube (CRT) screen . If a manual data entry is required, the processing system further includes an input device, such as one or more of an alphanumeric input unit such as a keyboard, a pointer control device such as a mouse, etc. The terms storage device, storage subsystem or memory unit used in this document, if it is evident from the context and unless
fifteen explicitly stated otherwise, they also cover a storage system such as a hard disk drive. In some configurations, the processing system may include a sound output device and a network interface device.
Therefore, the storage subsystem includes a computer-readable medium that is configured with, for example encoded with, instructions, for example logic, for example software, that when executed by one
or more processors, cause one or more of the method steps described in this document to be carried out. The software can reside on the hard disk or it can also reside, completely or at least partially, in the memory, such as a RAM, and / or in the internal memory of the processor during the execution thereof by means of the computer system. Therefore, the memory and the processor that includes the memory also constitute a medium
25 computer readable in which there are coded instructions.
In addition, a computer-readable medium may form a computer program product or it may be included in a computer program product.
In alternative embodiments, the one or more processors function as a stand-alone device or may be connected, for example networked to another processor (s), in a networked implementation, the one
or more processors can operate in the capacity of a server or a client machine in a client-server network environment, or as a homologous machine in a distributed or peer-to-peer network environment. The term processing system encompasses all these possibilities unless explicitly excluded in this
35 document. The one or more processors can form a personal computer (PC), a multimedia playback device, a tablet PC, a decoder (STB), a personal digital assistant (PDA), a gaming machine, a cell phone, a device web, a network router, a switch or bridge, or any machine capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken by that machine.
It should be noted that although some diagrams only show a single processor and a single storage subsystem, for example a single memory that stores the logic that includes instructions, those skilled in the art will understand that many of the components described above are included, but not shown. or explicitly described so as not to obscure the inventive aspect. For example, although only a single one is illustrated
Four. Five machine, the term "machine" should also be considered to include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions for carrying out any one or more of the methodologies described in this document.
Therefore, one embodiment of each of the methods described in this document is in the form of a computer-readable medium configured with a set of instructions, for example, a computer program that when executed in one or more processors, for example one or more processors that are part of a multimedia device, causes the method steps to be carried out. Some embodiments take the form of logic itself. Therefore, as those skilled in the art will appreciate, the embodiments of the present invention can be realized as a method, an apparatus such as a special purpose apparatus, an apparatus such as a logical data processing system, for example included in a computer readable storage medium or a computer readable storage medium encoded with instructions, for example a computer readable storage medium configured as a computer program product. The computer-readable medium is configured with a set of instructions that, when executed by one or more processors, cause the method steps to be carried out. Accordingly, aspects of the present invention may take the form of a method, a completely hardware embodiment that includes several functional elements, where a functional element is understood as an element that performs a processing function. Each element of this type can be a hardware element, for example, special purpose hardware, or a processing system that includes a storage medium that includes instructions that when executed perform the function. Aspects of the present invention may take the form of an embodiment completely in software or an embodiment that combines aspects of software and hardware. In addition, the present invention may take the form of program logic, for example in a
computer-readable medium, for example a computer program on a computer-readable storage medium or the computer-readable medium configured with a computer-readable program code, for example a computer program product. It should be noted that in the case of special purpose hardware, defining the function of the hardware is sufficient to allow a person skilled in the art to write a functional description that
5 can be processed by programs that automatically determine a hardware description to generate hardware that performs the function. Therefore, the description of this document is sufficient to define such special purpose hardware.
Although the computer-readable medium is shown in an exemplary embodiment as a single medium, it should be noted that the term "medium" includes a single medium or multiple media (eg, multiple memories, a centralized or distributed database and / or cache memories and associated servers) that store the one or more instruction sets. A computer-readable medium can take many forms, including, but not limited to, non-volatile media and volatile media. Non-volatile media include, for example, optical, magnetic and magneto-optical discs. Volatile media includes dynamic memories, such as a memory
fifteen principal.
It should also be understood that the embodiments of the present invention are not limited to any particular implementation or programming technique, and that the invention can be implemented using any appropriate technique to implement the functionality described herein. In addition, the embodiments are not limited to any particular programming language or operating system.
Reference throughout this specification to "an embodiment" means that a particular property, structure or feature described in relation to the embodiment is included in at least one embodiment of the present invention. Therefore, not all occurrences of the expression "in one embodiment" in various parts of this
25 Descriptive report necessarily refers to the same embodiment, although they can do so. In addition, the particular properties, structures or features may be combined in any suitable manner, as will be apparent to one skilled in the art from this specification, in one or more embodiments.
Likewise, it should be appreciated that in the above description of exemplary embodiments of the invention, several features of the invention are sometimes grouped together in a single embodiment, figure or description thereof in order to simplify the disclosure and help the understanding of one. or more of the various inventive aspects. However, it should not be construed that this method of disclosure reflects the intention that the claimed invention requires more features than those expressly mentioned in each claim. However, as the following claims reflect, the inventive aspects do not reside in all the features
35 of a single embodiment described above. Therefore, the claims that follow the DESCRIPTION OF EXAMPLE EMBODIMENTS are expressly incorporated into this DESCRIPTION OF EXAMPLE EMBODIMENTS, where each claim itself represents a different embodiment of this invention.
In addition, although some embodiments described herein include some but not other features included in other embodiments, combinations of features of different embodiments are within the scope of the invention and form different embodiments, as those skilled in the art will understand. For example, in the following claims, any of the claimed embodiments can be used in any combination.
Four. Five In addition, some of the embodiments are described herein as a method or combination of elements of a method that can be implemented by a processor of a computer system or by other means that perform the function. Therefore, a processor with the instructions necessary to carry out such a method or element of a method forms a means to carry out the method or element of a method. In addition, an element, described herein, of an embodiment of apparatus is an example of a means for carrying out the function performed by the element in order to carry out the invention.
Numerous specific details are set forth in the description provided in this document. However, it should be understood that embodiments of the invention can be practiced without these specific details. In other cases, widely known methods, structures and techniques have not been shown in detail so as not to
55 obscure the understanding of this description.
As used herein, unless otherwise specified, the use of the ordinal adjectives "first," "second," "third," etc., to describe a common object, simply indicates that reference is made to different instances of similar objects and should not be considered to imply that the described objects must be in a given sequence, whether temporary, spatial, hierarchical or in any other way.
It should be appreciated that although the invention has been described in the context of the E-AC-3 standard, the invention is not limited to such contexts and can be used to decode data encoded by other methods that use techniques that are similar to the E-standard. AC-3 For example, embodiments of the invention can also be applied to decode encoded audio that is compatible with earlier versions of E-AC-3. Other embodiments may be applied to decode encoded audio that has been encoded according to the HE-AAC standard and to decode
encoded audio that is compatible with previous versions of HE-AAC. Other encoded streams can also be decoded advantageously using embodiments of the present invention.
In the following claims and in the description of this document, any one of the terms
5 “Understand”, “understood by” or “comprising” is an open term that means that it includes at least the elements / features listed after it, but not excluding others. Therefore, it should not be construed that the term "comprising", when used in the claims, limits the means, elements or steps listed therein. For example, the scope of the expression "a device comprising A and B" should not be limited to devices consisting only of elements A and B. One of the terms "include" or "that includes"
10 used in this document is also an open term that also means that it includes at least the elements / features that follow the term, but not excluding others. Therefore, including is synonymous with understanding.
Likewise, it should be noted that the term "coupled", when used in the claims, should not be construed only as limited to direct connections. The terms "coupled" and "connected", together with their derivatives, can be used. It should be understood that these terms are not synonyms. Therefore, the scope of the expression "a device A coupled to a device B" should not be limited to devices or systems in which an output of device A is directly connected to an input of device B. It means that there is a path between a A's output and a B's input, which can be a path that includes other devices or media. "Coupled" may mean that two or more elements are in direct physical or electrical contact or that two or more
twenty elements are not in direct contact with each other but still act together or interact with each other.
Thus, although what is considered to be the preferred embodiments of the invention has been described, those skilled in the art will recognize that additional modifications may be made therein, all these changes and modifications being within the scope of the invention. For example, any
25 The formulas provided above are simply representations of procedures that can be used. Functionality can be added or deleted in block diagrams and operations can be exchanged between functional elements. Steps can be added or deleted in the methods described within the scope of the present invention.
13 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13
81 members in 38 offices
Priority claims4
| Document | Office | Kind | Date |
|---|---|---|---|
| 305871P | United States of America | – | |
| 30587110 | United States of America | P | |
| 359763P | United States of America | – | |
| 35976310 | United States of America | P |
Members81
| Document | Office | Kind | |
|---|---|---|---|
| EP2360683A1 | European Patent Office (EPO) | A1 | |
| CA2757643A1 | Canada | A1 | |
| CA2794029A1 | Canada | A1 | |
| CA2794047A1 | Canada | A1 | |
| WO2011102967A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU2011218351A1 | Australia | A1 | |
| SG174552A1 | Singapore | A1 | |
| AP2011005900A0 | African Regional Intellectual Property Organization (ARIPO) | A0 | |
| TW201142826A | Taiwan Province of China | A | |
| MX2011010285A | Mexico | A | |
| IL215254A0 | Israel | A0 | |
| IL215254D0 | Israel | D0 | |
| US2012016680A1 | United States of America | A1 | |
| ECSP11011358A | Ecuador | A | |
| AR080183A1 | Argentina | A1 | |
| EA201171268A1 | Eurasian Patent Organization (EAPO) | A1 | |
| KR20120031937A | Republic of Korea | A | |
| CN102428514A | China | A | |
| MA33270B1 | Morocco | B1 | |
| NI201100175A | Nicaragua | A | |
| US8214223B2 | United States of America | B2 | |
| HK1160282A | Hong Kong, China | A | |
| HK1160282A1 | Hong Kong, China | A1 | |
| CO6501169A2 | Colombia | A2 | |
| PE20121261A1 | Peru | A1 | |
| US2012237039A1 | United States of America | A1 | |
| JP2012527021A | Japan | A | |
| AU2011218351B2 | Australia | B2 | |
| ZA201106950B | South Africa | B | |
| CA2757643C | Canada | C | |
| HK1170059A | Hong Kong, China | A | |
| HK1170059A1 | Hong Kong, China | A1 | |
| UA101262C2 | Ukraine | C2 | |
| AU2013201583A1 | Australia | A1 | |
| TN2011000541A1 | Tunisia | A1 | |
| KR20130055033A | Republic of Korea | A | |
| CN102428514B | China | B | |
| IL227701A0 | Israel | A0 | |
| IL227701D0 | Israel | D0 | |
| IL227702A0 | Israel | A0 | |
| IL227702D0 | Israel | D0 | |
| IL215254A | Israel | A | |
| KR101327194B1 | Republic of Korea | B1 | |
| CN103400581A | China | A | |
| EP2698789A2 | European Patent Office (EPO) | A2 | |
| GT201100246A | Guatemala | A | |
| EP2360683B1 | European Patent Office (EPO) | B1 | |
| EP2698789A3 | European Patent Office (EPO) | A3 | |
| GEP20146086B | Georgia | B | |
| JP5501449B2 | Japan | B2 | |
| PT2360683E | Portugal | E | |
| ES2467290T3This record | Spain | T3 | |
| DK2360683T3 | Denmark | T3 | |
| TWI443646B | Taiwan Province of China | B | |
| HRP20140506T1 | Croatia | T1 | |
| SI2360683T1 | Slovenia | T1 | |
| JP2014146040A | Japan | A | |
| NZ595739A | New Zealand | A | |
| PL2360683T3 | Poland | T3 | |
| AR089918A2 | Argentina | A2 | |
| US8868433B2 | United States of America | B2 | |
| RS53336B | Serbia | B | |
| TW201443876A | Taiwan Province of China | A | |
| ME01880B | Montenegro | B | |
| IL227701A | Israel | A | |
| HN2011002584A | Honduras | A | |
| IL227702A | Israel | A | |
| AP3147A | African Regional Intellectual Property Organization (ARIPO) | A | |
| AU2013201583B2 | Australia | B2 | |
| US2016035355A1 | United States of America | A1 | |
| JP5863858B2 | Japan | B2 | |
| US9311921B2 | United States of America | B2 | |
| BRPI1105248A2 | Brazil | A2 | |
| CN103400581B | China | B | |
| MY157229A | Malaysia | A | |
| TWI557723B | Taiwan Province of China | B | |
| EA025020B1 | Eurasian Patent Organization (EAPO) | B1 | |
| EP2698789B1 | European Patent Office (EPO) | B1 | |
| KR101707125B1 | Republic of Korea | B1 | |
| CA2794029C | Canada | C | |
| BRPI1105248B1 | Brazil | B1 |
Numbers
- Publication
- 2467290
- Application
- 11154910
Titles2
- Spanish
- Descodificación de audio usando un mezclado descendente eficaz
- English
- Audio decoding using efficient downstream mixing
Classification
- CPC, 8
- G10L19/008
- H04S3/008
- G10L19/02
- G10L19/06
- G10L19/167
- G10L19/24
- H04R5/02
- G10L19/022
- IPC, 2
- G10L19 008
- H04S3 00