Audio decoder and decoding method using efficient downmixing
Abstract
A method, an apparatus, a computer-readable storage medium configured with instructions for carrying out a method, and an arrangement for carrying out the method. The method must decode audio information that includes Nn channels in Mm decoded audio channels, which includes unpacking metadata and unpacking and decoding frequency domain and mantissa exponent information; the determination of transformation coefficients from the unpacked and decoded information of frequency domain mantissa and mantissa; the inverse transformation of frequency domain information; and in the case M <N, the audio channel reduction combination according to the audio channel reduction combination information, where the audio channel reduction combination is carried out efficiently.

Term
No projected expiry on record.
- Priority
- Filed
- Published
- Today
13 claims: 8 independent, 5 dependent
- 1CLAIMS REIVINDICACIONES Habiendo así especialmente descripto y determinado la naturaleza de la presente invención y la forma como la misma ha de ser llevada a la práctica, se declara reivindicar como de propiedad y derecho exclusivo:Having thus specially described and determined the nature of the present invention and the manner in which it is to be put into practice, it is claimed to claim ownership and exclusive right: 1. A method for the operation of an audio decoder to decode audio information that includes encoded blocks of Nn audio information channels, to form decoded audio information that includes Mm decoded audio channels, where M> 1, where n is the number of 1. Un método para la operación de un decodificador de audio para decodificar información de audio que incluye bloques codificados de N.n canales de información de audio, para formar información de audio decodificada que incluye M.m canales de audio decodificado, donde M>1, donde n es el número de 10 Low frequency effect channels in the encoded audio information, and m is the number of low frequency effect channels in the decoded audio information, where the method characterized by the fact that it comprises: 10 canales de efectos de baja frecuencia en la información de audio codificada, y m es el número de canales de efectos de baja frecuencia en la información de audio decodificada, donde el método caracterizado por el hecho de que comprende: la aceptación de la información de audio que incluye bloques de N.n canales de información de audio codificada, codificada por un método de the acceptance of audio information that includes blocks of Nn channels of encoded audio information, encoded by a method of 15 codificación, donde el método de codificación incluye la transformación de N.n canales de información de audio digital, y la formación y el empaque de información de exponente de dominio de frecuencia y mantisa;y la decodificación de la información de audio aceptada, donde la decodificación incluye: fifteen coding, where the coding method includes the transformation of Nn channels of digital audio information, and the formation and packaging of frequency domain and mantissa exponent information;and decoding of accepted audio information, where decoding includes: 20 el desempaque y la decodificación de la información de exponente de dominio de frecuencia y mantisa;twenty unpacking and decoding of frequency domain and mantissa exponent information;la determinación de coeficientes de transformación a partir de la información de exponente de dominio de frecuencia y mantisa desempacada y decodificada;the determination of transformation coefficients from the frequency domain exponent information and unpacked and decoded mantissa;25 the inverse transformation of frequency domain information and the application of additional processing, in order to determine samples of audio information;Y 25 la transformación inversa de la información de dominio de frecuencia y la aplicación de procesamiento adicional, a fin de determinar muestras de información de audio;y 100 the downstream mix in the time domain of at least some blocks of the audio information samples determined according to the downstream mix for the case M <N;100 la mezcla descendente en el dominio de tiempo de por lo menos algunos bloques de las muestras de información de audio determinadas de acuerdo con la mezcla descendente para el caso M<N;en donde el método incluye la identificación (835) de uno o más canales no 5 contributivos de los N.n canales de entrada, donde un canal no contributivo es aquel que no contribuye a los M.m canales, y en donde el método no requiere realizar una transformada inversa de los datos del dominio de la frecuencia y no necesita la aplicación de procesos posteriores sobre los uno o más canales no contributivos identificados. where the method includes the identification (835) of one or more non-contributory channels of the Nn input channels, where a non-contributory channel is one that does not contribute to the Mm channels, and where the method does not require a transform Inverse of the frequency domain data and does not need the application of subsequent processes on the one or more identified non-contributory channels. 10 10
- 4El método de acuerdo con cualquiera de las reivindicaciones precedentes, caracterizado por el hecho de que el decodificador (200) utiliza por lo menos un procesador x86 cuyo conjunto de instrucciones incluyen instrucciones de vector que comprenden múltiples extensiones de información de Four. The method according to any of the preceding claims, characterized in that the decoder (200) uses at least one x86 processor whose instruction set includes vector instructions comprising multiple extensions of information of 25 a single instruction in execution (SSE) comprising instructions in vector, and where the downstream mix the execution of instructions in vector in at least one of one or more of the x86 processors. 25 una sola instrucción en ejecución (SSE) que comprende instrucciones en vector, y donde la mezcla descendente la ejecución de instrucciones en vector en por lo menos uno de uno o más de los procesadores x86. 101 101
- 8A computer readable storage medium, characterized by the fact that it stores decoding instructions that 8. Un medio de almacenamiento legible por computadora, caracterizado por el hecho de que almacena instrucciones de decodificación que 5 when executed by one or more processors of a processing system, they cause the processing system to carry out the method according to any of the preceding claims. 5 cuando son ejecutadas por uno o más procesadores de un sistema de procesamiento, hacen que el sistema de procesamiento lleve a cabo el método de acuerdo con cualquiera de las reivindicaciones precedentes.
- 1010 used for the processing of audio information for the decoding of audio information that includes blocks encoded from Nn channels of audio information, in order to form decoded audio information that includes Mm channels of decoded audio, M £ 1, where n is the number of low frequency effect channels in the encoded audio information, and m 10 utiliza para el procesamiento de información de audio para la decodificación de la información de audio que incluye bloques codificados de N.n canales de información de audio, a fin de formar información de audio decodificada que incluye M.m canales de audio decodificado, M£1, donde n es el número de canales de efectos de baja frecuencia en la información de audio codificada, y m
- 1115 es el número de canales de efectos de baja frecuencia en la información de audio decodificada. fifteen is the number of channels of low frequency effects in decoded audio information. 10. An arrangement (1200) for carrying out the method according to any of claims 1 to 7, characterized in that it is configured for decoding audio information that 10. Una disposición (1200) para para llevar a cabo el método de acuerdo con cualquiera de las reivindicaciones 1 a 7, caracterizada por el hecho de que está configurada para la decodificación de información de audio que
- 1220 incluye N.n canales de información de audio codificada a fin de formar información de audio decodificada que incluye M.m canales de audio decodificado, M£1, donde n es el número de canales de efectos de baja frecuencia en la información de audio codificada, y m es el número de canales de efectos de baja frecuencia en la información de audio decodificada, donde la twenty includes Nn channels of encoded audio information in order to form decoded audio information that includes Mm channels of decoded audio, M £ 1, where n is the number of low frequency effect channels in the encoded audio information, and m is the number of low frequency effect channels in the decoded audio information, where the
- 1325 provision includes:25 disposición comprende: one or more processors;and a storage subsystem coupled to one or more processors;uno o más procesadores;y un subsistema de almacenamiento acoplado a uno o más procesadores;103 where the arrangement is configured to accept the audio information that includes Nn blocks of encoded audio information channels, encoded by an encoding method, where the encoding method includes the transformation of Nn digital audio information channels, and training Y 103 donde la disposición está configurada para aceptar la información de audio que incluye N.n bloques de canales de información de audio codificada, codificada por un método de codificación, donde el método de codificación incluye la transformación de N.n canales de información de audio digital, y la formación y 5 The packaging of exponent information and frequency domain mantissa. 5 el empaque de Información de exponente y mantisa de dominio de frecuencia. p. de: DOLBY LABORATORIES LICENSING CORPORA INTERNATIONAL AB p. from: DOLBY LABORATORIES LICENSING CORPORA INTERNATIONAL AB Desempaque de información BSl Para bloque= 1 a B (el nro. de bloques) Unpacking BSl information For block = 1 to B (the number of blocks) Desempaque de Información establecida Guardado de punteros de exponentes empacados Para canal=1 a N (el nro. de canales codificados) Unpacking Established Information Saving pointers of packed exponents For channel = 1 to N (the number of coded channels) Desempaque de exponentes Para banda= 1 a L (el nro. de bandas) Unpacking of exponents For band = 1 to L (the number of bands) Bit allocation computation Unpacking mantissa Cómputo de asignación de bits Desempaque de mantisas Desempaque de canal de acoplamiento (guardado de punteros) Unpacking the coupling channel (saving pointers) Graduado de mantlsas/desecho de acoplamiento Desnormalización de mantisas por exponentes Cómputo de transformación Inversa a dominio de ventana Combin. reduc. canal de audio a nro. apropiado M de canales de salida Para canal= 1 a M (el nro. de canales de salida) Graduate of mantlsas / coupling waste Normalization of mantissa by exponents Computation of transformation Reverse to window domain Combin. reduction audio channel to no. appropriate M of output channels For channel = 1 to M (the number of output channels) Ventana y superp. de adición con un buffer de demora Window and superp. addition with a delay buffer Copla de valores de buffer de comb.de reduc. a buffer de demora Copy of buffer values of comb.de reduc. to delay buffer FIG. 1 (ARTE PREVIO) FIG. 1 (PRIOR ART) 100 100 FRAME AC-3 / E-AC-3 TRAMA AC-3/E-AC-3 FRAME AC-3 / E-AC-3 TRAMA AC-3/E-AC-3 FIG. 2B FIG. 2B FIG. 2A FIG. 2A FRAME AC-3 / E-AC-3 TRAMA AC-3/E-AC-3 Hasta 7.1 canales de PCM Up to 7.1 PCM channels FRAME AC-3 / E-AC-3 TRAMA AC-3/E-AC-3 Hasta 5.1 canales de PCM Up to 5.1 PCM channels FIG. 2 C FIG. 2C FIG. 2D FIG. 2D Decodificación extremo inicial /‘Decod.extremo inicial primer pasa*/ Initial Extreme Decoding /'Decox. Initial End First Pass * / For block = 0 to (B-1 number of blocks) Para bloque = 0 a (B-1 nro. de bloques) Desempaque de inform. establecida Inform unpacking established For channel = 0 to N-1 (N = no. Of encoded channels) Para canal= 0 a N-1 (N=nro. de canales codificados) Guardado de puntero de corriente de bits de exponentesempacados) Desempaque de exponentes Save pointer of current exponent bits packaged) Unpack exponents Guardado de puntero de corriente de bits de mantisas empacadas Saving bit stream pointer from packed mantissa Bit allocation computation Cómputo de asignación de bits Omitir mantisas sobre la base de asignación de bits /*Decod. extremo inicial segundo pase*/ Skip mantissa based on bit allocation / * Decod. initial end second pass * / For channel = 0 to N-1 (N = no. Of encoded channels) Para canal= 0 a N-1 (N=nro. de canales codificados) For bJoque = 0 to B-1 B = no. of blocks) / * unpacking * / Para bJoque= 0 a B-1 B= nro. de bloques) /*desempaque*/ Bitstream pointer load saved from expos, packet Unpacking exponents Computation of bit allocation Carga de puntero de corriente de bits guardado de expon, empac. Desempaque de exponentes Cómputo ae asignación de bits Bit stream pointer load saved from mant. empa. Carga de puntero de corriente de bits guardado de mant. empa. Desempaque de mantisas /*decocl·*/ Unwrap mantissa / * decocl · * / I do decouple standard / improved (amplitude only) Realiz. desacopl. norma/mejorado (solo amplitud) Generación de banda de extensión espectral ) Spectral extension band generation) I sent, inform. Expose, and mantis. int memory to ext. Transí, inform. expon, y mantis. de memoria int. a ext. FIG. 3 frame AC-3 / E-AC-3 FIG. 3 trama AC-3/E-AC-3 FIG. 4 FIG. 4 Información de metadatos y trama de audio, información de bloque de audio Metadata and audio frame information, audio block information Decodificación extremo finaí: ca. Finai extreme decoding: ca. Pare bloque = Oa B-1 (B -- bloques por trama) t Stop block = Oa B-1 (B - blocks per frame) t Transfer in all the blocks of a channel from external memory Transferencia en toóos los bloques de un canal desde memoria externa For channel - 0 to NI and LFE if n = l | Nn = no. Of coded channels) í Para canal - 0 a N I y LFE si n=l |Nn = nro.de canales codificados) í Aplicación de control de rango dinámico, normalización de diálogo y variación de ganancia Desnormalización de mantisas por exponentes Computo de transformación inversa Ventana / superposición de adición con buffer de demora Realización de procesamiento de ruido previo transitono Combinación de reducción de canales de audío a canales de salida apropiados 1 Application of dynamic range control, dialogue normalization and gain variation Normalization of mantras by exponents Reverse transformation compute Window / addition overlay with delay buffer Performing pre-noise processing processing Transono Combination of reduction of audio channels to output channels appropriate 1 Dynamic range control module (dialog normalization, dynamic range control) Módulo control rango dinámico (normalización de diálogo, control rango dinámico) Transformación Transformation Procesamiento rudio previo transitorio Transitional prior Russian processing Adtc overlay window. Ventana-superposición de adtc. Combin.reduc. TD audio channel Combin.reduc. canal audio TD FIG. 5 A FIG. 5 a 520 % Final extreme decoding: 520 % Decodlficación extremo final: For block · - Oa 0-1 (B - blocks per frame) { Para bloque ·- Oa 0-1 (B - bloques por trama) { Transfer in all the blocks of an external memory channel Evaluation of whether to reduce the audio channel FOocomb reduc. audio channel TD Yes comb. reduction audio channel FD;Transferencia en todos tes bloques de un canal de memoria externa Evaluación de si comb reduccanal audio FOocomb reduc. canal audio TD Si comb. reduc. canal audio FD;< < For channel «0 to ΝΊ and LFE if π · = 1 i Para canal « 0 a ΝΊ y LFE si π ·= 1 i Aplic, control de rango dinámico, nermafiz de diálogo, pero deshabito. variación ganancia Aplic, dynamic range control, dialogue nermaphiz, but uninhabited. gain variation Demantlement mantises by exponents Comb. FD audio channel reduction Desnormafización mantisas por exponentes Comb. reduc.canal audio FD Procesam. difte. bloque transición luego de bloque somet a combn. reduc. canal audroTD Tratamiento de cualquier canal apagado saliente } ) Processing diff. transition block after block subject to combn. reduction audroTD channel Treatment of any outgoing off channels}) De lo contrario,% Combin. reduc. canales audio TD { Otherwise,% Combin. reduction TD audio channels { For channel = 0a N-1 and LFE if n = T (But only for channels in combination. Audio channel reduction) Para canal = 0a N-1 y LFE si n=T (Pero solo para canales en combin. reduc. canal audio) Different process block transtc.after block submit combin.reduc audio channel FD Stack.control dynamic range, normalization dialog, but uninhabited. Variation gain Normalization mantises by exponents Proceso difnte. bloque transtc.luego de bloque someta combin.reduc canal audio FD Apile.control rango dinámico, normalización diálogo, pero deshabito. variación ganancia Desnormalización mantisas por exponentes For channel = 0a N-1 and LFE if n = 1 (But «Mo for channels in combination. Audio channel reduction;N = M if audio channel reduction comb FD) { Para canal = 0a N-1 y LFE si n =1 (Pero «Mo para canales en combin. reduc canal audio;N=M si comb. reduc.canal audio FD) { Cómputo transformación inversa Ventana/su^serposkión de adición con buffer demora Si combin. reduc.canal audio TD f Reverse transformation computation Window / its ^ serposkión of addition with buffer delay If combined. audio channel reduction TD f Combination prior noise processing. audio channel reduction to appropriate output channels Realización procesamiento ruido previo transitorio Combin. reduc.canal audio hasta canales de salida apropiados Dynamic range control module (dialog normalization, dynamic unce control) Combination method selection module. reduction audio channel Módulo control rango dinámico (normalización diálogo, control unce dinámica)Módulo selección método combin. reduc. canal audio Módulo selec.transic.comb reduc. FI1 Select module transic.comb reduc. FI1 Comb. reduc. i L) (incluye módulo lógica transic. comb. reduc. TD) Comb. reduction i L) (includes logic module transic. comb. reduction TD) Transformación Transformation Proces. ruido previo transit. Process. pre transit noise Ventana - stiperp. adición Window - stiperp. addition Comb. recfuc. TD audio channel Comb. recfuc. canal audio TD FIG. 5 B FIG. 5 B Exponentes Mantisas Metadatos Exponents Mantisas Metadata I ί I ί I I PCM Muestras PCM Samples FIG. 6 r0 ~ Logka séléccTlon mltóBo FIG. 6 r0 ~Logka séléccTón mltóBo FIG. 8 FIG. 8 FIG. 9 / * Comb reduc.TD according to inform.comb. reduction * / Si (new inform.comb.reduc. - old inform.comb.reduc.) (execute SSE instructions comb. reduction using comb. FIG. 9 /* Comb reduc.TD de acuerdo con inform.comb. reduc.*/ Si (nueva inform.comb.reduc. - vieja inform.comb.reduc.) ( ejecutar instrucciones SSE comb. reduc. usando inform. comb. reduc.;de lo contrario, ( on the contrary, ( apagado cruzado vieja inform. usando comb. reduc. vent. usando inform. comb. reduc. apag. cruz.;cross off old inform. using comb reduction vent. using inform. comb. reduction off cross.;1101 1101 FIG. 11 FIG. eleven Processing system Sistema de procesamiento 1200 1200 1207 network interfaces Interfaces 1207 de red F 2Q3 x86 processor Procesador f 2Q3 x86 --— iW --—iW Dispositivos de audio l/O L / O audio devices 1205 1205 Collective bar subsystem. Subsistema barra colect. Storage device storage subsystem .;^ 11 memory and possibly other storage devices. Subsistema de almacenamiento de dispositivos de almacén.;^11 memoria y posiblemente otros dispositivos de almacén. -Ί7ΪΓ programas ¡nform -Ί7ΪΓ programs nform Análisis inform. Computer Analysis plot trama 1221 1221 1213 1213 Cartografiador canal_1227 Cartographer channel_1227
Independent claims8
326 paragraphs in 6 sections, as filed
Invention Patent
On
AUDIO DECODER AND DECODING METHOD USING EFFICIENT AUDIO CHANNEL REDUCTION (DOWNMIXING) COMBINATION
226338 APP / fhcr
Requested by:
DOLB AND LABORA TORIES LICENSING CORPORA TION * DOLBY INTERNAL TIONAL AB.
neAidetOeA eti 100 Potrero Avenue San Francisco, California 94103 EU USA * Atlas Comp / ex,
Africa Building Hoogoorddreef 9, NETHERLANDS
DIVISIONAL DELA APPLICATION ACT N ° P11 Ol 00457 * AR 080183 As presented on 02/15/2011 ® AUDIO DECODER AND DECODING METHOD USING
COMBINATION OF AUDIO CHANNEL REDUCTION (DOWNMIXING)
EFFICIENT
FIELD OF THE INVENTION
The present invention generally relates to the processing of audio signals.
BACKGROUND
Compression of digital audio information has become an important technique in the audio industry. New formats have been introduced that allow the reproduction of high quality audio, without the need for the high bandwidth of information that would be required using traditional techniques. The AC-3 and, more recently, enhanced AC-3 (EAC-3) encoding technology has been adopted by the Advanced Television Systems Committee (ATSC) as the standard of audio services for High Definition Television (HDTV) in the United States. E-AC-3 has also found applications in consumer media (digital video disc) and direct satellite transmission. E-AC-3 is an example of perception coding, and provides the coding of multiple digital audio channels in a stream or time series of encoded audio bits and metadata.
There is an interest in efficient decoding of a stream of encoded audio bits. For example, the battery life of portable devices is mainly limited by the energy consumption of its main processing unit. The energy consumption of a processing unit is closely related to the complexity of calculating its tasks. Consequently, reducing the average complexity of computations of a
226,338 portable audio processing system should extend the battery life of that system.
The term x86, as is commonly understood by the person skilled in the art, refers to a family of instruction group architectures for processors whose origins date back to the Intel 8086 processor. As a consequence of the. ubiquity of the x86 instruction group architecture, there is also an interest in the efficient decoding of an audio bit stream encoded in a processor or in a processing system that has an x86 instruction group architecture. Many decoder implementations are of a general nature, while others are specifically designed for embedded processors. New processors, such as AMD's Geode and the new Intel Atom, are examples of 32-bit and 64-bit designs that use the x86 instruction set and are being used on small portable devices.
BRIEF DESCRIPTION OF THE DRAWINGS.
FIG. 1 shows pseudocode 100 for instructions that, when executed, carry out a typical AC-3 decoding process.
FIGS. 2A-2D show, in the form of a simplified block diagram, some different decoder configurations that can conveniently use one or more common modules.
FIG. 3 shows the pseudocode and a simplified block diagram of an embodiment of a start-end decoding module.
FIG. 4 shows a simplified data flow diagram for the operation of an embodiment of a start-end decoding module.
FIG. 5A shows the pseudocode and a simplified block diagram of an embodiment of an end-end decoding module.
FIG. 5B shows the pseudo code and a simplified block diagram of another embodiment of an end-end decoding module.
FIG. 6 shows a simplified data flow diagram for the operation of an embodiment of an end-end decoding module.
FIG. 7 shows a simplified Information flow diagram for the operation of another embodiment of an end-end decoding module.
FIG. 8 shows a flow chart of a processing embodiment for an end-end decoding module such as that set forth in FIG. 7.
FIG. 9 shows an example of processing five blocks that
It includes the combination of audlo channel reduction from 5.1 to 2.0 using an embodiment of the present invention, in the case of a non-overlapping transformation that includes the combination of audio channel reduction of '5.1 to 2.0.
- FIG. 10 shows another example of five-block processing that includes the combination of audio channel reduction from 5.1 to 2.0 using an embodiment of the present invention, in the case of an overlay transformation.
FIG. 11 shows a simplified pseudocode for a combination of time domain audio channel reduction combination.
FIG. 12 shows a simplified block diagram of an embodiment of a processing system that includes at least one processor and can perform decoding, which includes one or more features of the present invention.
DESCRIPTION OF EXEMPLARY EMBODIMENTS
Review.
The embodiments of the present invention include a method, an apparatus, and logic encoded in one or more computer-readable tangible means, for carrying out actions.
Particular embodiments include a method for the operation of an audio decoder to decode audio information that includes encoded blocks of Nn channels of audio information, to form decoded audio information that includes Mm channels of decoded audio, M ^ 1, where n is the number of low frequency effect channels in the encoded audio information, and m is the number of low frequency effect channels in the decoded audio information. The method comprises the acceptance of audio information that includes blocks of Nn channels of encoded audio information, encoded by an encoding method that includes the transformation of Nn channels of digital audio information, and the formation and packaging of information of exponent of frequency domain and mantissa; and decoding of accepted audio information. Decoding includes: unpacking and decoding of frequency domain and mantissa exponent information; the determination of transformation coefficients from the frequency domain exponent information and unpacked and decoded mantissa; the inverse transformation of frequency domain information and the application of additional processing, in order to determine samples of audio information; and the time domain audio channel reduction combination of at least some blocks of the audio information samples determined in accordance with the audio channel reduction combination information for the case M <N. At least one of A1, B1 and C1 is true:
A1 means that decoding includes determining, block by block, whether to apply frequency domain audio channel reduction combination or time domain audlo channel reduction combination, and if the combination application is determined for a particular block of frequency domain audio channel reduction, the application of frequency domain audio channel reduction combination application for the particular block;
B1 means that the time domain audlo channel reduction combination includes evaluating whether the audio channel reduction combination information is changed with respect to previously used audio channel reduction combination information, and, if it is changed , the cross-off application to determine the combination information of cross-off audio channel reduction combination, and the time domain audio channel reduction combination according to the cross-off audlo channel reduction combination information, and if not changed, the direct time domain audio channel reduction combination according with the combination information of audio channel reduction; Y
C1 means that the method includes the identification of one or more non-contributing channels of the input channels Nn, where a non-contributing channel is a channel that does not contribute to the Mm channels; and that the method does not perform the inverse transformation of frequency domain information; and the application of additional processing on one or more identified non-contributing channels.
Particular embodiments of the invention include a computer-readable storage medium, which stores decoding instructions which, when executed by one or more processors of a processing system, causes the processing system to perform the decoding of audio information. which includes encoded blocks of Nn channels of audio information, to form decoded audio information that includes Mm channels of decoded audio, M¿: 1, where n is the number of low frequency effect channels in the encoded audio information, and m is the number of low frequency effect channels in the decoded audio information. Decoding instructions include: instructions that, when executed, produce acceptance of audio information that includes blocks of channels Nn of encoded audio information, encoded by an encoding method, where the encoding method includes the transformation of channels Nn of digital audio information, and the formation and packaging of frequency domain and mantissa exponent information; and instructions that, when executed, produce the decoding of the accepted audio information. The instructions that when executed produce decoding include: instructions that, when executed, produce the unpacking and decoding of frequency domain and mantissa exponent information; instructions that, when executed, produce the determination of transformation coefficients from the unpacked and decoded information of frequency domain mantissa and mantissa; instructions that, when executed, produce the inverse transformation of frequency domain information and the application of additional processing to determine samples of audio information; and instructions that, when executed, produce the evaluation of whether M <N and instructions that, when executed, produce the time domain audio channel reduction combination of at least some blocks of the audio information samples determined according to the combination information of audio channel reduction if M <N. At least one of A2, B2 and C2 is true:
Α2 means that the Instructions that when executed produce the decoding include instructions that when executed produce the determination, block by block, of whether to apply frequency domain audio channel reduction combination or domain audio channel reduction combination of time, and instructions that when executed produce the frequency domain audio channel combination combination application if the frequency domain audio channel combination combination application is determined for a particular block;
B2 means that the time domain audio channel reduction combination includes evaluating whether the audio channel reduction combination information is changed with respect to previously used audio channel reduction combination information, and, if it is changed , the cross-off application to determine the combination information of cross-off audio channel reduction combination, and the time domain audio channel reduction combination according to the cross-off audio channel reduction combination information, and if not changed, the direct time domain audio channel reduction combination according with the combination information of audio channel reduction; Y
C2 means that the instructions that when executed produce decoding include the identification of one or more non-contributing channels of the input channels Nn, where a non-contributing channel is a channel that does not contribute to the Mm channels, and that the method does not carries out the inverse transformation of the frequency domain information; and the application of additional processing in one or more identified non-contributing channels.
j
Particular embodiments include an apparatus for processing audio information for decoding audio information that includes encoded blocks of Nn channels of audio information, in order to form decoded audio information that includes Mm channels of decoded audio, M ^ 1, where n is the number of low frequency effect channels in the encoded audio information, and m is the number of low frequency effect channels in the decoded audio information. The apparatus comprises: a means for accepting audio information that includes blocks of Nn channels of encoded audio information, encoded by an encoding method, wherein the coding method includes the transformation of Nn channels of digital audio information, and the training and packaging of frequency domain and mantissa exponent information; and a means for decoding the accepted audio information. The means for decoding includes: a means for unpacking and decoding the frequency domain and mantissa exponent information; a means for the determination of the transformation coefficients, based on the unpacked and decoded information of frequency domain mantissa and mantissa; a means for the inverse transformation of frequency domain information and for the application of additional processing in order to determine samples of audio information; and a means for the time domain audio channel reduction combination of at least some blocks of the audio information samples determined in accordance with the audio channel reduction combination information for the case M <N. At least one of A3, B3 and C3 is true:
A3 means that the means for decoding includes a means for the block-by-block determination of whether to apply frequency domain audio channel reduction combination or time domain audio channel reduction combination, and a means for The frequency domain audio channel combination combination application, where the means for the frequency domain audio channel reduction combination application applies frequency domain audio channel reduction combination for the particular block, if the channel reduction combination application is determined for a particular block audio frequency domain;
B3 means that the means for the time domain audio channel reduction combination carries out the evaluation of whether the audio channel reduction combination information is changed with respect to previously combined audio channel reduction information. used, and if changed, apply cross-off in order to determine the cross-off audio channel reduction combination information and perform the time domain audio channel reduction combination according to the cross-off audio channel reduction combination information ; and if it is not changed, apply the time domain audio channel reduction combination directly according to the audio channel reduction combination information; Y
C3 means that the apparatus includes a means for the identification of one or more non-contributing channels of the input channels Nn, where a non-contributing channel is a channel that does not contribute to the Mm channels, and that the apparatus does not carry out the reverse transformation of frequency domain information; and the application of additional processing in one or more identified non-contributing channels.
Particular embodiments include an apparatus for processing audio information that includes Nn channels of encoded audio information in order to form decoded audio information that includes channels
Mm decoded audio, M> 1, n = 0 or 1 which is the number of low frequency effect channels in the encoded audio information, and m = 0 or 1 which is the number of low frequency effect channels in the decoded audio information. The apparatus comprises: a means for accepting audio information that includes Nn channels of encoded audio information, encoded by a coding method, wherein the coding method comprises the transformation of Nn channels of Digital Audio Information, such that Reverse transformation and additional processing can recover time domain samples without aliasing errors; the formation and packaging of frequency domain and mantissa exponent information; and the formation and packaging of metadata related to the frequency domain and mantissa exponent information, where the metadata optionally includes metadata related to the processing of transient prior noise; and a means for decoding the accepted audio information. The means for decoding comprises: one or more means for decoding the start end (front end) and one or more means for decoding end end (back end). The means for decoding the starting end includes a means for unpacking the metadata, for unpacking and for decoding the frequency domain and mantissa exponent information. The means for end-end decoding includes a means for determining the transformation coefficients from the unpacked and decoded information of frequency domain mantissa and mantissa; for the inverse transformation of frequency domain information; for the application of window and overlay-aggregate operations to determine samples of audio information; for the application of any transient pre-noise processing decoding required in accordance with the metadata related to the transient pre-noise processing; and for the time domain audio channel reduction combination according to the audio channel reduction combination information, where the audio channel reduction combination is configured for the domain audio channel reduction combination of at least some blocks of information according to the combination information of audio channel reduction in the case M <N. At least one of A4, B4 and 4C is true:
A4 means that the means for end-end decoding includes a means for the block-by-block determination of whether to apply frequency domain audio channel reduction combination or time domain audio channel reduction combination, and a means for the application of frequency domain audio channel reduction combination, where the means for the frequency domain audio reduction combination application applies frequency domain audio channel reduction combination for the particular block, if the combination reduction application of the frequency reduction is determined for a particular block frequency domain audio channels;
B4 means that the means for the time domain audio channel reduction combination performs the evaluation of whether the audio channel reduction combination information is changed with respect to the audio channel reduction combination information. previously used, and if changed, apply cross-off in order to determine the cross-off audio channel reduction combination information and perform the time domain audio channel reduction combination according to the cross-off audio channel reduction combination information, and if it is not changed, directly apply the time domain audio channel reduction combination according to the
Combination information of audio channel reduction; Y
C4 means that the apparatus includes a means for the identification of one or more non-contributing channels of the input channels Nn, where a non-contributing channel is a channel that does not contribute to the Mm channels, and that the means for end decoding final does not carry out the reverse transformation of the frequency domain information; and the application of additional processing in one or more identified non-contributing channels.
Particular embodiments include a system for decoding audio information that includes Nn channels of encoded audio information in order to form decoded audio information that includes Mm channels of decoded audio, M ^ 1, where n is the number of channels of Low frequency effects on encoded audio information, and m is the number of channels of low frequency effects on decoded audio information. The system comprises: one or more processors; and a storage subsystem coupled to one or more processors. The system must accept the audio information that includes blocks of Nn channels of encoded audio information, encoded by an encoding method, where the encoding method includes the transformation of Nn channels of Digital Audio Information, and the formation and packaging of frequency domain and mantissa exponent information; and in addition, you must decode the accepted audio information, which includes: unpacking and decoding the frequency domain and mantissa exponent Information; the determination of the transformation coefficients from the unpacked and decoded information of frequency domain and mantissa exponent; the reverse transformation of the frequency domain information and the application of additional processing in order to determine samples of audio information; and the time domain audio channel reduction combination of at least some blocks of the Audio Information samples determined in accordance with the audio channel reduction combination information for the case M <N. At least one of A5, B5 and C5 is true:
A5 means that decoding includes the determination, block by block, of whether to apply frequency domain audio channel reduction combination or time domain audio channel reduction combination, and if the application is determined for a particular block the application Combination of frequency domain audio channel reduction, the application of frequency domain audio channel reduction combination for the particular block;
B5 means that the time domain audio channel reduction combination includes the evaluation of whether the audio channel reduction combination information is changed with respect to the previously used audio channel reduction combination information, and whether is changed the cross-off application in order to determine the cross-off audio channel reduction combination information and the realization of time domain audio channel combination according to the Combined Audio Channel Reduction Combination Information crusade; and if not changed, the direct realization of time domain audio channel reduction combination according to the Audio Channel Reduction Combination Information; Y
C5 means that the method includes the identification of one or more non-contributing channels of the input channels Nn, where a non-contributing channel is a channel that does not contribute to the Mm channels, and that the method does not perform the inverse transformation of frequency domain information; and the application of additional processing in one or more identified non-contributing channels.
In some versions of the system embodiment, the accepted audio information is presented in the form of a bit stream of encoded information frames, and the storage subsystem is configured with instructions that, when executed by one or more of the Processor processors produce the decoding of the accepted audio information .
Some versions of the system embodiment include one or more 10 subsystems that are networked through a network connection, where each subsystem includes at least one processor.
In some embodiments in which A1, A2, A3, A4 or A5 is true, the determination of whether to apply frequency domain audio channel reduction combination or time domain audio channel reduction combination includes the determination of if there is any transient prior noise processing, and the determination of whether any of the N channels has a different type of block, so that the frequency domain audio channel reduction combination is applied only for a block that has the same type of block on the N channels, without prior transient noise processing, and
M <N.
In some embodiments in which A1, A2, A3, A4 or A5 is true, and where the transformation in the coding method uses an overlay transformation, and the additional processing includes the application of window and overlay-addition operations, in order to determine samples of audio information, (i) the application of frequency domain audio channel reduction combination for the particular block includes the determination of whether the audio channel reduction combination for the previous block was by domain domain audio channel reduction combination of time, and if the combination of audio channel reduction for the previous block was by combination of time domain audio channel reduction, the application of time domain audio channel reduction combination (or combination of audio channel reduction in a pseudo time domain) to the information of the previous block that must be superimposed with the decoded information of the particular block; and (i) the time domain audio channel combination combination application for a particular block includes the determination of whether the audio channel reduction combination for the previous block was by combination of audio channel reduction of frequency domain, and if the combination of audio channel reduction for the previous block was by combination of frequency domain audio channel reduction, the processing of the particular block differently from the case in which the combination of audio channel reduction for the previous block was not by combination of frequency domain audio channel reduction.
In some embodiments in which B1, B2, B3, B4 or B5 is true, at least one x86 processor is used whose instruction set includes vector instructions comprising multiple information extensions of a single direct current instruction (SSE, according to its acronym in English), and the time domain audio channel reduction combination includes the execution of vector instructions in at least one of one or more of the x86 processors.
In some embodiments in which C1, C2, C3, C4 or C5 is true, n = 1 and m = 0, so that the inverse transformation and the additional processing application are not carried out in the low frequency effect channel . In addition, in some embodiments in which C is true, the audio information that includes encoded blocks comprises information that defines the combination of audio channel reduction, and where the identification of one or more non-contributing channels uses the information that defines the combination of audio channel reduction. Also, in some embodiments in which C is true, the identification of one or more non-contributing channels also includes the identification of whether one or more channels have an insignificant amount of content in relation to one or more additional channels, where one channel has an insignificant amount of content in relation to another channel if its energy or absolute level is at least 15 dB lower than that of the other channel. For some cases, a channel has an insignificant amount of content in relation to another channel if its energy or its absolute level is at least 18 dB lower than that of another channel, while for other applications, a channel has an insignificant amount of content in relation to another channel if its energy or its absolute level is at least 25 dB lower than that of another channel.
In some embodiments, the encoded audio information is encoded according to one of the group of standards consisting of the AC-3 standard, the E-AC-3 standard, a retroactive standard compatible with the E-AC3 standard, the MPEG standard -2 AAC and the HE-AAC standard.
In some embodiments of the Invention, the transformation in the coding method uses an overlay transformation, and the additional processing includes the application of window and overlay operations in order to determine samples of audio information.
In some embodiments of the invention, the coding method includes the formation and packaging of metadata related to the frequency domain and mantissa exponent information, where the metadata optionally includes metadata related to the processing of transient prior noise and with the combination of audio channel reduction.
Particular embodiments may provide all, some or none of these aspects, features or advantages. Particular embodiments may provide one or more additional aspects, features or advantages, one or more of which may be apparent without difficulty to the person skilled in the art, from the figures, descriptions and claims of the present application.
Decoding of a coded stream.
Embodiments of the present invention are described for audio decoding that has been encoded in accordance with the Extended AC-3 (EAC-3) standard, in an encoded bit stream. The E-AC-3 standard and the previous AC-3 standard are described in detail in the Advanced Television Systems Committee, Inc., (ATSC): “Digital Audio Compresslon Standard (AC-3, E-AC-3),” Revision B, Document A / 52B, June 14, 2005 (“Digital Audlo Compression Standard (AC-3, E-AC-3)”, Revision B, Document A / 52B, June 4, 2005), retrieved on 1 December 2009 on the Internet at www<sup>TO</sup>point<sup>TO</sup>atsc<sup>TO</sup>point<sup>TO</sup>org / standards / a_52b<sup>TO</sup>point<sup>TO</sup>pdf (where <sup>TO</sup>point<sup>TO</sup> denotes the dot sign (".") in the actual Internet address). However, the invention is not limited to the decoding of a bit stream encoded in E-AC-3, and can be applied to a decoder and for decoding a bit stream encoded according to another encoding method, since methods for said decoding, decoding apparatus, systems for carrying out said decoding, to computer programs that, when executed, cause one or more processors to carry out said decoding, and to tangible storage media in which said computer programs are stored. For example, the embodiments of the present invention are also applicable to audio decoding that has been encoded according to the MPEG-2 AAC (ISO / IEC 13818-7) and MPEG-4 Audio (ISO / IEC) audio standards.
14496-3). The MPEG-4 Audio standard includes both High Efficiency coding
AAC version 1 (HE-AAC v1) as the High Efficiency coding AAC version 2 (HE-AAC v2), jointly referred to in the present application, HEAAC.
AC-3 and E-AC-3 are also known as DOLBY ® DIGITAL and DOLBY ® DIGITAL PLUS. A version of HE-AAC that incorporates some additional compatible enhancements is also known as DOLBY © PULSE. These are registered trademarks of Dolby Laboratories Licensing Corporation, the assignee of the present invention, and may be registered in one or more jurisdictions. E-AC-3 is compatible with AC-3 and includes additional functionality.
The x86 architecture.
The term x86 is commonly understood by those skilled in the art by referring to a family of processor instruction group architectures whose origins date back to the Intel 8086 processor. The architecture has been implemented in processors of companies such as Intel, Cyrix, AMD, VIA, and many others. In general, it is understood that the term implies binary compatibility with the 32-bit instruction group of the Intel 80386 processor. Currently (early 2010), the x86 architecture is ubiquitous between desktop computers and laptops, as in a growing majority of servers and workstations. A large number of computer programs support the platform, which includes operating systems such as MS-DOS, Windows, Linux, BSD, Solaris and Mac OS X.
In accordance with the present application, the term "x86" means an x86 processor instruction group architecture that also supports an SSE, an English acronym of: single instruction multiple information instruction group extension (SIMD, according to its acronym in English). SSE is an extension of a single instruction multiple information instruction group (SIMD) of the original x86 architecture, introduced in 1999 in the Intel's Pentium III series processors, and now, common in x86 architectures manufactured by many vendors.
AC-3 and E-AC-3 bit currents.
An AC-3 bit stream of a multi-channel audio signal is composed of frames, which represent a constant time interval of 1536 pulse code modulated samples (PCM) of the audio signal a through all encoded channels. Up to five main channels are provided and, optionally, a low frequency effects (LFE) channel called ".1"; that is, up to 5.1 audio channels are provided. Each frame has an established size, which depends only on the sample rate and the encoded data rate.
In summary, AC-3 coding includes the use of an overlay transformation — the modified separate cosine transformation (MDCT) with a window derived from Kaiser Bessel (KBD) with 50 Overlay% — in order to convert time information to frequency information. The frequency information is perceptually encoded so as to compress the information to form a stream of compressed bits of frames each including encoded audio information and metadata. Each AC-3 frame is an independent entity, which does not share information with previous frames except the transformation overlay inherent in the MDCT used to convert time information into frequency information.
At the beginning of each AC-3 frame are the SI (Synchronization Information) and BSI (Bit Current Information) fields. The
J SI and BSI fields describe the bitstream configuration, which includes the sample rate, the data rate, the number of encoded channels, and several other system level elements. There are also two CRC words (cyclic redundancy code) per frame, one at the beginning and one at the end, which provide a means of error detection.
Within each frame are six audio blocks, each representing 256 PCM samples per encoded channel of audio information. The audio block contains the block change flags, coupling coordinates, exponents, bit allocation parameters and mantissa. It is allowed to share information within a frame, so that the information present in Block 0 can be reused in subsequent blocks.
An optional auxiliary information field is located at the end of the frame. This field allows system designers to embed private control or status information in the AC-3 bit stream for broad system transmission.
E-AC-3 preserves the AC-3 frame structure of six 256 coefficient transformations, and at the same time, allows shorter frames composed of one, two and three 256 coefficient transformation blocks. This allows audio transport at data rates greater than 640 kbps (kllobit per second). Each E-AC-3 frame includes metadata and audio information.
E-AC-3 allows a significantly larger number of channels than the AC-3 5.1; in particular, E-AC-3 allows the transport of 6.1 and 7.1 common audio today, and the transport of at least 13.1 channels to support, for example, audio soundtracks of multiple future channels. Additional channels beyond 5.1 are obtained by associating the bit stream of the main audio program with up to eight additional dependent subcurrents, all of which are multiplied into an E-AC-3 bit stream. This allows the main audio program to transport the 5.1-channel format of AC-3, while the capacity of additional channels comes from the dependent bit streams. This means that a 5.1-channel version and the various combinations of conventional audio channels are always available, and that encoding artifacts induced by matrix subtraction are eliminated through the use of a channel substitution process.
In addition, multiple program support is available through the ability to transport seven additional independent audio streams, each with possible associated dependent subcurrents, in order to increase the channel transport of each program, beyond 5.1 channels
AC-3 uses a relatively short transformation and simple scalar quantification to perceptually encode audio material. E-AC-3, while supporting AC-3, provides improved spectral resolution, improved quantization and improved coding. With E-AC-3, the coding efficiency has increased with respect to that of AC-3, in order to allow the beneficial use of lower data rates. This is achieved using an improved filter bank, in order to convert time information into frequency information, improved quantification, improved channel coupling, spectral extension and a technique called transient preprocessing noise processing (TPNP). .
In addition to the overlay transformation MDCT to convert Time Information to Frequency Information, E-AC-3 uses a hybrid adaptive transformation (AHT) for stationary audio signals. The AHT includes the MDCT with the Kaiser Bessel derived window (KBD) overlay, followed, for stationary signals, of a secondary block transformation in the form of a separate non-overlapping, non-window Type II cosine transformation (DCT, according to its English acronym). The AHT thus adds a second stage of DCT after the existing AC-3 MDCT / KBD filter bank when audio with stationary characteristics is presented to convert the six 256-coefficient transformation blocks into a single transformation block. hybrid of 1536 coefficients, with higher frequency resolution. This higher frequency resolution is combined with 6-dimensional vector quantification (VQ) and gain adaptation quantification (GAQ) to improve coding efficiency for some signals. , for example, the "hard to encode" signals. VQ is used to efficiently encode frequency bands that require lower accuracy, while GAQ provides greater efficiency when higher accuracy quantification is required.
The improved coding efficiency is also obtained by using channel coupling with phase preservation. This method expands on the AC-3 channel coupling method of using a high frequency mono compound channel that reconstitutes the high frequency portion of each decoding channel. The addition of encoder-controlled phase and processing information, of spectral amplitude information sent in the bit stream improves the fidelity of this process, so that the mono composite channel can be extended at frequencies lower than previously possible. This decreases the encoded effective bandwidth, and consequently increases the coding efficiency.
E-AC-3 also includes spectral extension. The spectral extension includes the replacement of higher frequency transformation coefficients with lower frequency spectral segments translated in frequency. The spectral characteristics of the transferred segments coincide with the original complete spectral modulation of the transformation coefficients, and in addition, the complete torsion of noise components shaped with the less frequently transferred spectral segments.
E-AC-3 includes a low frequency effects channel (LFE). This is an optional single channel of limited (<120 Hz) bandwidth, which is intended to be played at a level of +10 dB with respect to full bandwidth channels. The optional LFE channel allows the provision of high sound pressure levels for low frequency sounds. Other coding standards, for example, AC-3 and HE-AAC, also include an optional LFE channel.
An additional technique for improving audio quality at low data rates is the use of transient prior noise processing, described in more detail below.
AC-3 decoding.
In typical implementations of AC-3 decoders, in order to keep the decoder and memory latency requirements as low as possible, each AC-3 frame is decoded in a series of nested circuits.
A first step establishes the frame alignment. This means finding the AC-3 synchronization word, and then confirming that the CRC error detection words do not indicate errors. Once frame synchronization is found, the BSl information is unpacked in order to determine important frame information such as the number of encoded channels. One of the channels may be an LFE channel. The number of coded channels is denoted in this application as Nn, where n is the number of LFE channels, and N is the number of main channels. In the coding standards currently used, n = 0 or 1. In the future, there may be cases where n> 1.
The next stage in decoding is unpacking each of the six audio blocks. In order to minimize the memory requirements of the information pulse modulated buffers (PCM), the audio blocks are unpacked one at a time. At the end of each block period, PCM results, in many implementations, are copied to output buffers, which for real-time operation on a hardware decoder, typically are double or circular buffers for direct interrupt access. by a digital to analog converter (DAC, according to its acronym in English).
The AC-3 decoder audio block processing can be divided into two distinct stages, referred to in this application as the input and output processing. The input processing includes any maneuver of coded and unpacked bit stream channels. The output processing refers mainly to the window and overlay stages of the inverse MDCT transformation.
This distinction is made because the number of main output channels, in this application, called M> 1, generated by an AC-3 decoder, does not necessarily coincide with the number of main input channels, in this application, called N, N> 1 encoded in the bit stream, where typical, although not necessarily, N> M. By using the audio channel reduction combination, a decoder can accept a bit stream with any amount of N of encoded channels, and produce an arbitrary amount M, M> 1, of output channels. Note that, in general, the number of output channels is denoted Mm in this application, where M is the number of main channels, and m is the amount of LFE output channels. In current applications, m = 0 or 1. It may be possible to have m> 1 in the future.
Note that, in the audio channel reduction combination, not all encoded channels are included in the output channels. For example, in a combination of audio channel reduction from 5.1 to stereo, the LFE channel information is usually discarded. Consequently, in some combinations of audlo channels, n = 1 and m = 0, that is, there is no LFE output channel.
FIG. 1 sample. Pseudocode 100 for Instructions, which when executed, carry out a typical AC-3 decoding process.
The input processing in the AC-3 decoding typically begins when the decoder unpacks the established audio block information, which is a set of parameters and flags located at the beginning of the audlo block. This established information includes items such as block change flags, coupling information, exponents, and bit allocation parameters. The term "established information" refers to the fact that word sizes for these bitstream elements are known a priori, and therefore, a variable length decoding process is not required to retrieve those elements.
The exporters form the largest individual field in the established Information region, since they include all the exponents of each encoded channel. According to the coding mode, in AC-3, there can be as many as one exponent per mantissa, up to 253 mantissa per channel. Instead of unpacking all these exponents into local memory, many decoder implementations keep pointers for exponent fields, and unpack them when necessary, one channel at a time.
Once the established information is unpacked, many known AC-3 decoders begin processing each encoded channel.
First, the exponents for the given channel are unpacked from the input frame. Next, a bit allocation calculation is typically performed, which takes the exponents and the bit allocation parameters and computes the word sizes for each unpacked mantissa. Mantisas are then typically unpacked from the input frame. The mantises are graduated to provide the appropriate dynamic range control, and if necessary, to undo the coupling operation, and then, they are denormalized by the exponents. Finally, an inverse transformation is computed in order to determine the information prior to the superposition-addition, information in what is called the "window domain", and the results are subjected to the combination of reduction of audio channels in the buffers of combination of reduction of appropriate audio channels, for subsequent output processing.
In some implementations, the exponents for the individual channel are unpacked in a 256-sample long buffer, called the "MDCT buffer". These exponents are then grouped into as many as 50 bands for bit allocation purposes. The number of exponents in each band increases towards higher audio frequencies, following approximately a logarithmic division that models the psychoacoustic critical bands.
For each of these bit allocation bands, the exponents and the bit allocation parameters are combined in order to generate a mantissa word size for each mantissa in said band. These word sizes are stored in a 24 sample long band buffer, where the widest bit allocation band consists of 24 frequency bins.
Once the word sizes have been computed, the corresponding mantissa are unpacked from the input frame and stored in place again in the band buffer. These mantissa are graduated and denormalized by the corresponding exponent, and written, for example, written on site again in the MDCT buffer. After all bands have been processed, and all mantras are unpacked, any remaining location in the MDCT buffer is typically written with zeros.
A reverse transformation is performed, for example, on site, in the MDCT buffer. The output of this processing, the window domain information, can then be subjected to the audio channel reduction combination in the appropriate audio channel reduction combination buffers, in accordance with the channel reduction combination parameters audio, determined according to metadata, for example, brought from the predefined information according to the metadata.
Once the input processing is complete, and the audio channel reduction combination buffers have been fully generated with window domain audio channel combination information, the decoder can perform the output processing . For each output channel, a combination buffer of audio channel reduction and its corresponding half-block delay buffer, 128 samples long, are subjected to the window operation and combined to produce 256 PCM output samples. In a sound support sound system that includes a decoder and one or more DACs, these samples are rounded to the DAC word width and copied to the output buffer. Once this is done, half of the audio channel reduction combination buffer is then copied to its corresponding delay buffer, to provide the 50% overlay information necessary for the proper reconstruction of the next audio block.
E-AC-3 decoding.
Particular embodiments of the present invention include an operation method of an audio decoder for decoding audio information that includes an amount, denoted Nn, of encoded audio information channels, for example, an audio decoder of E- AC-3 for decoding of encoded audio information of E-AC-3 to form decoded audio information that includes Mm channels of decoded audio, n = 0o1, m = 0o1, and M £ 1. The expression n = 1 indicates an input LFE channel; m = 1 indicates an output LFE channel; M <N indicates combination of audio channel reduction; M> N Indicates combination of audio channel increase.
The method includes the acceptance of the audio information that includes Nn channels of encoded audio information, the encoding by the coding method, for example, by an encoding method that includes the transformation using an overlay transformation of N channels of Information of digital audio, training and packaging of frequency domain and mantissa exponent Information, and the formation and packaging of metadata related to the frequency domain and mantissa exponent information, where the metadata optionally includes metadata related to the processing of prior transient noise, for example, by an E-AC-3 coding method .
Some embodiments described in this application are designed to accept encoded audio information, encoded in accordance with the EAC-3 standard or in accordance with a retroactive standard compatible with the E-AC-3 standard, and may include more than 5 encoded main channels .
As will be described in more detail below, the method includes decoding of the Accepted Audio Information, where decoding j
includes: unpacking of metadata and unpacking and decoding of frequency domain and mantissa exponent information; the determination of the transformation coefficients from the unpacked and decoded information of frequency domain and mantissa exponent; the inverse transformation of frequency domain information; the application of window operations and addition overlay, in order to determine samples of audio information; the application of any transient pre-noise processing decoding required in accordance with the metadata related to the transient pre-noise processing; and in the case of M <N, the audio channel reduction combination according to the audio channel reduction combination information. The audio channel reduction combination Includes the evaluation of whether the audio channel reduction combination information is changed with respect to the previously used audio channel combination information, and if it is changed, the cross-off application in order to determine the cross-off audio channel reduction combination information and perform the audio channel reduction combination according to the cross-off audio channel reduction combination information; and if it is not changed, the combination of direct audio channel reduction, according to the combination information of audio channel reduction.
In some embodiments of the present invention, the decoder uses at least one x86 processor that executes SSE, single instruction multiple information extension (SIMD) instructions in direct current, which includes vector instructions. In said embodiments, the audio channel reduction combination includes the execution of vector instructions in at least one of one or more of the x86 processors.
In some embodiments of the present invention, the E-AC-3 audio decoding method, which could be AC-3 audio, is divided into operation modules that can be applied more. at once, that is, appeared more than once in different decoder implementations. In the case of a method that includes decoding, the decoding is divided into a group of start-end decoding operations (EDF) and a group of end-end decoding operations (BED, according to its English acronym). As will be detailed below, the start-end decoding operations include unpacking and decoding frequency domain exponent and mantissa information from a frame of an AC-3 or E-AC-3 bit stream in information unpacked and decoded of frequency domain and mantissa exponent for the frame, and the attached metadata of the frame. The end-end decoding operations include the determination of the transformation coefficients, inverse transformation of the determined transformation coefficients, the application of window and addition overlay operations, the application of any required transient pre-noise processing decoding, and the audio channel reduction combination application, in the case that there are fewer output channels than the channels encoded in the bit stream.
Some embodiments of the present invention include a computer-readable storage medium, which stores instructions that, when executed by one or more processors of a processing system, cause the processing system to perform the decoding of audio information that includes Nn channels of encoded audio information, in order to form decoded audio information that includes Mm channels of decoded audio, M £ 1. In current standards, n = 0 or 1 and m =
O or 1, although the invention is not limited thereto. The instructions include instructions that, when executed, result in the acceptance of audio information that includes Nn channels of encoded audio information, encoded by an encoding method, for example, AC-3 or E-AC-3. The instructions also include instructions that, when executed, produce the decoding of the accepted audio information.
In some of these embodiments, the accepted audio information is presented in the form of an AC-3 or E-AC-3 bit stream of encoded information frames. The instructions that, when executed, produce the decoding of the accepted audio information, are divided into a group of reusable instruction modules, which include a start-end decoding module (EDF) and an end-end decoding module (BED) The start-end decoding module includes instructions that, when executed, perform the unpacking and decoding of the frequency domain exponent information and maintaining a frame of the bit stream, in unpacked and decoded exponent information frequency domain and mantissa for the frame, and the attached metadata of the frame. The end-end decoding module includes instructions that, when executed, produce the determination of the transformation coefficients, the inverse transformation, the application of window operations and addition overlay, the application of any transient pre-noise processing decoding required, and the audio channel reduction combination application, in the case that there are fewer output channels than encoded input channels.
FIGS. 2A-2D show, in the form of simplified block diagrams, some different decoder configurations that can conveniently use one or more common modules. FIG. 2A shows a simplified block diagram of an exemplary decoder of E-AC-3 200, for 5.1 audio encoded by AC-3 or E-AC-3. Naturally, the use of the term "block", when referring to the blocks in a block diagram, is not the same as an audio information block, where the latter term refers to a quantity of audio information. The decoder 200 includes a start-end decoder module (EDF) 201 that must accept AC-3 or E-AC-3 frames and carry out, frame by frame, unpacking the frame metadata, and decoding of the audio information of the frame in exponent information of frequency domain and mantissa. The decoder 200 also includes an end-end decoder module (BED) 203, which accepts the frequency domain exponent and mantissa information from the start-up end decoding module 201 and decodes it up to 5.1 channels of PCM Audio Information.
The decomposition of the decoder into a Home end decoding module and an end end decoding module is a design choice, not a necessary fractionation. Said fractionation does provide benefits of having common modules in several alternate configurations. The EDF module may be common to such alternate configurations, and many configurations have in common the unpacking of frame metadata and decoding of frame audio information in frequency domain and mantissa exponent information carried out by an EDF module.
As an example of an alternate configuration, FIG. 2B shows a simplified block diagram of an E-AC-3 210 decoder / converter for 5.1 audio encoded by E-AC-3, which decodes 5.1 audio encoded by both AC-3 and E-AC-3, and also , converts an EAC-3 encoded frame of up to 5.1 audio channels into an AC-3 encoded frame of up to
5.1 channels The decoder / converter 210 includes a start-end decoder module (EDF) 201 that accepts frames of AC-3 or E-AC-3 and performs, frame by frame, unpacking the frame metadata and the decoding of frame audio information in frequency domain and mantissa exponent information. The decoder / converter 210 also includes an end-end decoder module (BED) 203, which is the same as the BED module 203, or similar, of the decoder 200, and which accepts the frequency domain and mantissa exponent information from the starter decoder module 201, and decodes it up to 5.1 channels of PGM audio information. The decoder / converter 210 also includes a metadata converter module 205, which converts metadata, and an end-end coding module 207, which accepts the frequency domain exponent and mantissa information of the start-end decoding module 201 and encodes the information as an AC-3 frame of up to 5.1 channels of audio information, at a maximum data rate of not more than 640 kbps possible with AC-3.
As an example of an alternate configuration, FIG. 2C shows a simplified block diagram of an E-AC-3 decoder that decodes an AC-3 frame of up to 5.1 channels of encoded audio, and also decodes an E-AC-3 encoded frame of up to 7.1 channels of Audio. The decoder 220 includes a Frame Information analysis module 221 that unpacks the BSI information and identifies the frames and types of frames, and provides the frames for the appropriate start-end decoder elements. In a typical implementation that includes one or more processors and memory in which instructions are stored that, when executed, carry out the functionality of the modules, multiple occurrences of a start-end decoding module and multiple occurrences of an end-end decoding module. In some embodiments of an E-AC-3 decoder, the unpacking functionality of BSl is separated from the starting end decoding module to look at the BSl information. This provides common modules per used in various alternate implementations. FIG. 2C shows a simplified block diagram of a decoder with said architecture, suitable for up to 7.1 channels of audio information. FIG. 2D shows a simplified block diagram of a 5.1 240 decoder, with said architecture. The decoder 240 includes a frame information analysis module 241, a start end decoding module 243, and an end end decoding module 245. These FED and BED modules can be of a structure similar to the FED modules and BED used in the architecture of FIG. 2 C.
Returning to FIG. 2C, the Frame Information analysis module 221 provides the information of a frame encoded by AC-3 / E-AC3 independent of up to 5.1 channels, to a start-end decoding module 223 that accepts AC-3 frames or E-AC-3 and carries out, frame by frame, the unpacking of frame metadata and decoding of frame audio information in frequency domain and mantissa exponent information. The frequency domain and mantissa exponent information is accepted by an end-end decoding module 225 which is the same as the BED module 203, or similar, of the decoder 200, and which accepts the frequency domain exponent information and Mantissa from the start end decoding module 223, and decodes the information to 5.1 channels of PCM audio information. Any frame encoded by dependent AC-3 / EAC3, additional channel information is provided to another start-end decoding module 227, which is similar to the other EDF module, and thus unpacks the frame metadata and decodes it. the audio information of the plot in exponent information of frequency domain and mantissa. An end-end decoding module 229 accepts the information from the EDF module 227 and decodes the information in PCM audio information of any additional channel. A PCM 231 channel mapping module is used to combine the decoded information of the respective BED modules to provide up to 7.1 channels of PCM information.
If there are more than 5 coded main channels, that is, in the case where N> 5, for example, there are 7.1 coded channels, the coded bitstream includes an independent frame of up to 5.1 coded channels, and at least one frame dependent on coded information In embodiments of computer programs for said case, for example, embodiments comprising a computer-readable medium that stores instructions for execution, the instructions are arranged as a plurality of decoding modules.
5.1 channels, where each 5.1 channel decoding module includes a respective appearance of a start end decoding module and a respective appearance of an end end decoding module. The plurality of 5.1 channel decoding modules includes a first 5.1 channel decoding module which, when executed, produces decoding of the independent frame, and one or more additional channel decoding modules for each respective dependent frame. In some of these embodiments, the instructions include an instruction frame information analysis module that, when executed, produces the unpacking of the Bit Current Information (BSI) field of each frame, to identify the frames and types of frames, and provides the frames
Identified to the appropriate appearance of the starter decoder module; and an instruction channel mapping module, which when executed, and in the case of N> 5, produces the combination of the decoded information of the respective end-end decoding modules to form the main N channels of decoded information.
A method for the operation of an AC-3 / E-AC-3 dual converter decoder.
An embodiment of the invention is in the form of a dual converter decoder (DDC, according to its acronym in English) which decodes two input bit streams of AC-3 / E-AC-3, designated "main" and "associated" , with up to 5.1 channels each, in PCM audio, and in the case of conversion, converts the main audio bit stream of E-AC-3 to AC-3, and in the case of decoding, decodes the current of main bits, and if present, the associated bit stream. The dual converter decoder optionally combines the two PCM outputs using the combination of metadata extracted from the associated audio bit stream.
An embodiment of the dual converter decoder performs a method of operating a decoder, to carry out the processes included in the decoding and / or conversion of up to two input bit streams of AC3 / E-AC-3. Another embodiment is presented in the form of a tangible storage medium that has instructions, for example, computer program instructions, which when executed by one or more processors of a processing system, causes the processing system to carry out the processes included in the decoding and / or conversion of up to two input bit streams of AC-3 / E-AC-3.
An embodiment of the AC-3 / E-AC-3 dual converter decoder has six subcomponents, some of which include common subcomponents.
The modules are:
Decoder-converter: The decoder-converter is configured when it is executed to decode an input bit stream of AC3 / E-AC-3 (up to 5.1 channels) into PCM audio, and / or convert the input bit stream of E -AC-3 in AC-3. The decoder-converter has three main subcomponents, and can implement an embodiment 210 set forth in FIG. 2B above. The main subcomponents are:
Start-end decoder: The EDF module is configured, when executed, to decode a frame of an AC3 / E-AC-3 bit stream in Raw Frequency Domain Audio Information and its attached metadata.
End-end decoder: The BED module is configured, when executed, to complete the rest of the decoding process that was initiated by the FED module. In particular, the BED module decodes the audio information (in mantissa and exponent format) in PCM audio information.
End-end encoder: The end-end encoder module is configured, when executed to encode an AC-3 frame using six blocks of audio information from the EDF. The end-end encoder module is also configured, when executed, to synchronize, resolve and convert E-AC-3 metadata into Dolby Digital metadata using a Metadata converter module Included.
5.1 Decoder: The 5.1 decoder module is configured when it is executed to decode an input bit stream of AC3 / E-AC-3 (up to 5.1 channels) in PCM audio. The 5.1 decoder also, optionally, outputs combination metadata for use by an external application to mix two streams of AC-3 / E-AC-3 bits. The decoder module includes two main subcomponents: a module
EDF, as described in this application before, and a module
BED, as described in this application before. A block diagram of an example 5.1 decoder is shown in FIG. 2D.
Frame information: The Frame Information module is configured 5 when it is executed to analyze an AC-3 / E-AC-3 frame and unpack its Bitstream Information. CRC control is performed in the frame, as part of the unpacking process.
Buffers descriptors: The buffers descriptor module contains descriptions of buffers of AC-3, E-AC-3 and PCM, and functions for buffer operations.
Sample index converter: The sample index converter module is optional, and is configured, when executed, to increase the amount of PCM audio samples by a factor of two.
External mixer: The external mixer module is optional, and is configured when it is executed to mix a main audio program and an associated audio program into a single output audio program using combination metadata provided in the associated audio program .
Starter decoder module design.
The start-end decoder module decodes information in accordance with AC-3 methods, and in accordance with additional decoding aspects of E-AC-3, including decoding AHT information for stationary signals, coupling Improved channel E-AC-3 and spectral extension.
In the case of an embodiment in the form of a tangible storage medium, the startup end decoding module comprises computer program instructions stored in a tangible storage medium, which, when executed by one or more processors of a system processing, produce the actions described in the details provided in this request for the operation of the start-end decoding module. In a physical support implementation, the start-end decoder module includes elements that are configured in the operation, to perform the actions described in the details provided in this request, for the operation of the start-end decoder module. .
In AC-3 decoding, block by block decoding is possible. With E-AC-3, the first audio block — the 0 audio block of a frame includes the AHT mantissa of the 6 blocks. Consequently, block-by-block decoding is typically not used, but instead, several blocks are processed at once. The processing of real information, however, carried out, of course, in each block.
In one embodiment, in order to use a uniform decoding / architecture method of a decoder, regardless of the use of AHT, the EDF module carries out two passes, channel by channel. A first pass includes unpacking block-by-block metadata and saving pointers for places where exponent and mantissa information is stored, and a second pass includes the use of pointers stored for packed exponents and mantissa, and the unpacking and decoding of exponent and mantissa information, channel by channel.
FIG. 3 shows a simplified block diagram of an embodiment of a start-end decoder module, for example, implemented as a set of instructions stored in a memory, which when executed, produce EDF processing. FIG. 3 it also shows the pseudocode for instructions for a first pass of the two-pass start decoder module 300, as well as the pseudocode for instructions for the second pass of the two-pass start end decoder module. The EDF module includes the following modules, where each one includes instructions, where some of these instructions are defining, in terms of defining structures and parameters:
Channel: The channel module defines structures for the representation of an audio channel in memory, and provides instructions for unpacking and decoding an audio channel of a bit stream of AC-3 or E-AC3.
Bit allocation: The bit allocation module provides instructions for calculating the masking curve, and for calculating the bit allocation for the encoded information.
Continuous current or bit time series operations: The continuous bit stream operation module provides instructions for unpacking information from an AC-3 or E-AC-3 bit stream.
Exponents: The exponent module defines structures for the representation of exponents in memory, and provides configured instructions when executed to unpack and decode exponents of an AC-3 or E-AC-3 bit stream.
Exponents and mantissa: The exponent and mantissa module defines structures for the representation of exponents and mantissa in memory, and provides configured instructions when executed, to unpack and decode exponents and mantissa of a bit stream of AC-3 or E- AC-3
Matrix: The matrix module provides configured instructions when executed to support the dematrization of matrix channels.
Auxiliary Information: The Auxiliary Information module defines auxiliary information structures used in the EDF module to carry out the EDF processing.
Mantissa: The mantissa module defines structures for the representation of mantissa in memory, and provides configured instructions when executed, to unpack and decode mantissa of a bit stream of AC-3 or E-AC-3.
Hybrid adaptation transformation (AHT, according to its acronym in English):
The AHT module provides configured instructions when executed, to unpack and decode hybrid transformation information adapting an E-AC-3 bit stream.
Audio frame: The audio frame module defines structures for the representation of an audio frame in memory, and provides configured instructions when executed, to unpack and decode an audio frame of an AC-3 bit stream or E-AC-3.
Improved coupling: The enhanced coupling module defines structures for the representation of an improved coupling channel in memory, and provides configured instructions when executed, to unpack and decode an improved coupling channel of an AC-3 bit stream or E-AC-3. The enhanced coupling extends the traditional coupling into a stream of E-AC-3 bits by providing phase and chaos information.
Audio block: The audio block module defines structures for the representation of an audio block in memory, and provides configured instructions when executed to unpack and decode an audio block of an AC-3 or E bit stream. -AC-3.
Spectral extension: The spectral extension module provides support for decoding of spectral extension in a bit stream of
E-AC-3.
Coupling: The coupling module defines structures for the representation of a coupling channel in memory, and provides configured instructions when executed, to unpack and decode a coupling channel of a stream of bits of AC-3 or E-AC- 3.
FIG. 4 shows a simplified information flow diagram for the operation of an embodiment of the start-end decoding module 300 of FIG. 3, which describes the manner in which the pseudo code elements and submodules set forth in FIG. 3 cooperate to carry out the functions of a Home end decoding module. The term functional element means an element that performs a processing function. Each of said elements can be a physical support element, or a processing system and a storage medium that includes instructions that, when executed, carry out the function. A bit stream unpacking functional element 403 accepts an AC-3 / EAC-3 frame and generates bit allocation parameters for an AHT 405 standard and / or bit allocation functional element, which produces additional information for the unpacking the bit stream in order to ultimately generate exponent and mantissa information for an improved standard / decouple functional element Included 407. The functional element 407 generates exponent and mantissa information for a functional rematrization element included 409 in order to carry out any necessary rematrization. The functional element 409 generates exponent and mantissa information for an included functional element of spectral extension decoding 411 in order to carry out any necessary spectral extension. The functional elements 407 a
Π
<img file="AR089918A2_D0001.tif" />
411 they use information obtained by the unpacking operation of the functional element 403. The result of the start end decoding is the exponent and mantissa information, as well as additional unpacked audlo frame parameters and audio block parameters.
With reference in more detail to the first pass and second pass pseudocode shown in FIG. 3, the first pass instructions are configured, when executed, for unpacking metadata from an AC-3 / E-AC-3 frame. In particular, the first pass includes unpacking the BSI Information and unpacking the audio frame information. For each block, Starting with block 0 to block 5 (for 6 blocks per frame), the established information is unpacked, and for each channel, a pointer of the exponents packed in the bit stream is saved, the exponents are unpacked , and the position in the bit stream in which the packed mantissa resides is stored. The bit allocation is computed, and, based on the bit allocation, the mantissa can be omitted.
The second pass instructions are configured, when executed, to decode the audio information of a frame, in order to form mantissa and exponent information. For each block, starting with block 0, the unpacking includes loading the saved pointer of the packed exponents, and unpacking the exponents pointed in that way; the computation of bit allocation; the load of the saved pointer of packed mantissa, and the unpacking of the mantissa pointed in that way. The decoding includes the performance of decoupling of standard and improved and the generation of the spectral extension bands, and in order to be independent of other modules, the transfer of the resulting Information to a memory, for example, an external memory to the internal memory of the pass, so that the resulting information can be accessed by other modules, for example, the BED module. This memory, for reasons of convenience, is called "external memory," although, as will be apparent to those skilled in the art, it can be part of a single memory structure used for all modules.
In some embodiments, for the unpacking of exponents, the exponents unpacked during the first pass are not saved, in order to minimize memory transfers. If AHT is in use for a channel, the exponents are unpacked from block 0 and copied to five other blocks, numbered 1 to 5. If AHT is not in use for a channel, the pointers of the packed exponents are saved. If the channel exponent strategy is the reuse of exponents, the exponents are unpacked again using the saved pointers.
In some embodiments, for unpacking coupling mantissa, if the AHT is used for the decoupling channel, the six AHT coupling channel mantle blocks are unpacked in block 0, and regenerated on screen for each channel that is a coupled channel, in order to produce uncorrelated screening. If the AHT is not used for the coupling channel, the pointers of the coupling mantras are saved. These saved pointers are used to redepack the coupling mantras for each channel that is a channel coupled in a given block.
Design of the end-end decoding module.
The end-end decoding module (BED) operates to take the frequency domain exponent information and mantissa and decode it into PCM audio information. PCM audio information is presented on the basis of selected user modes, dynamic range compression and audio channel reduction combination modes.
In some embodiments in which the Home end decoding module stores exponent and mantissa information in a memory - which we call external memory - separated from the working memory of the Start end module, the BED module uses the processing block by block frame, in order to minimize the requirements of the combination of audio channel reduction and delay buffer, and to be compatible with the output of the boot end module, use external memory transfers to access the exponent information and mantissa for processing.
In the case of an embodiment in the form of a tangible storage medium, the end-end decoding module comprises Computer program instructions stored in a tangible storage medium which, when executed by one or more processors of a processing system , produce the actions described in the details provided in this application for the operation of the end-end decoding module. In a physical support implementation, the end-end decoding module includes elements that are configured in the operation to perform the actions described in the details provided in this application for the operation of the end-end decoding module.
FIG. 5A shows a simplified block diagram of an embodiment of an end-end decoding module 500 implemented as a set of instructions stored in a memory, which when executed, produces BED processing. FIG. 5A also shows the pseudocode for instructions for the end-end decoding module 500. The BED 500 module includes the following modules, where each one includes Instructions, where some of these Instructions are of definition:
Dynamic range control: The dynamic range control module provides instructions, which when executed, perform functions for the control of the dynamic range of the decoded signal, which includes the gain variation application, and the control application of rank
Transformation: The transformation module provides instructions, which when executed, produce the inverse transformations, which include the production of a separate modified modified cosine transformation (IMDCT), which includes the pre-rotation production used to calculate the reverse DCT transformation, post-rotation production, used to calculate the reverse DCT transformation, -and the determination of the inverse fast Fourier transformation (IFFT, according to its acronym in English).
Transient pre-noise processing: The transient pre-noise processing module provides instructions, which when executed, carry out the transient pre-noise processing.
Window and addition overlay: The window and overlay module with delay buffer provides instructions, which when executed, perform the window and addition overlay operation to reconstruct output samples from the samples Inverse transforms.
Time domain audio channel reduction (TD) combination: The TD audio channel combination module provides instructions, which when executed, produce the audio channel reduction combination in the time domain, as necessary to obtain a smaller number of channels.
FIG. 6 shows a simplified information flow diagram for the operation of an embodiment of the end-end decoding module 500 of FIG. 5A, which describes the manner in which the code elements and submodules set forth in FIG. 5A cooperate to carry out the functions of an end-end decoding module. A gain control functional element 603 accepts exponent and mantissa information from the start end decoding module 300, and applies any required dynamic range control, dialog normalization, and gain variation, according to the metadata. The resulting exponent and mantissa information is accepted by a functional mantissa denormalization element by. exponents 605, which generates the transformation coefficients for the inverse transformation. A functional transformation-inverse element 607 applies the IMDCT to the transformation coefficients, in order to generate time samples that are subjected to previous window operations and addition overlay. Said preaddition overlay time domain samples are called "time pseudo-domain" samples in this application, and these samples are found in what is called in this application the time-domain pseudo-domain. These are accepted by a functional element of window operation and addition overlay 609, which generates PCM samples by applying window operations and addition overlay to pseudo-domain time samples. Any necessary transient pre-noise processing is applied by a functional transient pre-noise processing element 611 according to the metadata. If specified, for example, in metadata or otherwise, the subsequent PCM samples resulting from subsequent transient preprocessing are subjected to the combination of audio channel reduction until the Mm number of output channels of PCM samples is obtained, by a functional element of 613 audio channel reduction combination.
With reference again to FIG. 5A, the pseudocode for the processing of the BED module includes, for each information block, the transfer of mantissa and exponent information for the blocks of a channel from the external memory, and, for each channel: the application of any control of dynamic range required, dialogue normalization and gain variation, according to metadata; the denormalization of mantissa by exponents, to generate the transformation coefficients, for the inverse transformation; the computation of an IMDCT for the transformation coefficients, in order to generate pseudo-time samples; the application of window operations and superposition of addition to the pseudo-domain time samples; the application of any prior transient noise processing according to the metadata; and, if required, the submission to the time domain audio channel reduction combination until the amount of Mm of output channels from PCM samples is obtained.
The decoding embodiments set forth in FIG. 5A include making such gain adjustments when dialog normalization offsets are applied according to the metadata, and the application of dynamic range control gain factors according to the metadata. Making such gain adjustments at the stage that the information is provided in the form of mantissa and exponent in the frequency domain is convenient. The gain changes may vary depending on the time, and such gain changes made in the frequency domain achieve smooth cross-offs once the inverse and window transformation and addition overlay operations have occurred.
Processing of prior transitory noise.
The encoding and decoding of E-AC-3 were designed to operate and provide better audio quality at lower data rates, than in AC-3. At lower data rates, the quality of the encoded audio audio can be negatively affected, especially, for relatively difficult transient material. This effect on audio quality is mainly due to the limited amount of data bits available to accurately encode these types of signals. Transient coding artifacts are displayed as a reduction in the definition of the transient signal, as is the “transient pre-noise” artifact, which erases audible noise throughout the entire coding window, due to quantization errors of coding.
As described above, and in FIGS. 5 and 6, the BED provides the prior transient noise processing. EAC-3 encoding Includes transient pre-noise processing coding, to reduce transient pre-noise artifacts that can be introduced when transient-containing audio is encoded, by replacing the appropriate audio segment with synthesized audio using the audio located before of the previous transitory noise. The audio is processed using the time grading synthesis, so that its duration is increased so that it is of an appropriate length to replace the audio that contains the transient pre-noise. The audio synthesis buffer is analyzed using audio scene analysis and maximum similarity processing, and then the time is graded so that its duration is increased sufficiently to replace the audio containing the transient pre-noise. The longer synthesized audio is used to replace the transient pre-noise and is subjected to cross-off with the existing transient pre-noise, just before the transient location, in order to ensure a smooth transition from the synthesized audio to the audio information originally coded. Using the transient pre-noise processing, the length of the transient pre-noise can be drastically reduced or eliminated, even for the case where block switching is disabled.
In an embodiment of an E-AC-3 encoder, time grading synthesis analysis and processing for the transient prior noise processing tool are performed on time domain information in order to determine metadata information, for example , time graduation parameters. The metadata information is accepted by the decoder, along with the encoded bit stream. The transmitted transient pre-noise metadata is used to perform time domain processing on the decoded audio, in order to reduce or eliminate the transient pre-noise introduced by the low bit rate audio coding at low data rates.
The E-AC-3 encoder performs the time gradation synthesis analysis and determines the time grading parameters, based on the audio content, for each transient detected. The time grading parameters are transmitted as additional metadata, along with the encoded audio information.
In an E-AC-3 decoder, the optimal time graduation parameters provided in the E-AC-3 metadata are accepted as part of the accepted E-AC-3 metadata for use in prior noise processing transient. The decoder performs the audio buffer splicing and cross-off using the transmitted time grading parameters obtained from the E-AC-3 metadata.
By utilizing the optimal timing information and applying it with the appropriate cross-off processing, the transient pre-noise introduced by the low bit rate audio coding can be drastically reduced or eliminated in decoding.
Therefore, the transient prior noise processing writes about the previous noise with an audio segment that most closely resembles the original content. Transient pre-noise processing instructions, when executed, maintain a four-block delay buffer for use in copying. The instructions for the processing of transient pre-noise, when executed, in the case where overwriting occurs, produce a cross-off input and output over the overwritten previous noise.
Combination of audio channel reduction (downmixin).
Denoted by Nn is the number of channels encoded in the bit stream E-AC-3, where N is the number of main channels, and n = 0 or 1, is the number of LFE channels. Often, it is desired to submit the main N channels to the combination of audio channel reduction to obtain a smaller amount, denoted M, of main output channels. The combination of audio channel reduction from N to M channels, M <N, is sustained by the embodiments of the present invention. The combination of audio channel augmentation is also possible, in which case, M> N.
Therefore, in the more general implementation, - the audio decoder embodiments operate to decode audio information that includes Nn channels of encoded audio information, to decode audio information that includes Mm channels of encoded audio, and M> 1 , where n, m indicate the amount of LFE channels at the input, output, respectively. The combination of audio channel reduction is the case M <N, and according to a parameter of combination coefficients of audio channel reduction, it is included in the case M <N.
Combination of frequency domain audio channel reduction compared to time domain.
The combination of audio channel reduction can be performed entirely in the frequency domain, before the inverse transformation, in the time domain after the inverse transformation, although, in the case of the processing of addition overlay blocks, before of window operations and addition overlay, or in the time domain after window operation and addition overlay.
The frequency domain audio (FD) channel reduction combination is much more efficient than the time domain audio channel reduction combination. Its efficiency arises, for example, from the fact that any processing step after the audio channel reduction combination step is only performed on the remaining amount of channels, which is generally less after the channel reduction combination audio Consequently, the computational complexity of all processing steps after the audio channel reduction combination step reduces at least the ratio of input channels to output channels.
As an example, consider a combination of audio channel reduction from 5.0 channels to stereo. In this case, the computational complexity of any subsequent processing step will be reduced by approximately a factor of 5/2 = 2.5.
The time domain audio (TD) channel reduction combination is used in typical E-AC-3 decoders and in the embodiments described above, and which are illustrated in FIGS. 5A and 6. There are three main reasons why typical E-AC-3 decoders use the time domain audio channel reduction combination:
Channels with different types of blocks.
According to the audio content to be encoded, an encoder of
E-AC-3 can select between two types of different blocks - short block and long block - to segment the audio information. The harmonic, slow-change audio information is typically segmented and encoded using long blocks, while the transient signals are segmented and encoded into short blocks. Consequently, the frequency domain representation of the short blocks and that of the long blocks is inherently different, and cannot be combined in a combination operation of frequency domain audio channel reduction.
Only after the specific block type coding steps are undone in the decoder, can the channels be mixed. Therefore, in the case of changed block transformations, a different partial inverse transformation process is used, and the results of the two different transformations cannot be combined directly until just before the window stage.
However, methods for first converting the short length transformation information into the longer frequency domain information are known, in which case, the combination of audlo channel reduction can be carried out in the frequency domain. Nevertheless; In most of the known decoder implementations, the audio channel reduction combination is carried out after the Inverse transformation according to the audio channel reduction combination coefficients.
Combination of increasing audio channels.
If the number of main output channels is greater than the amount of main input channels, M> N, a time domain combination approach is beneficial, as this moves the step of combining audio channel increase towards the end of processing, to reduce the amount of channels in processing.
TPNP
The blocks that undergo the processing of prior transient noise (TPNP, according to its acronym in English) cannot be subjected to the combination of reduction of audio channels in the frequency domain, since TPNP operates in the time domain. TPNP requires a record of up to four blocks of PCM information (1024 samples), which must be present for the channel on which TPNP is applied. The change to the time domain audio channel reduction combination is therefore necessary to complete the PCM information history and to effect the previous noise substitution.
Combination of hybrid audio channel reduction using combination of audio channel reduction of both frequency domain and time domain.
The inventors recognize that the channels in most encoded audio signals use the same type of block for more than 90% of the time. This means that the most efficient frequency domain audio channel reduction combination would work for more than 90% of the information in typical encoded audio, assuming there is no TPNP. The remaining 10% or less would require a combination of time domain audio channel reduction, as occurs in the typical E-AC-3 decoders of the prior art.
The embodiments of the present invention include selection logic of audio channel reduction combination method in order to determine, block by block, the method of combining audio channel reduction to be applied, and both combination reduction logic of time domain audio channels, such as frequency domain audio channel reduction combination logic, to apply the particular audio channel reduction combination method, as appropriate. Therefore, one method embodiment includes the determination, block by block, of whether to apply frequency domain audio channel reduction combination or time domain audio channel reduction combination. The logic selection method of audio channel reduction combination operates to determine whether to apply frequency domain audio channel reduction combination or time domain audio channel reduction combination, and includes the determination of if there is any prior transient noise processing, and the determination of whether any of the N channels has a different type of block. The selection logic determines that a frequency domain audio channel reduction combination should be applied only to a block that has the same type of block on the N channels, which does not have transient prior noise processing, and M <N.
FIG. 5B shows a simplified block diagram of an embodiment of an end-end decoder module 520 implemented as a set of instructions stored in a memory, which when executed, produces BED processing. FIG. 5B also shows the pseudo code for instructions for the end-end decoding module 520. The BED 520 module includes the<sup>-</sup>modules shown in FIG. 5A that only use the time domain audio channel reduction combination, and the following additional modules, where each includes instructions, where some of these instructions are definition:
The audio channel reduction combination method selection module, which controls (i) the change of the block type; (¡I) if there is no true combination of audio channel reduction (M <N), but instead, combination of audio channel increase, and (iii) if the block is subject to TPNP, and if none of These are true, the combination selection of frequency domain audio channel reduction. This module carries out the determination, block by block, of whether to apply frequency domain audio channel reduction combination or time domain audio channel reduction combination.
The combination module of frequency domain audio channel reduction, which performs, after the distortion of the mantises by the exponents, the combination of frequency domain audio channel reduction. Note that the frequency domain audio channel reduction combination module also includes a time domain to frequency domain transition logic module, which controls whether the preceding block used the domain audio channel reduction combination. of time, in which case, the block is maneuvered differently, as described in more detail below. In addition, the transition logic module also deals with the processing steps associated with certain recurring events not regularly, for example, program changes such as outgoing shutdown channels.
The FD to TD audio channel reduction combination transition logic module, which controls whether the preceding block used frequency domain audio carriage combination combination, in which case, the block is handled differently as Describe in more detail below. In addition, the transition logic module also deals with the processing steps associated with certain recurring events not regularly, for example, program changes such as outgoing shutdown channels.
Likewise, the modules presented in FIG. 5A could behave differently in embodiments that include the hybrid audio channel reduction combination, that is, the FD and TD audio channel reduction combination, in accordance with one or more conditions for the current block.
With reference to the pseudocode of FIG. 5B, some embodiments of the end-end decoding method include, after transferring the information of a block frame from the external memory, evaluating whether reduction combination of FD audio channels or combination of reduction of TD audio channels is applied. For the FD audio channel reduction combination, for each channel, the method includes (i) the application of dynamic range control and dialog normalization, although, as described below, disabling the gain variation; (¡I) the denormalization of mantissa by exponents; (iii) the realization of the FD audio channel reduction combination; and (iv) assess whether there are outgoing shutdown channels, or if the previous block was subjected to the audio channel reduction combination by time domain audio channel reduction combination, in which case, the processing takes Perform differently as described in more detail below. For the case of TD audio channel reduction combination, and also, for the information submitted to FD audio channel combination combination, the process includes, for each channel: (I) the process differently from the blocks, to be subjected to the TD audio channel reduction combination, in the case that the previous block was subjected to the FD audio channel reduction combination, and in addition, the maneuver of any program change; (ii) the determination of the inverse transformation; (iii) performing the window operation and addition overlay; and in the case of the TD audio channel reduction combination, (iv) the realization of any TPNP and audio channel reduction combination until the appropriate output channel is obtained.
FIG. 7 shows a simple information flow chart. The block
701 corresponds to the logic selection method of audio channel reduction combination that evaluates the three conditions: block type change,
TPNP, or combination of increasing audio channels, and if any condition is true, directs the data flow to a combination branch of reduction of
L · TD 721 audio channels included in the FD 723 audio channel reduction combination transition logic to process differently a block that appears immediately following a block processed by the channel reduction combination FD audio; program change processing, and in 725, mantissa denormalization by exponents. The data flow after block 721 is processed by common processing block 731. If the audio logic reduction combination method selection test block 701 determines that the block is for the FD audio channel reduction combination, the data stream is branched to the combination reduction processing of 711 FD audio channels, which includes a combination process of frequency domain audio channel reduction 713 that disables gain variation, and for each channel, denormalises mantises by exponents and performs the FD audio channel reduction combination, and a TD 715 audio channel combination transition logic block, to determine if the previous block was processed by combination of reduction of TD audio channels, and to process said block differently, and in addition, to detect and maneuver any program change, such as outgoing shutdown channels. The data flow after the TD 715 audio channel reduction combination transition block is from the same common processing block 731.
The common processing block 731 includes the inverse transformation and any time domain processing. Additional time domain processing includes the undoing of the gain variation, and the window processing and addition overlay. If the block is from the TD 721 audio channel reduction combination block, the additional time domain processing also includes any TPNP processing and time domain audlo channel reduction combination.
FIG. 8 shows a flow chart of a processing embodiment for an end-end decoding module such as that set forth in FIG. 7. The flowchart is divided as follows, where the same reference numbers are used as in FIG. 7, for similar respective functional data flow blocks: a section of selection logic of the audio channel reduction combination method 701 in which a logical flag FD dmx is used to indicate when 1 said reduction combination of Frequency domain audlo channels are used for the block; a TD 721 audio channel reduction combination logic section that includes an FD audio channel reduction combination transition logic and a program change logic section 723 for differently processing a block that appears immediately following a block processed by FD audio channel reduction combination and program change processing, and a section to denormalize the mantissa by exponents for each input channel. The data flow after block 721 is processed by a common processing section 731. If the audio channel reduction combination method selection logic block 701 determines that the block is for the FD audio channel reduction combination, the data stream branches to the reduction combination processing section. of FD 711 audio channels, which includes a combination process of frequency domain audio channel reduction that disables gain variation, and for each channel, denormalises mantises by exponents and performs the FD audio channel reduction combination, and a TD 715 audio channel reduction combination transition section, in order to determine, for each channel of the previous block , if there is an outgoing channel shutdown or if the previous block was processed by the TD audio channel reduction combination, and to process said block differently. The data flow after the TD 715 audio channel reduction combination transition section goes to the same common processing logic section 731. The common processing logic section 731 includes, for each channel, the reverse transformation and any additional time domain processing. Additional time domain processing includes the undoing of the gain variation, and the window processing and addition overlay. If FD_dmx is 0, indicating the combination of TD audio channel reduction, the additional time domain processing in 731 also includes any TPNP processing and time domain audio channel reduction combination.
Note that after the FD audio channel reduction combination, in the TD 715 audio channel combination transition logic section, at 817, the number of input channels N is set to be the same than the amount of output channels M, so that the rest of the processing, for example, the processing in the common processing logic section 731, It is carried out only on the information subject to a combination of audio channel reduction. This reduces the amount of computing. Naturally, the time domain audio channel reduction combination of the previous block information when there is a transition of a block that was subjected to the TD audio channel reduction combination - said audio channel reduction combination TD set forth as 819 in section 715 - is carried out in all of said input channels N involved in the audio channel reduction combination.
Transition maneuver.
In decoding, it is necessary to have smooth transitions between the audio blocks. E-AC-3 and many other coding methods use a folded transformation, for example, a 50% overlap MDCT. Consequently, when a current block is processed, there is 50% overlap with the previous block, and in addition, there will be 50% overlap with the next block in the time domain. Some embodiments of the present invention use addition overlay logic, which includes an addition overlay buffer. When a present block is processed, the addition overlay buffer contains information from the previous audio block. Because it is necessary to have smooth transitions between the audio blocks, logic is included to maneuver differently the transitions of the TD audio channel combination to the FD audio channel reduction combination, and the combination of reduction of FD audio channels to the combination of TD audio channel reduction.
FIG. 9 shows an example of five block processing, denoted as block k, k + 1 ..... k + 4 of five-channel audio that includes, as is common: left, center, right, left surround and right surround channels , denoted L, C, R, LS and RS, respectively, and the combination of audio channel reduction to a stereo combination using the formula:
Left output denoted L-aC + bL + cLS, and
Right output denoted R '= aC + bR + cRS.
FIG. 9 assumes that a transformation without overlap is used. Each rectangle represents the audio contents of a block. The horizontal axes from left to right represent the blocks k, ..., k + 4, and the vertical axes from top to bottom represent the progress of information decoding.
Suppose block k is processed by the TD audio channel reduction combination, blocks k + 1 and k + 2 are processed by the FD audio channel reduction combination, and blocks k + 3 and k +4, by the combination of TD audio channel reduction. As can be seen, for each of the TD audio channel reduction combination blocks, the audio channel reduction combination does not occur until after the time domain audio channel reduction combination merges to the bottom. , after which, the contents are the channels subjected to a combination of reduction of audio channels L 'and R', while for the block subjected to a combination of reduction of audio channels of FD, Left and right channels in the frequency domain are already subject to the audio channel reduction combination after the frequency domain audio channel reduction combination, and the C, LS and RS channel information is ignored. Because there is no overlap between the blocks, no special case maneuver is required when switching from the TD audio channel reduction combination to the FD audio channel reduction combination, or the reduction reduction combination. FD audio channels to the TD audio channel reduction combination.
FIG. 10 describes the case of 50% overlap transformations.
Assume that the addition overlay is carried out by the addition overlay decoding using an addition overlay buffer. In this diagram, when the information block is shown as two triangles, the lower left triangle is information in the addition overlay buffer of the previous block, while the upper right triangle shows the information of the current block.
Transition maneuver for a combination transition from TD audio channel reduction to FD audio channel reduction combination.
Consider the block k + 1 which is an FD audio channel reduction combination block that immediately follows a TD audio channel reduction combination block. After the TD audio channel reduction combination, the addition overlay buffer contains the information of L, C, R, LS and RS of the last block to be included for this block. In addition, the contribution of the current block of k + 1 is included, already subject to the FD audio channel reduction combination. In order to properly determine the information of PCM subjected to a combination of audio channel reduction for output, both the information of the present block and the information of the previous block must be included. For this purpose, the information in the previous block must be expelled, and because it has not yet undergone the audio channel reduction combination, it must be subjected to the audio channel reduction combination in the time domain. The two contributions must be added in order to determine the PCM information submitted to the audlo channel reduction combination for output. This processing is included in the TD 715 audio channel reduction combination transition logic of FIGS. 7 and 8, and by the code in the TD audio channel reduction combination transition logic included in the FD audio channel reduction combination module set forth in FIG. 5B. The processing performed therein is summarized in the TD 715 audio channel reduction combination transition logic section of FIG. 8. In more detail, the transition maneuver for a transition from TD audio channel reduction combination to FD audio channel reduction combination includes:
• The expulsion of overlapping buffers by feeding zeroes in the addition overlay logic, and performing window and adding overlay operations. The copy of the output ejected from the addition overlay logic. This is the PCM information of the previous block of the particular channel before the audio channel reduction combination. The overlay buffer now contains zeros.
• The time domain audio channel reduction combination of the PCM information of the overlay buffers, to generate the PCM information of the TD audio channel reduction combination of the previous block.
• The combination of frequency domain audio channel reduction of the new information of the current block. The realization of the inverse transformation and the feeding of new information after the combination of reduction of FD audio channels and the inverse transformation in addition overlay logic. Performing window operations and addition overlay, and so on with the new information in order to generate PCM information of the FD audio channel reduction combination of the current block.
• The addition of the PCM information of the TD audio channel reduction combination and the FD audio channel reduction combination in order to generate the PCM output.
Note that in an alternative embodiment, assuming there was no TPNP in the previous block, the information in the addition overlay buffers is subjected to the audio channel reduction combination, and then an addition overlay operation is performed on the Output channels already subject to the combination of audio channel reduction. This avoids the need to perform an addition overlay operation for each previous block channel. In addition, as described above for the decoding of AC-3, when an audio channel reduction combination buffer is used and its corresponding half-block delay buffer of 128 samples long is used and subjected to operation of window and combines to produce 256 PCM output samples, the audio channel reduction combination operation is simpler, since the delay buffer only has 128 samples, instead of 256. This aspect reduces the peak computational complexity that is inherent in transition processing. Therefore, in some embodiments, for a particular block that is subjected to the FD audio channel reduction combination following a block whose information was subjected to the TD audio channel reduction combination, the processing of Transition includes the application of audlo channel reduction combination in the time pseudo-domain to the information of the previous block that must be superimposed with the decoded information of the particular block.
Transition maneuver for a transition from a combination of reduction of FD audio channels to a combination of reduction of TD audio channels.
Consider the block k + 3 which is a TD audio channel reduction combination block that immediately follows an FD k + 2 audlo channel combination block. Because the previous block was an FD audlo channel reduction combination block, the addition overlay buffer in the previous stages, for example, before the TD audio channel reduction combination, contains the information undergoing a combination of audio channel reduction on the left and right channels, and does not contain information on the other channels. The contributions of the current block are not subject to audio channel reduction combination until after the TD audio channel reduction combination. In order to properly determine the information of PCM subjected to a combination of audio channel reduction for output, both the information of the present block and the information of the previous block must be included. For this, it is necessary to eject the information from the previous block. The information of the present block must be subjected to the combination of reduction of audio channels in the time domain, and must be added to the inverse transformation information that was expelled, in order to determine the PCM information submitted to the combination of reduction of audio channels for output. This processing is included in the FD 723 audio channel reduction combination transition logic of FIGS. 7 and 8, and by code in the FD audio channel reduction combination transition logic module set forth in FIG. 5B. The processing performed therein is summarized in the FD 723 audio channel reduction combination transition logic section of FIG. 8. In more detail, assuming that there are output PCM buffers for each output channel, the transition maneuver for a combination transition from reduction of FD audio channels to combination reduction of TD audio channels includes:
• The expulsion of the overlapping buffers by feeding zeroes in the addition overlay logic, and performing window operation and addition overlay. The copy of the output in the output PCM buffer. The ejected information is the PCM information of the FD audio channel reduction combination of the previous block. The overlay buffer now contains zeros.
• Performing the inverse transformation of the new information of the current block, in order to generate information prior to the audio channel reduction combination of the current block. Feeding this new time domain information (after transformation) into the addition overlay logic.
• Performing window operation and addition overlay, TPNP if any, and combination of TD audio channel reduction with the new information of the current block, so as to generate PCM information of the channel reduction combination of TD audio of the current block.
• The addition of the PCM information of the TD audio channel reduction combination and the FD audio channel reduction combination in order to generate the PCM output.
In addition to the transitions of the time domain audio channel reduction combination to the frequency domain audio channel reduction combination, program changes are maneuvered in the audio channel reduction combination transition logic of time domain and program change maneuver. Newly emerging channels are automatically included in the combination of audio channel reduction, and consequently, do not need any special treatment. Channels that are no longer present in the new program must be subjected to outgoing shutdown. This is carried out, as shown in section 715 in FIG. 8, in the case of the FD audio channel reduction combination, by ejecting the overlay buffers from the shutdown channels. The expulsion is carried out by feeding zeroes in the logic of addition overlay and the performance of window operation and addition overlay.
Note that the flow chart set forth, and in some embodiments, the frequency domain audio channel reduction combination logic section 711, includes disabling the optional gain variation feature for all channels that are part of the combination of frequency domain audio channel reduction. Channels can have different parameters of gain variation, which would induce different graduation of the spectral coefficients of a channel, in order to avoid a combination of audio channel reduction.
In an alternative implementation, the FD 711 audio channel reduction combination logic section is modified, so that the minimum of all gains is used to perform the gain variation for a channel subject to the reduction combination of audio channels (frequency domain).
Combination of reduction of time domain audio channels with change of combination coefficients of reduction of audio channels and need for explicit cross-off.
The combination of audio channel reduction can create several problems. Different audio channel reduction combination equations are demanded in different circumstances, and therefore, it may be necessary to dynamically change the audio channel reduction combination coefficients based on the signal conditions. Metadata parameters can be obtained that allow to adapt the combination coefficients of reduction of audio channels to obtain optimal results.
Therefore, the combination coefficients of audio channel reduction may change as a function of time. When there is a change from a first parameter of combination coefficients of reduction of audio channels to a second parameter of combination coefficients, the information must be subjected to cross-off from the first parameter to the second parameter.
When the audio channel reduction combination is performed in the frequency domain, and in addition, in many decoder implementations, for example, in an AC-3 decoder of the prior art, as shown in FIG. 1, the audio channel reduction combination is performed before window operations and addition overlay. The advantage of carrying out the combination of audio channel reduction in the frequency domain, or in the time domain before the window and addition overlay is that there is inherent cross-off as a result of the addition overlay operations. Therefore, in many known AC-3 decoders and decoding methods in which the combination of audio channel reduction takes place in the window domain after inverse transformation, or in the frequency domain in the Combination implementations of hybrid audio channel reduction, there is no explicit cross-off operation.
In the case of the combination of time domain audio channel reduction and transient prior noise processing (TPNP), there would be a delay of a block in decoding of transient prior noise processing caused by program change issues, for example, in a 7.1 decoder. Accordingly, in embodiments of the present invention where the combination of audio channel reduction in the time domain is performed and TPNP is used, the time domain audio channel reduction combination is carried out after window operations and addition overlay. The processing order in the case of the time domain audio channel reduction combination used is: performing the inverse transformation, for example, MDCT, performing window operations and adding overlay, performing any decoding of prior transient noise processing (without delay), and then, the combination of time domain audio channel reduction.
In such a case, the time domain audio channel reduction combination requires the cross-off of the previous and current audio channel reduction combination information, for example, the audio channel reduction combination coefficients or the Combination tables of audio channel reduction, in order to ensure that no change in the combination coefficients is not smoothed.
One option is to carry out the cross-off operation in order to compute the resulting coefficient. Denoted by c [z] the combination coefficient to be used, where i denotes the time index of 256 time domain samples, so that the range is / = 0 ..... 255. Denoted by w<sup>2</sup>[z '] · a positive window function so that w<sup>2</sup>[z '] + w<sup>2</sup>[255-z '] = l for i = 0, ..., 255. Denoted by c<sub>v</sub>¡<sub>and</sub>j<sub>or</sub> the combination coefficient prior to the update, and by
New updated combination coefficient. The cross-off operation to apply is:
c [z] = w<sup>2</sup> [z] · cnuem + w<sup>2</sup> [255 - z] cviej0 for i = 0 ..... 255.
After each pass through the coefficient cross-off operation, the old coefficients are updated with the new one, as c .. <- c old new
In the next pass, if the coefficients are not updated,
ΦΊ = w<sup>2</sup>M · cnuew + w<sup>2</sup>[255 - z] · new = c<sub>mtem</sub> .
In other words, the influence of the old coefficient parameter has completely disappeared.
The inventors noted that in many audio streams and audio channel reduction combination situations, the combination coefficients often do not change. In order to improve the performance of the time domain audio channel reduction combination process, the embodiments of the time domain audio channel combination module include the test to assess whether the combination coefficients of time reduction combination Audio channels have changed from their previous value, and if they have not done so, performing the audio channel reduction combination, or else, if they have changed, performing cross-off of the combination coefficients according to a preselected positive window function. In one embodiment, the window function is the same window function as that used in window operations and addition overlay. In another embodiment, a different window function is used.
FIG. 11 shows the simplified pseudocode for a combination embodiment of audio channel reduction. The decoder for said embodiment uses at least one x86 processor that executes SSE vector instructions. The audio channel reduction combination includes the evaluation of whether the new audio channel reduction combination information is changed with respect to the old audio channel reduction combination information. If it is, the audio channel reduction combination includes the execution of SSE vector instructions on at least one of one or more x86 processors, and the audio channel reduction combination using the unchanged reduction combination information. of audio channels that includes the execution of at least one SSE vector instruction. Otherwise, if the new audio channel reduction combination information is changed with respect to the old audio channel reduction combination information, the method includes determining the information subject to audio channel reduction combination. with cross-off through the cross-off operation.
Exclusion of unnecessary processing information.
In some audio channel reduction combination situations, there is at least one channel that does not contribute to the output of the audio channel reduction combination. For example, in many cases of audio channel reduction combination from 5.1 to stereo audio, the LFE channel is not included, so that the audio channel reduction combination is 5.1 to 2.0. The exclusion of the LFE channel from the audio channel reduction combination may be inherent in the encoding format, as is the case for AC-3, or it may be controlled by metadata, as is the case for E-AC-3. In E-AC-3, the Ifemixlevcode parameter determines whether the LFE channel is included in the audlo channel reduction combination. When the Ifemixlevcode parameter is 0, the LFE channel is not included in the audlo channel reduction combination.
Remember that the audlo channel reduction combination can be carried out in the frequency domain ^ in the pseudo-time domain after the Inverse transformation but before the window operation and addition overlay, or in the time domain, then of the inverse transformation and after the window operation and addition overlay. The combination of pure time domain audio channel reduction is carried out in many known E-AC-3 decoders, and in some embodiments of the present invention, and is convenient; for example, due to the presence of TPNP, the combination of pseudo-domain audlo channel reduction is carried out in many AC-3 decoders and in some embodiments of the present invention, and is convenient because the operation Addition overlay provides inherent cross-off that is convenient when audio channel reduction combination coefficients change, and the combination of frequency domain audio channel reduction is carried out in some embodiments of the present invention when conditions permit.
As described in the present application, the frequency domain audio channel reduction combination is the most efficient audio channel reduction combination method, since it minimizes the amount of inverse transformation and window operations and overlay of Addition required to produce a 2-channel output of a 5.1-channel input. In some embodiments of the present invention, when the FD audio channel reduction combination is carried out, for example, in FIG. 8, in the FD 711 audio channel reduction combination circuit section in the circuit that begins with element 813, ends with 814 and increases by 815 to the next channel, those channels not included in the channel reduction combination Audio are excluded in processing.
The combination of audio channel reduction in the pseudo-time domain after the inverse transformation but before the window operation and addition overlay, or in the time domain, after the inverse transformation and the window and overlay operation In addition, it is less computationally efficient than in frequency domain. In many current decoders, such as the current AC-3 decoders, the combination of audio channel reduction takes place in the pseudo-time domain. The reverse transformation operation is carried out independently of the audio channel reduction combination operation, for example, in separate modules. The reverse transformation in said decoders is carried out on all input channels. This is relatively inefficient from the computational point of view, because, in the case that the LFE channel is not included, the inverse transformation is still carried out for this channel. This unnecessary processing is significant, because, even when the LFE channel is limited in bandwidth, the application of the Inverse transformation to the LFE channel requires both computation and the application of the inverse transformation to any full bandwidth channel . The inventors recognized this inefficiency. Some embodiments of the present invention include the identification of one or more non-contributing channels of the input channels Nn, where, a non-contributing channel is a channel that does not contribute to the output channels Mm of decoded audio. In some embodiments, the identification uses information, for example, metadata that defines the combination of audio channel reduction. In the example of a combination of reduction of audio channels from 5.1 to 2.0, the LFE channel is thus identified as a non-contributing channel. Some embodiments of the invention include performing a frequency transformation in time on each channel that contributes to the output channels Mm, and not performing any frequency transformation in time on each identified channel that does not contribute to the output signal. Mm In the example from 5.1 to 2.0 in which the LFE channel does not contribute to the combination of audio channel reduction, the reverse transformation, for example, an IMCDT, is only carried out on the five full bandwidth channels, so that the reverse transformation portion is carried out with approximately 16% reduction of the computational resources required for the 5.1 channels. Because the IMDCT is a significant source of computational complexity in the decoding method, this reduction can be significant.
In many current decoders, such as current EAC-3 decoders, the combination of audio channel reduction takes place in the time domain. The reverse transformation operation and the addition overlay operations are carried out before any TPNP and before the audio channel reduction combination, independent of the audio channel reduction combination operation, for example, in modules separated. Reverse transformation and addition and overlay window operations on these decoders are carried out on all input channels. This is relatively inefficient from the computational point of view, since, in the case that the LFE channel is not included, the inverse transformation and the o.peraclón of window / addition overlay are still carried out for this channel. This unnecessary processing is significant, since, even if the LFE channel is limited in bandwidth, the application of the inverse transformation and the channel addition overlay
LFE requires both computation and the application of reverse transformation and the addition of overlap to any full bandwidth channel. In some embodiments of the present invention, the audio channel reduction combination is carried out in the time domain, and in other embodiments, the audio channel reduction combination may be carried out in the time domain of agreement. with the result of the application of the logic selection method of audio channel reduction combination. Some embodiments of the present invention in which the TD audio channel reduction combination is used include the Identification of one or more non-contributing channels of the Nn input channels. In some embodiments, the identification uses information, for example, metadata. that define the combination of audio channel reduction. In the example of a combination of audio channel reduction from 5.1 to 2.0, the LFE channel is identified as a non-contributing channel. Some embodiments of the invention include the realization of an inverse transformation, that is, the frequency transformation in time, in each channel contributing to the output channels Mm, and the non-realization of any frequency transformation in time and other processing of time domain on each identified channel that does not contribute to the Mm channel signal In the example from 5.1 to 2.0 in which the LFE channel does not contribute to the combination of audio channel reduction, the reverse transformation, for example, an IMCDT, the addition overlay and the TPNP are only carried out in the five Full bandwidth channels, so that the inverse transformation and the addition / overlay window operation portions are carried out with approximately 16% reduction of the computational resources required for the 5.1 channels. In the flowchart of FIG. 8, in the common processing logic section 731, a feature of some embodiments includes that the processing in the circuit that begins with element 833, continues to 834, and that includes the increment to the next channel element 835 is carried out for all channels except non-contributing channels. This happens inherently for a block that is subjected to the FD audio channel reduction combination.
Although in some embodiments the LFE is a non-contributing channel, that is, it is not included in the output channels subjected to the combination of audio channel reduction, as is common in AC-3 and E-AC-3, in other embodiments, a different LFE channel is also, or instead, a non-contributing channel that is not included in the output subjected to the audio channel reduction combination. Some embodiments of the invention include the control of such conditions for the identification of one or more channels, if any, which are non-contributing in terms of said channel not being included in the audio channel reduction combination, and, in the case of the time domain audio channel reduction combination, non-processing through reverse transformation and window operations and addition overlay for any identified non-contributing channel.
Li
For example, in AC-3 and E-AC-3, there are certain conditions in which the surround channels and / or the central channel are not included in the output channels subjected to the combination of audio channel reduction. These conditions are defined by metadata included in the encoded bitstream that takes predefined values. Metadata, for example, may include information that defines the combination of audio channel reduction, which includes mixing level parameters.
Next, some such examples of said mixing level parameters are described, for illustrative purposes in the case of E-AC3. In the combination of audio channel reduction to stereo in E-AC-3, two types of audio channel reduction combination are provided: the combination of audio channel reduction to a pair of LtRt matrix surround encoded stereo, and the combination of reducing audio channels to a conventional stereo signal, LoRo. The stereo signal subjected to the audio channel reduction combination (LoRo, or LtRt) can be additionally mixed to mono. A 3-bit surround combination level code of LtRt denoted Itrtsurmixlev, and a 3-bit combination combination level code of LoRo denoted lorosurmixlev indicate the nominal combination level of audio channel reduction combination of the surround channels with respect to the Left and right channels in a combination of audio channel reduction from LtRt, or LoRo, respectively. A binary value '111' indicates a combination level of audio channel reduction of 0, that is, - °° dB. The 3-bit center combination level codes of LtRt and LoRo denoted Itrtcmixlev, lorocmixlev indicate the nominal combination level of audio channel reduction of the central channel with respect to the left and right channels in a combination of channel reduction audio of LtRt and LoRo, respectively. A binary value '111' indicates a combination level of audio channel reduction of 0, that is, - °° dB.
There are conditions in which the surround channels are not included in the output channels subjected to the audlo channel reduction combination. In E-AC-3, these conditions are identified by metadata. These conditions include cases where surmixlev = Ί0 '(AC-3 only), Itrtsurmixlev = Ί11', and lorosurmixlev = Ί11 '. For these conditions, in some embodiments, a decoder includes the use of combination level metadata for the identification that said metadata indicates that the enveloping channels are not included in the audlo channel reduction combination, and non-processing of the envelope channels through the inverse transformation and the addition / overlay window stages. In addition, there are conditions in which the center channel is not included in the output channels subjected to the audio channel reduction combination, identified by ltrtcm¡xlev = -11Γ, Iorocmixlev == 'l11', For these conditions, In some embodiments, a decoder includes the use of combination level metadata to identify that said metadata indicates that the center channel is not included in the audio channel reduction combination, and non-processing of the center channel through the inverse transformation and the addition / overlay window / stages.
In some embodiments, the identification of one or more non-contributing channels depends on the content. By way of an example, the Identification includes the Identification of whether one or more channels have an insignificant amount of content in relation to one or more additional channels. A measure of the amount of content is used. In one embodiment, the measure of the amount of content is energy, while in another embodiment, the measure of the amount of content is the absolute level. The Identification includes the comparison of the difference in the measure of the amount of content between pairs of channels at a set threshold. By way of example, in one embodiment, the identification of one or more non-contributing channels includes the assessment of whether the amount of envelope content of a block is less than the amount of content of each start channel at least by a threshold of fixation, in order to assess whether the surround channel is a non-contributing channel.
Ideally, the threshold is selected to be as low as possible, without introducing noticeable artifacts in the version of the signal subjected to the audio channel reduction combination, in order to maximize the identification of channels as non-contributors, in order to reduce the amount of computing required, while minimizing the loss of quality. In some embodiments, different thresholds are provided for different decoding applications, where the choice of the threshold for a particular decoding application represents an acceptable balance between the quality of the audio channel reduction combination (upper thresholds) and the reduction of the computational complexity (lower thresholds) for the specific application.
In some embodiments of the present invention, one channel is considered insignificant with respect to another channel if its energy or absolute level is at least 15 dB lower than that of the other channel. Ideally, one channel is insignificant with respect to another channel if its energy or absolute level is at least 25 dB lower than that of another channel.
The use of a threshold for the difference between two channels denoted A and B that is equivalent to 25 dB is approximately equivalent to saying that the level of the sum of the absolute values of the two channels is within 0.5 dB of the level of the dominant channel. That is, if channel A is at -6 dBFS (dB in relation to the full scale) and channel B is at -31 dBFS, the sum of the absolute values of channel A and channel B will be approximately -5.5 dBFS , or about 0.5 dB higher than the level of channel A.
If the audio is of relatively low quality, and for low-cost applications, it may be acceptable to sacrifice quality to reduce complexity, the threshold could be less than 25 dB. In one example, a threshold of 18 dB is used. In that case, the sum of the two channels can be found within about 1 dB of the channel level with the highest level. This may be audible in certain cases, although it should not be too objectionable. In another embodiment, a threshold of 15 dB is used, in which case, the sum of the two channels is within 1.5 dB of the dominant channel level.
In some embodiments, several thresholds are used, for example, 15dB, 18dBy25dB.
Note that although the Identification of the non-contributing channels for AC-3 and E-AC-3 is described in this application beforehand, the Non-contributing Channel Identification feature of the invention is not limited to such formats. Other formats, for example, also provide information, such as metadata regarding the combination of audio channel reduction that is useful for identifying one or more non-contributing channels. Both MPEG-2 AAC (ISO / IEC 13818-7) and MPEG-4 Audio (ISO / IEC 14496-3) are capable of transmitting what is referred to by the standard as a “combination reduction coefficient of audio channels of matrix. Some embodiments of the invention for decoding said formats use this coefficient to construct a stereo or mono signal from 3/2, that is, a Left, Center, Right, Surround Left, Surround Right signal. The combination coefficient of matrix audio channel reduction determines the way in which the surround channels are mixed with the front channels to construct the stereo or mono output. Four values of the matrix audio channel reduction combination coefficient are possible according to each of these standards, one of which is 0. A value of 0 ensures that the surround channels are not included in the combination reduction of audio channels Some embodiments of the MPEG-2 AAC decoder or MPEG-4 Audio decoder of the invention include generating a combination of stereo or mono audio channel reduction from a 3/2 signal, using the combination coefficients of reduction of audio channels indicated in the bit stream, and also include the identification of a non-contributing channel by a. Combination coefficient of matrix audio channel reduction of 0, in which case, inverse transformation and window processing / addition overlay is not carried out.
FIG. 12 shows a simplified block diagram of an embodiment of a processing system 1200 that includes at least one processor 1203. In this example, an x86 processor is shown whose instruction group includes SSE vector instructions. In addition, a busbar subsystem 1205 is shown in simplified blocks by which the various components of the processing system are coupled. The processing system includes a storage subsystem 1211 coupled to the processors, for example, by means of the busbar subsystem 1205, where the storage subsystem 1211 has one or more storage devices, which include at least one memory and in some embodiments, one or more additional storage devices, such as magnetic or optical storage components. Some embodiments also include at least one network interface 1207, and an audio input / output subsystem 1209 that can accept PCM information and that includes one or more DACs to convert PCM information into electric waveforms to conduct a group of speakers or headphones. Other elements may also be included in the processing system, which will be apparent to those skilled in the art, and which are not set forth in FIG. 12 for reasons of simplicity.
The storage subsystem 1211 includes instructions 1213 which, when executed in the processing system, causes the processing system to carry out the decoding of audio information including Nn channels of encoded audio information, for example, EAC information. 3 to form decoded audio information that includes Mm channels of decoded audio, M ^ l, and in the case of the combination of audio channel reduction, M <Ñ. For today's known coding formats, n = 0 or 1 and m = 0 or 1, although the invention is not limited to these. In some embodiments, instructions 1211 are divided into modules. Other instructions (other computer programs) 1215 are also typically included in the storage subsystem. The exposed embodiment includes the following modules in instructions 1211: two decoder modules: a 5.1-channel independent frame decoder module 1223, which includes a start end decoding module 1231 and an end end decoding module 1233 ; a 1225 dependent frame decoder module that includes a Start end decoding module 1235 and an end end decoding module 1237; an instruction frame information analysis module 1221, which when executed, produces the unpacking of the Bit Current Information (BSl) field information of each frame in order to identify the frames and types of frames and provide frames identified to appropriate occurrences of starter decoder module 1231 or 1235; and an instruction channel mapping module 1227, which when executed, and in the case N> 5, produce the combination of the decoded information of the respective end-end decoding modules in order to form the Nn channels of decoded information .
The embodiments of alternate processing systems may include one or more processors coupled by at least one network connection, that is, they are distributed. That is, one or more of the modules may be in other processing systems coupled to a main processing system by a network connection. Such alternate embodiments will be apparent to those skilled in the art. Consequently, in some embodiments, the system comprises one or more subsystems that are networked through a network connection, where each subsystem includes at least one processor.
Therefore, the processing system of FIG. 12 forms an embodiment of an apparatus for processing audio information that includes Nn channels of encoded audio information in order to form decoded audio information that includes Mm channels of decoded audio, M> 1, in the case of the combination of audio channel reduction, M <N, and for the combination of audio channel increase, M> N. While for current standards, n = 0 or 1 and m = 0 or 1, other embodiments are possible. The apparatus includes several functional elements functionally expressed as means for carrying out a function. A functional element means an element that performs a processing function. Each of said elements may be a physical support element, for example, special purpose physical support, or a processing system that includes a storage medium comprising instructions that, when executed, perform the function. The apparatus of FIG. 12 includes means for accepting audio information that includes N channels of encoded audio information, encoded by an encoding method, for example, an E-AC-3 encoding method, and more generally, an encoding method that comprises the transformation using the N-channel transformation of digital audio information channels, to form and package frequency domain and mantissa exponent information, and form and pack metadata related to the frequency domain and mantissa exponent information, where the metadata optionally includes metadata related to the processing of transient prior noise.
The apparatus includes means for decoding the accepted audio information.
In some embodiments, the means for decoding include means for unpacking the metadata and means for unpacking and for decoding the frequency domain and mantissa exponent information; means for determining the transformation coefficients of the unpacked and decoded information of frequency domain and mantissa exponent; means for the inverse transformation of frequency domain information; means for applying window operations and addition overlay to determine samples of audio information; means for the application of any transient pre-noise processing decoding required in accordance with the metadata related to the processing of transient pre-noise; and means for the combination of TD audio channel reduction according to the combination information of audio channel reduction. The means for the TD audio channel reduction combination, in the case M <N, performs the audio channel reduction combination according to the audio channel reduction combination information, which includes, in some embodiments , the evaluation of whether the audio channel reduction combination information is changed with respect to the previously used audio channel reduction combination information, and if it is changed, The cross-off application in order to determine the combination information of audio channel reduction with cross-off, performing the combination of audio channel reduction according to the combination information of reduction of audio channels of cross-off , and if it is not changed, the direct realization of the audio channel reduction combination according to the audio channel reduction combination information.
Some embodiments include means for the evaluation of a block, in terms of the use of the TD audio channel reduction combination or FD audio channel combination combination , and means for the FD audio channel reduction combination , activated if the means for the evaluation of a block in terms of the use of the TD audio channel reduction combination or the FD audio channel combination combination evaluate the FD audio channel reduction combination, including means for the transition transition processing combination of audio channels from TD to FD. Such embodiments further include means for the transition processing of FD to TD audio channel reduction combination. The operation of these elements is as described in the present application.
In some embodiments, the apparatus includes means for the identification of one or more non-contributing channels of the input channels Nn, where a non-contributing channel is a channel that does not contribute to the channels Mm The apparatus does not perform the inverse transformation of frequency domain information and the application of additional processing, such as TPNP or addition overlay on one or more identified non-contributing channels.
In some embodiments, the apparatus includes at least one x86 processor whose instruction group includes vector instructions comprising multiple information extensions of a single instruction in continuous comment (SSE). The means for combining the reduction of audio channels in the operation executes vector instructions in at least one of one or more x86 processors.
Alternative devices to those set forth in FIG. 12. For example, one or more of the elements may be implemented by physical support devices, while others may be implemented through the operation of an x86 processor. Such variations will be apparent to those skilled in the art.
In some embodiments of the apparatus, the means for decoding include one or more means for starting end decoding and one or more means for end-end decoding. The means for decoding the starting end includes the means for unpacking the metadata and the means for unpacking and decoding the frequency domain and mantissa exponent information. The means for end-end decoding includes the means for evaluating a block in terms of the use of the TD audio channel reduction combination or the FD audio channel reduction combination; the means for the FD audio channel reduction combination, which includes the means for the audio channel reduction combination transition processing from TD to FD, the means for the FD to TD transition processing processing, and the means for determining the transformation coefficients from the unpacked and decoded information of frequency domain mantissa and mantissa; for the inverse transformation of frequency domain information; for the application of window operations and addition overlay to determine samples of audio information; for the application of any transient pre-noise processing decoding required in accordance with the metadata related to the transient pre-noise processing; and for the time domain audio channel reduction combination in accordance with the audio channel reduction combination information. The time domain audio channel reduction combination, in the case M <N, performs the audio channel reduction combination according to the audio channel reduction combination information, which includes, in some embodiments. , the evaluation of whether the audio channel reduction combination information is changed with respect to previously used audio channel reduction combination information, and if it is changed, the cross-off application in order to determine Combination information of cross-channel audio reduction combination, and the realization of the combination of audio channel reduction according to the combination information of cross-channel audio reduction combination ; and if it is not changed, performing the audio channel reduction combination according to the Audio Channel Reduction Combination Information.
For the processing of E-AC-3 information from more than 5.1 channels of encoded information, the means for decoding include multiple occurrences of the medium for the Start-end decoding and the means for the end-end decoding, which include a first means for decoding start end and a first means for decoding end end, for decoding the independent frame of up to 5.1 channels; a second means for the start-end decoding and a second means for the end-end decoding, for the decoding of one or more frames dependent on Information. The apparatus also includes means for unpacking Bit Current Information field information in order to identify the frames and types of frames and provide the identified frames to appropriate start-end decoding means, and means for the combination of the decoded information of respective means for end-end decoding in order to form the N channels of decoded information.
Note that while E-AC-3 and other coding methods use an addition overlay transformation, and in the reverse transformation, include window and addition overlay operations, it is known that other forms of transformations operating from such that reverse transformation and additional processing can recover time domain samples without aliasing errors. Therefore, the invention is not limited to addition overlap transformations, and each time the inverse transformation of frequency domain information and the operation of window and addition overlay are mentioned in order to determine time domain samples , art experts will understand that, in general, These operations can be established as “reverse transformation of frequency domain information and the application of additional processing for the determination of audio information samples.
While the terms exponent and mantissa are used throughout the description because they are terms used in AC-3 and E-AC-3, other coding formats may use other terms, for example, graduation factors and coefficients spectral in the case of HE-AAC, and the use of the terms exponent and mantissa does not limit the scope of the invention to formats that use the terms exponent and mantissa.
Unless specifically stated otherwise, as will be apparent from the following description, it is appreciated that throughout the specification, descriptions using terms such as "processing," "computation / computation," " calculation ”,“ determination ”,“ generation ”or the like refers to the action or processes of a physical support element, for example, a computer or a computer system, a processing system or a similar electronic computing device, which manipulates or transforms information represented as physical, such as electronics, quantities, into other Information equally represented as physical quantities.
Similarly, the term "processor" may refer to any device or portion of a device that processes electronic information, for example, from records or memory, in order to transform said electronic information into other electronic information that, for example, may stored in cash registers or memory. A "processing system", a "computer", a "computing machine" or a "computing platform" may include one or more processors.
Note that when describing a method that includes several elements, for example, several steps, no order of such elements is involved, for example, steps, unless specifically stated.
In some embodiments, a computer-readable storage medium is configured, for example, encoded, with, for example, storage instructions that when executed by one or more processors of a processing system such as a digital signal processing device or A subsystem that includes at least one processor element and a storage subsystem, performs a method as described in this application. Note that, in the above description, when it is established that the instructions are configured, when they are executed to carry out a process, it should be understood that this means that the instructions, when executed, cause one or more processors to operate so that a physical support apparatus, for example, the processing system, carries out the process.
The methodologies described in this application, in some embodiments, can be performed by one or more processors that accept logic, instructions encoded in one or more computer-readable media. When executed by one or more of the processors, the instructions carry out at least one of the methods described in this application. It includes any processor capable of executing a set of instructions (successive or otherwise) that specify actions to be taken. Consequently, an example is a typical processing system that includes one or more processors. Each processor may include one or more of a CPU ("central processing unit") or a similar element, a graphics processing unit (GPU), and / or a DSP unit ( English acronym for: "digital signal processor") programmable. The processing system also includes a storage subsystem with at least one storage medium, which may include embedded memory in a semiconductor device, or a separate memory subsystem that includes RAM (English acronym for "random access memory" ) main and / or a static RAM, and / or ROM (English acronym for "read-only memory"), and also cache memory. The storage subsystem may further include one or more additional storage devices, such as magnetic, optical or other solid state storage devices. A busbar subsystem can be included for communication between the components. The processing system can also be a distributed processing system with processors coupled by a network, for example, by means of network interface devices or wireless interface devices. If the processing system requires a screen, said screen may include, for example, a liquid crystal display (LCD), an organic light emitting screen (OLED), or a cathode ray tube (CRT) display. If manual data entry is required, the processing system also includes an input device; such as one or more of an alphanumeric input unit such as a keyboard, a pointer control device, such as a mouse, etc. The term storage device, storage subsystem or memory unit as used in this application, if it is clear from the context and unless explicitly stated otherwise, also encompasses a storage system such as a disk drive. The processing system, in some configurations, may include a sound output device and a network interface device.
The storage subsystem, accordingly, includes a computer-readable medium that is configured, for example, encoded with instructions, for example, logic such as computer programs that, when executed by one or more processors, carry out one or more of the method steps described in this application. The computer program can be found on the hard disk, or it can also be found completely or at least partially, within the memory such as RAM, or within the internal memory of the processor during its execution by the computer system. Consequently, memory and the processor that includes memory also constitute a computer-readable medium in which instructions are encoded.
Also, a computer-readable medium can form a computer program product, or it can be included in a computer program product.
In alternative embodiments, one or more processors operate as a single device or may be connected, for example, in a network with other processors, in a network deployment, where the processors can operate as a server or a client machine in a server-client network environment, or as a machine attached in a distributed network environment or attached to attached. The term processing system encompasses all such possibilities, unless explicitly excluded in this application. One or more processors may form a personal computer (PC), a media playback device, a tablet PC, an analog or digital television (STB) signal decoding box, a personal digital assistant (PDA, according to its acronym) in English), a game machine, a cell phone, an Internet device, a network router, switch or bridge, or any machine capable of executing a set of instructions (successive or otherwise) that specify actions to be taken by the machine.
Note that while some diagrams only show a single processor and a single storage subsystem, for example, a single memory that stores the logic that includes instructions, those skilled in the art will understand that many of the components described above are included. , although they are not explicitly shown or described so as not to obscure the aspect of the invention. For example, while only a single machine is illustrated, the term “machine will also be considered to include any group of machines that individually or jointly execute a group (or multiple groups) of instructions in order to perform one or more of the described methodologies in this application.
Therefore, an embodiment of each of the methods described in this application is presented in the form of a computer-readable medium configured with a group of instructions, for example, a computer program, which, when executed in one or more processors, for example, one or more processors that are part of a media device, carry out the method steps. Some embodiments are presented in the form of logic itself. Consequently, as will be appreciated by those skilled in the art, the embodiments of the present invention can be realized as a method, an apparatus such as a special purpose apparatus, an apparatus such as an information processing system, logic, by for example, represented on a computer readable storage medium, or a computer readable storage medium that is encoded with instructions, for example, a computer readable storage medium configured as a computer program product. The computer-readable medium is configured with a group of instructions that, when executed by one or more processors, carry out the method steps. Accordingly, aspects of the present invention may take the form of a method, an embodiment of physical support entirely, which includes several functional elements, where a functional element means an element that performs a processing function. Each of said elements may be a physical support element, for example, special purpose physical support, or a processing system that includes a storage medium comprising instructions that, when executed, perform the function. Aspects of the present invention may take the form of an entire computer program embodiment, or an embodiment that combines aspects of the computer program, or software, and physical support. Also, the present invention may take the form of program logic, for example, in a computer-readable medium, for example, a computer program in a computer-readable storage medium, or the computer-readable medium configured with computer code. computer readable program, for example, a computer program product. Note that, in the case of special purpose physical support, the definition of the function of the physical support is sufficient to allow the person skilled in the art to write a functional description that can be processed by programs that then automatically determine the description of the physical support for the generation of the function by physical support. Consequently, the description in this application is sufficient to define said special purpose physical support.
While the computer-readable medium is shown in an exemplary embodiment as a single medium, the term "medium" must be construed to include a single medium or multiple media (eg, multiple memories, a centralized or distributed database, or associated caches and servers) that store one or more instruction groups. A computer-readable medium can take many forms, including, without limitation, non-volatile media and volatile media. Non-volatile media include, for example, optical, magnetic, and magneto-optical discs. Volatile media includes dynamic memory, such as main memory.
It will be further understood that the embodiments of the present invention are not limited to any particular implementation or programming technique, and that the invention can be implemented using any appropriate technique to implement the functionality described in this application. In addition, the embodiments are not limited to any particular programming language or operating system.
Reference throughout the specification to "an embodiment means that a particular feature, structure or feature described in relation to the embodiment is included in at least one embodiment of the present invention. Consequently, the appearance of the phrase "in one embodiment" in various places throughout this specification does not necessarily always refer to the same embodiment, although it may do so.
In addition, the particular features, structures or features may be combined in any suitable manner, as will be apparent from the person skilled in the art of this invention, in one or more embodiments.
Likewise, it should be appreciated that in the above description of the exemplary embodiments of the invention, various features of the invention are sometimes grouped into a single embodiment, figure, or description of the invention for the purpose of profiling the invention and assisting in the understanding of one or more of the various aspects of the invention. However, this method of description should not be construed as reflecting an intention that the claimed invention requires more features than those expressly cited in each claim. Instead, as the following claims reflect, aspects of the invention are sustained in less than in all features of a single embodiment disclosed above. Therefore, the claims that follow the DESCRIPTION OF EXEMPLARY EMBODIMENTS are expressly incorporated into this DESCRIPTION OF EXEMPLARY EMBODIMENTS, wherein each claim itself represents a separate embodiment of this invention.
In addition, while some embodiments described in this application include some, but not other features included in other embodiments, combinations of features of different embodiments are intended to be within the scope of the invention, and form different embodiments, as will be understood by the art experts. For example, in the following claims, any of the claimed embodiments can be used in any combination.
Moreover, some of the embodiments are described in this application as a method or a combination of elements of a method that can be implemented by a processor of a computer system, or by another method to carry out the function. Accordingly, a processor with the instructions necessary to carry out said method or element of a method forms a means to carry out the method or element of a method.
Likewise, an element described in this application for an embodiment of an apparatus is an example of a means for carrying out the function performed by the element for the purpose of carrying out the invention.
In the description provided in this application, numerous specific details are set forth. However, it is understood that embodiments of the invention can be brought to the. practice without these specific details. In other cases, well-known methods, structures and techniques have not been shown in detail, so as not to obscure the understanding of this description.
According to this request, unless otherwise specified, the use of the ordinal adjectives "first," "second," "third, etc., to describe a common object, merely indicates that reference is made to different occurrences of Similar objects, and it is not intended to imply that the objects described must be in a certain sequence, whether temporary, spatial, in classification or in any other way.
It should be noted that while the Invention has been described in the context of E-AC-3, the Invention is not limited to that context, and can be used for decoding information encoded by other methods that use techniques that have certain similarity with E-AC-3. For example, the embodiments of the invention are also applicable for decoding encoded audio that is retroactively compatible with E-AC-3. Other embodiments are applicable for decoding of encoded audio that is encoded in accordance with the HE-AAC standard, and for decoding of encoded audio that is retroactively compatible with HE-AAC. Other encoded streams can also be conveniently decoded using embodiments of the present invention.
All United States patents, patent applications of
United States and international patent applications (PCT) designating
The United States cited in this application is incorporated into this application by reference. In the event that the rules or statutes on patents do not allow the Incorporation as a reference of material that itself incorporates information by way of reference, the incorporation as a reference, of the material of this application excludes any Information Incorporated as a reference in said material incorporated as a reference, unless such information is explicitly incorporated in this application by way of reference.
Any description of prior art in this specification should not be considered in any way an admission that said prior art is widely and publicly known, or is part of the general knowledge in the field.
In the claims that follow and in the description of this application, any of the terms "comprising", "composed of" or "comprising" is an open term which means that it includes at least the elements / features that follow, although without exclusion of others. Consequently, the term "comprising", when used in the claims, should not be construed as limiting the means, elements or steps mentioned below. For example, the scope of the expression “a device that includes A and B” should not be limited to devices consisting only of elements A and B. Any of the terms “that includes” or “includes” as used in this Application is also an open term that also means that it includes at least the elements / features that follow the term, although without the exclusion of others. Therefore, the term “that includes” is synonymous with “that includes”, and means “that includes”.
It should also be noted that the term "coupled", when used in the claims, should not be construed as limiting direct connections only. The terms "coupled" and "connected, together with their derivatives, can be used. It should be understood that these terms are not proposed as synonyms for each other. Consequently, the scope of the expression "a device A coupled to a device B" should not be limited to devices or systems where a device output A is directly connected to a device input B. It means that there is a path between an output of A and a B input that can be a path that includes other devices or media. The term “coupled can mean that two or more elements are either in direct physical or electrical contact, or that two or more elements are not in direct contact with each other, although they still cooperate or interact with each other.
Accordingly, while what is believed to be the preferred embodiments of the invention have been described, those skilled in the art will recognize that different and additional modifications may be made in such embodiments without departing from the spirit of the invention, and it is intended to claim all such changes and modifications that are within the scope of the invention. For example, any formula provided above is merely representative of procedures that can be used. Functionalities can be added to block diagrams, or they can be eliminated, and operations can be exchanged between functional elements. Steps can be added to the methods described, or eliminated, within the scope of the present invention.
Contents6
14 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14
73 members in 38 offices
Priority claims10
| Document | Office | Kind | Date |
|---|---|---|---|
| 30587110 | United States of America | P | |
| 30587110 | United States of America | P | |
| 61305871 | United States of America | – | |
| 35976310 | United States of America | P | |
| 35976310 | United States of America | P | |
| 61359763 | United States of America | – | |
| 61305871 | – | – | – |
| 61359763 | – | – | – |
| US20100305871P | – | – | – |
| US20100359763P | – | – | – |
Members73
| Document | Office | Kind | |
|---|---|---|---|
| EP2360683A1 | European Patent Office (EPO) | A1 | |
| CA2757643A1 | Canada | A1 | |
| CA2794029A1 | Canada | A1 | |
| CA2794047A1 | Canada | A1 | |
| WO2011102967A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU2011218351A1 | Australia | A1 | |
| SG174552A1 | Singapore | A1 | |
| TW201142826A | Taiwan Province of China | A | |
| MX2011010285A | Mexico | A | |
| IL215254D0 | Israel | D0 | |
| US2012016680A1 | United States of America | A1 | |
| ECSP11011358A | Ecuador | A | |
| AR080183A1 | Argentina | A1 | |
| EA201171268A1 | Eurasian Patent Organization (EAPO) | A1 | |
| KR20120031937A | Republic of Korea | A | |
| CN102428514A | China | A | |
| MA33270B1 | Morocco | B1 | |
| NI201100175A | Nicaragua | A | |
| US8214223B2 | United States of America | B2 | |
| HK1160282A1 | Hong Kong, China | A1 | |
| CO6501169A2 | Colombia | A2 | |
| PE20121261A1 | Peru | A1 | |
| US2012237039A1 | United States of America | A1 | |
| JP2012527021A | Japan | A | |
| AU2011218351B2 | Australia | B2 | |
| ZA201106950B | South Africa | B | |
| CA2757643C | Canada | C | |
| HK1170059A1 | Hong Kong, China | A1 | |
| UA101262C2 | Ukraine | C2 | |
| TN2011000541A1 | Tunisia | A1 | |
| KR20130055033A | Republic of Korea | A | |
| CN102428514B | China | B | |
| IL227701D0 | Israel | D0 | |
| IL227702D0 | Israel | D0 | |
| IL215254A | Israel | A | |
| KR101327194B1 | Republic of Korea | B1 | |
| CN103400581A | China | A | |
| EP2698789A2 | European Patent Office (EPO) | A2 | |
| GT201100246A | Guatemala | A | |
| EP2360683B1 | European Patent Office (EPO) | B1 | |
| EP2698789A3 | European Patent Office (EPO) | A3 | |
| GEP20146086B | Georgia | B | |
| JP5501449B2 | Japan | B2 | |
| PT2360683E | Portugal | E | |
| ES2467290T3 | Spain | T3 | |
| DK2360683T3 | Denmark | T3 | |
| TWI443646B | Taiwan Province of China | B | |
| HRP20140506T1 | Croatia | T1 | |
| SI2360683T1 | Slovenia | T1 | |
| JP2014146040A | Japan | A | |
| NZ595739A | New Zealand | A | |
| PL2360683T3 | Poland | T3 | |
| AR089918A2This record | Argentina | A2 | |
| US8868433B2 | United States of America | B2 | |
| RS53336B | Serbia | B | |
| TW201443876A | Taiwan Province of China | A | |
| ME01880B | Montenegro | B | |
| IL227701A | Israel | A | |
| HN2011002584A | Honduras | A | |
| IL227702A | Israel | A | |
| AP3147A | African Regional Intellectual Property Organization (ARIPO) | A | |
| US2016035355A1 | United States of America | A1 | |
| JP5863858B2 | Japan | B2 | |
| US9311921B2 | United States of America | B2 | |
| BRPI1105248A2 | Brazil | A2 | |
| CN103400581B | China | B | |
| MY157229A | Malaysia | A | |
| TWI557723B | Taiwan Province of China | B | |
| EA025020B1 | Eurasian Patent Organization (EAPO) | B1 | |
| EP2698789B1 | European Patent Office (EPO) | B1 | |
| KR101707125B1 | Republic of Korea | B1 | |
| CA2794029C | Canada | C | |
| BRPI1105248B1 | Brazil | B1 |
2 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Grant, registrationFG | FG | |
| Grant, registrationFG | FG |
Numbers
- Publication
- 089918
- Publication, DOCDB
- 089918
- Publication, EPODOC
- AR089918
- Application
- 100367
- Application, DOCDB
- P130100367
- Application, EPODOC
- AR2013P100367
Titles2
- Spanish
- DECODIFICADOR DE AUDIO Y METODO DE DECODIFICACION USANDO COMBINACION DE REDUCCION DE CANALES DE AUDIO (DOWNMIXING) EFICIENTE
- English
- AUDIO DECODER AND DECODING METHOD USING EFFICIENT AUDIO CHANNEL REDUCTION (DOWNMIXING) COMBINATION
Classification
- CPC, 8
- G10L19/008
- H04S3/008
- G10L19/02
- G10L19/06
- G10L19/167
- G10L19/24
- H04R5/02
- G10L19/022
- IPC, 1
- H03M13 00