Individual channel temporal envelope shaping for binaural cue coding schemes and the like
Abstract
Method for encoding audio channels, the method comprising: generate two or more indication codes for one or more audio channels, in which at least one indication code is an envelope indication code generated by the characterization of a temporary envelope in one of the one or more audio channels, wherein the one or more indication codes further comprise one or more interchannel correlation codes (ICC), interchannel level difference code (ICLD) and interchannel time difference codes (ICTD), in which a first time resolution associated with the envelope indication code is finer than a second time resolution associated with the other indication code (s) and in which the time envelope is characterized for the corresponding audio channel in a time domain or individually for different signal subbands of the corresponding audio channel in a subband domain; and transmit the two or more indication codes.

Term
Term ended
Projected expiry passed 7 September 2025, 1 year ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
41 claims: 5 independent, 36 dependent
- 1ES 2 323 275 T3 REIVINDICACIONES 1. Método para codificar canales de audio, comprendiendo el método:generar dos o más códigos de indicación para uno o más canales de audio, en el que al menos un código de indicación es un código de indicación de envolvente generado por la caracterización de una envolvente temporal en uno de los uno o más canales de audio, en el que el uno o más códigos de indicación comprenden además uno o más de códigos de correlación intercanal (ICC), código de diferencia de nivel intercanal (ICLD) y códigos de diferencia de tiempo intercanal (ICTD), en el que una primera resolución de tiempo asociada con el código de indicación de envolvente es más fina que una segunda resolución de tiempo asociada con el (los) otro(s) código(s) de indicación y en el que la envolvente temporal se caracteriza para el canal de audio correspondiente en un dominio de tiempo o individualmente para diferentes subbandas de señal del canal de audio correspondiente en un dominio de subbanda;y transmitir los dos o más códigos de indicación.
- 2Método según la reivindicación 1, que comprende además transmitir E canal(es) de audio transmitido(s) correspondiente^) al uno o más canales de audio, siendo E 1.
- 3Método según la reivindicación 2, en el que:el uno o más canales de audio comprenden C canales de audio de entrada, siendo C E;y los C canales de entrada se mezclan descendentemente para generar el (los) E canal(es) transmitido(s).
- 4Método según la reivindicación 1, en el que los dos o más códigos de indicación se transmiten para permitir que un descodificador efectúe la conformación de envolvente durante la descodificación del (de los) E canal(es) transmitido(s) basándose en los dos o más códigos de indicación, en el que el (los) E canal(es) de audio transmitido(s) corresponde(n) al uno o más canales de audio, siendo E 1.
- 5Método según la reivindicación 4, en el que la conformación de envolvente ajusta una envolvente temporal de una señal sintetizada generada por el descodificador para coincidir con la envolvente temporal caracterizada.
- 6Método según la reivindicación 1, en el que la envolvente temporal se caracteriza solamente para frecuencias especificadas del canal de audio correspondiente.
- 7Método según la reivindicación 1, en el que la envolvente temporal se caracteriza solamente para frecuencias del canal de audio correspondiente por encima de una frecuencia de corte especificada.
- 8Método según la reivindicación 1, en el que el dominio de subbanda corresponde a un filtro de espejo en cuadratura (QMF).
- 9Método según la reivindicación 1, que comprende además determinar si se habilita o deshabilita la caracterización.
- 10Método según la reivindicación 9, que comprende además generar y transmitir una bandera de habilitación/deshabilitación basándose en la determinación de instruir a un descodificador si implementar o no la conformación de envolvente durante la descodificación del (los) E canal(es) transmitido(s) correspondiente(s) al uno o más canales de audio, siendo E 1.
- 11Método según la reivindicación 9, caracterizado porque la determinación está basada en analizar un canal de audio para detectar transitorios en el canal de audio de tal manera que la caracterización se habilita si se detecta la presencia de un transitorio.
- 12Método según la reivindicación 1, en el que la etapa de generar el código de indicación de envolvente incluye elevar al cuadrado (1006) o formar una magnitud y filtrar paso bajo (1008) muestras de señal del canal de audio o señales de subbanda del canal de audio con el fin de caracterizar la envolvente temporal.
- 13Método según la reivindicación 1 ó 12, en el que la etapa de generación comprende además la etapa de parametrizar, cuantificar y codificar una envolvente temporal estimada.
- 14Aparato para codificar canales de audio, comprendiendo el aparato:medios para generar dos o más códigos de indicación para uno o más canales de audio, en el que al menos un código de indicación es un código de indicación de envolvente generado mediante la caracterización de una envolvente temporal en uno de los uno o más canales de audio, en el que los dos o más códigos de indicación comprenden además uno o más de códigos de correlación intercanal (ICC), códigos de diferencia de nivel ES 2 323 275 T3 intercanal (ICLD) y códigos de diferencia de tiempo intercanal (ICTD), en el que una primera resolución de tiempo asociada con el código de indicación de envolvente es más fina que una segunda resolución de tiempo asociada con el (los) otro(s) código(s) de indicación, y en el que la envolvente temporal se caracteriza para el canal de audio correspondiente en un dominio de tiempo o individualmente para diferentes subbandas de señal del canal de audio correspondiente en un dominio de subbanda;y medios para transmitir información acerca de los dos o más códigos de indicación.
- 15Aparato según la reivindicación 14, en el que el aparato es operativo para codificar C canales de audio de entrada para generar E canal(es) de audio transmitido(s), en el que los medios de generación comprenden un analizador de envolvente adaptado para caracterizar la envolvente temporal de entrada de al menos uno de los C canales de entrada, en el que los medios de generación comprenden además un estimador de código adaptado para generar códigos de indicación para dos o más de los C canales de entrada y en el que el aparato comprende además un mezclador descendente adaptado para mezclar descendentemente los C canales de entrada para generar el (los) E canal(es) transmitido(s), siendo C E 1, en el que los medios de transmisión están adaptados para transmitir información acerca de los dos o más códigos de indicación para permitir que un descodificador efectúe la síntesis y conformación de envolvente durante la descodificación del (los) E canal(es) transmitido(s).
- 16Aparato según la reivindicación 15, en el que:el aparato es un sistema seleccionado del grupo que consiste en un grabador de vídeo digital, un grabador de audio digital, un ordenador, un transmisor de satélite, un transmisor de cable, un transmisor de difusión terrestre, un sistema de entretenimiento en casa y un sistema de cine, y el sistema comprende el analizador de envolvente, el estimador de código y el mezclador descendente.
- 17Producto de programa informático que tiene código de programa, en el que, cuando el código de programa se ejecuta por una máquina, la máquina implementa un método según la reivindicación 1.
- 18Flujo de bits de audio codificado, que tiene:dos o más códigos de indicación generados para uno o más canales de audio, en el que al menos un código de indicación es un código de indicación de envolvente generado mediante la caracterización de una envolvente temporal en uno de los uno o más canales de audio, en el que los dos o más códigos de indicación comprenden además uno o más de códigos de correlación intercanal (ICC), código de diferencia de nivel intercanal (ICLD) y códigos de diferencia de tiempo intercanal (ICTD), en el que una primera resolución de tiempo asociada con el código de indicación de envolvente es más fina que una segunda resolución de tiempo asociada con el (los) otro(s) código(s) de indicación, y en el que la envolvente temporal se caracteriza para el canal de audio correspondiente en un dominio de tiempo o individualmente para diferentes subbandas de señal del canal de audio correspondiente en un dominio de subbanda, y los dos o más códigos de indicación y E canal(es) de audio transmitido(s) que corresponden al uno o más canales de audio, siendo E 1, se codifican en el flujo de bits de audio codificado.
- 19Flujo de bits de audio codificado según la reivindicación 18, que comprende además E canal(es) de audio transmitido(s), en el que:el (los) E canal(es) de audio transmitido(s) corresponde(n) al uno o más canales de audio.
- 20Método para descodificar E canal(es) de audio transmitido(s), para generar C canales de audio de reproducción, siendo C E 1, comprendiendo el método:recibir códigos de indicación correspondientes al (a los) E canal(es) transmitido(s), en el que los códigos de indicación comprenden un código de indicación de envolvente correspondiente a una envolvente temporal caracterizada de un canal de audio correspondiente al (a los) E canal(es) transmitido(s), en el que los dos o más códigos de indicación comprenden además uno o más de códigos de correlación intercanal (ICC), códigos de diferencia de nivel intercanal (ICLD) y códigos de diferencia de tiempo intercanal (ICTD), en el que una primera resolución de tiempo asociada con el código de indicación de envolvente es más fina que una segunda resolución de tiempo asociada con el (los) otro(s) código(s) de indicación;ES 2 323 275 T3 mezclar ascendentemente uno o más del (de los) E canal(es) transmitido(s) para generar uno o más canales mezclados ascendentemente;y sintetizar uno o más de los C canales de reproducción mediante la aplicación de los códigos de indicación a uno o más canales mezclados ascendentemente, en el que el código de indicación de envolvente se aplica a un canal mezcla ascendentemente o una señal sintetizada para ajustar una envolvente temporal de la señal sintetizada basándose en la envolvente temporal caracterizada mediante ajuste a escala de muestras de señal en el dominio de tiempo o en el dominio de subbanda utilizando un factor de ajuste a escala tal que la envolvente temporal ajustada coincida con la envolvente temporal caracterizada.
- 21Método según la reivindicación 20, en el que el código de indicación de envolvente corresponde a una envolvente temporal caracterizada en un canal de entrada original usado para generar el (los) E canal(es) transmitido(s).
- 22Método según la reivindicación 20, en el que la síntesis comprende síntesis de ICC de reverberación tardía.
- 23Método según la reivindicación 21, en el que la envolvente temporal de la señal sintetizada se ajusta antes de la síntesis de ICLD.
- 24Método según la reivindicación 20, en el que:la envolvente temporal de la señal sintetizada está caracterizada;y la envolvente temporal de la señal sintetizada se ajusta basándose tanto en la envolvente temporal caracterizada correspondiente al código de indicación de envolvente como en la envolvente temporal caracterizada de la señal sintetizada.
- 25Método según la reivindicación 24, en el que:se genera una función de ajuste a escala basándose en la envolvente temporal caracterizada correspondiente al código de indicación de envolvente y la envolvente temporal caracterizada de la señal sintetizada;y la función de ajuste a escala se aplica a la señal sintetizada.
- 26Método según la reivindicación 20, que comprende además ajustar un canal transmitido basándose en la envolvente temporal caracterizada para generar un canal aplanado, en el que la mezcla ascendente y la síntesis se aplican al canal aplanado para generar un canal de reproducción correspondiente.
- 27Método según la reivindicación 20, que comprende además ajustar un canal mezcla ascendentemente basándose en la envolvente temporal caracterizada para generar un canal aplanado, en el que la síntesis se aplica al canal aplanado para generar un canal de reproducción correspondiente.
- 28Método según la reivindicación 20, en el que la envolvente temporal de la señal sintetizada se ajusta solamente para frecuencias especificadas.
- 29Método según la reivindicación 28, en el que la envolvente temporal de la señal sintetizada se ajusta solamente para frecuencias por encima de una frecuencia de corte especificada.
- 30Método según la reivindicación 20, en el que las envolventes temporales se ajustan individualmente para diferentes subbandas de señal en la señal sintetizada.
- 31Método según la reivindicación 20, en el que un dominio de subbanda corresponde a un QMF.
- 32Método según la reivindicación 20, en el que la envolvente temporal de la señal sintetizada se ajusta en un dominio de tiempo.
- 33Método según la reivindicación 20, que comprende además determinar si se habilita o deshabilita el ajuste de la envolvente temporal de la señal sintetizada.
- 34Método según la reivindicación 33, caracterizado porque la determinación está basada en una bandera de habilitación/deshabilitación generada por un codificador de audio que generó el (los) E canal(es) transmitido(s).
- 35Método según la reivindicación 33, en el que la determinación está basada en el análisis del (los) E canal(es) transmitido(s) para detectar transitorios de tal manera que el ajuste se habilita si se detecta la presencia de un transitorio. ES 2 323 275 T3
- 36Método según la reivindicación 20, que comprende además:caracterizar una envolvente temporal de un canal transmitido;y determinar si se usa (1) la envolvente temporal caracterizada correspondiente al código de indicación de envolvente o (2) la envolvente temporal caracterizada del canal transmitido para ajustar la envolvente temporal de la señal sintetizada.
- 37Método según la reivindicación 20, en el que la potencia dentro de una ventana especificada de la señal sintetizada después del ajuste de la envolvente temporal es igual a la potencia dentro de una ventana correspondiente de la señal sintetizada antes del ajuste.
- 38Método según la reivindicación 37, en el que la ventana especificada corresponde a una ventana de síntesis asociada con uno o más códigos de indicación sin envolvente.
- 39Aparato para descodificar E canal(es) de audio transmitido(s) para generar C canales de audio de reproducción, siendo C E 1, comprendiendo el aparato:medios para recibir códigos de indicación correspondientes al (a los) E canal(es) transmitido(s), en el que los códigos de indicación comprenden un código de indicación de envolvente correspondiente a una envolvente temporal caracterizada de un canal de audio correspondiente al (a los) E canales transmitido(s), en el que los dos o más códigos de indicación comprenden además uno o más de códigos de correlación intercanal (ICC), códigos de diferencia de nivel intercanal (ICLD) y códigos de diferencia de tiempo intercanal (ICTD), en el que una primera resolución de tiempo asociada con el código de indicación de envolvente es más fina que una segunda resolución de tiempo asociada con el (los) otro(s) código(s) de indicación;medios para mezclar ascendentemente uno o más de los E canales transmitidos para generar uno o más canales mezclados ascendentemente;y medios para sintetizar uno o más de los C canales de reproducción mediante la aplicación de los códigos de indicación al uno o más canales mezclados ascendentemente, en el que el código de indicación de envolvente se aplica a un canal mezcla ascendentemente o una señal sintetizada para ajustar una envolvente temporal de la señal sintetizada basándose en la envolvente temporal caracterizada mediante el ajuste a escala de muestras de señal en el dominio de tiempo o en el dominio de subbanda utilizando un factor de ajuste a escala tal que la envolvente temporal ajustada coincida sustancialmente con la envolvente temporal caracterizada.
- 40Aparato según la reivindicación 39, en el que:el aparato es un sistema seleccionado a partir del grupo que consiste en un reproductor de vídeo digital, un reproductor de audio digital, un ordenador, un receptor de satélite, un receptor de cable, un receptor de difusión terrestre, un sistema de entretenimiento en casa y un sistema de cine;y el sistema comprende el receptor, el mezclador ascendente, el sintetizador y el ajustador de envolvente.
- 41Producto de programa informático que tiene código de programa, en el que, cuando el código de programa se ejecuta por una máquina, la máquina implementa el método de descodificación según la reivindicación 20.
Independent claims41
194 paragraphs in 11 sections, as filed
ES 2 323 275 T3
DESCRIPTION
Individual channel time envelope shaping for binaural indication coding schemes and the like.
Background of the invention
The content of this application is related to the content of the following US application publications:
Or US 2003/0026441;
or US 2003/0035553;
or US 2003/0219130;
or US 2003/0236583;
or US 2009/0180579;
or US 2005/0058304;
or US 2005/0157883; I US 2006/0085200.
The content of this application is also related to the content described in the following documents:
o F. Baumgarte and C. Faller, “Binaural Cue Coding - Part I: Psychoacoustic fundamentals and design principles”, IEEE Trans. on Speech and Audio Proc., vol. 11, no.6, November 2003;
o C. Faller and F. Baumgarte, “Binaural Cue Coding - Part II: Schemes and applications”, IEEE Trans. on Speech and Audio Proc., vol. 11, no.6, November 2003; I C. Faller, “Coding of spatial audio compatible with different playback formats”, Preprint 17th Conv. Aud. Eng. Soc., October 2004.
Field of the invention
The present invention relates to the encoding of audio signals and the subsequent synthesis of auditory scenes from the encoded audio data.
Description of Related Art
When a person hears an audio signal (that is, sounds) generated by a particular audio source, the audio signal will commonly arrive at the person's left and right ears at two different times and at two different audio levels (for example, decibels), in which these different times and levels are a function of the differences in the paths through which the audio signal travels to reach the left and right ears, respectively. The person's brain interprets these differences in time and level to give the person the perception that the received audio signal is being generated by an audio source located in a particular position (for example, direction and distance) with respect to the person. An auditory scene is the net effect of a person simultaneously listening to audio signals generated by one or more different audio sources located at one or more different positions relative to the person.
The existence of this processing by the brain can be used to synthesize auditory scenes, in which audio signals from one or more different audio sources are intentionally modified to generate left and right audio signals that give the perception that different audio sources audio are located at different positions with respect to the listener.
Figure 1 shows a high-level block diagram of the conventional binaural signal synthesizer 100, which converts a single audio source signal (eg, a mono signal) into the left and right audio signals of a binaural signal, defining a binaural signal like the two signals received at a listener's eardrums. In addition to the audio source signal, the synthesizer 100 receives a set of spatial indications corresponding to the desired position of the audio source relative to the listener. In typical implementations, the set of spatial indications comprises an interchannel level difference (ICLD) value (which identifies the difference in audio level between the left and right audio signals as received by the left ears.
ES 2 323 275 T3 do and right, respectively) and an inter-channel time difference (ICTD) value (which identifies the arrival time difference between the left and right audio signals as received by the left and right ears, respectively). Additionally or alternatively, some synthesis techniques involve modeling a direction-dependent transfer function for sound from the signal source to the eardrums, also referred to as the head-related transfer function (HRTF). See, for example, J. Blauert, The Psychophysics of Human Sound Localization, MIT Press, 1983.
Using the binaural signal synthesizer 100 of FIG. 1, the mono audio signal generated by a single sound source can be processed such that, when listening through headphones, the sound source is spatially positioned by applying an appropriate set of cues. (eg ICLD, ICTD and / or HRTF) to generate the audio signal for each ear. See, for example, DR Begault, 3-D Sound for Virtual Reality and Multimedia, Academic Press, Cambridge, MA, 1994.
The binaural signal synthesizer 100 of FIG. 1 generates the simplest type of auditory scenes: those that have a single audio source positioned relative to the listener. More complex auditory scenes comprising two or more audio sources located at different positions relative to the listener can be generated using an auditory scene synthesizer that is implemented essentially using multiple instances of the binaural signal synthesizer, with each binaural signal synthesizer instance generating the binaural signal corresponding to a different audio source. Since each different audio source has a different location relative to the listener, a different set of spatial cues is used to generate the binaural audio signal for each different audio source.
An object of the present invention is to provide an improved concept for audio coding.
This object is achieved by a method for encoding according to claim 1, an apparatus for encoding according to claim 14, a computer program product according to claim 17, an encoded audio bitstream according to claim 18, a method for decoding according to claim 20, a decoding apparatus according to claim 39 and a computer program product according to claim 41. US 5,812,971 (HERRE) discloses a method of intensity stereo coding of multichannel audio signals using time envelope shaping.
Article 6447 of the AES Convention, J. Herre et. al, "The Reference Model Architecture for MPEG Spatial Audio Coding", discloses a way in which you do not have a totally discrete representation of multichannel sound, but you have a compatible stereo transmission rate only slightly higher than the usual rates. used for mono / stereo sound. Specifically, OTT and TTT elements are used which are based on level difference parameters and interchannel coherent / cross-correlation parameters representing time / frequency varying coherence or cross-correlation between two input channels.
The technical publication "Parametric Coding of Spatial Audio", C. Faller, Proceedings of the 7th International Conference on Digital Audio Effect, Naples, Italy, October 5, 2004, pages 151-156, discloses BCC technology. BCC represents stereo or multichannel audio signals as a single or more downmixed audio channels plus side information. The lateral information contains the interchannel cues inherent in the original audio signal that are relevant to the perception of properties of the auditory spatial image. The relationship between interchannel indications and attributes of the audio-auditory spatial image is discussed.
Brief description of the figures
Other aspects, elements, and advantages of the present invention will become more fully apparent from the following detailed description, the appended claims, and the accompanying drawings in which the same reference numerals identify similar or identical elements.
Figure 1 shows a high-level block diagram of the conventional binaural signal synthesizer;
Figure 2 is a block diagram of a generic Binaural Indication Coding (BCC) audio processing system;
Figure 3 shows a block diagram of a down mixer that can be used for the down mixer of Figure 2;
Figure 4 shows a block diagram of a BCC synthesizer that can be used for the decoder of Figure 2;
Figure 5 shows a block diagram of the BCC estimator of Figure 2 according to an embodiment of the present invention;
Figure 6 illustrates ICTD and ICLD data generation for five channel audio;
Figure 7 illustrates the generation of ICC data for five channel audio;
ES 2 323 275 T3 Figure 8 shows a block diagram of an implementation of the BCC synthesizer of Figure 4 that can be used in a BCC decoder to generate a stereo or multichannel audio signal given an individual transmitted sum signal s (n) plus the spatial indications;
Figure 9 illustrates how ICTD and ICLD are modified within a subband as a function of frequency;
Figure 10 shows a block diagram of the time domain processing that is added to a BCC encoder, such as the encoder of Figure 2, in accordance with one embodiment of the present invention;
Figure 11 illustrates an exemplary time domain application of TP processing in the context of the BCC synthesizer of Figure 4;
Figure 13 shows a block diagram of frequency domain processing that is added to a BCC encoder, such as the encoder of Figure 2, in accordance with an alternative embodiment of the present invention;
Figure 14 illustrates an exemplary frequency domain application of TP processing in the context of the BCC synthesizer of Figure 4;
Figure 15 shows a block diagram of the frequency domain processing that is added to a BCC encoder, such as the encoder of Figure 2, according to another alternative embodiment of the present invention;
Figure 16 illustrates another exemplary frequency domain application of TP processing in the context of the BCC synthesizer of Figure 4;
Figures 17 (a) - (c) show block diagrams of possible implementations of the TPAs of Figures 15 and 16 and ITP and TP of Figure 16; and Figures 18 (a) and (b) illustrate two exemplary modes of operation of the control block of Figure 16.
Detailed description
In binaural indication coding (BCC), an encoder encodes C input audio channels to generate E transmitted audio channels, with C> E> 1. In particular, two or more of the C input channels are provided in one frequency domain and one or more indication codes are generated for each of one or more different frequency bands on the two or more input channels in the domain. of frequency. In addition, the C input channels are downmixed to generate the E transmitted channels. In some downmix implementations, at least one of the E transmitted channels is based on two or more of the C input channels and at least one of the E transmitted channels is based on only one of the C input channels.
In one embodiment, a BCC encoder has two or more filter banks, a code estimator, and a down-mixer. The two or more filter banks convert two or more of the C input channels from a time domain to a frequency domain. The code estimator generates one or more indication codes for each of one or more different frequency bands on the two or more converted input channels. The downmixer downmixes the C input channels to generate the E transmitted channels, with C> E> 1.
In BCC decoding, E transmitted audio channels are decoded to generate C playback audio channels. In particular, for each of one or more different frequency bands, one or more of the E transmitted channels are upmixed in a frequency domain to generate two or more of the C reproduction channels in the frequency domain. , where C> E> 1. One or more indication codes are applied to each of the one or more different frequency bands in the two or more reproduction channels in the frequency domain to generate two or more modified channels, and the two or more modified channels are converted from the frequency domain to a time domain. In some upmix implementations, at least one of the C playback channels is based on at least one of the E transmitted channels and at least one cue code, and at least one of the C playback channels is based on only one. only of the E channels transmitted and independent of any indication code.
In one embodiment, a BCC decoder has an up-mixer, a synthesizer, and one or more reverse filter banks. For each of one or more different frequency bands, the upmixer upmixes one or more of the E channels transmitted in a frequency domain to generate two or more of the C playback channels in the frequency domain, where C> E> 1. The synthesizer applies one or more cue codes to each of the one or more different frequency bands on the two or more reproduction channels in the frequency domain to generate two or more modified channels. The one or more inverse filter banks convert the two or more modified channels from the frequency domain to a time domain.
Depending on the particular implementation, a given playback channel may be based on a single transmitted channel, rather than a combination of two or more transmitted channels. For example, when there is only one transmitted channel, each of the C playback channels is based on that transmitted channel. In this situations,
ES 2 323 275 T3 the upmix corresponds to copying the corresponding transmitted channel. As such, for applications where there is only one transmitted channel, the up-mixer can be implemented using a replicator that copies the transmitted channel for each playback channel.
BCC encoders and / or decoders can be incorporated into various systems or applications including, for example, digital video recorders / players, digital audio recorders / players, computers, satellite transmitters / receivers, cable transmitters / receivers, terrestrial broadcast transmitters / receivers, home entertainment systems and theater systems.
Generic BCC Processing
FIG. 2 is a block diagram of a generic binaural indication coding (BCC) audio processing system 200 comprising an encoder 202 and a decoder 204. Encoder 202 includes down mixer 206 and estimator 208 BCC.
Down mixer 206 converts C input audio channels x, (n) into E transmitted audio channels yi (n), where C> E> 1. In this specification, the signals expressed using the variable n are signals in the time domain, while the signals expressed using the variable k are signals in the frequency domain. Depending on the particular implementation, downmixing can be implemented in either the time domain or the frequency domain. The BCC estimator 208 generates BCC codes from the C input audio channels and transmits these BCC codes as either in-band or out-of-band side information with respect to the E transmitted audio channels. Typical BCC codes include one or more interchannel time difference (ICTD), interchannel level difference (ICLD), and interchannel correlation (ICC) data estimated between certain pairs of input channels as a function of frequency and time. The particular implementation will determine between which particular pairs of input channels the BCC codes are estimated.
ICC data corresponds to the coherence of a binaural signal, which is related to the perceived width of the audio source. The wider the audio source, the lower the coherence between the left and right channels of the resulting binaural signal. For example, the coherence of the binaural signal for an orchestra spread across an auditorium stage is typically lower than the coherence of the binaural signal for a single solo violin. In general, an audio signal with lower coherence is usually perceived as more spread out in the auditory space. As such, ICC data typically refers to the apparent font width and the degree to which the listener is involved. See, for example, J. Blauert, The Psychophysics of Human Sound Localization, MIT Press, 1983.
Depending on the particular application, the E transmitted audio channels and the corresponding BCC codes may be transmitted directly to decoder 204 or stored in some appropriate type of storage device for later access by decoder 204. Depending on the situation, the term " transmission ”can refer either to direct transmission to a decoder or to storage for later provision to a decoder. In either case, the decoder 204 receives the transmitted audio channels and side information and performs an upmix and BCC synthesis using the BCC codes to convert the E transmitted audio channels into more than E (usually, although not necessarily C) playback audio channels x, (n) for audio playback. Depending on the particular implementation, upmixing can be done in either the time domain or the frequency domain.
In addition to the BCC processing shown in Figure 2, a generic BCC audio processing system may include additional encoding and decoding steps, to further compress the audio signals in the encoder and then decompress the audio signals in the decoder, respectively. These audio codecs can be based on conventional audio compression / decompression techniques, such as those based on pulse code modulation (PCM), differential PCM (DPCM), or adaptive DPCM (ADPCM).
When the down-mixer 206 generates a single summation signal (i.e., E = 1), the BCC encoding can represent multichannel audio signals at a bit rate only slightly higher than that required to represent an audio signal. monkey. This is so because the estimated ICTD, ICLD, and ICC data between a pair of channels contains approximately two orders of magnitude less information than an audio waveform.
Not only is the low bit rate of the BCC encoded interesting, but also its backward compatibility aspect. A single transmitted sum signal corresponds to a mono downmix of the original stereo or multichannel signal. For receivers that do not support stereo or multichannel sound reproduction, listening to the transmitted sum signal is a valid method of presenting audio material on low-profile mono reproduction equipment. Accordingly, BCC encoding can also be used to enhance existing services involving the delivery of mono audio material to multi-channel audio. For example, mono audio radio broadcast systems can be enhanced for stereo or multichannel reproduction if the BCC side information can be embedded in the existing broadcast channel. Analog capabilities exist when multichannel audio is down-mixed into two summation signals that correspond to stereo audio.
ES 2 323 275 T3
BCC processes audio signals with a certain time and frequency resolution. The frequency resolution used is largely motivated by the frequency resolution of the human auditory system. Psychoacoustics suggests that spatial perception is most likely based on a critical band representation of the acoustic input signal. This frequency resolution is considered using an invertible filter bank (for example, based on a fast Fourier transform (FFT) or a quadrature mirror filter (QMF)) with subbands with bandwidths equal to or proportional to the bandwidth. critical of the human auditory system.
Generic downmix
In preferred implementations, the transmitted sum signal (s) contain (s) all signal components of the input audio signal. The goal is for each signal component to be fully maintained. The simple addition of the input audio channels often results in amplification or attenuation of the signal components. In other words, the power of the signal components in a "simple" sum is often larger or smaller than the sum of the power of the corresponding signal component of each channel. A downmixing technique can be used that equalizes the sum signal, such that the power of the signal components in the sum signal is approximately the same as the corresponding power on all input channels.
Figure 3 shows a block diagram of a down-mixer 300 that can be used for the down-mixer 206 of Figure 2 in accordance with certain implementations of the BCC system 200. Down mixer 300 has filter bank 302 (FB) for each input channel x, (n), down mix block 304, optional scale / delay block 306, and inverse FB 308 (IFB) for each scrambled channel yfn).
Each filter bank 302 converts each frame (eg 20 ms) of a corresponding digital input channel xfn) in the time domain to a set of input coefficients xfk) in the frequency domain. The downmix block 304 downmixes each subband of C corresponding input coefficients into a corresponding subband of E coefficients in the frequency domain downmixed. Equation (1) represents the down-mixing of the k-th subband of input coefficients (x<sub>1</sub>(k), x<sub>2</sub>(k) ,,, Xc (k)) to generate the k-th subband of descendingly mixed coefficients (y (k), and<sub>2</sub>(k), ..., and<sub>AND</sub>(k)) as follows:
<td> |------------------------------------------------------------------ 1_________________________________________________</td><td></td><td> 1------------------------------------------------------------------ 1_________________________________________________</td>
<td></td><td></td><td>x<sub>c</sub>(k)</td>
where D<sub>EC</sub> is a real-valued C by E downmix matrix.
The optional scale / delay block 306 comprises a set of multipliers 310, each of which multiplies a down-mixed coefficient and<sub>1</sub>(k) corresponding by a scaling factor e<sub>1 </sub>(k) to generate a corresponding scaled coefficient yfk). The motivation for the scaling operation is equivalent to generalized EQ for downmix with arbitrary weighting factors for each channel. If the input channels are independent, then the power P<sub>yi (k)</sub> of the down-mixed signal in each subband is given by equation (2) as follows:
<td>Py, (k) Pydk)</td><td></td><td>Pxgk) Px<sub>2</sub>w</td>
<td>-PyríD.</td><td></td><td>_Px<sub>c</sub>(k) _</td>
(2) where D<sub>EC</sub> is obtained by squaring each matrix element in matrix D<sub>EC</sub> downmixing of C by E and P<sub>X1W</sub> is the power of subband k of input channel i.
ES 2 323 275 T3
If the subbands are not independent, then the power values P<sub>y1 (k)</sub> of the downmixed signal will be larger or smaller than that calculated using equation (2), due to signal amplifications or cancellations when the signal components are in phase or out of phase, respectively. To prevent this, the downmix operation of equation (1) is applied in subbands followed by the scaling operation of multipliers 310. The scaling factors e<sub>1</sub> (k) (1 <i <E) can be obtained using equation (3) as follows:
<img file="ES2323275T3_D0001.tif" />
where P<sub>yi (k)</sub> is the subband power calculated by equation (2) and P<sub>yi (k)</sub> is the power of the corresponding downmixed subband signal yfk).
In addition to or instead of providing the optional scaling, the scaling / delay block 306 may optionally apply delays to the signals.
Each inverse filter bank 308 converts a set of scaled coefficients and<sub>1</sub>(k) corresponding in the frequency domain in a frame of a digital transmitted channel and<sub>1</sub>(n) corresponding.
Although Figure 3 shows all of the C input channels converted to the frequency domain for subsequent downmixing, in alternative implementations one or more (but less than C-1) of the C input channels could skip part or all of the processing shown in Figure 3 and transmitted as an equivalent number of unmodified audio channels. Depending on the particular implementation, these unmodified audio channels may or may not be used by the BCC estimator 208 of FIG. 2 in generating the transmitted BCC codes.
In a down-mixer 300 implementation that generates a single sum signal y (n), E = 1 and the x signals<sub>c</sub>(k) of each subband of each input channel C are added and then multiplied by a factor e (k), according to equation (4) as follows:
y (k) = e (k) £ x<sub>c</sub>(k) <sup>e = l</sup> (4) the factor e (k) is given by equation (5) as follows:
<img file="ES2323275T3_D0002.tif" />
where P ^ (k) is a time estimate of the power of x<sub>c</sub> (k) at the time index k, and P<sub>x</sub> (k) is a time estimate of the power of £<sup>C</sup> , x<sub>c</sub>(k). The equalized subbands are transformed back to the time domain resulting in the sum signal y (n) which is transmitted to the decoder BCC.
Generic BCC synthesis
Figure 4 shows a block diagram of a 400 BCC synthesizer that can be used by the decoder 204 of Figure 2 according to certain implementations of the 200 BCC system. The 400 BCC synthesizer has a 402 filter bank for each transmitted channel and<sub>1</sub> (n), an upmix block 404, delays 406, multipliers 408, correlation block 410, and a reverse filter bank 412 for each playback channel x<sub>1</sub>(n).
Each filter bank 402 converts each frame of a corresponding digital transmitted channel and end) in the time domain into a set of input coefficients and<sub>1</sub>(k) in the frequency domain. The upmix block 404 upmixes each subband of corresponding transmitted channel coefficients E into a
ES 2 323 275 T3 corresponding subband of C coefficients in the frequency domain mixed up. Equation (4) represents the upmixing of the k-th subband of transmitted channel coefficients (yfk), and<sub>2</sub>(k),., and<sub>AND</sub>(k)) to generate the k-th subband of up-mixed coefficients (S<sub>1</sub>(k), S<sub>2</sub>(k), ..., s<sub>C</sub>(k)) as follows:
<img file="ES2323275T3_D0003.tif" />
where u<sub>EC</sub> is a real-valued E by C upmix matrix. Uplifting in the frequency domain allows you to upmix individually on each different subband.
Each delay 406 applies a delay value dfk) based on a corresponding BCC code for ICTD data to ensure that the desired ICTD values appear between certain pairs of reproduction channels. Each multiplier 408 applies a scaling factor fk) based on a corresponding BCC code for ICLD data to ensure that the desired ICLD values appear between certain pairs of playback channels. Correlation block 410 performs a decorrelation operation A based on corresponding BCC codes for ICC data to ensure that the desired ICC values appear between certain pairs of reproduction channels. A further description of the operations of the mapping block 410 can be found in US 2003/0219130.
Synthesis of ICLD values may be less problematic than synthesis of ICTD and ICC values, since ICLD synthesis merely involves scaling of subband signals. Since ICL cues are the most commonly used directional cues, it is usually more important that ICLD values approximate those of the original audio signal. As such, ICLD data could be estimated across all channel pairs. The scaling factors afk) (1 <i <C) for each subband are preferably chosen such that the subband power of each reproduction channel approaches the corresponding power of the original input audio channel.
One goal may be to apply relatively few signal modifications to synthesize ICTD and ICC values. As such, the BCC data might not include ICTD and ICC values for all channel pairs. In that case, the 400 BCC synthesizer would synthesize ICTD and ICC values only between certain pairs of channels.
Each inverse filter bank 412 converts a set of corresponding synthesized coefficients xfk) in the frequency domain into a frame of a corresponding digital reproduction channel xfn).
Although Figure 4 shows all of the E transmitted channels converted to the frequency domain for subsequent upmixing and BCC processing, in alternative implementations, one or more (but not all) of the E transmitted channels could skip some or all of the processing shown in Figure 4. For example, one or more of the transmitted channels may be unmodified channels that are not upmixed. In addition to being one or more of the C playback channels, these unmodified channels could, in turn, but need not be used as reference channels to which BCC processing is applied to synthesize one or more of the other playback channels. reproduction. Either way, such unmodified channels can be delayed to compensate for the processing time involved in upmixing and / or BCC processing used to generate the rest of the playback channels.
Note that although Figure 4 shows C reproduction channels synthesized from E transmitted channels, where C was also the number of original input channels, the BCC synthesis is not limited to the number of reproduction channels. In general, the number of reproduction channels can be any number of channels, including numbers greater or less than C and possibly even situations where the number of reproduction channels is equal to or less than the number of transmitted channels.
"Perceptually relevant differences" between audio channels
Assuming a single summation signal, BCC synthesizes a stereo or multichannel audio signal in such a way that ICTD, ICLD, and ICC approximate the corresponding indications of the original audio signal. The role of ICTD, ICLD, and ICC with respect to auditory spatial imaging attributes is discussed below.
Knowledge about spatial hearing implies that for an auditory event, ICTD and ICC are related to the perceived direction. When considering room binaural impulsive responses (BRIRs) from a source, there is a relationship between the width of the auditory event and how the listener becomes involved and the estimated ICC data for early and late parts of the BRIRs. However, the relationship between ICC and these properties for general signals (and not just BRIRs) is not direct.
ES 2 323 275 T3
Stereo and multichannel audio signals usually contain a complex mix of simultaneously active source signals superimposed by the reflected signal components resulting from recording in close quarters or added by the recording technician to artificially create a spatial impression. Signals from different sources and their reflections occupy different regions in the time-frequency plane. This is reflected by ICTD, ICLD, and ICC, which vary with time and frequency. In this case, the relationship between instantaneous ICTD, ICLD and ICC and auditory event directions and spatial impression is not obvious. The strategy of certain BCC embodiments is to blindly synthesize these cues, such that they approximate the corresponding cues of the original audio signal.
Filter banks with subbands of bandwidths equal to twice the Equivalent Rectangular Bandwidth (ERB) are used. Informal listening reveals that BCC audio quality does not noticeably improve when a higher frequency resolution is chosen. A lower frequency resolution may be desirable as it results in fewer ICTD, ICLD and ICC values that need to be transmitted to the decoder and thus a lower bit rate.
With regard to time resolution, ICTD, ICLD and ICC are normally considered at regular time intervals. High performance is achieved when ICTD, ICLD, and ICC are considered approximately every 4 to 16 ms. Note that unless the prompts are considered at very short time intervals, the precedence effect is not directly considered. Assuming a classical lead-lag pair of sound stimuli, if lead and lag fall in a time interval in which only one set of cues is synthesized, then lead location dominance is not considered. Despite this, BCC gets audio quality reflected in an average MUSHRA score of about 87 (ie "excellent" audio quality) on average and up to nearly 100 for certain audio signals.
The perceptually small difference frequently obtained between the reference signal and the synthesized signal implies that indications related to a wide range of auditory spatial image attributes are implicitly considered when synthesizing ICTD, ICLD and ICC at regular time intervals. Here are some arguments for how ICTD, ICLD, and ICC can be related to a range of auditory spatial image attributes.
Estimation of spatial indications
The following describes how ICTD, ICLD, and ICC are estimated. The bit rate for the transmission of these spatial indications (quantized and encoded) can be only a few kb / s and therefore, with BCC, it is possible to transmit stereo and multichannel audio signals at close bit rates to that required for a single audio channel.
Figure 5 shows a block diagram of the 208 BCC estimator of Figure 2, according to one embodiment of the present invention. The BCC estimator 208 comprises filter banks 502 (FB), which can be the same as the filter banks 302 of Figure 3 and the estimation block 504, which generates ICTD, ICLD and ICC spatial indications for each different frequency subband generated by filter banks 502.
ICTD, ICLD and ICC estimation for stereo signals
The following measures are used for ICTD, ICLD, and ICC for subband signals x<sub>1</sub>(k) and x<sub>2</sub>(k) corresponding two-channel audio (for example stereo):
o ICTD [samples]:
<img file="ES2323275T3_D0004.tif" />
with a time estimate of the normalized cross-correlation function given by equation (8) as follows:
<img file="ES2323275T3_D0005.tif" />
ES 2 323 275 T3 where
<img file="ES2323275T3_D0006.tif" />
and P ^ (d, k) is a time estimate of the mean of x<sub>1</sub>(kd<sub>1</sub>) x<sub>2</sub>(kd<sub>2</sub>). o ICLD [dB]:
A £,<sub>2</sub>(?) = 101og<sub>10</sub>
<img file="ES2323275T3_D0007.tif" />
Λ (10) or ICC:
c<sub>]2</sub> (k) = max | Φ<sub>Ι2</sub> (d, k) \ (11)
Note that the absolute value of the normalized cross-correlation is considered and c<sub>12</sub>(k) has an interval of
[0,1].
ICTD, ICLD and ICC estimation for multichannel audio signals
When there are more than two input channels, it is normally sufficient to define ICTD and ICLD between a reference channel (for example, channel number 1) and the other channels, as illustrated in figure 6 for the case of C = 5 channels , in which r<sub>1 C</sub>(k) and AL<sub>Í2</sub> (k) denote ICTD and ICLD, respectively, between reference channel 1 and channel c.
In contrast to ICTD and ICLD, ICC typically has more degrees of freedom. The ICC as defined can have different values among all possible input channel pairs. For C channels, there are C (C-1) / 2 possible channel pairs; for example for 5 channels there are 10 pairs of channels as illustrated in figure 7 (a). However, such a scheme requires that, for each subband at each time index, the ICC values of C (C-1) / 2 are estimated and transmitted, resulting in high computational complexity and high bit rate.
Alternatively, for each subband, ICTD and ICLD determine the direction in which the auditory event of the corresponding signal component is provided in the subband. A single ICC parameter per subband can therefore be used to describe the overall coherence between all audio channels. Good results can be obtained by estimating and transmitting ICC indications only between the two channels with the highest energy in each subband at each time index. This is illustrated in Figure 7 (b), in which for time instants k-1 and k, the channel pairs (3,4) and (1,2) are the strongest, respectively. A heuristic rule can be used to determine ICC between the other channel pairs.
Synthesis of spatial indications
Figure 8 shows a block diagram of an implementation of the 400 BCC synthesizer of Figure 4 that can be used in a BCC decoder to generate a stereo or multichannel audio signal given an individual transmitted sum signal s (n) plus spatial indications. . The sum signal s (n) is decomposed into subbands, where s (k) denotes one such subband. To generate the corresponding subbands of each of the output channels, delays d are applied<sub>c</sub>, scaling factors a<sub>c</sub>, and filters h<sub>c</sub> to the corresponding subband of the sum signal. (For simplicity of notation, the time index k is ignored in delays, scaling factors, and filters.) ICTDs are synthesized by imposing lags, ICLD by scaling, and ICC by applying decorrelation filters. The processing shown in Figure 8 is applied independently to each subband.
ES 2 323 275 T3
ICTD synthesis
The delays d<sub>c</sub> are determined from the ICTD T<sub>1 C</sub>(k) according to equation (12) as follows:
<img file="ES2323275T3_D0008.tif" />
The delay for the reference channel d<sub>1</sub> is calculated in such a way that the maximum magnitude of the delays d<sub>c</sub> is minimized. The less the subband signals are modified, the less danger there is of artifacts. If the subband sampling rate does not provide high enough time resolution for ICTD synthesis, delays can be imposed more precisely using appropriate all-pass filters.
ICLD synthesis
In order for the output subband signals to have desired ICLDs AL<sub>12</sub>(k) between channel c and reference channel 1, the gain factors a<sub>c</sub> must satisfy equation (13) as follows:
¿MU ík <sub>=</sub> io 20 <sup>to</sup>' ,(13)
Additionally, the output subbands are preferably normalized, such that the sum of the power of all the output channels is equal to the power of the input sum signal. Since the total original signal power in each subband is conserved in the summation signal, this normalization results in the absolute subband power for each output channel approaching the corresponding power of the encoder input audio signal. original. Given these restrictions, the scaling factors ac are given by equation (14) as follows:
<img file="ES2323275T3_D0009.tif" />
ICC synthesis
In certain embodiments, the goal of ICC synthesis is to reduce the correlation between subbands after delays and scaling have been applied, without affecting ICTD and ICLD. This can be achieved by designing the filters h<sub>c</sub> in Figure 8 such that ICTD and ICLD are effectively modified as a function of frequency such that the average variation is zero in each subband (auditory critical band).
Figure 9 illustrates how ICTD and ICLD are modified within a subband as a function of frequency. The amplitude of the ICTD and ICLD modification determines the degree of decorrelation and is controlled for ICC. Note that ICTD is smoothly modified (as in Figure 9 (a)), while ICLD is randomly modified (as in Figure 9 (b)). They could be modified ICLD as smoothly as ICTD, but this would result in more coloration of the resulting audio signals.
Another method of synthesizing ICC, particularly suitable for multichannel ICC synthesis, is described in more detail in C. Faller, "Parametric multi-channel audio coding: Synthesis of coherence cues", IEEE Trans. on Speech and Audio Proc., 2003.
Based on time and frequency, specific amounts of artificial late reverb are added to each of the output channels to obtain a desired ICC. Additionally, spectral modification can be applied such that the spectral envelope of the resulting signal approximates the spectral envelope of the original audio signal.
Other related and unrelated ICC synthesis techniques for stereo signals (or audio channel pairs) have been presented in E. Schuijers, W. Oomen, B. den Brinker, and J. Breebaart, “Advances in parametric coding for high quality audio ”, In Preprint 114<sup>th</sup> Conv. Aud. Eng. Soc., March 2003 and J. Engdegard, H. Purnhagen, J. Roden, and L. Liljeryd, "Synthetic ambience in parametric stereo coding," in Preprint 117<sup>th</sup> Conv. Aud. Eng. Soc., May 2004.
ES 2 323 275 T3
BCC from C to E
As described above, BCC can be implemented with more than one transmission channel. A variation of BCC has been described that represents C channels of audio not as a single (transmitted) channel, but as E channels, denoted BCC from C to E. There are (at least) two motivations for BCC from C to E:
◦ BCC with one broadcast channel provides a compatible backward path to upgrade existing mono systems for stereo or multichannel audio playback. Upgraded systems transmit the BCC downmixed sum signal over the existing mono infrastructure, while additionally transmitting the BCC side information. BCC from C to E is applicable to channel C audio backward compatible encoding.
◦ BCC from C to E introduces scalability in terms of different degrees of reduction in the number of transmitted channels. It is expected that the more channels of audio that are transmitted, the better the audio quality.
Signal processing details for BCCs from C to E, such as how to define ICTD, ICLD and ICC indications, are described in US 2005/0157883.
Individual channel shaping
In certain embodiments, both BCC with a transmission channel and BCC from C to E involve algorithms for the synthesis of ICTD, ICLD, and / or ICC. Usually, it is sufficient to synthesize the ICTD, ICLD, and / or ICC indications approximately every 4 to 30 ms. However, the perceptual phenomenon of precedence effect implies that there are specific instants in time when the human auditory system evaluates cues at a higher time resolution (eg, every 1 to 10 ms).
A single static filter bank cannot commonly provide sufficiently high frequency resolution appropriate for most instants in time, while providing sufficiently high time resolution at instants in time when the precedence effect becomes effective.
Certain embodiments of the present invention are directed to a system that uses relatively low time resolution ICTD, ICLD, and / or ICC synthesis, while adding additional processing to address time points when time resolution is required. highest. Additionally, in certain embodiments, the system eliminates the need for signal adaptive windowing technology that is usually difficult to integrate into the fabric of a system. In certain embodiments, the temporal envelopes of one or more of the input audio channels of the original encoder are estimated. This can be done, for example, directly by analyzing the time structure of the signal or examining the autocorrelation of the signal spectrum with respect to frequency. Both approaches will be further developed in subsequent implementation examples. The information contained in these envelopes is transmitted to the decoder (as envelope indication codes) if it is perceptually required and advantageous.
In certain embodiments, the decoder applies some processing to impose these desired temporary envelopes on its output audio channels:
OR This can be achieved by TP processing, eg manipulation of the signal envelope by multiplying the samples in the time domain of the signal with a time varying amplitude modifying function. Similar processing can be applied to spectral / subband samples if the time resolution of the subbands is high enough (at the cost of coarse frequency resolution).
Or Alternatively, a convolution / filtering of the spectral representation of the signal with respect to frequency can be used in a manner analogous to that used in the prior art in order to shape the quantization noise of a low bit rate audio encoder. bits or to enhance intensity stereo coded signals. This is preferred if the filter bank has a high frequency resolution and therefore a rather low time resolution. For the convolution / filtration approach:
O The envelope shaping method extends from intensity stereo to multi-channel CaE coding.
O The technique comprises an adjustment in which the envelope formation is controlled by parametric information (eg, binary flags) generated by the encoder, but is actually carried out using sets of decoder-derived filter coefficients.
OR In another setting, sets of filter coefficients are transmitted from the encoder, for example only when perceptually necessary and / or beneficial.
The same is also true for the time domain / subband domain approach. Accordingly, criteria (eg, transient detection and an estimated tonality value) can be entered to further control the transmission of envelope information.
ES 2 323 275 T3
There may be situations where it is beneficial to disable TP processing in order to avoid potential artifacts. Just in case, it is a good strategy to leave temporary processing disabled by default (that is, BCC would operate according to a conventional BCC scheme). Additional processing is enabled only when higher channel time resolution is expected to produce improvement, for example, when the precedence effect is expected to become active.
As noted above, this enable / disable control can be accomplished by transient detection. That is, if a transient is detected, then TP processing is enabled. The precedence effect is the most effective for transients. Transient detection can be used in advance to efficiently shape not only individual transients but also the signal components shortly before and after the transient. Possible ways to detect transients include:
O Observe the temporal envelope of the BCC encoder input signals or transmitted BCC sum signal (s). If there is a sudden increase in energy, then a transient has occurred.
O Examine the linear predictive coding (LPC) gain as estimated at the encoder or decoder. If the LPC prediction gain exceeds a certain threshold, then the signal can be assumed to be transient or highly fluctuating. The LPC analysis is calculated on the autocorrelation of the spectrum.
Additionally, to prevent possible artifacts in the tonal signals, TP processing is preferably not applied when the tonality of the transmitted sum signal (s) is high.
According to certain embodiments of the present invention, the temporal envelopes of the individual original audio channels are estimated in a BCC encoder in order to enable a BCC decoder to generate output channels with temporal envelopes similar (or perceptually similar) to those of the original audio channels. Certain embodiments of the present invention address the phenomenon of the precedence effect. Certain embodiments of the present invention involve the transmission of envelope indication codes in addition to the other BCC codes such as ICLD, ICTD and / or ICC, as part of the BCC side information.
In certain embodiments of the present invention, the time resolution for time envelope indications is finer than the time resolution of other BCC codes (eg, ICLD, ICTD, ICC). This allows the envelope shaping to be performed within the time period provided by a synthesis window that corresponds to the length of a block of an input channel for which the other BCC codes are derived.
Implementation examples
Figure 10 shows a block diagram of the time domain processing that is added to a BCC encoder, such as the encoder 202 of Figure 2, in accordance with one embodiment of the present invention. As shown in Figure 10 (a), each temporal processing analyzer (TPA) 1002 estimates the temporal envelope of an original input channel x<sub>c</sub>(n) different, although in general any of one or more of the input channels can be analyzed.
Figure 10 (b) shows a block diagram of a possible time domain based implementation of TPA 1002 in which the input signal samples are squared (1006) and then low pass filtered (1008) to characterize the temporal envelope of the input signal. In alternative embodiments, the time envelope can be estimated using an autocorrelation / LPC method or with other methods, eg, using a Hilbert transform.
Block 1004 of Figure 10 (a) parameterizes, quantizes, and encodes the estimated temporal envelopes prior to their transmission as temporal processing information (TP) (i.e., envelope indication codes) that is included in the side information of the figure 2.
In one embodiment, a detector (not shown) within block 1004 determines whether TP processing in the decoder will improve audio quality, such that block 1004 transmits TP side information only during those instants of time when the Audio quality will improve through TP processing.
Figure 11 illustrates an exemplary time domain application of TP processing in the context of the BCC synthesizer (400) of Figure 4. In this embodiment, there is a single transmitted sum signal s (n), C signals base are generated by replicating that summation signal and the envelope shaping is applied individually to different synthesized channels. In alternative embodiments, the order of delays, scaling, and other processing may be different. Furthermore, in alternative embodiments, the envelope shaping is not restricted to the processing of each channel independently. This is especially true for convolution / filtering-based implementations that take advantage of coherence over frequency bands to derive information regarding the fine temporal structure of the signal.
ES 2 323 275 T3
In FIG. 11 (a), decoding block 1102 recovers time envelope signals a for each output channel from the transmitted TP side information received from the BCC encoder; and each TP block 1104 applies the corresponding envelope information to shape the envelope of the output channel.
Figure 11 (b) shows a block diagram of a possible time-domain-based implementation of TP 1104 in which the synthesized signal samples are squared (1106) and then low-pass filtered (1108) to characterize the time envelope b of the synthesized channel. A scaling factor (for example, sqrt (a / b)) is generated (1110) and then (1112) is applied to the synthesized channel to generate an output channel that has a time envelope substantially equal to that of the input channel. corresponding original entry.
In alternative implementations 1002 of TPA of Figure 10 and TP 1104 of Figure 11, the temporal envelopes are characterized using magnitude operations rather than squaring the signal samples. In such implementations, the ratio a / b can be used as the scaling factor without having to apply the square root operation.
Although the scaling operation of Figure 11 (c) corresponds to a time domain-based implementation of TP processing, TP processing (as well as inverse TP processing (ITP) and TPA) can also be implemented using signals in the frequency domain, as in the embodiment of Figures 16-17 (described later herein). As such, for the purposes of this specification, the term "scaling function" is to be construed as encompassing either time domain operations or frequency domain operations, such as filtering operations of the Figures 17 (b) and (c).
In general, each TP 1104 is preferably designed in such a way that it does not modify the signal strength (that is, energy). Depending on the particular implementation, this signal power may be a short-time average signal power on each channel, for example based on the total signal power per channel in the time period defined by the synthesis window or some other measure. of appropriate power. As such, scaling for ICLD synthesis (eg, using multipliers 408) can be applied before or after envelope shaping.
Since full-band scaling of the BCC output signals can result in artifacts, the envelope shaping could only be applied at specified frequencies, for example frequencies greater than a certain cutoff frequency.<sub>TP</sub> (for example, 500 Hz). Note that the frequency range for analysis (TPA) may differ from the frequency range for synthesis (TP).
Figures 12 (a) and (b) show possible implementations of TPA 1002 of Figure 10 and TP 1104 of Figure 11 in which envelope shaping is applied only at frequencies higher than the cutoff frequency fTP. In particular, Figure 12 (a) shows the addition of the high pass filter 1202, which filters frequencies lower than fTP before the time envelope characterization. Figure 12 (b) shows the addition of two-band filter bank 1204 having a cutoff frequency f<sub>TP</sub> between the two subbands, in which only the high frequency part is temporarily shaped. The two-band reverse filter bank 1206 then recombines the low-frequency portion with the temporarily-shaped high-frequency portion to generate the output channel.
Figure 13 shows a block diagram of the frequency domain processing that is added to a BCC encoder, such as the encoder 202 of Figure 2, in accordance with an alternative embodiment of the present invention. As shown in Figure 13 (a), the processing of each TPA 1302 is applied individually in a different subband, in which each filter bank (FB) is the same as the corresponding FB 302 of Figure 3 and the block 1304 is a subband implementation analogous to block 1004 of FIG. 10. In alternative implementations, the subbands for TPA processing may differ from the BCC subbands. As shown in Figure 13 (b), the TPA 1302 can be implemented analogously to the TPA 1002 of Figure 10.
Figure 14 illustrates an exemplary frequency domain application of TP processing in the context of the 400 BCC synthesizer of Figure 4. Decoding block 1402 is analogous to decoding block 1102 of Figure 11, and each TP 1404 is a subband implementation analogous to each TP 1104 of Figure 11, as shown in Figure 14 (b).
Figure 15 shows a block diagram of the frequency domain processing that is added to a BCC encoder, such as the encoder 202 of Figure 2, in accordance with another alternative embodiment of the present invention. This scheme has the following setting: The envelope information for each input channel is derived by LPC calculation through the frequency (1502), parameterized (1504), quantized (1506) and encoded in the bit stream (1508) by the encoder. Figure 17 (a) illustrates an example implementation of the TPA 1502 of Figure 15. The side information to be transmitted to the multichannel synthesizer (decoder) could be the LPC filter coefficients calculated by an autocorrelation method, the resulting reflection coefficients or line spectrum pairs, etc., in order to maintain the rate data transmission of small side information, parameters derived from, for example, the LPC prediction gain as "transients present / not present" binary flags.
ES 2 323 275 T3
Figure 16 illustrates another exemplary frequency domain application of TP processing in the context of the 400 BCC synthesizer of Figure 4. The encoding processing of Figure 15 and decoder processing of Figure 16 may be implemented to form a corresponding pair of an encoder / decoder configuration. Decoding block 1602 is analogous to decoding block 1402 of Figure 14, and each TP 1604 is analogous to each TP 1404 of Figure 14. In this multichannel synthesizer, the transmitted TP side information is decoded and used to control the envelope shaping of individual channels. However, in addition, the synthesizer includes an Envelope Characterizer Stage (TPA) 1606 for the analysis of transmitted summation signals, an Inverse TP (ITP) 1608 to "flatten" the temporal envelope of each base signal, in which the envelope adjusters (TP) 1604 impose a modified envelope on each output channel. Depending on the particular implementation, ITP can be applied either before or after the upmix. In detail, this is done using the convolution / filtering approach in which the envelope shaping is obtained by applying LPC-based filters over the spectrum through frequency as illustrated in Figures 17 (a), (b ) and (c) for TPA, ITP and TP processing, respectively. In FIG. 16, control block 1610 determines whether or not envelope shaping is to be implemented and, if so, whether it will be based on (1) transmitted TP side information or (2) characterized envelope data. locally from TPA 1606.
Figures 18 (a) and (b) illustrate two exemplary modes of operation of the control block 1610 of Figure 16. In the implementation of Figure 18 (a), a set of filter coefficients is transmitted to the decoder and the shaping Envelope by convolution / filtering is done based on the transmitted coefficients. If the encoder detects that the shaping of transients is not beneficial, then no filter data is sent and the filters are disabled (shown in Fig. 18 (a) by switching to a set of unit filter coefficients "[1,0. ..] ”).
In the implementation of Figure 18 (b), only one "transient / non-transient flag" is transmitted for each channel and this flag is used to turn shaping on or off based on the sets of filter coefficients calculated from the downmix signals transmitted in the decoder.
Additional alternative embodiments
Although the present invention has been described in the context of BCC coding schemes in which there is a single summation signal, the present invention can also be implemented in the context of BCC coding schemes that have two or more summation signals. In this case, the temporal envelope for each different "base" summation signal can be estimated before the application of BCC synthesis, and different BCC output channels can be generated based on different temporal envelopes, depending on which summation signals were used for synthesize the different output channels. An output channel that is synthesized from two or more different summing channels could be generated based on an effective time envelope that takes into account (eg, by weighted averaging) the relative effects of the constituent summing channels.
Although the present invention has been described in the context of BCC coding schemes involving ICTD, ICLD, and ICC codes, the present invention can also be implemented in the context of other BCC coding schemes involving only one or two of these three types. of codes (for example, ICLD and ICC, but not ICTD) and / or one or more types of additional codes. Furthermore, the BCC synthesis processing sequence and envelope conformation may vary in different implementations. For example, when envelope shaping is applied to signals in the frequency domain, as in Figures 14 and 16, envelope shaping could alternatively be implemented after ICTD synthesis (in those embodiments employing ICTD synthesis), but before of ICLD synthesis. In other embodiments, the envelope shaping could be applied to up-mixed signals before any other BCC synthesis is applied.
Although the present invention has been described in the context of BCC encoders that generate envelope indication codes from the original input channels, in alternative embodiments, the envelope indication codes could be generated from downmixed channels corresponding to the original input channels. This would allow the implementation of a processor (for example, a separate envelope indication encoder) that could (1) accept the output of a BCC encoder that generates the downmixed channels and certain BCC codes (for example, ICLD, ICTD and / or ICC) and (2) characterize the temporal envelope (s) of one or more of the downmixed channels to add envelope indication codes to the BCC side information.
Although the present invention has been described in the context of BCC coding schemes in which the envelope indication codes are transmitted with one or more audio channels (i.e., the E channels transmitted) along with other BCC codes, in embodiments Alternatively, the envelope indication codes could be transmitted, either alone or with other BCC codes, to a location (e.g., a decoder or a storage device) that already has the transmitted channels and possibly other BCC codes.
Although the present invention has been described in the context of BCC coding schemes, the present invention can also be implemented in the context of other audio processing systems in which audio signals are decorrelated or other audio processing that needs to decorrelate signals. .
ES 2 323 275 T3
Although the present invention has been described in the context of implementations in which the encoder receives the input audio signal in the time domain and generates the audio signals transmitted in the time domain and the decoder receives the audio signals transmitted in the time domain. the time domain and generates playback audio signals in the time domain, the present invention is not limited in this way. For example, in other implementations, any one or more of the input, transmitted, and playback audio signals could be represented in a frequency domain.
BCC encoders and / or decoders can be used in conjunction with or incorporated into a variety of different applications or systems, including systems for television or electronic music distribution, cinemas, broadcast, streaming, and / or reception. These include systems for encoding / decoding transmissions via, for example, terrestrial, satellite, cable, internet, intranet, or physical media (for example, compact discs, digital versatile disks, semiconductor chips, hard drives, memory cards and the like). BCC encoders and / or decoders may also be used in games and gaming systems, including, for example, interactive software products designed to interact with a user for entertainment (action, role-playing games, strategy, adventure, simulations, racing , sports, recreational, card and board games) and / or educational that can be published for multiple machines, platforms or media. In addition, BCC encoders and / or decoders can be incorporated into audio recorders / players or CD-ROM / DVD systems. BCC encoders and / or decoders can also be incorporated into PC software applications incorporating digital decoding (eg, player, decoder) and software applications incorporating digital encoding capabilities (eg, encoder, ripper ("ripper"), recoder and music managers).
The present invention may be implemented as circuit-based processes, including possible implementations such as a single integrated circuit (such as an ASIC or FPGA), a multi-chip module, a single card, or a multi-card circuit pack. . As will be apparent to the person skilled in the art, various functions of the circuit elements can also be implemented as processing steps in a software program. Such software can be used for example in a digital signal processor, microcontroller or general purpose computer.
The present invention can be carried out in the form of methods and apparatus for practicing these methods. The present invention can also be realized in the form of program code implemented on tangible media, such as floppy disks, CD-ROMs, hard drives or any other machine-readable storage medium, in which, when the program code is loaded in and run by a machine, such as a computer, the machine becomes an apparatus for practicing the invention. The present invention can also be realized in the form of a program code, for example, either stored on a storage medium, loaded into and / or executed by a machine, or transmitted by some transmission medium or carrier, such as lines or electrical wiring, by means of optical fibers or through electromagnetic radiation, in which, when the program code is loaded into and executed by a machine, such as a computer, the machine becomes an apparatus for practicing the invention. When implemented in a general purpose or multipurpose processor, the program code segments combine with the processor to provide a single device that operates in a manner analogous to specific logic circuits.
It will be further understood that those skilled in the art may make various changes to the details, materials, and arrangements of the parts that have been described and illustrated in order to explain the nature of this invention, without departing from the scope of the invention as stated. expressed in the following claims.
Although the steps in the following method claims, if any, are cited in a particular sequence with corresponding labeling, unless the mentions in the claims otherwise imply a particular sequence to implement some or all of these steps, it is not those steps are necessarily envisaged to be limited to being implemented in that particular sequence.
Contents11
24 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24
35 members in 21 offices
Priority claims10
| Document | Office | Kind | Date |
|---|---|---|---|
| 20040620480P | United States of America | – | |
| 62048004 | United States of America | P | |
| 62048004 | United States of America | P | |
| 20040006482 | United States of America | – | |
| 648204 | United States of America | A | |
| 648204 | United States of America | A | |
| 620480P05792350 | – | – | – |
| 6482 | – | – | – |
| US20040006482 | – | – | – |
| US20040620480P | – | – | – |
Members35
| Document | Office | Kind | |
|---|---|---|---|
| US2006083385A1 | United States of America | A1 | |
| AU2005299068A1 | Australia | A1 | |
| CA2582485A1 | Canada | A1 | |
| WO2006045371A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW200628001A | Taiwan Province of China | A | |
| NO20071493L | Norway | L | |
| KR20070061872A | Republic of Korea | A | |
| EP1803117A1 | European Patent Office (EPO) | A1 | |
| MX2007004726A | Mexico | A | |
| IL182236A0 | Israel | A0 | |
| CN101044551A | China | A | |
| HK1106861A1 | Hong Kong, China | A1 | |
| JP2008517333A | Japan | A | |
| BRPI0516405A | Brazil | A | |
| BRPI0516405A | Brazil | A | |
| AU2005299068B2 | Australia | B2 | |
| RU2339088C1 | Russian Federation | C1 | |
| EP1803117B1 | European Patent Office (EPO) | B1 | |
| AT424606T | Austria | T | |
| ATE424606T1 | Austria | T1 | |
| DE602005013103D1 | Germany | D1 | |
| PT1803117E | Portugal | E | |
| DK1803117T3 | Denmark | T3 | |
| ES2323275T3This record | Spain | T3 | |
| PL1803117T3 | Poland | T3 | |
| KR100924576B1 | Republic of Korea | B1 | |
| TWI318079B | Taiwan Province of China | B | |
| US7720230B2 | United States of America | B2 | |
| JP4664371B2 | Japan | B2 | |
| IL182236A | Israel | A | |
| CN101044551B | China | B | |
| CA2582485C | Canada | C | |
| NO338919B1 | Norway | B1 | |
| BRPI0516405A8 | Brazil | A8 | |
| BRPI0516405B1 | Brazil | B1 |
Numbers
- Publication
- 2323275
- Publication, DOCDB
- 2323275
- Publication, EPODOC
- ES2323275T
- Application
- 5792350
- Application, DOCDB
- 05792350
- Application, EPODOC
- ES20050792350T
Titles2
- Spanish
- CONFORMACION DE ENVOLVENTE TEMPORAL DE CANAL INDIVIDUAL PARA ESQUEMAS DE CODIFICACION DE INDICACION BINAURAL Y SIMILARES.
- English
- INDIVIDUAL CHANNEL TEMPORARY ENVELOPE CONFORMATION FOR BINAURAL AND SIMILAR INDICATION CODING SCHEMES.
Classification
- CPC, 3
- G10L19/008
- G10L19/02
- H03M7/30
- IPC, 2
- G10L19 02
- G10L19 00