Individual channel temporal envelope shaping for binaural cue coding schemes and the like.
Abstract
At an audio encoder, cue codes are generated for one or more audio channels, wherein an envelope cue code is generated by characterizing a temporal envelope in an audio channel. At an audio decoder, E transmitted audio channel(s) are decoded to generate C playback audio channels, where C>=Eo1. Received cue codes include an envelope cue code corresponding to a characterized temporal envelope of an audio channel corresponding to the transmitted channel(s). One or more transmitted channel(s) are upmixed to generate one or more upmixed channels. One or more playback channels are synthesized by applying the cue codes to the one or more upmixedchannels, wherein the envelope cue code is applied to an upmixed channel or a synthesized signal to adjust a temporal envelope of the synthesized signal based on the characterized temporal envelope such that the adjusted temporal envelope substantially matches the characterized temporal envelope.

Term
Term ended
Expired 7 September 2025, 1 year ago.
- Priority
- Filed
- Granted
- Expired
- Today
36 claims: 12 independent, 24 dependent
- 1REIVINDICACIONES 1. Un método para la codificación de canales de audio, el método está caracterizado porque comprende:generar uno o más códigos de indicación para uno o 5 más canales de audio, en donde por lo menos un código de indicación es un código de indicación de envolvente generado por la caracterización de una envolvente temporal en uno de los uno o más canales de audio, en donde el uno o más códigos de indicación comprenden además uno o más de códigos de 10 correlación de intercanal (ICC), código de diferencia de nivel intercanal (ICLD), y códigos de diferencia de tiempo intercanal (ICTD), en donde una primera resolución de tiempo asociada con el código de indicación de envolvente es más fina que una segunda resolución en el tiempo asociada con los otros códigos 15 de indicación y en donde la envolvente temporal es caracterizada por el canal de audio correspondiente en un dominio de tiempo o individualmente para diferentes sub-bandas de señal del canales de audio correspondiente en un dominio de sub-banda;y 20 transmitir el uno o más códigos de indicación.
- 2El método de conformidad con la reivindicación 1, caracterizado porgue comprende además transmitir E canal(es) de audio transmitidos correspondientes al uno o más canales de audio, en donde E 1. 25
- 3El método de conformidad con la reivindicación 2, caracterizado porque:el uno o más canales de audio comprenden C canales de audio de entrada, en donde OE;y los C canales de entrada son mezclados descendentemente para generar el (los) E canal(es) transmitidos.
- 4El método de conformidad con la reivindicación 1, caracterizado porque el uno o más códigos de indicación son transmitidos para permitir que un descodificador efectúe la formación de envolvente durante la descodificación de (los) E canal(es) transmitidos en base al uno o más códigos de indicación, en donde el (los) E canal(es) de audio transmitidos corresponden al uno o más canales de audio, en donde E 1.
- 5El método de conformidad con la reivindicación 4, caracterizado porque la formación de envolvente ajusta una envolvente temporal de una señal sintetizada generada por el descodificador para coincidir sustancialmente con la envolvente temporal caracterizada.
- 6El método de conformidad con la reivindicación 1, caracterizado porque la envolvente temporal es caracterizada solamente para frecuencias especificadas del canal de audio correspondiente.
- 7El método de conformidad con la reivindicación
- 88, caracterizado porque la envolvente temporal es caracterizada solamente para frecuencias del canal de audio correspondiente por encima de una frecuencia de corte especificada. 8. El método de conformidad con la reivindicación 10, caracterizado porque el dominio de sub-banda corresponde a un filtro de espejo de cuadratura (QMF). 5 9. El método de conformidad con la reivindicación 1, caracterizado porque comprende además determinar si se habilita o deshabilita la caracterización. 10. El método de conformidad con la reivindicación 9, caracterizado porque comprende además generar y transmitir 10 una bandera de habilitación/deshabilitación en base a la determinación de instruir a un descodificador si implementar o no la formación de envolvente durante la descodificación del (los) E canal (es) transmitidos correspondiente al uno o más canales de audio, en donde E 1. 15 11. El método de conformidad con la reivindicación
- 99, caracterizado porque la determinación está basada en análisis de un canal de audio para detectar transitorios en el canal de audio de tal manera que la caracterización es habilitada si se detecta la presencia de un transitorio. 20 12. El método de conformidad con la reivindicación 1, caracterizado porque la etapa de generación del código de indicación de envolvente incluye elevar al cuadrado o formar una magnitud y filtración de paso de bajos de muestras de señal del canal de audio o de las señales de sub-banda del canal de 25 audio con el fin de caracterizar la envolvente temporal. 13. El método de conformidad con la reivindicación 1 o 12, caracterizado porque la etapa de generación comprende además la etapa de parametrización, cuantificación y codificación de una envolvente temporal estimada. 14 . Un aparato para la codificación de canales de audio, el aparato está caracterizado porque comprende:medios para generar uno o más códigos de indicación para uno o más canales de audio, en donde porque por lo menos un código de indicación es un código de indicación de envolvente generado mediante la caracterización de una envolvente temporal en uno de los uno o más canales de audio, en donde los uno o más códigos de indicación comprende además uno o más de códigos de correlación de intercanal (ICC), códigos de diferencia de nivel de intercanal (ICLD), y códigos de diferencia de tiempo de intercanal (ICTD), en donde una primera resolución de tiempo asociada con el código de indicación de envolvente es más fina que una segunda resolución en el tiempo asociada con el (los) otros códigos de indicación y en donde la envolvente temporal es caracterizada para el canal de audio correspondiente en un dominio de tiempo o individualmente para diferentes sub-bandas de señal del canal de audio correspondiente en un dominio de sub-banda;y medios para transmisión del uno o más códigos de indicación. 15. Un aparato para la codificación de C de audio de entrada para generar E canal(es) de audio transmitidos, el aparato está caracterizado porque comprende: un analizador de envolvente para caracterizar una envolvente temporal de entrada de por lo menos uno de los C 5 canales de entrada;un estimador de código adaptado para generar códigos de indicación para dos o más de los C canales de entrada, en donde los uno o más códigos de indicación comprenden además uno o más de códigos de correlación de intercanal (ICC), código de
- 1010 diferencia de nivel de intercanal (ICLD), y códigos de diferencia de tiempo de intercanal (ICTD), en donde una primera resolución en el tiempo asociada con el código de indicación de envolvente es más fina que una segunda resolución en el tiempo asociada con el (los) otro(s) código(s) de indicación y en
- 1115 donde la envolvente temporal es caracterizada para el canal de audio correspondiente en un dominio de tiempo o individualmente para diferentes sub-bandas de señal del canal de audio correspondiente en un dominio de sub-banda, y un mezclador descendente adaptado para mezclar 20 descendentemente los C canales de entrada para generar el (los) E canal(es) transmitidos, en donde OE 1, en donde el aparato está adaptado para transmitir información acerca de los códigos de indicación y la envolvente temporal de entrada caracterizada para permitir que un descodificador efectúe la síntesis y 25 formación de envolvente durante la descodificación del (los) E canal(es) transmitidos.
- 1216. El aparato de conformidad con la reivindicación 15, caracterizado porque:el aparato es un sistema seleccionado del grupo que 5 consiste de una grabadora de video digital, una grabadora de audio digital, una computadora, un transmisor de satélite, un transmisor de cable, un transmisor de difusión terrestre, un sistema de entretenimiento en casa y un sistema de cine, y el sistema comprende el analizador de envolvente, ;10 estimador de código y mezclador descendente.
- 1317. Un medio que se puede leer por la máquina que tiene codificado en el mismo código de programas, caracterizado porque, cuando el código de programa es ejecutado por la máquina, la máquina implementa el método de conformidad con la 15 reivindicación 1.
- 1418. Una corriente de bits de audio codificada, caracterizada porque tiene:uno o más códigos de indicación generados para uno o más canales de audio, en donde por lo menos un código de 20 indicación es un código de indicación de envolvente generado mediante la caracterización de una envolvente temporal en uno de los uno o más canales de audio, en donde porque el uno o más códigos de indicación comprenden además uno o más de códigos de correlación de intercanal (ICC), código de diferencia de nivel 25 de intercanal (ICLD), y códigos de diferencia de tiempo de intercanal (ICTD), en donde una primera resolución en el tiempo asociada con el código de indicación de envolvente es más fina que una segunda resolución en el tiempo asociada con el (los) otro(s) código(s) de indicación y en donde la envolvente temporal es caracterizada para el canal de audio correspondiente en un dominio de tiempo o individualmente para diferentes sub-bandas de señal del canal de audio correspondiente en un dominio de sub-banda, y + el uno o más códigos de indicación y E canal (es) de audio transmitidos corresponden al uno o más canales de audio, en donde E 1, son codificados a la corriente de bits de audio codificada.
- 1519. Una corriente de bits de audio codificada, que comprende uno o más códigos de indicación y E canal (es) de audio transmitidos, caracterizada porque:el uno o más códigos de indicación son generados para uno o más canales de audio, en donde por lo menos un código de indicación es un código de indicación de envolvente generado mediante la caracterización de una envolvente temporal en uno de los uno o más canales de audio, en donde el uno o más códigos de indicación comprenden además uno o más códigos de correlación de intercanal (ICC), código de diferencia de nivel de intercanal (ICLD), y códigos de diferencia de tiempo de intercanal (ICTD), en donde una primera resolución en el tiempo asociada con el código de indicación de envolvente es más fina que una segunda resolución en el tiempo asociada con el (los) otro(s) código(s) de indicación, y en donde la envolvente temporal está caracterizada para el canal de audio correspondiente en un dominio de tiempo o individualmente para 5 diferentes sub-bandas de señal del canal de audio correspondiente en un dominio de sub-banda;y el (los) E canal (es) de audio transmitidos corresponden al uno o más canales de audio.
- 1620. Un método para la descodificación de E canal(es) 10 de audio transmitidos, para generar C canales de audio de reproducción, en donde C .E 1, el método está caracterizado porgue comprende:recibir códigos de indicación correspondiente al (los) E canal(es) transmitidos, en donde los códigos de 15 indicación comprenden un código de indicación de envolvente correspondiente a una envolvente temporal caracterizada de un canal de audio correspondiente al (los) E canal(es) transmitidos, en donde el uno o más códigos de indicación comprenden además uno o más de códigos de correlación de 20 intercanal (ICC), códigos de diferencia de nivel de intercanal (ICLD), y códigos de diferencia de tiempo de intercanal (ICTD), en donde una primera resolución de tiempo asociada con el código de indicación de envolvente es más fina gue una segunda resolución en el tiempo asociada con el (los) otro(s) código(s) 25 de indicación;mezclar ascendentemente uno o más del (los) E canal(es) transmitidos para generar uno o más canales mezclados ascendentemente;y sintetizar uno o más de los C canales de reproducción 5 mediante la aplicación de los códigos de indicación a uno o más canales mezclados ascendentemente, en donde el código de indicación de envolvente es aplicado a un canal mezclado ascendentemente o una señal sintetizada para ajustar una envolvente temporal de la señal sintetizada en base a la 10 envolvente temporal caracterizada mediante escalamiento de dominio de tiempo o muestras de señal de dominio de sub-banda utilizando un factor de escalamiento, de tal manera que la envolvente temporal ajustada coincide sustancialmente con la envolvente temporal caracterizada. 15 21. El método de conformidad con la reivindicación 20 ++++, caracterizado porque el código de indicación de envolvente corresponde a una envolvente temporal caracterizada en un canal de entrada original usado para generar el (los) E canal(es) transmitidos. 20 22. El método de conformidad con la reivindicación 20, caracterizado porque la síntesis comprende síntesis de ICC de reverberación tardía. 23. El método de conformidad con la reivindicación
- 1721, caracterizado porque la envolvente temporal de la señal 25 sintetizada es ajustada antes de la síntesis de ICLD.
- 1824. El método de conformidad con la reivindicación 20, caracterizado porque:la envolvente temporal de la señal sintetizada es caracterizada;y 5 la envolvente temporal de la señal sintetizada es ajustada en base tanto a la envolvente temporal caracterizada correspondiente al código de indicación de envolvente como la envolvente temporal caracterizada de la señal sintetizada.
- 1925. El método de conformidad con la reivindicación 10 24, caracterizado porque:una función de escalamiento es generada en base a la envolvente temporal caracterizada correspondiente al código de indicación de envolvente y la envolvente temporal caracterizada de la señal sintetizada;y 15 la función de escalamiento es aplicada a la señal sintetizada.
- 2026. El método de conformidad con la reivindicación 20, caracterizado porgue comprende además ajustar un canal transmitido en base a la envolvente temporal caracterizada para 20 generar un canal aplanado, en donde la mezcla ascendente y la síntesis son aplicados al canal aplanado para generar un canal de reproducción correspondiente.
- 2127. El método de conformidad con la reivindicación 20, caracterizado porque comprende además ajustar un canal 25 mezclado ascendentemente en base a la envolvente temporal caracterizada para generar un canal aplanado, en donde la síntesis es aplicada al canal aplanado para generar un canal de reproducción correspondiente.
- 2228. El método de conformidad con la reivindicación 5 20, caracterizado porque la envolvente temporal de la señal sintetizada es ajustadas solamente para frecuencias especificadas.
- 2329. El método de conformidad con la reivindicación 28, caracterizado porque la envolvente temporal de la señal 10 sintetizada es ajustada solamente para frecuencias mayores que una frecuencia de corte especificada.
- 2430. El método de conformidad con la reivindicación 20, caracterizado porque las envolventes temporales son ajustadas individualmente para diferentes sub-bandas de señal 15 en la señal sintetizada.
- 2531. El método de conformidad con la reivindicación 20, caracterizado porque un dominio de sub-banda corresponde a un QMF.
- 2632. El método de conformidad con la reivindicación 20 20, caracterizado porque la envolvente temporal de la señal sintetizada es ajustada en un dominio de tiempo.
- 2733. El método de conformidad con la reivindicación 20, caracterizado porque comprende además determinar si se habilita o deshabilita el ajuste de la envolvente temporal de 25 la señal sintetizada.
- 2834. El método de conformidad con la reivindicación 33, caracterizado porque la determinación está basada en una bandera de habilitación/deshabilitación generada por un codificador de audio que generó el (los) E canal (es) transmitidos.
- 2935. El método de conformidad con la reivindicación 33, caracterizado porque la determinación está basada en el análisis del (los) E canal (es) transmitidos para detectar transitorios de tal manera que el ajuste se habilita si se detecta la presencia de un transitorio.
- 3036. El método de conformidad con la reivindicación 20, caracterizado porque comprende además:caracterización de una envolvente temporal de un canal transmitido;y determinar si se usa (1) la envolvente temporal caracterizada correspondiente al código de indicación de envolvente o (2) la envolvente temporal caracterizada del canal transmitido para ajustar la envolvente temporal de la señal sintetizada.
- 3137. El método de conformidad con la reivindicación 20, caracterizado porque la energía dentro de una ventana específica de la señal sintetizada después del ajuste de la envolvente temporal es sustancialmente igual a la energía de una ventana correspondiente de la señal sintetizada antes del ajuste.
- 3238. El método de conformidad con la reivindicación 37, caracterizado porque la ventana especificada corresponde a una ventana de síntesis asociada con uno o más códigos de indicación sin envolvente.
- 3339. Un aparato para descodificar E canal(es) de audio transmitidos para generar C canales de audio de reproducción, en donde C E 1, el aparato está caracterizado porque comprende:medios para recibir códigos de indicación correspondientes al (los) E canal(es) transmitidos, en donde los códigos de indicación comprenden un código de indicación de envolvente correspondiente a una envolvente temporal caracterizada de un canal de audio correspondiente al (los) E canales transmitidos, en donde el uno o más códigos de indicación comprenden además uno o más de códigos de correlación de intercanal (ICC), códigos de diferencia de nivel de intercanal (ICLD), y códigos de diferencia de tiempo de intercanal (ICTD), en donde una primera resolución de tiempo asociada con el código de indicación de envolvente es más fina que una segunda resolución de tiempo asociada con el (los) otro(s) código(s) de indicación;medios para mezclar ascendentemente uno o más de los E canales transmitidos para generar uno o más canales mezclados ascendentemente;y medios para sintetizar uno o más de los C canales de reproducción mediante la aplicación de los códigos de indicación al uno o más canales mezclados ascendentemente, en donde el código de indicación de envolvente es aplicado a un canal mezclado ascendentemente o una señal sintetizada para 5 ajustar una envolvente temporal de la señal sintetizada en base a la envolvente temporal caracterizada mediante el escalamiento de dominio de tiempo o muestras de señal de dominio de subbanda utilizando un factor de escalamiento, de tal manera que la envolvente temporal ajustada corresponde sustancialmente con 10 la envolvente temporal caracterizada.
- 3440. Un aparato para la descodificación de E canal(es) de audio transmitidos, para generar C canales de audio de reproducción, en donde C E 1, el aparato está caracterizado porgue comprende:15 un receptor adaptado para recibir códigos de indicación correspondiente al (los) E canal(es) transmitidos, en donde los códigos de indicación comprenden un código de indicación de envolvente correspondiente a una envolvente temporal caracterizada de un canal de audio correspondiente a 20 los E canales transmitidos, en donde el uno o más códigos de indicación comprenden además uno o más de códigos de correlación de intercanal (ICC), códigos de diferencia de nivel de intercanal (ICLD), y códigos de diferencia de tiempo de intercanal (ICTD), en donde una primera resolución de tiempo 25 asociado con el código de indicación de envolvente es más fina que una segunda resolución en el tiempo asociada con el (los) otro(s) código(s) de indicación;un mezclador ascendente adaptado para mezclar ascendentemente uno o más de los E canales transmitidos para generar uno o más canales mezclados ascendentemente;y un sintetizador adaptado para sintetizar uno o más de los C canales de reproducción mediante la aplicación de los códigos de indicación al uno o más canales mezclados ascendentemente, en donde el código de indicación de envolvente es aplicado a un canal mezclado ascendentemente o una señal sintetizada para ajustar una envolvente temporal de la señal sintetizada en base a la envolvente temporal caracterizada mediante escalamiento de dominio de tiempo o muestras de señal de dominio de banda utilizando un factor de escalamiento de tal manera que la envolvente temporal ajustada coincide sustancialmente con la envolvente temporal caracterizada.
- 3541. El aparato de conformidad con la reivindicación 40, caracterizado porque:el aparato es un sistema seleccionado del grupo que consiste de un reproductor de video digital, un reproductor de audio digital, una computadora, un receptor de satélite, un receptor de cable, un receptor de difusión terrestre, un sistema de entretenimiento en casa, y un sistema de cine;y el sistema comprende el receptor, mezclador ascendente, sintetizador, y ajustador de envolvente.
- 3642. Un medio que se puede leer por la máquina, que tiene codificado en el mismo códigos de programa, caracterizado porque, cuando el código de programa es ejecutado por una máquina, la máquina implementa el método para descodificación 5 de conformidad con la reivindicación 20.
Independent claims36
203 paragraphs in 10 sections, as filed
(54) Title: INDIVIDUAL CHANNEL FORMATION FOR SCHEMES OF BCC AND THE LIKE.
(54) Title: INDIVIDUAL CHANNEL TEMPORAL ENVELOPE SHAPING FOR BINAURAL CUE CODING SCHEMES AND THE LIKE.
(57) Summary
Individual channel formation for BCC schemes and the like is described. In an audio encoder, cue codes are generated for one or more audio channels, where an envelope cue code is generated by characterizing a temporal envelope on an audio channel. In an audio decoder, E channel (s) of audio transmitted with decoded to generate C channels of playback audio, where C> E.1. The received indication codes include an envelope indication code corresponding to a characterized temporal envelope of an audio channel corresponding to the transmitted channel (s). One or more transmitted channel (s) is (are) upmixed to generate one or more upmixed channels. One or more playback channels are synthesized by applying the cue codes to the one or more upmixed channels, wherein the envelope cue code is applied to a downmixed channel or synthesized signal to set a temporal envelope of the synthesized signal, based on the characterized time envelope, such that the adjusted time envelope corresponds substantially to the characterized time envelope.
(57) Abstract
At an audio encoder, cue codes are generated for one or more audio channels, wherein an envelope cue code is generated by characterizing a temporal envelope in an audio channel. At an audio decoder, E transmitted audio channel (s) are decoded to generate C playback audio channels, where C> = Eo1. Received cue codes include an envelope cue code corresponding to a characterized temporal envelope of an audio channel corresponding to the transmitted channel (s). One or more transmitted channel (s) are upmixed to generate one or more upmixed channels. One or more playback channels are synthesized by applying the cue codes to the one or more upmixedchannels, wherein the envelope cue code is applied to an upmixed channel or a synthesized signal to adjust a temporal envelope of the synthesized signal based on the characterized temporal envelope such that the adjusted temporal envelope substantially matches the characterized temporal envelope.
I
INDIVIDUAL CHANNEL FORMATION FOR SCHEMES OF BCC AND THE LIKE
FIELD OF THE INVENTION
The present invention is concerned with the encoding of audio signals and the subsequent synthesis of auditory scenes from the encoded audio data.
BACKGROUND OF THE INVENTION
When a person hears an audio signal (that is, sounds) generated by a particular audio source, the audio signal will commonly arrive at the person's left and right ears at two different times and at two different audio levels (for say, decibles), where these different times and levels are functions of the differences in the paths through which the audio signal travels to reach the left and right ears, respectively. The person's brain interprets these differences in time and level to give the person the perception that the received audio signal is generated by an audio source located at a particular position (for example, direction and distance) in relation to the person. . An auditory scene is the net effect of the person simultaneously listening to audio signals generated by one or more different audio sources located at one or more different positions in relation to the person.
The existence of this processing by the brain can be used to synthesize auditory scenes, where audio signals from one or more different audio sources are purposely modified to generate left and right audio signals that give the perception that the different sources audio are located in different positions in relation to the person.
Figure 1 shows a high-level blogging diagram of the conventional binaural signal synthesizer 100, 10 that converts a single audio source signal (for example, a single signal) to the left and right audio signals of a binaural signal, in where it is defined that a binaural signal is the two signals received at the user's eardrums.
In addition to the audio source signal, the synthesizer 100 receives a set of spatial indications corresponding to the desired position of the audio source relative to the user. In typical implementations, the set of spatial cues comprises an inter-channel level difference (ICLD) value (which identifies the difference in audio level between the left and right audio signals as received by the left and right ears. right, respectively) and an interchannel time difference (ICTD) value (which identifies the difference in arrival time between the left and right audio signals as they are received in the left and right ears, respectively).
Additionally or alternatively, some synthesis techniques involve modeling a direction-dependent transfer function for the sound from the signal source to the eardrums, also referred to as the head-related transfer function (HRTF). See, for example, J. Blauert, The Psychophysics of Human Sound Localization, MIT Press, 1983, the teachings of which are incorporated herein by reference.
By using the binaural signal synthesizer 100 of Figure 1, the mono audio signal generated by a single sound source can be processed in such a way that when listened to in headphones, the sound source is spatially positioned by applying an appropriate set. of spatial cues (eg ICLD, ICTD and / or HRTF) to generate the audio signal for each ear. See, for example, DR Begault, 3-D Sound for Virtual Reality and Multimedia, Academic Press, Cambridge, Mass., 1994.
The binaural signal synthesizer 100 of FIG. 1 generates the simplest type of auditory scenes: those that have a single audio source positioned relative to the user. More complex auditory scenes comprising two or more audio sources located at different positions in relation to the user can be generated using an auditory scene synthesizer that is essentially implemented using multiple instances of the binaural signal synthesizer, where each instance of the binaural signal synthesizer Binaural signals generates the binaural signal corresponding to a different audio source. Since each different audio source has a different location in relation to the user, a different set of spatial indications is used to generate the binaural audio signal for each different audio source.
BRIEF DESCRIPTION OF THE INVENTION
According to one embodiment, the present invention is a machine-readable method, apparatus, and medium for encoding audio channels. One or more cue codes are generated and transmitted for one or more audio channels, wherein at least one cue code is an envelope cue code generated by characterizing a temporal envelope in the one or more audio channels. .
According to one embodiment, the present invention is an apparatus for encoding C input audio channels to generate E channel (s) of transmitted audio. The apparatus comprises an envelope analyzer, code estimator and down mixer. The envelope analyzer characterizes an input temporal envelope of at least one of the C input channels. The code estimator generates indication codes for two or more of the C input channels. The down-mixer downmixes the C input channels to generate the E transmitted channel (s), where OE.l, where the device transmits information about the indication codes and the input time envelope characterized to allow that a decoder performs synthesis and envelope formation during decoding of the transmitted E channel (s).
According to another embodiment, the present invention is an encoded audio bitstream generated by encoding audio channels, wherein one or more indication codes are generated for one or more audio channels, wherein at least one code Indication is an envelope indication code generated by characterizing a temporal envelope on one of the one or more audio channels. The one or more indication codes and E transmitted audio channel (s) corresponding to the one or more audio channels, wherein £ 31, are encoded to the encoded audio bitstream.
According to another embodiment, the present invention is an encoded audio bit stream comprising one or more indication codes and E channel (s) of transmitted audio. The one or more cue codes are generated for one or more audio channels, wherein at least one cue code is an envelope cue code generated by characterizing a temporal envelope on one of the one or more channels audio. The transmitted audio channel (s) correspond to one or more audio channels.
According to another embodiment, the present invention is a machine-readable method, apparatus and medium for decoding E transmitted audio channel (s) to generate C playback audio channels, where C> £ 31. Indication codes corresponding to the E channel (s) transmitted are received, wherein the indication codes comprise an envelope indication code corresponding to a characterized temporal envelope of an audio channel corresponding to the E channel ( es) transmitted. One or more of the transmitted E channel (s) are upmixed to generate one or more upmixed channels. One or more of the C playback channels are synthesized by applying the cue codes to one or more upmixed channels, wherein the envelope cue code is applied to a downmixed channel or synthesized signal to set an envelope. time of the synthesized signal based on the characterized time envelope, such that the adjusted temporal envelope substantially corresponds to the characterized temporal envelope.
BRIEF DESCRIPTION OF THE FIGURES
<img file="MX2007004726A_D0001.tif" />
Other aspects, elements, and advantages of the present invention will become more fully apparent from the following detailed description, the appended claims, and the accompanying figures in which like reference numerals identify similar or identical elements.
Figure 1 shows a high-level block diagram of the conventional binaural signal synthesizer;
Figure 2 is a block diagram of a generic binaural cue coding (BCC) audio processing system;
Figure 3 shows a block diagram of a down mixer that can be used for the down mixer of Figure 2;
Figure 4 shows a block diagram of a BCC synthesizer that can be used for the decoder of Figure 2;
Figure 5 shows a block diagram of the BCC estimator of Figure 2 in accordance with one embodiment of the present invention;
Figure 6 illustrates the ICTD data generation and
ICLD for five channel audio;
Figure 7 illustrates the generation of ICC data for five channel audio;
Figure 8 shows a block diagram of an implementation of the BCC synthesizer of Figure 4 that can be used in a BCC decoder to generate a stereophonic or multichannel audio signal given an individual transmitted sum signal s (n) plus the spatial indications;
Figure 9 illustrates how ICTD and ICLD are varied within a sub-band as a function of frequency;
Figure 10 shows a block diagram of time domain processing that is added to a BCC encoder, such as the encoder of Figure 2, in accordance with one embodiment of the present invention;
Figure 11 illustrates an exemplary time domain application of TP processing in the context of the BCC synthesizer of Figure 4;
Figures 12 (a) and (b) show possible implementations of the TPA of Figure 10 and the TP of Figure 11, 15 respectively, where envelope shaping is applied only at frequencies higher than the cutoff frequency frp!
Figure 13 shows a block diagram of the frequency domain processing that is added to a BCC encoder, such as the encoder of Figure 2, in accordance with an alternative embodiment of the present invention;
Figure 14 illustrates an exemplary frequency domain application of TP processing in the context of the BCC synthesizer of Figure 4;
Figure 15 shows a block diagram of the frequency domain processing that is added to a BCC encoder, such as the encoder of Figure 2, in accordance with another alternative embodiment of the present invention;
Figure 16 illustrates another exemplary frequency domain application of TP precision in the context of the BCC synthesizer of Figure 4;
Figures 17 (a) - (c) show block diagrams of possible implementations of TPA of Figures 15 and 16 and ITP and TP of Figure 16; and
Figures 18 (a) and (b) illustrate two exemplary modes of operation of the control block of Figure 16.
DETAILED DESCRIPTION OF THE INVENTION
In binaural indication coding (BCC), an encoder encodes C input audio channels to generate E transmitted audio channels, where OE ^ l. In particular, two or more of the C input channels are provided in one frequency domain and one or more indication codes are generated for each of one or more different frequency bands on the two or more input channels in the domain. of frequency. In addition, the C input channels are down-mixed to generate the E transmitted channels. In some downmix implementations, at least one of the E transmitted channels is based on two or more of
<img file="MX2007004726A_D0002.tif" />
the C input channels and at least one of the E transmitted channels is based on only one of the C input channels.
In one embodiment, a BCC encoder has two or more filter banks, a code decimator, and a down-mixer. The two or more filter banks convert two or more of the C input channels from a time domain to a frequency domain. The code estimator generates one or more indication codes for each of one or more different frequency bands on the two or more converted input channels. The downmixer downmixes the C input channels to generate the E transmitted channels, where C> E ^ 1.
In BCC decoding, E channels of transmitted audio are decoded to generate C channels of playback audio. In particular, for each of one or more different frequency bands, one or more of the E transmitted channels are upmixed in one frequency domain to generate two or more of the C reproduction channels in the frequency domain, where OE21. One or more indication codes are applied to each of the one or more different frequency bands in the two or more reproduction channels in the frequency domain to generate two or more modified channels and the two or more modified channels are converted from the frequency domain to a time domain.
II
In some upmix implementations, at least one of the C playback channels is based on at least one of the E transmitted channels and at least one cue code and at least one of the C playback channels is based on only one of the E channels transmitted and independent of any indication codes.
In one embodiment, a BCC decoder has an up-mixer, a synthesizer, and one or more reverse filter banks. For each of one or more different frequency bands, the upmixer upmixes one or more of the E channels transmitted in a frequency domain to generate two or more of the C playback channels in the frequency domain, where C > E> 1. The synthesizer applies one or more cue codes to each of the one or more different frequency bands on the two or more reproduction channels in the frequency domain to generate two or more modified channels. The one or more reverse filter banks convert the two or more modified frequency domain channels to a time domain.
Depending on the particular implementation, a given playback channel can be used on a single transmitted channel, rather than a combination of two or more transmitted channels. For example, when there is only one transmitted channel 25, each of the C playback channels is based on that transmitted channel. In these situations, upmixing corresponds to copying the corresponding broadcast channel. As such, for applications in which there is only one transmitted channel, the up-mixer can be implemented using a replicator that copies the transmitted channel for each playback channel.
BCC encoders and / or decoders can be incorporated into a number of systems or applications including, for example, digital video recorders / players, digital video recorders / players, computers, satellite transmitters / receivers, transmitters cable / receivers, terrestrial broadcast transmitters / receivers, home entertainment systems, and movie theater systems.
.
Generic BCC Processing
FIG. 2 is a blog diagram of a generic binaural cue coding (BCC) audio processing system 200 comprising an encoder 202 and 20 a decoder 204. Encoder 202 includes down mixer 206 and BCC estimator 208 .
Down mixer 206 converts C input audio channels Xi (n) to E transmitted audio channels yi (n), where C> EÚ1. In this specification, the signals expressed using variable n are time domain signals, while the expressed signals used using variable k are frequency domain signals. Depending on the particular implementation, downmixing can be implemented in either the time domain or the frequency domain. The BCC estimator 208 generates BCC codes from the C input audio channels and transmits those BCC codes as either in-band or out-of-band side information relative to the E transmitted audio channels. Typical BCC codes include one or more inter-channel time difference (ICTD), inter-channel level difference (ICLD), and inter-channel correlation data (ICC) estimated between certain pairs of input channels as a function of frequency and weather. The particular implementation will determine between which particular pairs of input channels, the BCC codes are estimated.
ICC data corresponds to the coherence of a binaural signal, which is related to the perceived width of the audio source. The wider the audio source, the lower the coherence between the left and right channels of the resulting binaural signal. For example, the coherence of the binaural signal corresponding to an orchestra spread across an auditorium stage is commonly lower than the coherence of the binaural signal corresponding to an individual violin performance solo. In general, an audio signal with lower coherence is usually perceived as more spread out in auditory space. As such, ICC data is commonly concerned with the apparent source width and degree of envelope of the listener. See, for example, J. Blauert, The Psychophysics of Human Sound Localization, MIT Press, 1983.
Depending on the particular application, the E transmitted audio channels and corresponding BCC codes may be transmitted directly to decoder 204 or stored in some appropriate type of storage device for subsequent access by the decoder.
204. Depending on the situation, the term transmission can refer to either direct transmission to a decoder or storage for subsequent provision to a decoder. In either case, the decoder 204 receives the transmitted audio channels and side information and performs BCC upmix and synthesis using the BCC codes to convert the E transmitted audio channels to more than E (commonly, but not necessarily C) playback audio channels f, (n) for audio playback. Depending on the particular implementation, upmixing can be done in either the time domain or the frequency domain.
In addition to BCC processing shown in figure
3, a generic BCC audio processing system may include additional encoding and decoding steps, 25 to further compress the audio signals in the encoder and then decompress the audio signals in the decoder, respectively. These audio codes can be based on conventional audio compression / decompression techniques, such as those based on pulse code modulation (PGM), differential PGM (DPCM) or adaptive DPCM (ADPCM).
When the downmixer 206 generates a single summing signal (that is, E = 1), the BCC encoding is capable of representing multichannel audio signals at a bit rate only slightly higher than that required to represent a signal of mono audio. This is because the ICTD, ICLD, and ICC data estimated between a channel pair contains approximately two orders of magnitude less information than an audio waveform.
Not only the low bit rate of the BCC encoding, but also its backward compatibility aspect is of interest. A single transmitted sum signal corresponds to a mono downmix of the original stereo or multichannel signal. For receivers that do not support stereo or multichannel sound reproduction, listening to the transmitted sum signal is a valid method of representing audio material on low-profile mono reproduction equipment. Therefore, BCC encoding can also be used to enhance existing services that involve the delivery of mono audio material to audio from
<img file="MX2007004726A_D0003.tif" />
multichannel. For example, monaural audio radio broadcast systems can be enhanced for stereo or multichannel reproduction if the BCC side information can be embedded into the existing broadcast channel. Analog capabilities exist when multichannel audio is down-mixed to two summing signals that correspond to stereo audio.
BCC processes audio signals with a certain time and frequency resolution. The frequency resolution used is largely driven by the frequency resolution of the human auditory system. Psycho-acoustics suggests that spatial perception is most likely based on a critical band representation of the acoustic band signal. This frequency resolution is considered when using an invertible filter bank (for example, based on a fast Fourier transform (FFT) or a quadrature mirror filter (QMF)) with sub-bands with bandwidths equal or proportional to the critical bandwidth of the human auditory system.
Generic Down Mix
In preferred implementations, the transmitted sum signal (s) contain (s) all signal components of the input audio signal. The goal is for each signal component to be fully maintained. Simply adding the input audio channels often results in amplification or attenuation of the signal components. In other words, the energy of the signal components in a simple sum is often larger or smaller than the sum of the corresponding signal component energy of each channel. A downmixing technique can be used that equalizes the summing signal such that the energy of the signal components in the summing signal is approximately the same as the corresponding energy in all input channels.
Figure 3 shows a block diagram of a down-mixer 300 that can be used for the down-mixer 206 of Figure 2 in accordance with certain implementations of the BCC system 200. The down-mixer 300 has a filter bank (FB) 302 for each input channel (n), a downmix block 304, an optional scaling / delay block 306, and an inverse FB (IFB) 308 for each encoded channel yi (n).
Each filter bank 302 converts each frame (eg, 20 ms) of a corresponding digital input channel Xj (n) in the time domain to a set of input coefficients x ^ k) in the frequency domain. Downmix block 304 downmixes each subband of C input coefficients corresponding to a corresponding subband 25 of E downmixed frequency domain coefficients. Equation (1) represents the down-mixing of the k-th sub-band of input coefficients (l<sub>}</sub>(k), l<sub>7</sub>_ (k), ..., l<sub>c</sub>(k)} to generate the k-th downwardly mixed coefficient sub-band (y ^ k), y, (k), ..., y<sub>AND</sub>(k)} as follows:
<td>y¿k) y ^ k)</td><td>~ θο?</td><td>\ (k) \ (k)</td>
<td>y ^) _</td><td></td><td>k<sub>c</sub>(k)</td>
(1) where D<sub>rf</sub> is a real-valued C by E downmix matrix.
The optional scaling / delay block 306 comprises a set of multipliers 310, each of which multiplies a corresponding down-mixed coefficient and<sub>t</sub>(k) by a scaling factor (k) to generate a corresponding scaled coefficient y ^ k). The motivation for the scaling operation is equivalent to the generalized EQ for the downmix with arbitrary weights for each channel. If the input channels are independent, then the energy p ..., of the downmixed signal in each sub-band is given by equation (2) as follows:
I
<td>Pipe) P ^ jk)</td><td></td><td>P ^<sub>k)</sub></td><td> , (2</td>
<td>Py ^ k) _</td><td></td><td>Pxcík ^</td><td></td>
where D., is derived by squaring each matrix element in matrix D<sub>C £</sub> downmix of C by E and p - ,,, is the energy of subband k of input channel i.
If the subbands are not independent, then the energy values p ^<sub>(k)</sub> of the downmix signal will be larger or smaller than that calculated using equation (2), due to applications or signal cancellations when the signal components are in phase or out of phase, respectively. To prevent this, the downmix operation of equation (1) is applied in sub-bands followed by the scaling operation of multipliers 310. The scaling factors e<sub>TO</sub>(k) (l # i # E) can be 15 derivatives using equation (3) as follows:
yPñw where p-<sub>(k}</sub> is the sub-band energy as calculated by equation (2) and p<sub>íw</sub> is the energy of the corresponding down-mixed subband signal y ^ k).
In addition to or instead of providing the optional scaling, the scaling / delay block 306 may optionally apply delays to the signals.
Each inverse filter band 308 converts a set of corresponding scaled coefficients y ^ k) in the frequency domain to one frame of a corresponding digital transmitted channel yi (n).
Although Figure 3 shows all of the C input channels being converted to the frequency domain for the subsequent downmix, in alternative implementations, one or more (but less than C-1) of the C input channels could be skewed somewhat or all the processing shown in figure 3 and be transmitted as an equivalent number of unmodified audio channels. Depending on the particular implementation, these unmodified audio channels may or may not be used by the BCC estimator 208 of FIG. 2 in generating the transmitted BCC codes.
In an implementation of downmix that 20 generates a single sum signal y (n), E = 1 and the signals x<sub>c</sub>(k) of each sub-band of each input channel C are added and then multiplied with a factor e (k), according to equation (4) as follows:
yW = e (k) ^ x<sub>c</sub>(k). (4) c = l the factor e (k) is given by equation (5) as follows:
Z ^, W e (k) - \, (5)
V a (^) where p. (k) is a short-time estimate of the energy x<sub>c</sub>(k) at the time index k, and / r (£) is a short-time estimate of the energy of Zt ^^) <sup>The</sup> Equalized subbands are transformed back to the time domain resulting in the sum y (n) that is transmitted to the BCC decoder.
Generic BCC Synthesis
Figure 4 shows a block day of a BCC 400 synthesizer that can be used by the decoder 204 of Figure 2 in accordance with certain implementations of the BCC 200 system. The BCC 400 synthesizer has a filter bank 402 for each transmitted channel Yi (n), an upmix block 404, delays 406, multipliers 408, correlation block 410 and an inverse filter bank 412 for each reproduction channel x ^ n).
Each filter bank 402 converts each frame of a corresponding digital transmitted channel yi (n) in the time domain to a set of input coefficients y \ k) in that of mixed coefficients follows:
<img file="MX2007004726A_D0004.tif" />
frequency domain. The upmix block 404 upmixes each subband of E transmitted channel coefficients corresponding to a corresponding subband of C upmixed frequency domain coefficients. Equation (4) represents the up-mixing of the k-th sub-band of transmitted channel coefficients and to generate the k-th sub-band upwardly (, (Λ), 5<sub>2</sub> (A), ..., í<sub>c</sub>(/ :)) as (6) a real-valued E by C upmix matrix. Performing upmix in the frequency domain allows upmixing to be applied individually in each different sub-band.
Each delay 406 applies a delay value di (k) based on a corresponding BCC code for ICTD data to ensure that the desired ICTD values appear between certain pairs of reproduction channels. Each multiplier 408 applies a scaling factor ai (k) based on a corresponding BCC code for ICLD data to ensure that the desired ICLD values appear between certain pairs of playback channels. The correlation block 410 performs
<img file="MX2007004726A_D0005.tif" />
an A de-correlation operation based on corresponding BCC codes for ICC data to ensure that the desired ICC values appear between certain pairs of reproduction channels. A further description of the 5 operations of the correlation block 410 can be found in US Patent Application No. 10 / 155,437, filed 05/24/02 as Baumgarte 2-10.
The synthesis of ICLD values can be less troublesome than the synthesis of ICTD and ICC values, since the synthesis of ICLD involves only the scaling of the sub-band signals. Since ICL cues are the most commonly used directional cues, it is usually more important than ICLD values approximate to those of the original audio signal. As such, the 15 ICLD data could be estimated across all channel pairs. The scaling factors ai (k) (l # i # C) for each sub-band are preferably chosen such that the sub-band energy of each reproduction channel approximates the corresponding energy of the original derived audio channel. .
One goal may be to apply relatively few signal modifications to synthesize ICTD and ICC values. As such, the BCC data might not include ICTD and ICC values for all channel pairs. In that case, the BCC synthesizer 400 would synthesize ICTD and ICC 25 values only between certain pairs of channels.
Each inverse filter bank 412 converts a set of corresponding synthesized coefficients x ^ k) in the frequency domain to one frame of a corresponding digital reproduction channel x, (n).
Although Figure 4 shows all E transmitted channels being converted to the frequency domain for subsequent upmixing and BCC processing, in alternative implementations, one or more (but not all) of the E transmitted channels could deviate from some or all of the processing shown in Figure 4. For example, one or more of the transmitted channels may be unmodified channels that are not upmixed. In addition to being one or more of the C playback channels, these unmodified channels could in turn not have to be used as reference channels to which processing is applied.
BCC to synthesize one or more of the other playback channels. Either way, such unmodified channels can be delayed to compensate for the processing time involved in upmixing and / or BCC processing used to generate the rest of the playback channels.
Note that although Figure 4 shows C playback channels being synthesized from E transmitted channels, where C was also the number of original input channels, BCC synthesis is not limited to that number of playback channels. In general, the number of playback channels can be any number of channels, including numbers greater or less than C and 5 possibly even situations where the number of playback channels is equal to or less than the number of channels transmitted. .
Perceptually relevant differences between audio channels
Assuming a single summation signal, BCC synthesizes a stereo or multichannel audio signal in such a way that ICTD, ICLD, and ICC approximate the corresponding indications of the original audio signal. In the following, the role of ICTD, ICLD, and ICC in relation to auditory spatial imaging attributes is discussed.
Knowledge about spatial hearing implies that for an auditory event, ICTD and ICC are related to the perceived direction. When considering binaural room impulse responses (BRIR) from a source, there is a relationship between the width of the auditory event and the listening envelope and estimated ICC data for premature and posterior parts of the BRIRs. However, the relationship between ICC and these properties for general signals (and not just BRIRs) is not direct.
Stereo and multichannel audio signals usually contain a complex mix of concurrently active source signals overlaid by the reflected signal components resulting from recording in close quarters or added by the recording technician to artificially create a spatial impression. Signals from different sources and their reflections occupy different regions in the time-frequency plane. This is reflected by ICT, ICLD and ICC that vary as a function of time and frequency. In this case, the relationship between instantaneous ICTD, ICLD, and ICC and auditory event addresses and spatial impression is not obvious. The strategy of certain BCC modalities is to blindly synthesize these cues, so that they approximate the corresponding cues of the original audio signal.
Filter banks with sub-bands of bandwidths equal to twice the equivalent rectangular bandwidth (ERB) are used. Informal listening reveals that BCC's audio quality does not noticeably improve when a higher frequency resolution is chosen. Lower frequency resolution may be desirable since it results in fewer ICTD, ICLD and ICC values that need to be transmitted to the decoder and thus in a lower bit rate.
<img file="MX2007004726A_D0006.tif" />
With regard to time resolution, ICTD, ICLD and ICC are commonly considered at regular time intervals. High performance is obtained when ICTD, ICLD and ICC are considered approximately every 4 to 16 ms. Note that, unless the indications are considered at very short time intervals, the precedence effect is not considered directly. Assuming a classical sound stimulus lead-lag pair if lead and lag fall to a time interval where only one set of 10 cues is synthesized, then the forward location dominance is not considered. Ά Despite this, BCC achieves reflected audio quality at an average MUSHRA score of about 87 (that is, excellent audio quality) on average and up to nearly 100 for certain 15 audio signals.
The frequently obtained perceptually small difference between the reference signal and the synthesized signal implies that indications related to a wide range of auditory spatial image attributes are implicitly considered when synthesizing ICTD, ICLD, and ICC at regular time intervals. In the following, some arguments are given as to how ICTD, ICLD and ICC can be related to a range of auditory spatial image attributes.
Estimation of spatial indications
The following describes how ICTD, ICLD, and ICC are estimated. The bit rate for the transmission of these spatial indications (quantized and encoded) can be only a few kb / s and thus, with BCC, it is possible to transmit stereo and multichannel audio signals at bit rates close to the one required for single channel of audio.
Figure 5 shows a block diagram of the BCC estimator 208 of Figure 2, in accordance with one embodiment of the present invention. The BCC estimator 208 comprises filter banks (FB) 502, which can be the same as the filter banks 302 of Figure 3 and the estimation block 504, which generates ICTD, ICLD and ICC spatial indications for each sub-band. of different frequency generated by filter banks 502.
ICTD, ICLD and ICC estimation for stereo signals
The following measurements are used for ICTD, ICLD and ICC for corresponding subband signals x<sub>}</sub>(k) and x<sub>2</sub>(k) two-channel audio (for example stereo):
ICTD [samples]:
ñ2 (^) <sup>= ar</sup>S<sup>rnaX</sup>{<sup>(I)</sup>|2(^^)} ' (<sup>7</sup>) d ''
<img file="MX2007004726A_D0007.tif" />
with a short-time estimate of the normalized cross-correlation function given by equation (8) as follows:
. (8)
P / kd ^ p. (k-dj where d<sub>}</sub> = max {-ú, 0} d<sub>2</sub> = max {í /, 0} and ρ<sub>: K</sub>- (d, k) is a short-time estimate of the mean of x<sub>t</sub>(kd<sub>t</sub>) x<sub>2</sub>(kd<sub>2</sub>) .
ICLD [dB]:
TO THE<sub>|2</sub>(Z :) = 10log<sub>l0</sub> 'Pjk / (10)
ICC:
<sub>C]</sub>^ k ^ max \ d> ^ d, k) \. (eleven)
Note that the absolute value of the normalized cross-correlation is considered and c<sub>l2</sub>(A) has an interval of
[0,1].
ICTD, ICLD, and ICC estimation for multichannel audio signals
When there are more than two input channels, it is usually sufficient to define ICTD and ICLD between a reference channel (for example channel number 1) and the other channels, as illustrated in figure 6 for the case of C = 5 channels, where r<sub>lf</sub>(k) and & L<sub>n</sub>(k) denote the ICTD and ICLD, respectively, between reference channel 1 and channel c.
In contrast to ICTD and ICLD, ICC commonly has more degrees of freedom. The ICC as defined can have different values between all possible input channel pairs. For C channels, there are C (Cl) / 2 possible channel pairs; for example for 5 channels there are 10 pairs of channels as illustrated in figure 7 (a). However, such a scheme requires that, for each subband at each time index, the C (C — 1) / 2 ICC values be estimated and transmitted, resulting in high computational complexity and high bit rate.
Alternatively, for each subband, ICTD and ICLD determine the direction in which the auditory event of the corresponding signal component is provided in the subband. A single ICC parameter per subband can then be used to describe the overall coherence between all audio channels. Good results can be obtained by estimating and transmitting ICC indications only between the two channels with the highest energy in each sub-band at each time index. This is illustrated in Figure 7 (b), where for time instants k-1 and k, the channel pairs (3,4) and (1,2) are stronger, respectively. A heuristic rule can be used to determine ICC between the other pairs of channels.
Synthesis of spatial indications
Figure 8 shows a block diagram of an implementation of the BCC synthesizer 400 of Figure 4 that can be used in a BCC decoder to generate a stereo or multi-channel audio signal given an individual transmitted sum signal s (n) plus the spatial indications. The sum signal s (n) is decomposed into sub-bands, where s (k) denotes one such sub-band. To generate the corresponding sub-bands of each of the output channels, delays d<sub>cl</sub> scale factors a<sub>c</sub>, and filters h<sub>c</sub> to the corresponding sub-band of the sum signal. (For simplicity of notation, the time index k is ignored in delays, scale factors, and filters.) ICTDs are synthesized by imposing delays, ICLD by scaling, and ICC by applying de-correlation filters. The processing shown in Figure 8 is applied independently to each subband.
ICTD synthesis
The delays d<sub>c</sub> are determined from the ICTD r, (£) according to equation (12) as follows:
- - (max<sub>2s / sr</sub> + min<sub>2s / sc</sub> η, (/ :)), r „(a) + í /,
The delay for the channel calculated such that the magnitude
2 <c <C.
reference di is maximum of the delays d<sub>c</sub> is minimized. The less the subband signals are modified, the less danger of artifacts being present. If the subband sampling rate does not provide high enough time resolution for ICTD synthesis, delays can be more precisely imposed by using appropriate all-step filters.
ICLD synthesis
In order for the output subband signals to have desired ICLDs AL<sub>l2</sub>(k) between channel c and reference channel 1, the gain factors a<sub>c</sub> must satisfy equation (13) as follows:
<sub>to</sub> ^ = 10 <sup>20</sup> ... (13) "I
Additionally, the output sub-bands are preferably normalized, such that the sum of the energy of all the output channels is equal to the energy of the input sum signal. Since the total original signal energy in each subband is preserved in the summing signal, this normalization results in the absolute subband energy for each output channel that approximates the
<img file="MX2007004726A_D0008.tif" />
corresponding power of the original encoder input audio signal. Given these constraints, the scale factors at<sub>c</sub> are given by equation (14) as follows:
<sup>C</sup> VIRUS '<sup>1</sup>'''<sup>0</sup>, c = l (14) 'otherwise.
ICC synthesis
In certain modalities, the goal of ICC synthesis is to reduce the correlation between the sub-bands after delays and scaling have been applied, without affecting ICTD and ICLD. This can be obtained by designing the filters h<sub>c</sub> in Figure 8 such that ICTD and ICLD are effectively varied as a function of frequency such that the average variation is zero in each sub-band (auditory critical band).
Figure 9 illustrates how ICTD and ICLD are varied within a sub-band as a function of frequency. The amplitude of the variation of ICTD and ICLD determines the degree of de-correlation and is controlled as a function of ICC. Note that ICTD is varied smoothly (as in Figure 9 (8a)), while ICLD is varied randomly (as in Figure 9 (b)). ICLD could be varied as smoothly as ICTD, but this would result in more coloration of the resulting 25 audio signals.
Another method for synthesizing ICC, particularly suitable for multichannel ICC synthesis, is described in more detail in Faller, Parametric multi-channel audio coding: Synthesis of coherence cues, IEEE Trans. on Speech and Audio 5 Proc., 2003, the teachings of which are incorporated herein by reference. As a function of time and frequency, specific amounts of artificial late reverb are added to each of the output channels to obtain a desired ICC. Additionally, spectral modification can be applied in such a way that the spectral envelope of the resulting signal approximates the spectral envelope of the original audio signal.
Other related and unrelated ICC synthesis techniques for stereo signals (or pairs of audio channels) 15 have been presented in E. Schuijers, W. Oomen, B. den Brinker, and J. Breebaart, Advances in parametric coding for high quality audio , in Preprint 114<sup>th</sup> Conv. Aud. Eng. Soc., March 2003 and J. Engdegard, H. Purnhagen, J. Roden, and L. Liljeryd, Synthetic ambience in parametric stereo coding, in 20 Preprint 117<sup>th</sup> Conv. Aud. Eng. Soc., May 2004, the teachings of both of which are incorporated herein by reference.
C to E BCC
As previously described, BCC can be implemented with more than one transmission channel. A variation of BCC has been described that represents C channels of audio not as a single channel (transmitted), but as E channels, denoted C through E BCC. There are (at least) two motivations for C to E BCC:
BCC with one broadcast channel provides a compatible backward path to upgrade existing monaural systems for stereo or multichannel audio playback. Upgraded systems transmit the BCC downmixed sum signal over the existing monaural infrastructure, while additionally transmitting the BCC side information. C to E BCC is applicable to C-channel audio backward compatible encoding of E-channel.
C to E BCC introduces scalability in terms of different degrees of reduction in the number of transmitted channels. It is expected that the more audio channels are transmitted, the better the audio quality.
Signal processing details for C to E BCC, such as how to define ICTD, ICLD, and ICC indications, are described in U.S. Patent Application Serial No. 10 / 762,100, filed 20/01/04 (Faller 13 -one).
Individual channel formation
In certain embodiments, both BCC with a transmission channel and C to E of BCC involve algorithms for the synthesis of ICTD, ICLD, and / or ICC. Usually, it is sufficient to synthesize the ICTD, ICLD, and / or ICC indications approximately every 4 to 30 ms. However, the perceptual phenomenon of precedence effect implies that there are specific instants in time when the human auditory system evaluates cues at a higher time resolution (eg, every 1 to 10 ms).
A single static filter bank cannot commonly provide sufficiently high frequency resolution appropriate for most instants in time, while providing sufficiently high time resolution at instants in time when the precedence effect becomes effective.
Certain embodiments of the present invention are concerned with a system that uses relatively low time resolution ICTD, ICLD, and / or ICC synthesis, while adding additional processing to address time points when higher time resolution is required. . Additionally, in certain embodiments, the system eliminates the need for adaptive signal window switching technology that is usually difficult to integrate into the fabric of a system. In certain embodiments, the temporal envelopes of one or more of the input audio channels of the original encoder are estimated. This can be done, for example directly by analyzing the time structure of the signal or by examining the autocorrelation of the signal spectrum with respect to frequency. Both procedures will be elaborated further in the subsequent implementation examples 5. The information contained in these envelopes is transmitted to the decoder (as envelope has codes) if it is perceptually required and advantageous.
In certain embodiments, the decoder applies some processing to impose these desired temporary envelopes 10 on its output audio channels:
This can be achieved by TP processing, eg, manipulating the signal envelope by multiplying the time domain samples of the signal with a time varying amplitude modifying function. Similar processing can be applied to spectral / subband samples if the time resolution of the subbands is high enough (at the cost of coarse frequency resolution).
Alternatively, a convolution / filtering of the spectral representation of the signal with respect to frequency can be used in a manner analogous to that used in the prior art for the purpose of forming the quantization noise - of a low speed audio encoder. bits or to enhance the intensity of stereo encoded signals. This is preferred if the filter bank has a high frequency resolution and therefore a rather low time resolution. For the convolution / filtration procedure:
The envelope formation method is extended from stereo intensity to multi-channel C to E coding.
The technique comprises a setting where the envelope formation is controlled by parametric information (eg, binary flags) generated by the encoder but is actually carried out using sets of decoder-derived filter coefficients.
In another setting, sets of filter coefficients are transmitted from the encoder, for example only when perceptually necessary and / or beneficial.
The same is also true for the time domain / sub-band domain procedure. Accordingly, criteria (eg, transient detection and an estimated tonality value) can be entered to further control the transmission of envelope information.
There may be situations when it is favorable to disable TP processing in order to avoid potential artifacts. In order to be on the safe side, it is a good strategy to leave temporary processing disabled by default (that is, BCC would operate according to a conventional BCC scheme). Additional processing is enabled only when higher temporal resolution of the channels is expected to produce improvement, for example, when the precedence effect is expected to become active.
As stated above, this enable / disable control can be obtained by transient detection. That is, if a transient is detected, then TP processing is enabled. The precedence object is most effective for transients. Transient detection can be used in advance 10 to effectively form not only individual transients but also the signal components briefly before and after the transient. Possible ways to detect transients include:
Observe the time envelope of the transmitted BCC encoder input signals or BCC sum signal (s). If there is a sudden increase in energy, then a transient occurred.
Examine the linear predictive coding (LPC) gain as estimated at the encoder or decoder. If the LPC prediction gain exceeds a certain threshold, then the signal can be assumed to be transient or highly fluctuating. The LPC analysis is calculated on the autocorrelation of the spectrum.
Additionally, to prevent possible artifacts in the tonal signals, the TP processing is preferably not applied when the tonality of the transmitted sum signal (s) is high.
In accordance with certain embodiments of the present invention, the temporal envelopes of the individual original audio channels are estimated in a BCC encoder in order to enable a BCC decoder to generate output channels with similar (or perceptually similar) temporal envelopes. ) to those of the original audio channels. Certain embodiments of the present invention address the precedence effect phenomenon. Certain embodiments of the present invention involve the transmission of envelope indication codes in addition to the other BCC codes such as ICLD, ICTD, and / or ICC, as part of the BCC side information.
In certain embodiments of the present invention, the time resolution for time envelope indications is finer than the time resolution of other BCC codes (eg, ICLD, ICTD, ICC). This allows the envelope formation to be performed within the time period stipulated by a synthesis window that corresponds to the length of a block of an input channel for which the other BCC codes are derived.
Implementation examples
Figure 10 shows a block diagram of time domain processing that is added to a BCC encoder, such as the encoder 202 of Figure 2, in accordance with one embodiment of the present invention. As shown in Figure 10 (a), each Temporal Processing Analyzer 5 (TPA) 1002 estimates the temporal envelope of a different original input channel x<sub>c</sub>(n), although in general any of one or more of the input channels can be analyzed.
Figure 10 (b) shows a block diagram of a possible time-domain-based implementation of TPA 1002 10 in which the input signal samples are squared (1006) and then low-pass filtered (1008 ) to characterize the time envelope of the input signal. In alternative embodiments, the time envelope can be estimated using an autocorrelation / LPC method or with other methods, for example using a Hilbert transform.
Block 1004 of FIG. 10 (a) parameterizes, quantizes, and encodes the estimated temporal envelopes prior to transmission as temporal processing information (TP) (i.e., envelope indication codes) that is included in the side information of figure 2.
In one embodiment, a detector (not shown) within block 1004 determines whether TP processing in the decoder will improve audio quality, such that block 1004 transmits TP side information only during those instants of time when the Audio quality will be improved by TP processing.
Figure 11 illustrates an exemplary time domain application of TP processing in the context of the BCC synthesizer 400 of Figure 4. In this embodiment, there is a single summation signal transmitted s (n), C base signals are generated by replicating that summing signal and the envelope formation is applied individually to different synthesized channels. In alternative embodiments, the order of delays, scaling, and other processing may be different. Furthermore, in alternative embodiments, the envelope formation is not restricted to the processing of each channel independently. This is especially true for convolution / filtering implementations that take advantage of coherence over frequency bands to derive information regarding the fine temporal structure of the signal.
In Fig. 11 (a), decoding block 1102 recovers time envelope signals a for each output channel of the transmitted TP side information from the BCC encoder and each TP block 1104 applies the corresponding envelope information to form the envelope of the output channel.
Figure 11 (b) shows a blog diagram of a possible time-domain based implementation of TP 1104 in which the synthesized signal samples are squared (1106) and then low-pass filtered (1108). to characterize the temporal envelope b of the synthesized channel. A scale factor (for example, sqrt (a / b)) (1110) 5 is generated and then applied (1112) to the synthesized channel to generate an output channel that has a time envelope substantially equal to that of the input channel corresponding original.
In alternative implementations of TPA 1002 of Figure 10 and TP 1104 of Figure 11, the temporal envelopes are characterized using magnitude operations rather than squaring the signal samples. In such implementations, the proportion a / b can be used as the scale factor without having to apply the square root operation.
Although the scaling operation of Figure 11 (c) corresponds to a time-domain-based implementation of TP processing, TP processing (also as TPA processing and reverse TP (ITP)) can also be implemented using frequency domain signals, as in the embodiment of Figures 16-17 (described later herein). As such, for the purposes of this specification, the term scaling function should be interpreted to cover either time domain operations or frequency domain operations, such as the filtering operations of Figures 17 (b) and (c). ).
In general, each TP 1104 is preferably designed in such a way that it does not modify the signal energy (ie, energy). Depending on the particular implementation, this signal energy may be a short-time average signal energy on each channel, for example based on the total signal energy per channel in the time period defined by the synthesis window or some other measure. of appropriate energy.
As such, scaling for ICLD synthesis (eg, using multipliers 408) can be applied before or after envelope formation.
Since full-band scaling of the BCC output signals can result in artifacts, the envelope shaping could only be applied at specified frequencies, for example frequencies greater than a certain cutoff frequency f<sub>TP</sub> (for example, 500 Hz). Note that the frequency range for analysis (TPA) may differ from the frequency range for synthesis (TP).
Figures 12 (a) and (b) show possible implementations of TPA 1002 of Figure 10 and TP 1104 of Figure 11 where envelope shaping is applied only at frequencies higher than the cutoff frequency f<sub>TP</sub>. In particular, Figure 12 (a) shows the addition of the high pass filter 1202, which filters frequencies lower than f<sub>TP </sub>before the temporal envelope characterization. Figure 12 (b) shows the addition of the two-band filter bank 1204 having a cutoff frequency f<sub>TP</sub> between the two sub-bands, 5 where only the high frequency part is temporarily formed. Then the reverse two-band filter bank 1206 recombined the low frequency part with the temporarily formed high frequency part to generate the output channel.
Figure 13 shows a block diagram of the frequency domain processing that is added to a BCC encoder, such as the encoder 202 of Figure 2, in accordance with an alternative embodiment of the present invention. As shown in figure 13 (a), the processing of each TPA 1302 is applied individually in a different sub-band, where each filter bank (FB) is the same as corresponding FB 302 of figure 3 and the block 1304 is a sub-band implementation analogous to block 1004 of FIG. 10. In alternative implementations, the sub-bands for TPA processing may differ from the BCC sub-bands. As shown in Figure 13 (b), the TPA 1302 can be implemented analogous to the TPA 1002 of Figure 10.
Figure 14 illustrates an exemplary frequency domain application of TP processing in the context of the BCC synthesizer 400 of Figure 4. Decode block 1402 is analogous to decode block 1102 of Figure 11, and each TP 1404 is a sub-band implementation analogous to each TP 1104 of Figure 11, as shown in Figure 14 (b).
Figure 15 shows a block diagram of the frequency domain processing that is added to a BCC encoder, such as the encoder 202 of Figure 2, in accordance with another alternative embodiment of the present invention. This scheme has the following setting: The envelope information for each input channel is derived by LPC calculation through frequency (1502), parameterized (1504), quantized (1506), and encoded to the bit stream (1508) using the encoder. Figure 17 (a) illustrates an example implementation of the TPA 1502 of Figure 15. The side information to be transmitted to the multichannel synthesizer (decoder) could be the LPC filter coefficients calculated by an autocorrelation method, the resulting reflection coefficients or line spectral pairs, etc., for the purpose of maintaining speed. of small side information data, parameters derived from, for example, LPC prediction gain as present / not present transient binary flags.
Figure 16 illustrates another exemplary frequency domain application of TP processing in the context of the BCC synthesizer 400 of Figure 4. The encoding processing of Figure 15 and decoder processing of Figure 16 can be implemented to form a corresponding pair of an encoder / decoder configuration. Decode block 1602 is analogous to decode block 1402 of FIG. 14, and each TP 1604 is analogous to each TP 1404 of FIG. 14. In this multichannel synthesizer, the transmitted TP side information is decoded and used to control the envelope formation of individual channels, however, in addition, the synthesizer includes an Envelope Characterizer (TPA) stage 1606 for analyzing signal signals. transmitted sum, an inverse TP (ITP) 1608 to flatten the temporal envelope of each base signal, wherein the envelope adjusters (TP) 1604 impose a modified envelope on each output channel. Depending on the particular implementation, ITP can be applied either before or after upmix. In detail, this is done using the convolution / filtration procedure where the envelope formation is obtained by applying filters based on LPC over the spectrum through frequency as illustrated in Figures 17 (a), (b ), and (c) for the processing of TPA, ITP, and TP, respectively. In Figure 16, the control block 1610 determines whether or not the envelope formation is to be implemented and if so, it will be based on (1) the transmitted TP side information or (2) the locally characterized envelope data from TPA 1606.
Figures 18 (a) and (b) illustrate two exemplary modes for operating the control block 1610 of Figure 16. In the implementation of Figure 18 (a), a set of filter coefficients is transmitted to the decoder and envelope formation by convolution / filtering is done based on the transmitted coefficients. If transient formation that is not beneficial is detected by the encoder, then no filter data is sent and the filters are disabled (shown in Figure 18 (a) by switching to a unit filter coefficient set [1,0 ...]).
In the implementation of Figure 18 (b), only one transient / non-transient flag is transmitted for each channel and this flag is used to activate or deactivate training based on the filter coefficient sets calculated from the signals. downmix signals transmitted on the decoder.
Additional alternative modalities
Although the present invention has been described in the context of BCC coding schemes in which there is a single summation signal, the present invention can also be implemented in the context of BCC coding schemes that have two or more summation signals. . In this case, the time envelope for each different base sum signal can be estimated before the application of BCC synthesis and different BCC output channels can be generated based on different time envelopes, depending on which sum signals were. used to synthesize the different output channels. An output channel that is synthesized from two or more different summation channels could be generated based on an effective time envelope that takes into account (eg, via weighted averaging) the relative effects of the constituent summation channels.
Although the present invention has been described in the context of BCC coding schemes involving ICTD, ICLD, and ICC codes, the present invention may also be implemented in the context of other BCC coding schemes involving only one or only two of these three types of codes (for example, ICLD and ICC, but not ICTD) and / or one or more types of additional codes. Furthermore, the BCC synthesis and envelope formation processing sequence may vary in different implementations. For example, when enveloping is applied to frequency domain signals, as in Figures 14 and 16, enveloping could alternatively be implemented after ICTD synthesis (in those 25 modalities that employ ICTD synthesis) , but before the ICLD analysis. In other embodiments, the envelope formation could be applied to up-mixed signals before any other BCC synthesis is applied.
Although the present invention has been described in the context of BCC coders that generate envelope indication codes from the original input channels, in alternative embodiments, the envelope indication codes could be generated from 10 downmixed channels. corresponding to the original input channels. This would allow the implementation of a processor (e.g. a separate envelope indication encoder) that could (1) accept the output of a BCC encoder that generates the downmixed channels and certain BCC codes (e.g. ICLD, ICTD , and / or ICC) and (2) characterize the temporal envelope (s) of one or more of the downmixed channels to add envelope indication codes to the BCC side information.
Although the present invention has been described in the context of BCC coding schemes in which the envelope indication codes are transmitted with one or more audio channels (i.e., the E channels transmitted) in conjunction with other BCC codes, In alternative embodiments, the envelope indication codes could be transmitted, either alone or with other BCC codes, to a location (e.g., a decoder or a storage device) that already has the transmitted channels and possibly other BCC codes.
Although the present invention has been described in the context of BCC coding schemes, the present invention may also be implemented in the context of other audio processing systems in which the audio signals are de-correlated or other audio processing. audio you need to de-correlate signals.
Although the present invention has been described in the context of implementations in which the encoder receives the input audio signal in the type domain and generates transmitted audio signals in the time domain and the decoder receives the transmitted audio signals in the time domain and generates playback audio signals in the time domain, the present invention is not limited in this way. For example, in other implementations, any one or more of the input playback audio signals 20 transmitted could be represented in a frequency domain.
BCC encoders and / or decoders can be used in conjunction with or incorporated into a variety of different applications or systems, including systems for television or electronic music distribution, cinemas, broadcast, stream, and / or reception. These include systems for encoding / decoding transmissions via for example terrestrial media, satellite, cable, internet, intranet or physical media (for example, compact discs, digital versatile discs, semi-conductor chips, hard drives, memory cards and the like) . BCC encoders and / or decoders can also be used in games and game systems that include, for example, interactive software or programming element products designed to interact with a user for entertainment (action, role playing, strategy, adventure , simulations, racing, sports, arcade, cards and board games) and / or education that can be published for multiple machines, platforms or media. In addition, BCC encoders and / or decoders can be incorporated into PC programming element applications that incorporate digital decoding (e.g., player, decoder) and programming element applications that incorporate digital encoding capabilities (e.g., encoder, decoder, recoder and consoles).
The present invention may be implemented as circuit-based processes, which include possible implementations such as a single integrated circuit (such as an ASIC or FPGA), a multi-chip module, a single card, or a circuit pack of multiple cards. As
<img file="MX2007004726A_D0009.tif" />
It will be apparent to one skilled in the art, various circuit element functions can also be implemented as processing steps in a program of programming elements. Such programming elements 5 can be used for example in a digital signal processor, microcontroller or general purpose computer.
The present invention can be implemented in the form of methods and apparatus for carrying out those methods. The present invention can also be implemented in the form of program code implemented on tangible media, such as floppy disks, CD-ROMs, hard drives or any other storage medium that can be read by the machine, where, when the code If the program is loaded to and performed by a machine, such as a computer, the machine 15 becomes an apparatus for carrying out the invention. The present invention may also be implemented in the form of program code, for example, whether stored on a storage medium, loaded to and / or executed by a machine, or transmitted on some transmission medium or carrier, such as wire or electrical wiring, via optical fibers or via electromagnetic radiation, where, when program code is loaded into and executed by a machine, such as a computer, the machine becomes an apparatus for practicing the invention.
When implemented in a general purpose or multipurpose processor, the program code segments combine with the processor to provide a single device that operates analogously to specific logic circuits.
It will be further understood that various changes in details, materials, and arrangements of parts that have been described and illustrated in order to explain the nature of this invention can be made by those skilled in the art without departing from the scope of the invention as expressed. in the following claims.
Although the steps in the following method claims, if any, are cited in a particular sequence with corresponding labeling, unless the claims citations otherwise imply a particular sequence to implement some or all of these steps, those steps are not necessarily intended to be limited to being implemented in that particular sequence.
Contents10
25 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25
34 members in 21 offices
Priority claims3
| Document | Office | Kind | Date |
|---|---|---|---|
| 62048004 | United States of America | P | |
| 648204 | United States of America | A | |
| 2005009618 | European Patent Office (EPO) | W |
Members34
| Document | Office | Kind | |
|---|---|---|---|
| US2006083385A1 | United States of America | A1 | |
| AU2005299068A1 | Australia | A1 | |
| CA2582485A1 | Canada | A1 | |
| WO2006045371A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW200628001A | Taiwan Province of China | A | |
| NO20071493L | Norway | L | |
| KR20070061872A | Republic of Korea | A | |
| EP1803117A1 | European Patent Office (EPO) | A1 | |
| MX2007004726AThis record | Mexico | A | |
| IL182236A0 | Israel | A0 | |
| CN101044551A | China | A | |
| HK1106861A1 | Hong Kong, China | A1 | |
| JP2008517333A | Japan | A | |
| BRPI0516405A | Brazil | A | |
| AU2005299068B2 | Australia | B2 | |
| RU2339088C1 | Russian Federation | C1 | |
| EP1803117B1 | European Patent Office (EPO) | B1 | |
| AT424606T | Austria | T | |
| ATE424606T1 | Austria | T1 | |
| DE602005013103D1 | Germany | D1 | |
| PT1803117E | Portugal | E | |
| DK1803117T3 | Denmark | T3 | |
| ES2323275T3 | Spain | T3 | |
| PL1803117T3 | Poland | T3 | |
| KR100924576B1 | Republic of Korea | B1 | |
| TWI318079B | Taiwan Province of China | B | |
| US7720230B2 | United States of America | B2 | |
| JP4664371B2 | Japan | B2 | |
| IL182236A | Israel | A | |
| CN101044551B | China | B | |
| CA2582485C | Canada | C | |
| NO338919B1 | Norway | B1 | |
| BRPI0516405A8 | Brazil | A8 | |
| BRPI0516405B1 | Brazil | B1 |
1 legal event, as the office reported them to INPADOC
Events
| Event | Code | |
|---|---|---|
| Grant or registrationFG | FG |
Numbers
- Application
- 2007004726
Titles2
- English
- INDIVIDUAL CHANNEL TEMPORAL ENVELOPE SHAPING FOR BINAURAL CUE CODING SCHEMES AND THE LIKE.
- Spanish
- FORMACION DE CANAL INDIVIDUAL PARA ESQUEMAS DE BCC Y LOS SEMEJANTES.
Classification
- CPC, 3
- G10L19/008
- G10L19/02
- H03M7/30
- IPC, 1
- G10L19 02