Parametric multi-channel audio representation
Abstract
A method for encoding a multi-channel audio signal, comprising at least two audio channels (RI, LI), such that the method comprises generating (1) an audio signal (SC "single channel") of a single channel, comprising a particular combination of the at least two audio channels (RI, LI), and encoding the single channel audio signal (SC) into a bit stream (EBS), as an audio signal from single coded channel (ESC), generate (2) information (INF) from the at least two audio channels (RI, LI), which allows to recover, with a required quality level, the multi-channel audio signal from the audio signal of single channel (SC) and information (INF), so that the generation (2) of the information comprises: - determine (2) a first portion of the information (P1), which consists of a single set of parameters (S1), determined for a first frequency zone (FR1) of the multi-channel audio signal, and encode it for the first portion of the information (P1) in the current bits (EBS), as a first coded portion of the information (EIN ¿encoded information), and - determining (2) a second portion of the information (P2) for a second frequency zone (FR2) of the multi-channel audio signal, such that the second frequency zone (FR2) is a portion of the first frequency zone (FR1), and encode the second portion of the information (P2) within the bit stream (EBS), as a second encoded portion of the information (EIN).

Term
Term ended
Projected expiry passed 22 April 2023, 3.4 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
20 claims: 3 independent, 17 dependent
- 1ES 2 268 340 T3 REIVINDICACIONES 1. Un método para codificar una señal de audio de múltiples canales, que comprende al menos dos canales de audio (RI, LI), de tal forma que el método comprende generar (1) una señal de audio (SC -“single channel”) de un único canal, que comprende una combinación particular de los al menos dos canales de audio (RI, LI), y codificar la señal de audio de canal único (SC) en una corriente de bits (EBS), como una señal de audio de canal único codificada (ESC), generar (2) información (INF) a partir de los al menos dos canales de audio (RI, LI), que permite recuperar, con un nivel de calidad requerido, la señal de audio de múltiples canales a partir de la señal de audio de canal único (SC) y de la información (INF), de tal modo que la generación (2) de la información comprende:- determinar (2) una primera porción de la información (P1), que consiste en un único conjunto de parámetros (S1), determinados para una primera zona de frecuencias (FR1) de la señal de audio de múltiples canales, y codificar la primera porción de la información (P1) en la corriente bits (EBS), como una primera porción codificada de la información (EIN -“encoded information”), y - determinar (2) una segunda porción de la información (P2) para una segunda zona de frecuencias (FR2) de la señal de audio de múltiples canales, de tal modo que la segunda zona de frecuencias (FR2) es una porción de la primera zona de frecuencias (FR1), y codificar la segunda porción de la información (P2) dentro de la corriente de bits (EBS), como una segunda porción codificada de la información (EIN).
- 2Un método para codificar una señal de audio de múltiples canales, de acuerdo con la reivindicación 1, que comprende adicionalmente:determinar únicamente (2) la segunda porción de la información (P2) para la segunda zona de frecuencias (FR2) de la señal de audio de múltiples canales en el caso de que una velocidad de bits de la señal de audio de múltiples canales codificada, que comprende la señal de audio de canal único (SC), la primera porción de la información (P1) y la segunda porción de la información (P2), no sea superior a una velocidad de bits máxima permisible (MBR).
- 3Un método de codificación de acuerdo con la reivindicación 1, caracterizado porque la información (INF) comprende conjuntos de parámetros (S1, S2,...), la primera porción (P1) comprende al menos un primero (S1) de los conjuntos de parámetros (S1, S2,...), y la segunda porción (P2) comprende al menos un segundo (S2) de los conjuntos de parámetros (S1, S2,...), de tal manera que cada conjunto de parámetros está asociado con una zona de frecuencias correspondiente (FR1, FR2,...).
- 4Un método de codificación de acuerdo con la reivindicación 3, caracterizado porque los conjuntos de parámetros comprenden al menos una indicación de localización (ILD, ITD, IPD, IC).
- 5Un método de codificación de acuerdo con la reivindicación 4, caracterizado porque la al menos una indicación de localización (ILD, ITD, IPD, IC) se selecciona de entre:una diferencia de niveles inter-auditivos o entre los dos oídos (ILD -“interaural level difference”), una diferencia de tiempos o de fases inter-auditivas, o entre los dos oídos (ITD -“interaural time difference”-, IPD -“interaural phase difference”), o una correlación transversal inter-auditiva, o entre los dos oídos (IC -“interaural cross-correlation”).
- 6Un método de codificación de acuerdo con la reivindicación 1 ó la reivindicación 2, caracterizado porque la primera zona de frecuencias (FR1) cubre una anchura banda completa (FBW -“full bandwidth”) de la señal de audio de múltiples canales.
- 7Un método de codificación de acuerdo con la reivindicación 1, caracterizado porque la primera zona de frecuencias (FR1) cubre sustancialmente una anchura de banda completa (FBW) de la señal de audio de múltiples canales, la segunda zona de frecuencias (FR2) cubre una porción de la anchura de banda completa (FBW), y por que determinar (2) la segunda porción de la información (P2) está destinada a determinar conjuntos de parámetros (S2, S3, ...) tanto para la segunda zona de frecuencias (FR2) como para el conjunto de zonas de frecuencias adicionales (FR3, FR4, FR5), de tal manera que la segunda zona de frecuencias (FR2) y el conjunto de zonas de frecuencias adicionales (FR3, FR4, FR5) cubren sustancialmente la anchura de banda completa (FBW), donde el conjunto de zonas de frecuencias adicionales (FR3, FR4, FR5) comprende al menos una zona de frecuencias adicional (FR3).
- 8Un método de codificación de acuerdo con la reivindicación 7, caracterizado porque la señal de audio de canal único (SC) y la primera porción (P1) de la información (INF) forman una capa de base de información (BL -“base layer”) que está siempre presente en la señal de audio de múltiples canales codificada (EBS), y porque el método comprende recibir (2) una velocidad de bits máxima permisible (MBR -“maximum bit rate”) de la señal de audio de múltiples canales codificada (EBS), de tal modo que la segunda porción de la información (P2) forma una capa de mejora de información (EL -“enhancement layer”) que es codificada únicamente si la velocidad de bits de la capa de base codificada (DL) y de la capa de mejora (EL) no es más alta que la velocidad de bits máxima permisible (MBR). ES 2 268 340 T3
- 9Un método de codificación de acuerdo con la reivindicación 3, caracterizado porque determinar (2) la primera porción de información (P1) en una trama particular (F2) de información codificada (EiN) comprende determinar (2) el primero de los conjuntos de parámetros (S1') contenido en la trama particular (F2), y codificar el primero de los conjuntos de parámetros (S1') basándose en el primero de los conjuntos de parámetros (S1) de una trama (F1) que precede a la trama particular (F2).
- 10Un método de codificación de acuerdo con la reivindicación 7, caracterizado porque determinar (2) la segunda porción de información (P2) contenida en una trama particular (F2) de la información codificada (EIN) comprende determinar (2) los conjuntos de parámetros (S2', S3', ...) de la segunda porción (P2) contenida en la trama particular (F2), y codificar los conjuntos de parámetros (S2', S3',...) de la segunda porción (P2) contenida en la trama particular (F2) basándose en los conjuntos de parámetros (S2, S3,...) de una trama (F1) que precede a la trama particular (F2).
- 11Un método de codificación de acuerdo con la reivindicación 7, caracterizado porque determinar (2) la segunda porción de información (P2) contenida en una trama particular (F2) de la información codificada (EIN) comprende determinar (2) los conjuntos de parámetros (S2', S3', ...) de la segunda porción (P2) contenida en la trama particular (F2), y codificar los conjuntos de parámetros (S2', S3',...) de la segunda porción (P2) contenida en la trama particular (F2) basándose en el primero de los conjuntos de parámetros (S1) de una trama (F1) que precede a la trama particular (F2).
- 12El método de codificación de acuerdo con una cualquiera de las reivindicaciones 9 a 11, caracterizado porque determinar (2) comprende calcular una diferencia entre los parámetros correspondientes de la trama particular (F2) y de la trama (F1) que precede a la trama particular (F2).
- 13Un codificador para codificar una señal de audio de múltiples canales que comprende al menos canales de audio (RI, LI), de tal modo que el codificador comprende:medios para generar (1) una señal de audio (SC -“single channel”) de un único canal, que comprende una combinación particular de los al menos dos canales de audio (RI, LI), medios para generar (2) información (INF) a partir de los al menos dos canales de audio (RI, LI), que permite recuperar, con un nivel de calidad requerido, la señal de audio de múltiples canales a partir de la señal de audio de canal único (SC) y de la información (INF), de tal modo que los medios para generar (2) la información comprenden: - medios para determinar (2) una primera porción de la información (P1), que consiste en un único conjunto de parámetros (S1), determinados para una primera zona de frecuencias (FR1) de la señal de audio de múltiples canales, y - medios para determinar (2) una segunda porción de la información (P2) para una segunda zona de frecuencias (FR2) de la señal de audio de múltiples canales, de tal modo que la segunda zona de frecuencias (FR2) es una porción de la primera zona de frecuencias (FR1).
- 14Un codificador para codificar una señal de audio de múltiples canales, de acuerdo con la reivindicación 13, que comprende adicionalmente medios para determinar (2) únicamente la segunda porción de la información (P2) para la segunda zona de frecuencias (FR2) de la señal de audio de múltiples canales, en el caso de que una velocidad de bits de la señal de audio de múltiples canales codificada, que comprende la señal de audio de canal único (SC), la primera porción de la información (P1) y la segunda porción de la información (P2), no sea superior a una velocidad de bits máxima permisible (MBR -“maximum bit rate”).
- 15Un aparato para suministrar una señal de audio, de tal modo que el aparato comprende:una entrada para recibir una señal de audio de múltiples canales, un codificador de acuerdo con la reivindicación 13 ó la reivindicación 14, destinado a codificar la señal de audio de múltiples canales con el fin de obtener una señal de audio de múltiples canales codificada, y una salida para suministrar la señal de audio de múltiples canales codificada.
- 16Una señal de audio de múltiples canales codificada, que comprende:una señal de audio (SC -“single channel”) de un único canal, que comprende una combinación particular de al menos dos canales de audio (RI, LI), e información (INF) procedente de los al menos dos canales de audio (RI, LI), lo que permite recuperar, con un nivel de calidad requerido, la señal de audio de múltiples canales a partir de la señal de audio de canal único (SC), y de la información (INF), de tal modo que la información comprende: - una primera porción de la información (P1), que consiste en un único conjunto de parámetros (S1) determinados para una primera zona de frecuencias (FR1) de la señal de audio de múltiples canales, y ES 2 268 340 T3 - una segunda porción de la información (P2) para una segunda zona de frecuencias (FR2) de la señal de audio de múltiples canales, de tal modo que la segunda zona de frecuencias (FR2) es una porción de la primera zona de frecuencias (FR1).
- 17Un medio de almacenamiento en el que se ha almacenado la señal de audio codificada de acuerdo con la reivindicación 16.
- 18Un método de descodificación de una señal de audio de múltiples canales codificada que se ha codificado de acuerdo con la reivindicación 16, de tal modo que el método de descodificación comprende:obtener (6, 7) una señal de audio de un único canal descodificada (SCO), que comprende una combinación particular de los al menos dos canales de audio (RI, LI), obtener (6, 8) información descodificada (INO) a partir de la información (INF), lo que permite recuperar la señal de audio de múltiples canales a partir de la señal de audio de canal único descodificada (SCO) y de la información descodificada (INO), de tal modo que la información descodificada (INO) comprende la primera porción de la información (P1) y la segunda porción de la información (P2), y aplicar (9), bien la primera porción de la información (P1) o bien la primera porción (P1) y la segunda porción de la información (P2) en la señal de audio de canal único (SCO) con el fin de generar una señal de audio de múltiples canales descodificada (LO, RO).
- 19Un descodificador para descodificar una señal de audio de múltiples canales codificada, la cual ha sido codificada de acuerdo con la reivindicación 16, de tal modo que el descodificador comprende:medios para obtener (6, 7) una señal de audio de un único canal descodificada (SCO), que comprende una combinación particular de los al menos dos canales de audio (RI, LI), medios para obtener (6, 8) información descodificada (INO) a partir de la información (INF), lo que permite recuperar la señal de audio de múltiples canales a partir de la señal de audio de canal único descodificada (SCO) y de la información descodificada (INO), de tal modo que la información descodificada (INO) comprende la primera porción de la información (P1) y la segunda porción de la información (P2), y medios para aplicar (9) la primera porción de la información (P1) y la segunda porción de la información (P2) en la señal de audio de canal único (SCO) con el fin de generar una señal de audio de múltiples canales descodificada (LO, RO).
- 20Un aparato para suministrar una señal de audio descodificada, de tal modo que el aparato comprende:una entrada para recibir una señal de audio de múltiples canales codificada, un descodificador de acuerdo con la reivindicación 19, destinado a descodificar la señal de audio de múltiples canales codificada, con el fin de obtener una señal de salida de múltiples canales, y una salida para suministrar o reproducir la señal de salida de múltiples canales.
Independent claims20
93 paragraphs in 7 sections, as filed
ES 2 268 340 T3
DESCRIPTION
Multi-channel parametric audio representation.
The invention relates to a method for encoding a multi-channel audio signal, to an encoder for encoding a multi-channel audio signal, to an apparatus for supplying an audio signal, to an encoded audio signal, to a medium. storage in which the encoded audio signal is stored, to a method for decoding an encoded audio signal, to a decoder for decoding an encoded audio signal, and to an apparatus for supplying a decoded audio signal.
EP-A-1107232 describes a parametric coding scheme for generating a representation of a stereo audio signal that is composed of a left channel signal and a right channel signal. In order to efficiently use the transmission bandwidth, this representation contains information concerning only a mono-auditory signal, or for a single ear, which is either the left channel signal or the right channel signal, and parametric information. The other stereo signal can be recovered based on the mono-auditory signal, together with the parametric information. The parametric information comprises location indications of the stereo audio signal, including intensity and phase characteristics of the left channel and the right channel.
The publication “Subband Coding of Stereophonic Digital Audio Signals”, by R. van der Waal, R. Veldhuis, Philips Reserch Laboratories, at the IEEE (Institute of Electrical Engineering and Electronics), 1991, vol. 2, pages 3.601-3.604 (ISBN: 0-7803-0003-3), describes a sub-band coding algorithm. In such sub-band coding algorithms, the frequency spectrum to be coded is divided into sub-bands that do not overlap. Encoding is done for each sub-band. Subband coding includes a rotational transformation.
Previous solutions that have been suggested in audio encoders to reduce the bit rate of stereo program material include intensity stereo and M / S stereo.
In the intensity stereo algorithm, high frequencies (typically above 5 kHz) are represented by a single audio signal (i.e. mono), combined with time-varying and dependent scale factors or intensity factors frequency, allowing you to recover a decoded audio signal that resembles the original stereo signal for these frequency zones. In the M / S algorithm, the signal is decomposed into a sum (or mean, or common) signal and a difference (or side, or uncommon) signal. This decomposition is sometimes combined with principle component analysis or with time-varying scale factors. These signals are then independently encoded, either by a transform encoder or by a subband encoder [both of which are waveform or waveform encoders]. The amount or magnitude of the information reduction that is achieved by this algorithm strongly depends on the spatial properties of the source signal. For example, if the source signal is mono-auditory, the difference signal is zero and can be discarded. However, if the correlation between the left and right audio signals is low (which is often the case for lower frequency areas), this scheme offers only a small bit rate reduction. For low-frequency areas, M / S coding generally provides significant merit.
Parametric descriptions of audio signals have been gaining interest in recent years, especially in the field of audio coding. It has been shown that the transmission of (quantized) parameters that describe audio signals requires only a small transmission capacity to re-synthesize a perceptually equal signal at the receiving end or terminal. However, current parametric audio encoders concentrate on encoding mono-auditory signals, and stereo signals are processed or treated as double mono signals.
It is an object of the invention to provide a multi-channel parametric audio system that is capable of scaling the quality of the encoded audio signal with the available bit rate, or of scaling the quality of the decoded audio signal. , with the complexity of the decoder or the available transmission bandwidth.
A first aspect of the invention provides a method for encoding a multi-channel audio signal, as claimed in claim 1. A second aspect of the invention provides an encoder for encoding a multi-channel audio signal, as claimed in claim 13. A third aspect of the invention provides an encoded audio signal as claimed in claim 16. A fourth aspect of the invention provides a storage medium in which the encoded signal is stored, and is claimed in claim 17. A fifth aspect of the invention provides a decoding method, as claimed in claim 18. A sixth Aspect of the invention provides a decoder for decoding an encoded audio signal, as claimed in claim 19. Advantageous embodiments are defined in the dependent claims.
In the method of encoding a multi-channel audio signal, according to the first aspect of the invention, a single-channel audio signal is generated. On the other hand, information is generated from the signal of the multi-channel audio signal, which allows the recovery, with a required level of quality, of the signal.
ES 2 268 340 T3 multi-channel audio from the single channel audio signal and information. Preferably, the information comprises sets of parameters, for example, as known from EP-A-1107232.
According to the first aspect of the invention, the information is generated by determining a first portion of the information for a first frequency area of the multi-channel audio signal, and determining a second portion of the information for a second frequency area. multi-channel audio signal. The second frequency zone is a portion of the first frequency zone and therefore constitutes a sub-interval or interval included in the first frequency zone. Now, two levels of quality in decoding are possible. For a low quality level of the decoded multi-channel audio signal, the decoder uses the encoded single channel audio signal, and the first portion of the information. For a higher quality level, the decoder uses the encoded single channel audio signal and both the first and second portions of the information. Of course, it is possible to select the decoding quality from a multiplicity of levels, if a multiplicity of information portions are present in such a way that each of them is associated with a different frequency range. For example, the first portion may comprise a single, determined set of parameters, with a frequency region covering the entire bandwidth of the multi-channel audio signal. And the second portion may comprise various sets of parameters, such that each set of parameters is determined by a sub-range or portion of the entire bandwidth. Together, the portions preferably cover the entire bandwidth.
This representation of the encoded audio signal allows the quality of the decoded audio signal to depend on the complexity of the decoder. For example, in a simple portable decoder a low complexity decoder can be used which has a low power consumption and is consequently capable of using only a part of the information. In a top-of-the-range application, a complex decoder is used that makes use of all the information available in the coded signal.
The quality of the decoded audio may also depend on the available transmission bandwidth. If the transmission bandwidth is high, then the decoder can decode all the available layers, since they are all transmitted. If the transmission bandwidth is low, then the transmitter may decide to transmit only a limited number of layers.
In one embodiment as defined in claim 2, the encoder receives a maximum allowable bit rate of the encoded multi-channel audio signal. This maximum allowable bit rate can be defined by the available bit rate of a transmission channel such as the Internet, or of a storage medium. In applications where the transmission bandwidth is variable and therefore the maximum allowable bit rate changes over time, it is important to be able to adapt to these fluctuations in the transmission bandwidth in order to avoid a very low quality of the decoded audio signal. Normally, the encoder encodes all available layers. It is decided in the transmitting terminal which layers are to be transmitted, depending on the capacity of the available channels. It is possible to do this with the encoder in the loop, but this is more complicated than separating or detaching some layers before transmission.
The encoder adds only the second portion of the information for the second frequency zone of the multi-channel audio signal to the encoded audio signal, in the case where a bit rate of the multi-channel audio signal encoded, comprising the single channel audio signal, and the first and second portions of the information are not higher than the maximum allowable bit rate. In this way, the second portion is not present in the encoded audio signal if the transmission bandwidth is not large enough to support the transmission of the second portion.
In one embodiment as defined in claim 3, the information comprises sets of parameters, such that each portion of the information is represented by one or more sets of parameters. The number of parameter sets depends on the number of frequency zones present in the information portions.
In one embodiment as defined in claim 4, the parameter sets comprise at least one of the location indications.
In one embodiment as defined in claim 6, the first frequency zone covers substantially the entire bandwidth of the multi-channel audio signal. In this way, a set of parameters is sufficient to provide the basic information that is required to decode the single channel audio signal into the multi-channel audio signal. In this way, a basic level of quality of the audio signal is guaranteed. The second frequency range covers part of the entire bandwidth. Thus, the second portion, when present in the encoded audio signal, improves the quality of the decoded audio signal in this frequency range.
In one embodiment as defined in claim 7, the second portion of the information comprises at least two frequency ranges that together cover substantially the entire bandwidth of the multi-channel audio signal. In this way, the improvement in quality provided by the second portion is present throughout the entire bandwidth.
ES 2 268 340 T3
In one embodiment as defined in claim 8, the base layer comprising the single channel audio signal and the first portion of the information is always present in the encoded audio signal. The enhancement layer comprising the second portion of the information is encoded only if the bit rate of the second audio signal does not exceed the maximum allowable bit rate. In this way, the quality of the decoded audio signal will depend on the maximum allowable bit rate. If the maximum allowable bit rate is too low to accommodate the enhancement layer, the decoded audio signal will be sourced from the base layer, resulting in a better quality of the decoded audio than will occur in the event that unpredictable parts of the encoded audio do not reach the decoder.
In the embodiments as defined in any one of claims 9-11, the portions of the information (usually containing sets of parameters, one set for each frequency band represented) contained in a subsequent frame are encoded based on the parameters of the previous plot. Typically this reduces the bit rate of the encoded portions of the information, because as a consequence of correlation the information contained in two successive frames will not differ substantially.
In embodiments as defined in claim 12, the difference between the parameters of two successive frames is encoded rather than the parameters themselves.
These and other aspects of the invention will become apparent from the embodiments described below, and will be clarified with reference thereto.
In the drawings:
Figure 1 shows a block diagram of a multi-channel encoder for stereo audio, Figure 2 shows a block diagram of a multi-channel decoder for stereo audio, Figure 3 shows a representation of the encoded data stream, the Figure 4 illustrates one embodiment of the frequency ranges according to the invention, Figure 5 shows another embodiment of the frequency ranges according to the invention, Figure 6 illustrates the determination of the parameter sets based on parameters of a previous frame, according to an embodiment of the invention, Figure 7 shows a set of parameters, Figure 8 shows the differential determination of the layer parameters base, and Figure 9 illustrates the differential determination of the parameters corresponding to a frequency zone of an enhancement layer.
Figure 1 shows a block diagram of a multi-channel encoder. The encoder receives a multi-channel audio signal which is displayed as RI, LI stereo signal, the encoder supplies the EBS encoded multi-channel audio signal.
The downstream mixer 1 combines the stereo signal or the stereo channels RI, LI into a single channel audio signal (also referred to as a mono-audio signal) SC. For example, the downstream mixer 1 can determine the average of the input audio signals RI, LI.
Encoder 2 encodes the SC mono-auditory signal to obtain an ESC encoded mono-auditory signal. Encoder 3 may be of a known type, for example an MPEG encoder (MPEG-LII, MPEG-LIII (mp3), or MPEG2-AAC).
The parameter determining circuit 2 determines the parameter sets S1, S2, ... that characterize the information INF, based on the input audio signals RI, LI. Optionally, the parameter determination circuit 2 receives the maximum allowable bit rate MBR ("maximum bit rate") in order to determine only the sets of parameters S1, S2, ..., which, once encoded by the encoder 4 parameters, together with the encoded mono-auditory signal ESC, do not exceed the maximum allowable MBR bit rate. Coded parameters are denoted by EIN.
The formatting device 5 combines the SC ("single channel") encoded mono-auditory signal and the EIN encoded parameters into a data stream of a desired format, in order to obtain the EBS encoded multi-channel audio signal.
The operation of the encoder is clarified in greater detail below, by way of example, with respect to one embodiment. LI, RI multi-channel audio signal is encoded into a single mono-audi4 signal
ES 2 268 340 T3 tiva SC (also further referred to as a single channel audio signal). The parameterization or quantization in parameters of spatial attributes of the multi-channel audio signals LI, RI is carried out by the parameter determination circuit 2. The parameters contain information about how to restore or restore the multi-channel audio signal LI, RI from the mono-auditory signal SC. The parameters are usually encoded by the parameter encoder 4, before combining them with the encoded single channel ESC ("encoded single channel"). In this way, for general audio coding applications, these parameters are transmitted or stored, combined with a single mono-auditory audio signal. The encoded and combined signal is the EBS encoded multi-channel audio signal. The transmission or storage capacity required to transmit or store the EBS encoded multi-channel audio signal is greatly reduced compared to audio encoders that independently process or handle the multiple channels. However, the original spatial impression is maintained by the INF information, which contains the (sets of) parameters.
In particular, the parametric description of multi-channel audio RI, LI is related to a bi-auditory (or two-ear) processing model that is aimed at describing the effective signal processing of the two-ear auditory system.
The model divides the incoming audio LI, RI into several band-limited signals, which are preferably linearly separated on an ERB rate scale. The bandwidth of these signals depends on the center frequency, following the ERB rate. Subsequently, the following properties of the incoming signals are preferably analyzed for each frequency band:
- the inter-auditory or inter-ear level difference, or ILD (“interaural level difference”), defined by the relative levels of the band-limited signal originating from the left and right ears,
- the inter-auditory or inter-ear time (or phase) difference, ITD (“interaural time difference”) (or IPD “interaural phase difference”), defined by the inter-ear delay (or phase shift) corresponding to the peak of the cross-ear correlation function, and
- the similarity (dissimilarity) of the waveforms that is not attributable to the ITDs or the ILDs, which can be quantified as a parameter by means of the maximum transversal correlation between ears, IC (for example, the value of the transversal correlation at the position of the maximum peak).
The sets S1, S2, ... of the three parameters, once established for each frequency band FR1, FR2, ..., vary over time. However, since the two-ear auditory system is very slow to process, the update rate for these properties is quite low (typically tens of milliseconds).
It can be assumed that the parameters that vary (slowly) with time are the only spatial signal properties available to the two-ear auditory system, and that, from these time- and frequency-dependent parameters, the auditory world Perceived is reconstructed by the higher levels of the auditory system.
Figure 2 shows a block diagram of a multi-channel decoder. The decoder receives the EBS encoded multi-channel audio signal and supplies the decoded multi-channel audio signal that it has recovered, which is displayed as an RO, LO stereo signal.
The format suppression device 6 recovers the encoded mono-auditory signal ESC 'and the encoded parameters EIN' from the EBS data stream. The decoder 7 decodes the encoded mono-auditory signal ESC 'to obtain the output mono-auditory signal SCO. The decoder 7 can be of any known type (of course, in correspondence with the encoder that has been used); for example, decoder 7 is an MPEG decoder. Decoder 8 decodes the encoded parameters EIN 'to obtain output parameters INO.
The demultiplexer 9 recovers the LO and RO output stereo audio signals by applying the parameter sets S1, S2, ... of the INO output parameters to the SCO output mono-audio signal.
Figure 3 shows a representation of the encoded data stream. For example, in each frame F1, F2, ..., the data packet begins with a header H, followed by the encoded mono-auditory signal ECS, now indicated by A, a first portion P1 of the encoded information EIN, a second portion P2 of the encoded information EIN, and a third portion P3 of the encoded information EIN.
If the frame F1, F2, ... comprises only the header H and the encoded mono-auditory signal ECS, only the mono-auditory signal SC is transmitted.
As described in EP-A-1107232, the entire frequency band in which the input audio signal takes place is divided into a plurality of frequency sub-bands, which together cover the band full frequency. In the terminology according to the invention, the multi-channel INF information is encoded in a plurality of parameter sets S1, S2, ..., one set for each frequency sub-band FR1, FR2, ... This plurality of parameter sets S1, S2, ... is encoded in the first portion P1
ES 2 268 340 T3 of the EIN encoded information. Thus, in order to transmit an entry-level quality multi-channel audio signal, the bit stream comprises the header H, the portion A, which is the encoded mono-audio signal, and the first portion P1.
In the bit stream according to one embodiment of the invention, the first portion P1 consists only of a single set of parameters S1. The unique set is determined for the full bandwidth FR1. This data stream, comprising header H and portions A and P1, provides a basic layer of quality, indicated by BL in Figure 3.
In order to support improved quality, additional P2, P3 portions of the EIN encoded information are present in the data stream. These additional portions form an EL enhancement layer. The bit stream may comprise a single additional portion P2 or more than 1 additional portion. The additional portion P2 preferably comprises a plurality of parameter sets S2, S3, ..., one set for each frequency sub-band FR2, FR3, ..., such that the frequency sub-bands FR2 , FR3 preferably cover the entire frequency band FR1. The improved quality may also be present in a step-by-step fashion, so that a first level of improvement is provided by the enhancement layer EL1, which comprises the first portion. And a second enhancement layer EL comprises the first enhancement layer EL1 and the second enhancement layer EL2, which comprises the portion P3.
The additional portion P2 may also comprise a single set S2 of parameters corresponding to a single frequency band FR2, which is a sub-band of the entire frequency band FR1. The additional portion P2 can also comprise a certain number of sets of parameters S2, S3, ... corresponding to the frequency bands FR2, FR3, ... which do not, together, cover the entire frequency band FR1.
The additional portion P3 preferably contains sets of parameters for frequency bands that subdivide at least one of the sub-bands of the additional portion P2.
This format of the bit stream according to the invention makes it possible to scale, in the transmission channel or in the decoder, the quality of the decoded audio signal, with the bit rate of the transmission channel, or with the complexity decoder decoder. For example, if the audio decoder is to have low power consumption, as is important in portable applications, the decoder may have low complexity and uses only the H, A and P1 portions. It would even be possible for the decoder to be able to carry out more complex operations with a higher power consumption, in the event that the user indicates that he wants a higher quality of the decoded audio.
It is also possible that the decoder is aware of the maximum allowable bit rate, MBR, that can be transmitted through the transmission channel or that can be stored on a storage medium. Now the encoder is able to decide how many additional portions P1, P2, ..., if any, fit within the maximum allowable MBR bit rate. The encoder encodes only these allowable portions P1, P2, ... of the bit stream.
Figure 4 shows an embodiment of the frequency ranges according to the invention. In this embodiment, the frequency band FR1 is equal to the full frequency band FBW ("full bandwidth") of the multi-channel audio signal LI, RI, and the frequency band FR2 is a sub-band of frequencies of the full bandwidth FBW.
If these are the only frequency ranges for which the parameter sets S1, S2, ... are determined, a single parameter set S1 is determined for the frequency band FR1 and is present in the portion P1, and it is determined a single set of parameters S2 for the frequency band FR2, and is present in the portion P2. Quality scale regulation is possible, either by using the P2 portion or by not using it.
Figure 5 shows another embodiment of the frequency ranges according to the invention. In this embodiment, the frequency band FR1 is again equal to the full bandwidth FBW, and the frequency sub-bands FR2 and FR3 together cover the full bandwidth FBW. Or, in other words, the frequency band FR1 is subdivided into the frequency sub-bands FR2 and FR3.
In the case that these are the only frequency ranges for which the parameter sets S1, S2, ... are determined, the portion P1 comprises a single parameter set S1, determined by the frequency band FR1, and the portion P2 comprises two sets of parameters S2 and S3, determined, respectively, by the frequency bands FR2 and FR3. Quality scaling is possible both by using the P2 portion and by not using it.
Figure 6 shows the determination of parameter sets based on parameters contained in a previous frame, according to an embodiment of the invention.
Figure 6 shows a data stream comprising, in each frame F1, F2, ..., the encoded information EIN, comprising the portion P1, which is a part of the base layer BL, and the portion P2, which forms the EL enhancement layer.
ES 2 268 340 T3
In frame F1, portion P1 comprises a single set of parameters S1 that are determined for the entire bandwidth FR1. The portion P2, by way of example, comprises four sets of parameters S2, S3, S4, S5 which are determined, respectively, for the frequency sub-bands FR2, FR3, FR4, FR5. The four frequency sub-bands FR2, FR3, FR4, FR5 sub-divide the frequency band FR1.
In frame F2, succeeding frame F1, portion P1 comprises a single set of parameters S1 'which are determined for the full bandwidth FR1 and form part of the base layer BL'. The portion P2 comprises four sets of parameters S2 ', S3', S4 ', S5' which are again determined, respectively, for the frequency subbands FR2, FR3, FR4, FR5 and which form the enhancement layer EL '.
It is possible to encode each of these sets of parameters S1, S2, ... for each of the frames F1, F2, ... separately. It is also possible to encode the parameter sets of the portion P2 with respect to the parameters of the portion P1. This is indicated by the arrows starting at S1 and ending at S2 through S5, in frame F1. Of course, this is also possible in other F2 frames, ... (not shown). In the same way, it is possible to encode the set of parameters S1 'with respect to S1. And finally, the parameter sets S2 ', S3', S4 ', S5' can be encoded with respect to the parameter sets S2, S3, S4, S5.
In this way, the bit rate of the EIN encoded information can be reduced as redundancy or correlation between sets of Si parameters is used.
Preferably, the new parameters of the new parameter sets S1 ', S2', S3 ', S4', S5 'are encoded as the difference between their value and the value of the parameters of the previous parameter sets S1, S2, S3 , S4, S5.
At uniform time intervals, at least the set of parameters S1 has to be encoded in an absolute and non-differential way, in order to prevent the errors from propagating too far.
Figure 7 shows a set of parameters. Each set of Si parameters can comprise one or more parameters. Typically, the parameters are location cues that provide information about the location of sound objects in the audio information. Typically, localization indications consist of the inter-auditory level difference, or between ears, ILD, the inter-auditory time or inter-auditory phase difference, ITD or IPD, and the inter-auditory cross-sectional correlation. auditory, or between ears, IC (“interaural cross-correlation”). More detailed information about these parameters is provided in Audio Engineering Society Convention Paper 5574, “Coding of bi-auditory, or two-ear, cues applied to stereo and multi-channel audio compression ”(“ Binaural Cue Coding Applied to Stereo and Multi-channel Audio Compression ”), presented at the 112<sup>to</sup> Convention, May 10-13, 2002 in Munich, Germany, by Christof Faller et al.
Figure 8 shows the differential determination of a base layer parameter. The horizontal axis indicates successive frames F1 to F5. The vertical axis shows the PVG value of a parameter from the S1 parameter set of the base layer BL (“base layer”). This parameter has the values A1 through A5 for frames F1 through F5, respectively. The contribution of this parameter to the bit rate of the EIN encoded information will decrease if the actual values A1 to A5 of the parameter are not encoded, but rather the smaller differences, D1, D2, ...
Figure 9 shows the differential determination of the parameters corresponding to a frequency zone of an enhancement layer. The horizontal axis indicates two successive frames F1 and F2. The vertical axis indicates the values of a particular parameter of the base layer BL and the enhancement layer EL. In this example, the base layer BL comprises the portion P1 of information INF with a single set of parameters, determined for the entire frequency range FBW, such that the particular parameter of the portion P1 has the value A1 for the frame F1 and A2 for frame F2. The enhancement layer EL comprises the information portion P2 INF with three sets of parameters determined for three respective frequency ranges FR2, FR3, FR4 which, together, fill the entire frequency range FBW. The three particular parameters (eg the parameter representing the ILD) have a value B11, b12, B13 in frame F1 and a value B21, B22, B23 in frame F2.
The contribution of these parameters to the bit rate of the encoded information EIN will be reduced if the true values B11 to B23 of the particular parameter are not encoded, but the differences D11, D12, ..., because these differences can be encoded more effectively than true values.
In summary, in a preferred embodiment according to the invention, it is proposed to organize the stereo parameter information INF in such a way that a base layer BL contains one of the parameter sets (preferably, the time / level difference and the correlation ) S1, which is determined for the full bandwidth FBW of the multi-channel audio signal LI, RI. The EL enhancement layer contains multiple sets of parameters S2, S3, ... corresponding to subsequent frequency intervals FR2, FR3 within the full bandwidth FBW. For the sake of bit rate efficiency, the parameter sets S2, S3, ... of the enhancement layer EL can be differentially encoded with respect to the parameter set S1 located in the base layer BL .
ES 2 268 340 T3
The INF information is encoded in a multilayer structured manner to allow scaling of decoding quality versus bit rate.
To conclude, in what follows below, a preferred embodiment according to the invention is elucidated, with respect to a program code and its explanation or clarification.
First of all, for all the subordinate frames or sub-frames (the portions P1, P2, ...) contained in the frames F1, F2, ..., the ESC data for the mono-auditory representation, or of single ear, SC, the EIN data for the stereo parameter set S1 for the full bandwidth FBW, and the stereo parameters S2, S3, ... for the frequency containers (or regions) FR2, FR3, .. .
The program code is shown on the left side, and a clarification of the program code being described is provided on the right side.
Code Description for (f = 0; f> nrof_frames; f + +) for all frames, do:
example_mono_frame (f) get data for monaural or single-ear signal representation (portion A of Figure 3) example_stereo_extension_layer_l (f) get full bandwidth of stereo parameters of data (the Pl portion) example_stereo_extension_layer_2 (f) get frequency containers of stereo data parameters (the P2 portion)
Second, depending on the value of the refresh_stereo bit, the stereo parameters for the full bandwidth are encoded absolutely (the true or true value is encoded), or the difference with the previous values is encoded. The following code is valid for the difference of inter-auditory levels, or between both ears, ILD.
Code example_stereo_extension_layer_l (f)
Description refresh stereo bit that denotes if the data has to be encoded in an absolute way or not if (refresh_stereo = l) if the data has to be encoded in an absolute way
ES 2 268 340 T3
<img file="ES2268340T3_D0001.tif" />
Third, depending on the value of the refresh_stereo bit, the stereo parameters for all frequency containers are encoded in an absolute way (the real or true value is encoded), or the difference is encoded with the corresponding parameters for the bandwidth complete. The following code is valid for the difference of inter-auditory levels, or between the two ears, ILD.
Code example_stereo_extensión_layer_2 (f) description if (refescar_stereo = l) if there is refreshment for (b = 0; b <nr of containers; b ++) for all frequency containers ild_container [f, b] encode the ild contained in that container with respect to the value global otherwise {
for (b = 0; b <nrof_containers; b ++) {ild_contenedor_dif [f, b] if there is no refresh for all containers, encode the ild inside a concrete container, with respect to the value inside that container from the previous frame
Where:
The expression "refresh_stereo" is an indicator that denotes whether or not the stereo parameters are to be refreshed (0: FALSE, 1 = TRUE).
ES 2 268 340 T3
The expression "ild_global [sf]" represents the Huffman-encoded absolute rendering level of the ILD for the entire frequency area for frame f.
The expression "ild_global_dif [f]" represents the Huffman-encoded relative rendering level of the ILD for the entire frequency area for frame f.
The expression "ild_container [f, b]" represents the Huffman encoded absolute rendering level of the ILD for frame f and container b.
The expression "ild_container_dif [f, b]" represents the relative Huffman-encoded rendering level of the ILD for frame f and container b.
It is to be appreciated that the aforementioned embodiments illustrate the invention rather than limit it, and that those skilled in the art will be able to design many alternative embodiments without departing from the scope of the accompanying claims.
Although the invention has been elucidated in the figures in relation to a stereo signal, the extension to an audio signal of more than two channels can be easily carried out by the skilled person.
In the claims, any reference symbols placed in parentheses are not to be construed as limiting the claim. The expression "comprising" does not exclude the presence of elements or steps other than those listed in a claim. The invention can be practiced by means of physical devices or hardware comprising several different elements, and by means of a suitably programmed computer. In the device claim that lists various means, several of these means can be realized by means of the same hardware element. The mere fact that certain measures are mentioned in the dependent claims distances from each other does not indicate that a combination of these measures cannot be used to advantage.
In sum, multi-channel audio signals are encoded into a mono-auditory, or single-ear, audio signal and information, allowing the multi-channel audio signal to be recovered from the mono audio signal. -hearing and information. The information is generated by determining a first portion of the information for a first frequency region of the multi-channel audio signal, and determining a second portion of the information for a second frequency region of the multi-channel audio signal. The second frequency zone is a portion of the first frequency zone and therefore constitutes a sub-interval of the first frequency zone. The information is structured in multiple layers, allowing for scaling of decoding quality versus bit rate.
Contents7
4 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4
83 members in 11 offices
Priority claims10
| Document | Office | Kind | Date |
|---|---|---|---|
| 02076588 | European Patent Office (EPO) | A | |
| 02076588 | European Patent Office (EPO) | A | |
| 20020076588 | European Patent Office (EPO) | – | |
| 02077869 | European Patent Office (EPO) | A | |
| 02077869 | European Patent Office (EPO) | A | |
| 20020077869 | European Patent Office (EPO) | – | |
| 02077869 | – | – | – |
| 0371259702076588 | – | – | – |
| EP20020076588 | – | – | – |
| EP20020077869 | – | – | – |
Members83
| Document | Office | Kind | |
|---|---|---|---|
| WO03090206A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO03090207A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO03090208A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU2003216682A1 | Australia | A1 | |
| AU2003216686A1 | Australia | A1 | |
| AU2003219426A1 | Australia | A1 | |
| WO2004036549A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU2003219428A1 | Australia | A1 | |
| BR0304540A | Brazil | A | |
| BR0304541A | Brazil | A | |
| BR0304542A | Brazil | A | |
| KR20040101552A | Republic of Korea | A | |
| KR20040102163A | Republic of Korea | A | |
| KR20040102164A | Republic of Korea | A | |
| EP1500082A1 | European Patent Office (EPO) | A1 | |
| EP1500083A1 | European Patent Office (EPO) | A1 | |
| EP1500084A1 | European Patent Office (EPO) | A1 | |
| KR20050049549A | Republic of Korea | A | |
| EP1554716A1 | European Patent Office (EPO) | A1 | |
| CN1647155A | China | A | |
| CN1647156A | China | A | |
| CN1647157A | China | A | |
| JP2005523479A | Japan | A | |
| JP2005523480A | Japan | A | |
| JP2005523624A | Japan | A | |
| US2005226426A1 | United States of America | A1 | |
| CN1689070A | China | A | |
| US2005254446A1 | United States of America | A1 | |
| JP2006503319A | Japan | A | |
| US2006100861A1 | United States of America | A1 | |
| EP1500083B1 | European Patent Office (EPO) | B1 | |
| AT332003T | Austria | T | |
| ATE332003T1 | Austria | T1 | |
| DE60306512D1 | Germany | D1 | |
| EP1500082B1 | European Patent Office (EPO) | B1 | |
| AT354161T | Austria | T | |
| ATE354161T1 | Austria | T1 | |
| ES2268340T3This record | Spain | T3 | |
| CN1307612C | China | C | |
| DE60311794D1 | Germany | D1 | |
| CN1312660C | China | C | |
| DE60306512T2 | Germany | T2 | |
| ES2280736T3 | Spain | T3 | |
| DE60311794T2 | Germany | T2 | |
| EP1500084B1 | European Patent Office (EPO) | B1 | |
| EP1881486A1 | European Patent Office (EPO) | A1 | |
| AT385025T | Austria | T | |
| ATE385025T1 | Austria | T1 | |
| DE60318835D1 | Germany | D1 | |
| ES2300567T3 | Spain | T3 | |
| US2008170711A1 | United States of America | A1 | |
| DE60318835T2 | Germany | T2 | |
| EP1881486B1 | European Patent Office (EPO) | B1 | |
| AT426235T | Austria | T | |
| ATE426235T1 | Austria | T1 | |
| DE60326782D1 | Germany | D1 | |
| ES2323294T3 | Spain | T3 | |
| JP2009271554A | Japan | A | |
| US2009287495A1 | United States of America | A1 | |
| JP4401173B2 | Japan | B2 | |
| KR20100039433A | Republic of Korea | A | |
| CN1647156B | China | B | |
| KR100978018B1 | Republic of Korea | B1 | |
| KR101016982B1 | Republic of Korea | B1 | |
| KR101021076B1 | Republic of Korea | B1 | |
| KR101021079B1 | Republic of Korea | B1 | |
| US7933415B2 | United States of America | B2 | |
| JP4714415B2 | Japan | B2 | |
| JP4714416B2 | Japan | B2 | |
| US2011166866A1 | United States of America | A1 | |
| JP2012161087A | Japan | A | |
| US8331572B2 | United States of America | B2 | |
| JP5101579B2 | Japan | B2 | |
| US8340302B2 | United States of America | B2 | |
| US2013094654A1 | United States of America | A1 | |
| US8498422B2 | United States of America | B2 | |
| JP5498525B2 | Japan | B2 | |
| US8798275B2 | United States of America | B2 | |
| US9137603B2 | United States of America | B2 | |
| BRPI0304541B1 | Brazil | B1 | |
| BRPI0304540B1 | Brazil | B1 | |
| BRPI0304542B1 | Brazil | B1 | |
| DE60311794C5 | Germany | C5 |
Numbers
- Publication
- 2268340
- Publication, DOCDB
- 2268340
- Publication, EPODOC
- ES2268340T
- Application
- 3712597
- Application, DOCDB
- 03712597
- Application, EPODOC
- ES20030712597T
Titles2
- Spanish
- REPRESENTACION DE AUDIO PARAMETRICO DE MULTIPLES CANALES.
- English
- REPRESENTATION OF PARAMETRIC AUDIO OF MULTIPLE CHANNELS.
Classification
- CPC, 6
- G10L19/008
- H04S3/008
- G10L19/24
- G10L19/0204
- H04S2420/03
- G10L19/02
- IPC, 5
- G10L19 008
- G10L19 02
- G10L19 24
- H03M7 30
- H04S3 00