Non-uniform parameter quantization for advanced coupling
Abstract
A method in an audio encoder (110) for the quantification of parameters related to the parametric spatial coding of audio signals, comprising: receiving at least a first parameter and a second parameter to be quantified; quantify (202) the first parameter based on a first scalar quantification scheme that has non-uniform stage sizes to obtain a first quantified parameter, wherein the non-uniform stage sizes are selected so that smaller stage sizes are used for intervals of the first parameter where human perception of sound is more sensitive, and larger stage sizes are used for intervals of the first parameter where human perception of sound is less sensitive; Unquantify (204) the first quantified parameter using the first scalar quantification scheme to obtain a first unquantified parameter that is an approximation of the first parameter; access a scaling function that maps the values of the first unquantified parameter into scaling factors that increase with the stage sizes corresponding to the values of the first unquantified parameter, and determine a scaling factor by submitting the first unqualified parameter to the function of scaling and quantifying (212) the second parameter based on the scaling factor and a second scalar quantification scheme having non-uniform stage sizes to obtain a second quantified parameter.

Term
8 yearsto projected expiry
Projected expiry 8 September 2034, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
15 claims: 10 independent, 5 dependent
- 1ES 2 645 839 T3 REIVINDICACIONES 1. Un procedimiento en un codificador (110) de audio para la cuantlflcaclón de parámetros relativos a la codificación espacial paramétrica de señales de audio, que comprende:recibir al menos un primer parámetro y un segundo parámetro a cuantificar;cuantificar (202) el primer parámetro basándose en un primer esquema de cuantificación escalar que tiene tamaños de etapa no uniformes para obtener un primer parámetro cuantificado, en donde los tamaños de etapa no uniformes se seleccionan de modo que menores tamaños de etapa se usan para intervalos del primer parámetro donde la percepción humana del sonido es más sensible, y tamaños de etapa más grandes se usan para intervalos del primer parámetro donde la percepción humana del sonido es menos sensible;descuantificar (204) el primer parámetro cuantificado utilizando el primer esquema de cuantificación escalar para obtener un primer parámetro descuantificado que es una aproximación del primer parámetro;acceder a una función de escalado que mapea los valores del primer parámetro descuantificado en factores de escalado que aumentan con los tamaños de etapa correspondientes a los valores del primer parámetro descuantificado, y determinar un factor de escalado al someter el primer parámetro descuantificado a la función de escalado;y cuantificar (212) el segundo parámetro basándose en el factor de escalado y un segundo esquema de cuantificación escalar que tiene tamaños de etapa no uniformes para obtener un segundo parámetro cuantificado.
- 2El procedimiento de la reivindicación 1, en donde la función de escalado es una función lineal por segmentos.
- 3El procedimiento de una cualquiera de las reivindicaciones precedentes, en donde la etapa de cuantificar el segundo parámetro basándose en el factor de escalado y el segundo esquema de cuantificación escalar comprende dividir el segundo parámetro por el factor de escalado antes de someter el segundo parámetro a cuantificación según el segundo esquema de cuantificación escalar.
- 4El procedimiento de una cualquiera de las reivindicaciones 1-2, en donde los tamaños de etapa no uniformes del segundo esquema de cuantificación escalar:se escalan mediante el factor de escalado antes de la cuantificación del segundo parámetro;y/o se aumentan con el valor del segundo parámetro.
- 5El procedimiento de una cualquiera de las reivindicaciones precedentes, en donde el primer esquema de cuantificación escalar:comprende más etapas de cuantificación que el segundo esquema de cuantificación escalar;y/o se construye compensando, duplicando y concatenando el segundo esquema de cuantificación escalar.
- 6Un codificador (110) de audio para la cuantificación de parámetros relativos a la codificación espacial paramétrica de señales de audio, que comprende:un componente receptor dispuesto para recibir al menos un primer parámetro y un segundo parámetro a cuantificar;un primer componente (202) de cuantificación dispuesto aguas abajo del componente receptor configurado para cuantificar el primer parámetro basándose en un primer esquema de cuantificación escalar que tiene tamaños de etapa no uniformes para obtener un primer parámetro cuantificado, en donde se seleccionan tamaños de etapa no uniformes de modo que los menores tamaños de etapa se usan para los intervalos del primer parámetro donde la percepción humana del sonido es más sensible, y tamaños de etapa más grandes se usan para los intervalos del primer parámetro donde la percepción humana del sonido es menos sensible;un componente (204) de descuantificación configurado para recibir el primer parámetro cuantificado desde el primer componente de cuantificación, y para descuantificar el primer parámetro cuantificado utilizando el primer esquema de cuantificación escalar para obtener un primer parámetro descuantificado que es una aproximación del primer parámetro;un componente de determinación del factor de escalado configurado para recibir el primer parámetro descuantificado, acceder a una función de escalado que mapea valores del primer parámetro descuantificado sobre factores de escalado que aumentan con los tamaños de etapa correspondientes a los valores del primer parámetro descuantificado, y determinar un factor de escalado al someter el primer parámetro descuantificado a la función de escalado;y un segundo componente (212) de cuantificación configurado para recibir el segundo parámetro y el factor de escalado, y cuantificar el segundo parámetro basándose en el factor de escalado y un segundo esquema de cuantificación escalar que tiene tamaños de etapa no uniformes para obtener un segundo parámetro cuantificado.
- 7Un procedimiento en un descodificador de audio (120) para la descuantificación de parámetros cuantificados ES 2 645 839 T3 relativos a la codificación espacial paramétrlca de señales de audlo, que comprende:recibir al menos un primer parámetro cuantificado y un segundo parámetro cuantificado;descuantificar (304) el primer parámetro cuantificado según un primer esquema de cuantificación escalar que tiene tamaños de etapa no uniformes para obtener un primer parámetro descuantificado, en donde los tamaños de etapa no uniformes se seleccionan de modo que menores tamaños de etapa se usan para intervalos del primer parámetro donde la percepción humana del sonido es más sensible, y mayores tamaños de etapas se usan para los intervalos del primer parámetro donde la percepción humana del sonido es menos sensible;acceder a una función de escalado que mapea valores del primer parámetro descuantificado en factores de escalado que aumentan con los tamaños de etapa correspondientes a los valores del primer parámetro descuantificado, y determinar un factor de escalado al someter el primer parámetro descuantificado a la función de escalado;y descuantificar (308) el segundo parámetro cuantificado basándose en el factor de escalado y un segundo esquema de cuantificación escalar que tiene tamaños de etapa no uniformes para obtener un segundo parámetro descuantificado.
- 8El procedimiento de la reivindicación 7, en donde la función de escalado es una función lineal por segmentos.
- 9El procedimiento de la reivindicación 7 o de la reivindicación 8, en donde la etapa de descuantificar el segundo parámetro basándose en el factor de escalado y el segundo esquema de cuantificación escalar comprende descuantificar el segundo parámetro cuantificado según el segundo esquema de cuantificación escalar y multiplicar el resultado del mismo por el factor de escalado.
- 10El procedimiento de la reivindicación 7 o de la reivindicación 8, en donde los tamaños de etapa no uniformes del segundo esquema de cuantificación escalar:se escalan por el factor de escalado antes de la descuantificación del segundo parámetro cuantificado;y/o se aumenta con el valor del segundo parámetro.
- 11El procedimiento de una cualquiera de las reivindicaciones 7 a 10, en donde el primer esquema de cuantificación escalar:comprende más etapas de cuantificación que el segundo esquema de cuantificación escalar;y/o se construye compensando, duplicando y concatenando el segundo esquema de cuantificación escalar.
- 12El procedimiento de una cualquiera de las reivindicaciones 7 a 11, o el procedimiento de una cualquiera de las reivindicaciones 1 a 5, en donde el tamaño de etapa más grande del primero y/o segundo esquema de cuantificación escalar es aproximadamente cuatro veces mayor que el tamaño de etapa más pequeño del primero y/o segundo esquema de cuantificación escalar.
- 13Un medio legible por ordenador que comprende instrucciones de códigos informáticos adaptadas para llevar a cabo el procedimiento de una cualquiera de las reivindicaciones 7 a 12 cuando se ejecuta mediante un dispositivo que tiene capacidad de procesamiento, o adaptado para llevar a cabo el procedimiento de una cualquiera de las reivindicaciones 1 a 5 cuando se ejecuta mediante un dispositivo que tiene capacidad de procesamiento.
- 14Un descodificador (120) de audio para la descuantificación de parámetros cuantificados relativos a la codificación espacial paramétrica de señales de audio, que comprende:un componente receptor configurado para recibir al menos un primer parámetro cuantificado y un segundo parámetro cuantificado;un primer componente (304) de descuantificación dispuesto aguas abajo del componente receptor y configurado para descuantificar el primer parámetro cuantificado según un primer esquema de cuantificación escalar que tiene tamaños de etapa no uniformes para obtener un primer parámetro descuantificado, en donde los tamaños de etapa no uniformes se seleccionan de manera que los menores tamaños de etapa se usan para los intervalos del primer parámetro donde la percepción humana del sonido es más sensible, y tamaños de etapa más grandes se usan para los intervalos del primer parámetro donde la percepción humana del sonido es menos sensible;un componente de determinación del factor de escalado configurado para recibir el primer parámetro descuantificado del primer componente de descuantificación, acceder a una función de escalado que mapea valores del primer parámetro descuantificado en factores de escalado que aumentan con los tamaños de etapa correspondientes a los valores del primer parámetro descuantificado, y determinar un factor de escalado al someter el primer parámetro descuantificado a la función de escalado;y un segundo componente de descuantificación (308) configurado para recibir el factor de escalado y el segundo parámetro cuantificado, y descuantificar el segundo parámetro cuantificado basándose en el factor de escalado y un segundo esquema de cuantificación escalar que tiene tamaños de etapa no uniformes para obtener un segundo parámetro descuantificado. ES 2 645 839 T3
- 15Un sistema de codificación/descodificación de audio que comprende un codificador de audio según la reivindicación 6 y un descodificador de audio según la reivindicación 14, en donde el codificador de audio está dispuesto para transmitir el primero y segundo parámetros cuantificados hasta el descodificador de audio.
Independent claims15
127 paragraphs in 7 sections, as filed
ES 2 645 839 T3
DESCRIPTION
Non-uniform parameter quantification for advanced coupling.
Cross reference to related applications
This application claims priority to United States Provisional Patent Application No. 61 / 877,166, filed September 12, 2013.
Technique field
The disclosure herein refers generally to audio coding. In particular, it relates to the perceptually optimized quantization of the parameters used in a system for the parametric spatial coding of audio signals.
Background
The performance of low-bit audio coding systems can be significantly improved for stereo signals when a Parametric Stereo (PS) coding tool is used. In such a system, a mono signal is quantized and transmitted, typically using an advanced audio encoder, and the stereo parameters are estimated and quantized at the encoder and added as side information to the bit stream. In the decoder, the stereo signal is reconstructed from the decoded mono signal with the help of stereo parameters.
There are several possible parametric stereo coding variants. Consequently, there are several types of encoders and, in addition to a mono downmix, they generate different stereo parameters that are integrated into the generated bit stream. The tools for such coding have also been standardized. An example of such a standard is MPEG-4 audio (IsO / IEC 14496-3).
The main idea behind audio coding systems, in general, and parametric stereo coding, in particular, and one of the various challenges in this field of art is to minimize the amount of information that has to be transferred in the stream. bits from an encoder to a decoder while still getting good audio quality. A high level of compression of the bitstream information can lead to unacceptable sound quality due to insufficient and complex calculation processes or because information has been lost in the compression process. On the other hand, a low level of compression of the bitstream information can lead to capacity problems that can also result in unacceptable sound quality.
Consequently, there is a need for better parametric stereo coding procedures.
The International Search Report published in connection with the present application cited the publication of United States Patent Application US No. 2007/0016416 A1, the '416 document, as a document that defines the general state of the art that does not it is considered to be of particular relevance. The '416 document discloses that a parameter that is a measure of a characteristic of a channel or a pair of channels with respect to another channel of a multichannel signal can be quantized more efficiently using a quantization rule that is generated based on a ratio of a channel or channel pair energy measure and a multichannel signal energy measure.
Brief description of the drawings
The example embodiments will now be described in greater detail and with reference to the accompanying drawings, in which:
Figure 1 discloses a block diagram of a parametric stereo encoding and decoding system according to an embodiment;
Figure 2 shows a block diagram related to the processing of stereo parameters in the coding portion of the parametric stereo coding system of Figure 1;
Figure 3 presents a block diagram related to the processing of stereo parameters in the decoding portion of the parametric stereo coding system of Figure 1;
Figure 4 shows the value of a scaling factor s as a function of one of the stereo parameters;
Figure 5 discloses non-uniform and uniform quantizers (fine and coarse) in the (a, b) plane, where a and b are stereo parameters; Y
Figure 6 presents a diagram showing the mean bit consumption with parametric stereo for examples of fine uniform, and approximate uniform quantization, compared to an approximate non-uniform fine and non-uniform quantization according to an exemplary embodiment.
ES 2 645 839 T3
Figure 7 discloses a block diagram of a parametric multi-channel coding and decoding system according to another embodiment.
All figures are schematic and generally only show parts that are necessary for clarification of the disclosure, while other parts may be omitted or simply suggested. Unless otherwise indicated, like reference numerals refer to like parts in different figures.
Detailed description
With the foregoing in mind, an object is to provide encoders, decoders, systems comprising encoders and decoders, and associated procedures that provide higher efficiency and quality of the encoded audio signal.
I. General Description - Encoder
According to a first aspect, the example embodiments propose encoding methods, encoders and computer products for encoding. The proposed processes, encoders, and software products may generally have the same features and benefits.
According to the example embodiments, a method is provided in an audio encoder for quantizing parameters relating to parametric spatial coding of audio signals, comprising: receiving at least a first parameter and a second parameter to be quantized; quantize the first parameter based on a first scalar quantization scheme having non-uniform stage sizes to obtain a first quantized parameter, wherein the non-uniform stage sizes are selected so that the smallest stage sizes are used for the intervals of the first parameter where human perception of sound is most sensitive, and the larger stage sizes are used for the intervals of the first parameter where human perception of sound is less sensitive; dequantizing the first quantized parameter using the first scalar quantization scheme to obtain a first dequantized parameter that is an approximation of the first parameter; access a scaling function that maps values of the first dequantized parameter to scaling factors that increase with the stage sizes corresponding to the values of the first dequantized parameter, and determine a scaling factor by subjecting the first dequantized parameter to the scaling function ; and quantizing the second parameter based on the scaling factor and a second scalar quantization scheme having non-uniform stage sizes to obtain a second quantized parameter.
The procedure is based on the understanding that human perception of sound is not homogeneous. Instead, it turns out that human perception of sound is higher for some characteristics of sound and lower for other characteristics of sound. This implies that human perception of sound is more sensitive for some parameter values related to parametric spatial coding of audio signals than for other values of this type. According to the procedure provided, such a first parameter is quantized in non-uniform stage sizes so that smaller stage sizes are used where human perception of sound is more sensitive and larger stage sizes are used where human perception of sound is less sensitive. By quantizing the use of such non-uniform stage size schemes, it is possible to reduce the average parametric stereo bit consumption without reducing the perceptible sound quality.
According to the embodiments, the scaling function of the method is a piecewise linear function.
According to the embodiments, the process step of quantizing the second parameter is based on the scaling factor and the second scalar quantization scheme comprises dividing the second parameter by the scaling factor before subjecting the second parameter to quantization according to the second scaling scheme. scalar quantification.
According to an alternative embodiment of the method, the non-uniform stage sizes of the second scalar quantization scheme are scaled by the scaling factor prior to quantization of the second parameter.
According to embodiments of the method, the non-uniform stage sizes of the second scalar quantization scheme increase with the value of the second parameter.
According to embodiments of the method, the first scalar quantization scheme comprises more quantization steps than the second scalar quantization scheme.
According to embodiments of the method, the first scalar quantization scheme is constructed by compensating, duplicating and concatenating the second scalar quantization scheme.
According to embodiments of the method, the largest stage size of the first and / or second scalar quantization scheme is approximately four times larger than the smallest stage size of the first and / or second scalar quantization scheme.
According to the example embodiments, a computer-readable medium is provided comprising instructions 3
ES 2 645 839 T3 with computer codes adapted to carry out any procedure of the first aspect when executed on a device having processing capacity.
According to the example embodiments, an audio encoder is provided for quantizing parameters relating to parametric spatial encoding of audio signals, comprising: a receiver component arranged to receive at least a first parameter and a second parameter to be quantized; a first quantization component arranged downstream of the receiver component configured to quantize the first parameter based on a first scalar quantization scheme having non-uniform stage sizes to obtain a first quantized parameter, wherein the non-uniform stage sizes are selected from so that the smaller stage sizes are used for the intervals of the first parameter where the human perception of sound is most sensitive, and larger stage sizes are used for the intervals of the first parameter where human perception of sound is less sensitive; a dequantization component configured to receive the first quantized parameter from the first quantization component, and to dequantize the first quantized parameter using the first scalar quantization scheme to obtain a first dequantized parameter that is an approximation of the first parameter; a component for determining the scaling factor configured to receive the first dequantized parameter, accessing a scaling function that maps the values of the first dequantized parameter into scaling factors that increase with the stage sizes corresponding to the values of the first dequantized parameter, and determining a scaling factor by subjecting the first dequantized parameter to the scaling function; and a second quantization component configured to receive the second parameter and the scaling factor, and quantizing the second parameter based on the scaling factor and a second scalar quantization scheme having non-uniform stage sizes to obtain a second quantized parameter.
II. Overview - Decoder
According to a second aspect, the example embodiments propose decoding methods, decoders and computer program products for decoding. The proposed procedures, decoders, and software products may have, in general, the same features and advantages.
The advantages over the features and configurations presented in the above encoder overview may be generally valid for the corresponding features and configurations of the decoder.
According to the example embodiments, a method is provided in an audio decoder for dequantizing quantized parameters relative to parametric spatial encoding of audio signals, comprising: receiving at least a first quantized parameter and a second quantized parameter; dequantize the first quantized parameter according to a first scalar quantization scheme having non-uniform stage sizes to obtain a first dequantized parameter, wherein the non-uniform stage sizes are selected so that the smallest stage sizes are used for the intervals of the first parameter where human perception of sound is most sensitive, and the largest stage sizes are used for the intervals of the first parameter where human perception of sound is less sensitive; access a scaling function that maps the values of the first unquantized parameter into scaling factors that increase with the stage sizes corresponding to the values of the first unquantized parameter, and determine a scaling factor by subjecting the first unquantized parameter to the function of scaled; and dequantizing the second quantized parameter based on the scaling factor and a second scalar quantization scheme having non-uniform stage sizes to obtain a second dequantized parameter.
According to example embodiments of the method, the scaling function is a piecewise linear function.
According to one embodiment, the step of dequantizing the second parameter based on the scaling factor and the second scalar quantization scheme comprises de-quantizing the second quantized parameter according to the second scalar quantization scheme and multiplying the result thereof by the scale factor.
According to an alternative embodiment, the non-uniform stage sizes of the second scalar quantization scheme are scaled by the scaling factor prior to dequantization of the second quantized parameter.
According to further embodiments, the non-uniform stage size of the second scalar quantization scheme increases with the value of the second parameter.
According to one embodiment, the first scalar quantization scheme comprises more quantization steps than the second scalar quantization scheme.
According to one embodiment, the first scalar quantization scheme is constructed by compensating, duplicating and concatenating the second scalar quantization scheme.
According to one embodiment, the largest stage size of the first and / or second scalar quantization scheme
ES 2 645 839 T3 is approximately four times larger than the smallest stage size of the first and / or second scalar quantization scheme.
According to example embodiments, a computer-readable medium is provided comprising computer code instructions adapted to carry out the procedure of any procedure of the second aspect when executed by a device having processing capability.
According to the example embodiments, there is provided an audio decoder for dequantizing quantized parameters relative to parametric spatial coding of audio signals, comprising: a receiver component configured to receive at least a first quantized parameter and a second quantized parameter; a first dequantization component arranged downstream of the receiver component and configured to dequantize the first quantized parameter according to a first scalar quantization scheme having non-uniform stage sizes to obtain a first dequantized parameter, where non-uniform stage sizes are selected such that smaller stage sizes are used for first parameter intervals where human perception of sound is most sensitive and larger stage sizes are used for first parameter intervals where perception human sound is less sensitive; a component for determining the scaling factor configured to receive the first dequantized parameter of the first dequantized component, accessing a scaling function that maps the values of the first dequantized parameter into scaling factors that increase with the stage sizes corresponding to the values of the first dequantized parameter, and determining a scaling factor by subjecting the first dequantized parameter to the scaling function; and a second dequantization component configured to receive the scaling factor and the second quantized parameter, and dequantize the second quantized parameter based on the scaling factor and a second scalar quantization scheme having non-uniform stage sizes to obtain a second parameter unquantified.
III. Overview - An audio encoding / decoding system
According to a third aspect, the example embodiments propose decoding / encoding systems comprising an encoder according to the first aspect and a decoder according to the second aspect.
Advantages over the features and configurations presented in the above encoder and decoder overview may generally apply to the corresponding features and configurations of the system.
According to example embodiments, such a system is provided wherein the audio encoder is arranged to transmit the first and second quantized parameters to the audio decoder.
IV. Example realizations
The description of this document discusses the perceptually optimized quantization of the parameters used in a system for parametric spatial coding of audio signals. In the examples considered below, the special case of parametric stereo coding for 2-channel signals is discussed. The same technique can also be used in multi-channel parametric coding, eg in a system operating in 5-3-5 mode. An example of an embodiment of such a system is presented in FIG. 7 and will be briefly discussed later. The example embodiments presented here refer to a simple non-uniform quantization that allows the reduction of the bit rate necessary to invoke these parameters without affecting the perceived audio quality, and further allows the continued use of established entropy coding techniques to scalar parameters (such as time difference or frequency difference encoding of Huffman encoding).
Figure 1 shows a block diagram of one embodiment of a parametric stereo encoding and decoding system 100 that is discussed here. A stereo signal comprising a left channel 101 (L) and a right channel 102 (R) is received by the encoder portion 110 of the system 100. The stereo signal is input to an Advanced Coupling (ACPL) encoder 112 which generates a mono downmix 103 (M) and stereo parameters a (referred to in figure 1 as 104a) and b (referred to in figure 1 as 104b) . In addition, the encoder part 110 comprises a downmix encoder 114 (DMX Enc) that transforms the mono downmix 103 into a bit stream 105, a stereo (Q) parameter quantization means 116 that generates a stereo parameter stream 106 quantized and a multiplexer 118 (MUX) that generates the final bit stream 108 which also comprises the quantized stereo parameters that are conveyed to the decoder portion 120. The decoder part 120 comprises a demultiplexer 122 (DE-MUX) that receives the final incoming bit stream 108 and regenerates the bit stream 105 and the quantized stereo parameter stream 106, a downmix decoder 124 (DMX Dec) that receives bit stream 105 and outputs a decoded mono downmix 103 '(M'), a stereo parameter dequantization means 126 (Q ') that receives a stream of quantized stereo parameters 106 and outputs dequantized stereo parameters to' 104a 'and b' 104b ', and finally the ACPL decoder 128 that receives the decoded mono downmix 103' and the unquantized stereo parameters 104a ', 104b' and transforms these incoming signals into reconstructed stereo signals 101 '(L') and 102 '(R').
ES 2 645 839 T3
From the incoming stereo signals 101 (L) and 102 (R), the ACPL encoder 112 calculates a mono downmix 103 (M) and a side signal (S) according to the following equations:
M = (L + R) / 2 (equation 1)
S = (L - R) / 2 (equation 2)
The stereo parameters a and b are calculated selectively in time and frequency, that is, for each time / frequency mosaic tile, typically with the help of a filter bank such as a QMF bank and using a non-uniform grouping of QMF bands to form a set of parameter bands according to a perceptual frequency scale.
In the ACPL decoder, the decoded mono downmix M 'together with the stereo parameters a', b 'and an uncorrelated version of M' (decorr (M ')) are used as input to reconstruct an approximation of the side signal according to the following equation:
S '= a' * M '+ b' * descorr (M ') (equation 3)
L 'and R' are then calculated as:
L '= M' + S '(equation 4)
R '= M' - S '(equation 5)
The pair of parameters (a, b) can be considered as a point in a two-dimensional plane (a, b). Parameters a, b are related to the perceived stereo image, where parameter a is mainly related to the position of the perceived sound source (e.g. left or right) and where parameter b is mainly related to size or the width of the perceived sound source (small and well localized or wide and ambient). Table 1 lists some typical examples of perceived stereo images and the corresponding values of parameters a, b.
Table 1
<td>Point</td><td>Parameter values</td><td>Signal description</td>
<td>Left</td><td>a = 1, b = 0</td><td>Signal fully pointed to the left side, that is, R = 0.</td>
<td>Center</td><td>a = 0, b = 0</td><td>Signal in the imaginary center, that is, L = R.</td>
<td>Right</td><td>a = -1, b = 0</td><td>Signal fully pointed to the right side, that is, L = 0.</td>
<td>Large</td><td>a = 0, b = 1</td><td>Broad signal, L and R are uncorrelated and have the same level.</td>
Note that b is never negative. It should also be noted that although b and the absolute value of a are often within the range of 0 to 1, they can also have absolute values greater than 1, for example, in the case of components with large phase shift in L and R, that is, when the correlation between L and R is negative.
The problem at hand now is to design a technique to quantize parameters a, b for transmission as side information in a parametric / spatial stereo coding system. A simple and straightforward approach to the prior art is to use uniform quantization and quantize a and b independently, that is, use two scalar quantizers. A typical quantization step size is delta = 0.1 for fine quantification or delta = 0.2 for rough quantification. The lower left and right panel of Figure 5 show the points in plane (a, b) that can be represented by such a quantization scheme for fine and coarse quantization. Typically, the quantized parameters a and b are independently entropy encoded, using time difference or frequency difference encoding in combination with Huffman encoding.
However, the present inventors have now realized that the performance (in a rate distortion sense) of parameter quantization can be improved through such a scalar quantization by taking perceptual aspects into account. In particular, the sensitivity of the human auditory system to small changes in parameter values (such as the error introduced by quantization) depends on the position in the (a, b) plane. Perceptual experiments investigating the audibility of small changes of this type or simply appreciable differences (JND) indicate that the JNDs for a and b are considerably higher.
ES 2 645 839 T3 small for sound sources with a perceived stereo image that is represented by the points (1, 0) and (-1, 0) in the (a, b) plane. Consequently, a uniform quantization of a and b may be too approximate (with audible artifacts) for the regions near (1, 0) and (-1, 0) and unnecessarily fine (causing an unnecessarily high side information bit rate) in other regions, such as around (0, 0) and (0, 1). Of course, it would be possible to consider a vector quantizer so that (a, b) achieves a joint and non-uniform quantization of the stereo parameters a and b. However, a vector quantizer is computationally more complex, and also the entropy encoding (by difference in time or frequency) would have to be adapted and would also become more complex.
Therefore, in this application a new non-uniform quantization scheme is introduced for parameters a and b. The non-uniform quantization scheme for a and b takes advantage of position-dependent JNDs (as a vector quantizer can do) but can be applied as a small modification of the prior art uniform and independent quantization of a and b. Furthermore, also the time difference or frequency difference entropy coding of the prior art can remain basically unchanged. Only Huffman codebooks need to be updated to reflect changes in index ranges and symbol probabilities.
The resulting quantization scheme is shown in Figures 2 and 3, where Figure 2 refers to the stereo parameter quantization means 116 of the encoder part 110 and Figure 3 refers to the stereo parameter dequantization means 126 from part 120 of the decoder. The stereo parameter quantization scheme begins by applying a non-uniform scalar quantization to the parameter a (referred to as 104a in Figure 2) in the Q means<sub>to</sub> quantization (referred to as 202 in Figure 2). The quantized parameter 106a is forwarded to the multiplexer 118. The quantized parameter is also dequantized directly in the Q means<sub>to</sub><sup>1</sup> dequantization (referred to as 204 in Figure 2) to the parameter a '. Since quantized parameter 106a is dequantized to aa '(referred to as 104a' in FIG. 3) also in decoder part 120, a 'will be identical in both encoder part 110 and decoder part 120 of system 100. Then, a 'is used to calculate a scaling factor s (carried out by the scaling means 206) which is used to make the quantization of b depend on the actual value of a. The parameter b (referred to 104b in figure 2) is divided by this scaling factor s (carried out by the inverting means 208 and the multiplication means 210) and then sent to another non-uniform scalar quantizer Qb (referred to 212 in FIG. 2) from which the quantized parameter 106b is forwarded. The process is partially reversed in the stereo parameter dequantization means 126 shown in FIG. 3. The incoming quantized parameters 106a and 106b are dequantized in the Q means.<sub>to</sub><sup>1 </sup>dequantization (referred to in Figure 3 as 304) and Qb<sup>-1</sup> (referred to in Figure 3 as 308) aa '(referred to as 104a' in Figure 3) and b 'previously divided by the scaling factor s in the part 110 of the encoder. The scaling means 306 determines the scaling factor s based on the unquantized parameter a '(104a) in the same way as the scaling means 206 in the encoder portion 110. The scaling factor is then multiplied by the result of the dequantization of the quantized parameter 106b in the multiplication means 310 and the dequantized parameter b 'is obtained (referred to as 104b' in FIG. 3). Consequently, the dequantization of a and the calculation of the scaling factor are applied in both the encoder part 110 and the decoder part 120, ensuring that exactly the same value of s is used to encode and decode b.
The non-uniform quantization of a and b is based on a simple non-uniform quantizer for values in the range 0 to 1 where the size of the quantization stage for values around 1 is approximately four times the size of the quantization stage for values around of 0, and where the size of the quantization stage increases with the value of the parameter. For example, the size of the quantization step can increase approximately linearly with the index that identifies the corresponding dequantized value. For a quantizer with 8 intervals (that is, 9 indications), the following values can be obtained, where the size of the quantization step is the difference between two neighboring dequantized values.
Table 2: Unquantified values within the interval from 0 to 1
<td>index</td><td>Value</td><td>index</td><td>Value</td>
<td> 0</td><td> 0</td><td> 5</td><td> 0,4844</td>
<td> 1</td><td> 0,0594</td><td> 6</td><td> 0,6375</td>
<td> 2</td><td> 0,1375</td><td> 7</td><td> 0,8094</td>
<td> 3</td><td> 0,2344</td><td> 8</td><td> 1,0000</td>
<td> 4</td><td> 0,3500</td><td></td><td></td>
This table is an example of a quantization scheme that could be used in the Qb quantification media<sup>-1 </sup>(referred to as 308 in figure 3). However, a larger range of values must be handled for parameter 7
ES 2 645 839 T3
to. An example of a quantification scheme of the means of quantification Q<sub>to</sub>-1 (referred to as 304 in Figure 3) could be constructed simply by doubling and concatenating the non-uniform quantifier intervals shown in Table 2 above to give a quantifier that can represent values in the range -2 to 2, where the size of the quantization stage for values around -2, 0 and 2 is approximately four times the size of the quantization stage for values around -1 and 1. The resulting values are shown in Table 3 below.
Table 3: Unquantified values within the interval from -2 to 2
<td>index</td><td>Value</td><td>index</td><td>Value</td>
<td> 0</td><td> -2,000</td><td> 17</td><td> 0,1906</td>
<td> 1</td><td> -1,8094</td><td> 18</td><td> 0,3625</td>
<td> 2</td><td> -1,6375</td><td> 19</td><td> 0,5156</td>
<td> 3</td><td> -1,4844</td><td> 20</td><td> 0,6500</td>
<td> 4</td><td> -1,3500</td><td> 21</td><td> 0,7656</td>
<td> 5</td><td> -1,2344</td><td> 22</td><td> 0,8625</td>
<td> 6</td><td> -1,1375</td><td> 23</td><td> 0,9406</td>
<td> 7</td><td> -1,0594</td><td> 24</td><td> 1,000</td>
<td> 8</td><td> -1,000</td><td> 25</td><td> 1,0594</td>
<td> 9</td><td> -0,9406</td><td> 26</td><td> 1,1375</td>
<td> 10</td><td> -0,8625</td><td> 27</td><td> 1,2344</td>
<td> 11</td><td> - 0,7656</td><td> 28</td><td> 1,3500</td>
<td> 12</td><td> - 0,6500</td><td> 29</td><td> 1,4844</td>
<td> 13</td><td> - 0,5156</td><td> 30</td><td> 1,6375</td>
<td> 14</td><td> - 0,3625</td><td> 31</td><td> 1,8094</td>
<td> 15</td><td> -0,1906</td><td> 32</td><td> 2,000</td>
<td> 16</td><td> 0</td><td></td><td></td>
Figure 4 shows the value of the scaling factor s as a function of a. It is a segment linear function, with s = 1 (that is, without scaling) for a = -1, a = 1 and s = 4 (4 times the closest quantization of b) for a = -2, a = 0, and a = 2. It is noted that the function of figure 4 is an example and that other functions of this type are theoretically possible. The same reasoning is applicable to quantification schemes.
The resulting non-uniform quantization of a and b is shown in the upper left panel of Figure 5, where each point in the plane (a, b) that can be represented by this quantizer is marked by a cross.
Around the most sensitive points (1, 0) and (-1, 0), the size of the quantization stage for both a and b is approximately 0.06, while it is approximately 0.2 for a and b around ( 0, 0). Consequently, the quantization steps are much more adapted to JND than those of a uniform scalar quantization of a and b.
If a more approximate quantization is desired, it is possible to simply leave every second dequantized value of the quantizers non-uniform, thereby doubling the sizes of the quantization steps. Table 4 shows the following approximate non-uniform quantizers for the parameter b and the non-uniform quantizers for the parameter a are obtained analogously to what has been shown above.
Table 4: Dequantized values for a more approximate quantification within the range of 0 to 1:
ES 2 645 839 T3
<td>index</td><td>Value</td>
<td> 0</td><td> 0</td>
<td> 1</td><td> 0,1375</td>
<td> 2</td><td> 0,3500</td>
<td> 3</td><td> 0,6375</td>
<td> 4</td><td> 1,000</td>
The scaling function shown in Figure 4 remains unchanged for the approximate quantifier, and the resulting approximate quantifier for (a, b) is shown in the upper right panel of Figure 5. Such coarse quantization may be desirable if the coding system is operated at very low target bit rates, where it may be advantageous to use the bits stored by the closest quantization of the stereo parameters to encode the M signal instead. downmix mono (referred to as 103 in Figure 1).
The difference in efficiency between a non-uniform and a uniform quantization of the stereo parameters a and b is shown in Figure 6. The differences are shown for a fine and a coarse quantization. The average consumption of bits per second corresponding to 11 hours of music is shown. It can be concluded from the figure that the bit consumption for non-uniform quantization is considerably lower than for uniform quantization. Furthermore, it can be concluded that a more approximate non-uniform quantization reduces the consumption of bits per second more than a more approximate uniform quantization does.
Finally, Figure 7 discloses a block diagram of an exemplary embodiment of a 5-3-5 parametric multi-channel encoding and decoding system 700. A multi-channel signal comprising a left front channel 701, a left surround channel 702, a center front channel 703, a right front channel 704 and a right surround channel 705 is received by the encoder portion 710 of the system 700. The signals from front left channel 701 and left surround channel 702 are input to a first Advanced Coupling (ACPL) encoder 712 which generates a left 706 downmix and stereo parameters aL (referred to as 708a) and bL (referred to as 708b ). Similarly, the signals from the front right channel 704 and the right surround channel 705 are input to a second Advanced Coupling (ACPL) encoder 713 which generates a 707 right downmix and stereo parameters aR (referred to as 709a) and bR. (referred to as 709b). In addition, the encoder portion 710 comprises a 3-channel downmix encoder 714 which transforms the signals from the left downmix 706, the center front channel 703 and the right downmix 707 to a bit stream 722, a first means 715 stereo parameter quantization that generates a first stream of 720 stereo parameters quantized based on the stereo parameters 708a and 708b, a second stereo parameter quantization means 716 that generates a second quantized stereo parameter stream 724 based on the stereo parameters 709a and 709b, and a multiplexer 730 that generates the final bit stream 735 that also comprises the quantized stereo parameters that are transported to part 740 of the decoder. The decoder portion 740 comprises a demultiplexer 742 that receives the final incoming bit stream 735 and regenerates the bit stream 722, the first quantized stereo parameter stream 720, and the second quantized stereo parameter stream 724. The first quantized stereo parameter stream 720 is received by dequantized first stereo parameter quantization means 745 708a 'and 708b'. The second stream of quantized stereo parameters 724 is received by the second stereo parameter dequantization means 746 which outputs the dequantized stereo parameters 709a 'and 709b'. The bit stream 722 is received by the 3 channel downmix decoder 744 which outputs the regenerated left down mix 706 ', the rebuilt center front channel 703' and the regenerated down mix 707 '. A first ACPL decoder 747 receives dequantized stereo parameters 708a 'and 708b', as well as a regenerated left down mix 706 'and outputs the reconstructed left front channel 701', and a reconstructed left surround channel 702 '. Similarly, a second ACPL decoder 748 receives the unquantized stereo parameters 709a ', 709b', and the regenerated right down mix 707 'and outputs the reconstructed right front channel 704' and reconstructed right surround channel 705 '.
Equivalents, extension, alternatives and observations
Additional embodiments of the present disclosure will become apparent to one of ordinary skill in the art after studying the foregoing description. Although the present description and drawings disclose embodiments and examples, the disclosure is not limited to these specific examples. Numerous modifications and variations can be made without departing from the scope of the present disclosure, which is defined by the appended claims. Any of the reference signs appearing in the claims should not be construed as limiting their scope.
ES 2 645 839 T3
Furthermore, variations in the disclosed embodiments can be understood and effected by the person skilled in the practice of the disclosure, based on a study of the drawings, the disclosure, and the appended claims. In the claims, the word comprising does not exclude other elements or steps, and the indefinite articles one or one does not exclude a plurality. The mere fact that certain measures are cited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage.
The systems and procedures disclosed earlier herein may be applied as software, firmware, hardware, or a combination thereof. In a hardware application, the division of tasks between the functional units referred to in the previous description does not necessarily correspond to the division into physical units; on the contrary, a physical component can have multiple functionalities, and a task can be carried out by several cooperating physical components. Certain or all components may be applied as software run by a digital signal processor or microprocessor, or applied as hardware or as an application-specific integrated circuit. Such software may be distributed on computer-readable media, which may comprise computer storage media (or non-transitory media) and communication media (or transitory media). As is well known to those skilled in the art, the term computer storage media includes both volatile and non-volatile, removable and non-removable media applied in any method or technology for storing information such as computer-readable instructions, data structures , program modules or other data. Computer storage media includes but is not limited to RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile discs (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, storage on magnetic disk or other magnetic disk storage devices, or any other medium that can be used to store the desired information and that can be accessed by a computer. In addition, it is well known to those of skill that communication media typically incorporate computer-readable instructions, data structures, program modules, or other data into a modulated data signal such as a carrier wave or other transport mechanism and include any medium. of information distribution.
Contents7
7 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7
55 members in 23 offices
Priority claims9
| Document | Office | Kind | Date |
|---|---|---|---|
| 201361877166 | United States of America | P | |
| 201361877166 | United States of America | P | |
| 201361877166P | United States of America | – | |
| 2014069040 | European Patent Office (EPO) | W | |
| 2014069040 | European Patent Office (EPO) | W | |
| 201361877166P | – | – | – |
| PCTEP2014069040 | – | – | – |
| US201361877166P | – | – | – |
| WO2014EP69040 | – | – | – |
Members55
| Document | Office | Kind | |
|---|---|---|---|
| CA2922256A1 | Canada | A1 | |
| WO2015036349A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW201521011A | Taiwan Province of China | A | |
| AU2014320538A1 | Australia | A1 | |
| SG11201601144WA | Singapore | A | |
| AR097618A1 | Argentina | A1 | |
| KR20160042113A | Republic of Korea | A | |
| IL244153A0 | Israel | A0 | |
| IL244153D0 | Israel | D0 | |
| CN105531763A | China | A | |
| MX2016002793A | Mexico | A | |
| EP3044788A1 | European Patent Office (EPO) | A1 | |
| US2016217800A1 | United States of America | A1 | |
| IL244153A | Israel | A | |
| JP2016531327A | Japan | A | |
| CL2016000571A1 | Chile | A1 | |
| HK1220037A | Hong Kong, China | A | |
| HK1220037A1 | Hong Kong, China | A1 | |
| TWI579831B | Taiwan Province of China | B | |
| US9672837B2 | United States of America | B2 | |
| EP3044788B1 | European Patent Office (EPO) | B1 | |
| CN105531763B | China | B | |
| US2017238208A1 | United States of America | A1 | |
| RU2628898C1 | Russian Federation | C1 | |
| KR101777631B1 | Republic of Korea | B1 | |
| JP6201057B2 | Japan | B2 | |
| AU2014320538B2 | Australia | B2 | |
| CA2922256C | Canada | C | |
| DK3044788T3 | Denmark | T3 | |
| ES2645839T3This record | Spain | T3 | |
| PL3044788T3 | Poland | T3 | |
| NO2996227T3 | Norway | T3 | |
| UA116482C2 | Ukraine | C2 | |
| EP3321932A1 | European Patent Office (EPO) | A1 | |
| MX356805B | Mexico | B | |
| US10057808B2 | United States of America | B2 | |
| HK1247432A | Hong Kong, China | A | |
| HK1247432A1 | Hong Kong, China | A1 | |
| US2018352475A1 | United States of America | A1 | |
| US10383003B2 | United States of America | B2 | |
| US2019320348A1 | United States of America | A1 | |
| US10694424B2 | United States of America | B2 | |
| EP3321932B1 | European Patent Office (EPO) | B1 | |
| US2020389815A1 | United States of America | A1 | |
| BR112016005192B1 | Brazil | B1 | |
| AR115819A2 | Argentina | A2 | |
| AR115820A2 | Argentina | A2 | |
| MY187124A | Malaysia | A | |
| US11297533B2 | United States of America | B2 | |
| US2022295347A1 | United States of America | A1 | |
| US11838798B2 | United States of America | B2 | |
| US2024155427A1 | United States of America | A1 | |
| MY204045A | Malaysia | A | |
| US12213004B2 | United States of America | B2 | |
| US2025261039A1 | United States of America | A1 |
Numbers
- Publication
- 2645839
- Publication, DOCDB
- 2645839
- Publication, EPODOC
- ES2645839T
- Application
- 14761831
- Application, DOCDB
- 14761831
- Application, EPODOC
- ES20140761831T
Titles2
- Spanish
- Cuantificación de parámetros no uniforme para acoplamiento avanzado
- English
- Non-uniform parameter quantification for advanced coupling
Classification
- CPC, 8
- G10L19/008
- H04W28/065
- G10L19/035
- H03M7/30
- H04S1/007
- H04S2420/03
- G10L19/038
- H04W28/0215
- IPC, 2
- G10L19 035
- G10L19 008