Non-uniform parameter quantization for advanced coupling
15 claims: 13 independent, 2 dependent
- 1Zastrzeżenia patentowe 1. Sposób w koderze (HO) audio do kwantyzacji parametrów dotyczących parametrycznego kodowania przestrzennego sygnałów audio obejmujący:odbieranie co najmniej pierwszego parametru i drugiego parametru podlegającego kwantyzacji;kwantyzowanie (202) pierwszego parametru na podstawie pierwszego skalarnego schematu kwantyzacji, mającego niejednolite wielkości kroków dla uzyskania kwantyzowanego pierwszego parametru, przy czym niejednolite wielkości kroków dobierane są tak, ze mniejsze wielkości kroków wykorzystywane są w przypadku zakresów pierwszego parametru, w przypadku którego ludzka percepcja dźwięku jest najbardziej wrażliwa, a większe wielkości kroków stosowane są w przypadku zakresów pierwszego parametru, w przypadku którego ludzka percepcja dźwięku jest mniej wrażliwa;dekwantyzację (204) pierwszego parametru przy wykorzystaniu pierwszego skalarnego schematu kwantyzacji, dla uzyskania dekwantyzowanego pierwszego parametru, będącego aproksymacją pierwszego parametru;uzyskanie dostępu do funkcji skalowania, która mapuje wartości dekwantyzowanego pierwszego parametru na współczynniki skalowania, które zwiększają wielkości kroków odpowiednio do wartości dekwantyzowanego pierwszego parametru i wyznaczanie współczynnika skalowania przez poddanie dekwantyzowanego pierwszego parametru funkcji skalowania;i kwantyzowanie (212) drugiego parametru na podstawie współczynnika skalowania i drugiego skalarnego schematu kwantyzacji, mającego niejednolite wielkości kroków, dla uzyskania kwantyzowanego drugiego parametru.
- 2Sposób według zastrz. 1, przy czym funkcja skalowania jest funkcją odcinkowoliniową.
- 3Sposób według dowolnego spośród powyższych zastrzeżeń, przy czym etap sposobu dotyczący kwantyzacji drugiego parametru opiera się na współczynniku skalowania, a drugi skalarny schemat kwantyzacji obejmuje dzielenie drugiego parametru przez współczynnik skalarny przed poddaniem drugiego parametru kwantyzacji według drugiego skalarnego schematu kwantyzacji.
- 4Sposób według zastrz. 1 - 2, przy czym niejednolite wielkości kroków drugiego skalarnego schematu kwantyzacji:są skalowane przez współczynnik skalowania przed kwantyzacją drugiego parametru;i/lub rosną wraz z wartością drugiego parametru.
- 5Sposób według dowolnego spośród powyższych zastrzeżeń, przy czym pierwszy skalarny schemat kwantyzacji:-16obejmuje więcej etapów kwantyzacji niż drugi skalarny etap kwantyzacji;i/lub skonstruowany jest przez przesunięcie, odbicie lustrzane i połączenie drugiego skalarnego schematu kwantyzacji.
- 6Koder audio (110) do kwantyzacji parametrów dotyczących parametrycznego kodowania przestrzennego sygnałów audio, obejmujący:komponent odbierający, zapewniony do odbierania co najmniej pierwszego parametru i drugiego parametru podlegającego kwantyzacji;pierwszy kwantyzujący komponent (202) zapewniony dalej za komponentem odbierającym skonfigurowanym do kwantyzowania pierwszego parametru na podstawie pierwszego skalarnego schematu kwantyzacji, mającego niejednolite wielkości kroków dla uzyskania kwantyzowanego pierwszego parametru, przy czym niejednolite wielkości kroków dobierane są tak, ze mniejsze wielkości kroków wykorzystywane są w przypadku zakresów pierwszego parametru, w przypadku którego ludzka percepcja dźwięku jest najbardziej wrażliwa, a większe wielkości kroków stosowane są w przypadku zakresów pierwszego parametru, w przypadku którego ludzka percepcja dźwięku jest mniej wrażliwa;komponent dekwantyzujący (204) skonfigurowany do odbierania pierwszego kwantyzowanego parametru z pierwszego komponentu kwantyzującego i do dekwantyzowania kwantyzowanego pierwszego parametru przy wykorzystaniu pierwszego skalarnego schematu kwantyzacji, dla uzyskania dekwantyzowanego pierwszego parametru, będącego aproksymacją pierwszego parametru;komponent wyznaczający współczynnik skalowania skonfigurowany do odbierania dekwanyzowanego pierwszego parametru, uzyskanie dostępu do funkcji skalowania, która mapuje wartości dekwantyzowanego pierwszego parametru na współczynniki skalowania, które zwiększają wielkości kroków odpowiednio do wartości dekwantyzowanego pierwszego parametru, i wyznaczanie współczynnika skalowania przez poddanie dekwantyzowanego pierwszego parametru funkcji skalowania;i drugi komponent kwantyzujący (212) skonfigurowany do odbierania drugiego parametru i współczynnika skalowania i kwantyzowania drugiego parametru na podstawie współczynnika skalowania i drugi skalarny schemat kwantyzacji, mający niejednolite wielkości kroków, dla uzyskania kwantyzowanego drugiego parametru.
- 7Sposób w dekoderze audio (120) dekwantyzacji kwantyzowanych parametrów dotyczących parametrycznego kodowania przestrzennego sygnałów audio, obejmujący:odbieranie co najmniej pierwszego kwantyzowanego parametru i drugiego kwantyzowanego parametru;-17dekwantyzowanie (304) kwantyzowanego pierwszego parametru na podstawie pierwszego skalarnego schematu kwantyzacji, mającego niejednolite wielkości kroków dla uzyskania dekwantyzowanego pierwszego parametru, przy czym niejednolite wielkości kroków dobierane są tak, że mniejsze wielkości kroków wykorzystywane są w przypadku zakresów pierwszego parametru, w przypadku którego ludzka percepcja dźwięku jest najbardziej wrażliwa, a większe wielkości kroków stosowane są w przypadku zakresów pierwszego parametru, w przypadku którego ludzka percepcja dźwięku jest mniej wrażliwa;uzyskanie dostępu do funkcji skalowania, która mapuje wartości dekwantyzowanego pierwszego parametru na współczynniki skalowania, które zwiększają wielkości kroków odpowiednio do wartości dekwantyzowanego pierwszego parametru, i wyznaczanie współczynnika skalowania przez poddanie dekwantyzowanego pierwszego parametru funkcji skalowania;i dekwantyzowanie (308) drugiego kwantyzowanego parametru na podstawie współczynnika skalowania i drugiego skalarnego schematu kwantyzacji, mającego niejednolite wielkości kroków, dla uzyskania dekwantyzowanego drugiego parametru.
- 8Sposób według zastrz. 7, przy czym funkcja skalowania jest funkcją odcinkowoliniową.
- 9Sposób według zastrz. 7 lub 8, przy czym etap dekwantyzacji drugiego parametru na podstawie współczynnika skalowania i drugiego skalarnego etapu kwantyzacji obejmuje dekwantyzację drugiego kwantyzowanego parametru według drugiego skalarnego schematu kwantyzacji i przemnożenie wyniku przez współczynnik skalowania.
- 10Sposób według zastrz. 7 lub 8, przy czym niejednolite wielkości kroków drugiego skalarnego schematu kwantyzacji:skalowane są przez współczynnik skalowania przed dekwantyzacją drugiego kwantyzowanego parametru;i/lub rosną wraz z wartością drugiego parametru.
- 11Sposób według zastrz. 7 - 10, przy czym pierwszy skalarny schemat kwantyzacji:obejmuje więcej etapów kwantyzacji niż drugi skalarny schemat kwantyzacji;i/lub skonstruowany jest przez przesunięcie, odbicie lustrzane i połączenie drugiego skalarnego schematu kwantyzacji.
- 12Sposób według zastrz. 7 - 11 lub sposób według zastrz. 1 - 5, przy czym największa wielkość kroku pierwszego i/lub drugiego skalarnego schematu kwantyzacji jest około -18cztery razy większa niż najmniejsza wielkość kroku pierwszego i/lub drugiego skalarnego schematu kwantyzacji.
- 13Nośnik odczytywany komputerowo obejmujący instrukcje kodu programowego dostosowane do realizacji sposobu według zastrz. 7 - 12, po wykonaniu przez urządzenie mające możliwości przetwarzające lub dostosowane do realizacji sposobu według zastrz. 1 - 5, po wykonaniu przez urządzenie mające możliwości przetwarzające.
- 14Dekoder audio (120) do dekwantyzacji kwantyzowanych parametrów dotyczących parametrycznego kodowania przestrzennego sygnałów audio, obejmujący:komponent odbierający skonfigurowany do odbierania co najmniej pierwszego kwantyzowanego parametru i drugiego kwantyzowanego parametru;pierwszy dekwantyzujący komponent (304) zapewniony dalej za komponentem odbierającym i skonfigurowany do dekwantyzowania pierwszego kwantyzowanego parametru na podstawie pierwszego skalarnego schematu kwantyzacji, mającego niejednolite wielkości kroków dla uzyskania dekwantyzowanego pierwszego parametru, przy czym niejednolite wielkości kroków dobierane są tak, ze mniejsze wielkości kroków wykorzystywane są w przypadku zakresów pierwszego parametru, w przypadku którego ludzka percepcja dźwięku jest najbardziej wrażliwa, a większe wielkości kroków stosowane są w przypadku zakresów pierwszego parametru, w przypadku którego ludzka percepcja dźwięku jest mniej wrażliwa;komponent wyznaczający współczynnik skalowania skonfigurowany do odbierania dekwanyzowanego pierwszego parametru z pierwszego komponentu dekwantyzującego, uzyskanie dostępu do funkcji skalowania, która mapuje wartości dekwantyzowanego pierwszego parametru na współczynniki skalowania, które zwiększają wielkości kroków odpowiednio do wartości dekwantyzowanego pierwszego parametru i wyznaczanie współczynnika skalowania przez poddanie dekwantyzowanego pierwszego parametru funkcji skalowania;i drugi komponent dekwantyzujący (308) skonfigurowany do odbierania współczynnika skalowania i drugiego kwantyzowanego parametru i dekwantyzowania drugiego kwantyzowanego parametru na podstawie współczynnika skalowania i drugi skalarny schemat kwantyzacji, mający niejednolite wielkości kroków, dla uzyskania dekwantyzowanego drugiego parametru.
- 15System kodowania/dekodowania audio zawierający koder audio według zastrz. 6 i dekoder audio według zastrz. 14, przy czym koder audio dostosowany jest do przesyłania pierwszego i drugiego parametru kwantyzowanego do dekodera audio. -19100 -20116 Fig. 2 Fig. 3 Fig. 4 b Niejednolity kwant zgrubny ACPL 2·--·—.—·-----·— --- 0 12 3 4 b Jednolity kwant zgrubny ACPL Jednolity kwant dokładny ACPL Fig. 5 Łft *» M7 ;.«0 r- -O m rst ' .'-«3 Fig. 7
Independent claims15
129 paragraphs, as filed
Description
References to related reports
[0001] This application claims the priority benefits of United States Provisional Patent Application No. 61 / 877,166, filed September 12, 2013.
Technical field
[0002] The invention relates generally to audio coding. It particularly relates to perceptually optimized quantization of parameters used in the system for parametric spatial coding of audio signals.
State of the art
[0003] The performance of a low bit audio coding system can be significantly improved for stereo signals when a parametric stereo (PS) coding tool is used. In this type of system, the mono signal is typically quantized and transferred using a state-of-the-art audio encoder, and the stereo parameters are estimated and quantized at the encoder and input as side information into the bitstream. In the decoder, the stereo signal is reconstructed from the decoded mono signal using stereo parameters.
[0004] There are several possible parametric stereo coding variants. Accordingly, there are several types of encoder and, in addition to the mono downmix, they generate various stereo parameters which are embedded in the integrated bitstream. The tools for this type of coding have also been standardized. An example of this type of standard is MPEG-4 Audio (ISO / IEC 14496-3).
[0005] The main aim of audio coding systems in general and parametric stereo coding in particular, and one of the several challenges in the art is to minimize the amount of information that has to be transmitted in the bitstream from encoder to decoder while ensuring good audio quality. A high degree of information compression in the data stream can lead to insufficient sound quality, both due to complex and insufficient computational methods, and also because information is lost during compression. A low level of information compression in the data stream may on the other hand lead to performance problems which may also translate into unacceptable audio quality.
[0006] Accordingly, there is a need for improved parametric stereo coding methods.
[0007] The International Search Report issued along with this application cites United States Application No. US 2007/0016416 A1, "document 416" as "a document defining a general art that is viewed as having no particular relevance." Document 416 discloses that a parameter that measures a characteristic channel or pair of channels relative to another channel of a multi-channel signal may be
Quantized more efficiently using a quantization rule that is generated from the association of a channel or channel pair energy measure and a multi-channel signal energy measure.
Brief description of the figures of the drawing
[0008] Preferred embodiments will be described below in more particular and with reference to the accompanying drawing figures, of which:
Fig. 1 shows a block diagram of a parametric stereo encoding and decoding system according to an embodiment;
Fig. 2 shows a block diagram relating to the processing of stereo parameters in the coding portion of the parametric stereo coding system shown in Fig. 1;
Fig. 3 shows a block diagram relating to the processing of stereo parameters in the decoding part of the parametric stereo coding system shown in Fig. 1;
Fig. 4 shows the value of the scaling factor s as a function of one of the stereo parameters;
Fig. 5 shows non-uniform and uniform quantizers (fine and coarse) in the plane (a, b), where a and b are stereo parameters; and
Fig. 6 is a diagram showing the average parametric stereo bit consumption related to the uniform, fine and uniform coarse quantization examples, as compared to non-uniform, fine and nonuniform coarse quantization according to a preferred embodiment.
Fig. 7 shows a block diagram of a parametric multi-channel coding and decoding system according to another preferred embodiment;
[0009] All figures are schematic and generally show only parts which are necessary to explain the disclosure, other parts may be omitted or merely suggested. The same numbers in the figures refer to the same parts unless otherwise indicated.
Detailed description
[0010] In view of the above, it is an object to provide encoders, decoders, systems including encoders and decoders as well as related methods, which ensure higher efficiency and quality of the encoded audio signal.
I. Summary - encoder
[0011] According to a first object, embodiments propose coding methods, coders and computer products for coding. The proposed methods, encoders, and computer program products may have substantially the same features and benefits.
[0012] According to embodiments, there is provided a method in an audio encoder for quantizing parameters related to parametric spatial coding of signals.
- 3audio, comprising: receiving at least a first parameter and a second quantizable parameter; quantizing the first parameter based on a first scalar quantization scheme, having non-uniform step sizes to obtain the quantized first parameter, with non-uniform step sizes being chosen such that smaller step sizes are used for the ranges of the first parameter where human perception of sound is most sensitive , and larger step sizes are used for the ranges of the first parameter, where the human perception of sound is less sensitive; dequantizing the first parameter using the first scalar quantization scheme to obtain a dequantized first parameter being an approximation of the first parameter; accessing a scaling function that maps the values of the dequantized first parameter to scaling factors that increase the step sizes according to the value of the dequantized first parameter, and determining a scaling factor by subjecting the dequantized first parameter of the scaling function; and quantizing the second parameter based on the scale factor and the second scalar quantization scheme having non-uniform step sizes to obtain the quantized second parameter.
[0013] The method is based on the understanding that the human perception of sound is not uniform. Instead, it turns out that human sound perception is higher for some characteristics and lower for others. This imposes that human perception of sound is more sensitive to some parameter values for parametric spatial coding of audio signals than for other such values. According to the provided method, a first such parameter is quantized with non-uniform step sizes, for example smaller step sizes which are used where human sound perception is most sensitive and larger step sizes which are used where human sound perception is less sensitive. Quantization using this type of non-uniform step size scheme allows the reduction of parametric average stereo bit consumption without reducing the perceptual quality of the sound.
[0014] According to embodiments, the scaling function of the method is a segment linear function.
[0015] According to embodiments, the method step of quantizing the second parameter is based on the scaling factor and the second scalar quantization scheme comprises dividing the second parameter by the scalar factor before subjecting the second quantizing parameter to the second scalar quantization scheme.
[0016] According to an alternative embodiment of the method, the non-uniform step sizes of the second scalar quantization scheme are scaled with a scaling factor before quantizing the second parameter.
[0017] According to the method embodiments, the non-uniform step sizes of the second scalar scalar quantization scheme increase with the value of the second parameter.
[0018] According to embodiments of the method, the first scalar quantization scheme comprises more quantization steps than the second scalar quantization scheme.
[0019] According to embodiments of the method, the first scalar quantization scheme is constructed by offsetting, mirroring, and combining a second scalar quantization scheme.
[0020] According to method embodiments, the largest step size of the first and / or second scalar quantization scheme is about four times greater than the smallest step size of the first and / or second scalar quantization scheme.
[0021] According to embodiments, a computer readable medium is provided including computer code instructions adapted to perform any method of the first item when implemented on a processing capable device.
[0022] According to embodiments, there is provided an audio encoder for quantizing parameters related to spatial parametric coding of audio signals, comprising: a receiving component provided for receiving at least a first parameter and a second quantized parameter; a first quantizing component provided downstream of the receiving component configured to quantize the first parameter based on the first scalar quantization scheme, having non-uniform step sizes to obtain the quantized first parameter, non-uniform step sizes being selected such that smaller step sizes are used for ranges of the first parameter. where human sound perception is most sensitive, and larger step sizes are used for ranges of the first parameter where human sound perception is less sensitive; a dequantization component configured to receive a first quantized parameter from the first quantization component and to dequantize the quantized first parameter using the first scalar quantization scheme to obtain a dequantized first parameter being an approximation of the first parameter; a scaling factor determining component configured to receive the dequantized first parameter, accessing a scaling function that maps the dequantized first parameter values to scaling factors that increase step sizes according to the dequantized first parameter value, and determining the scaling factor by subjecting the dequantized first parameter to the scaling function; and a second quantizing component configured to receive the second parameter and the scaling factor and quantize the second parameter based on the scale factor and the second
A scale quantization scheme having non-uniform step sizes to obtain a quantized second parameter.
II. Summary - decoder
[0023] According to a second object, embodiments propose decoding methods, decoders and decoding computer program products. The proposed methods, decoders and computer program products may have substantially the same features and benefits.
[0024] The features and configuration benefits presented in the general information about the encoder may be substantially important for the corresponding features and configuration of the decoder.
[0025] According to preferred embodiments, there is provided a method in an audio decoder for dequantizing quantized parameters related to parametric spatial coding of audio signals, comprising: receiving at least a first quantized parameter and a second quantized parameter; dequantizing the quantized first parameter based on a first scalar quantization scheme having non-uniform step sizes to derive the dequantized first parameter, where non-uniform step sizes are chosen such that smaller step sizes are used for ranges of the first parameter for which human perception of sound she is the most sensitive and larger step sizes are used for ranges of the first parameter where human sound perception is less sensitive; accessing a scaling function that maps the values of the dequantized first parameter to scaling factors that increase the step sizes according to the value of the dequantized first parameter, and determining a scaling factor by subjecting the dequantized first parameter of the scaling function; and dequantizing the second quantized parameter from the scaling factor and the second scalar quantization scheme having non-uniform step sizes to obtain the dequantized second parameter.
[0026] According to embodiments, the scaling function is a segmented linear function.
[0027] According to an embodiment, the step of dequantizing the second parameter from the scaling factor and the second scalar quantizing step comprises dequantizing the second quantized parameter according to the second scalar quantization scheme and multiplying the result by the scaling factor.
[0028] According to an alternative embodiment, the non-uniform step sizes of the second scalar quantization scheme are scaled with a scaling factor before the second quantized parameter is dequantized.
[0029] According to further embodiments, the non-uniform step sizes of the second scalar scalar quantization scheme increase with the value of the second parameter.
[0030] According to an embodiment, the first scalar quantization scheme comprises more quantization steps than the second scalar quantization scheme.
[0031] According to an embodiment, the first scalar quantization scheme is constructed by shifting, mirroring and combining the second scalar quantization scheme.
[0032] According to an embodiment, the largest step size of the first and / or second scalar quantization scheme is about four times greater than the smallest step size of the first and / or second scalar quantization scheme.
[0033] According to exemplary embodiments, a computer readable medium is provided including computer code instructions adapted to perform any method of the second subject when implemented on a processing capable device.
[0034] According to embodiments, there is provided an audio decoder for dequantizing quantized parameters related to parametric spatial coding of audio signals, comprising: a receiving component configured to receive at least a first quantized parameter and a second quantized parameter; a first dequantizing component provided downstream of the receiving component and configured to dequantize the first quantized parameter based on the first scalar quantization scheme having non-uniform step sizes to obtain the dequantized first parameter, whereby non-uniform step sizes are chosen such that smaller step sizes are used for the first quantization ranges. parameter, where human sound perception is most sensitive and larger step sizes are used for ranges of the first parameter where human sound perception is less sensitive; a scaling factor determining component configured to receive the dequantized first parameter from the first dequantizing component, accessing a scaling function that maps the values of the dequantized first parameter to scaling factors that increase step sizes according to the value of the dequantized first parameter, and determining the scaling factor by subjecting the dequantized first parameter scaling function parameter; and a second dequantizing component configured to receive the scaling factor and the second quantized parameter and dequantizing the second quantized parameter based on the scaling factor and a second scalar quantization scheme having non-uniform step sizes to obtain the dequantized second parameter.
III. Summary - Audio encoding / decoding system
[0035] According to a third object, embodiments propose decoding / coding systems comprising an encoder according to a first object and a decoder according to a second object.
[0036] The feature and configuration benefits disclosed in the general information about encoder and decoder may be substantially important for corresponding system features and configuration.
[0037] According to embodiments, there is provided a system in which an audio encoder is provided for transmitting the first and second quantized parameters to the audio decoder.
IV. Preferred Embodiments
[0038] The invention discloses perceptually optimized quantization of parameters used in the system for parametric spatial coding of audio signals. In the following examples, the special case of parametric stereo coding for 2-channel signals is discussed. The same technique can also be used for parametric multi-channel coding, e.g. in systems operating in the 5-3-5 mode. An embodiment of this type of system is shown in Fig. 7 and will be briefly discussed below. The embodiments presented herein relate to simple non-uniform quantization allowing bit rate reduction to collect such parameters without affecting perceived audio quality and allowing continuous use of fixed entropy coding techniques for scalar parameters (such as time difference or difference frequency coding followed by coding Huffman).
[0039] Fig. 1 shows a block diagram of a preferred embodiment of a stereo parametric coding and decoding system 100 discussed herein. A stereo signal comprising the left channel 101 (L) and the right channels 102 (R) is received by the encoder portion 110 of the system 100. The stereo signal is sent as an input signal to an ACPL (Advanced Coupling) encoder 112, generating a reduction in the number of channels. 103 mono (M) and stereo parameters a (denoted in Fig. 1 as 104a) and b (designated as 104b in FIG. 1). Additionally, the encoder portion 110 includes an encoder 114 for reducing the number of mono channels (DMXEnc) for converting the reduced number of mono channels 103 to a stream 105 bits, means 116 (Q) for quantizing a stereo parameter, generating a quantized parameter stream 106, and a multiplexer 118 (MUX) that generates the final bit stream 108, which also includes quantized stereo parameters that are transferred to decoder portion 120. The decoder portion 120 includes a demultiplexer 122 (DEMUX) which receives the incoming final bit stream 108 and re-generates the stream 105 bits and the quantized stream 106 stereo parameters, the decoder 124 down-channel (DMX Dec), which receives the stream 105 bits and generates the decoded reduced number channels 103 'mono (M<sup>1</sup>), stereo parameter dequantization means 126 (Q ') that receive the quantized stereo parameter stream 106 and generate quantized stereo parameters a' 104a 'and b' 104b 'and finally, an ACPO decoder 128 that receives a decoded reduced number of mono channels 103' and dequantized stereo parameters 104a 'and 104b' and converts these received signals into reconstructed stereo signals 101 '(L<sup>1</sup>) and 102 '(R.<sup>1</sup>).
[0040] Starting from the received stereo signals 101 (L) and 102 (R), the ACPL 112 encoder calculates the reduced number of mono (M) channels 103 and the side signal (S) according to the following equations:
M = (L + R) / 2 (equation 1)
S = (L - R) / 2 (equation 2)
[0041] The stereo parameters a and b are computed in a time and frequency selective manner, i.e. for each time / frequency tile, typically with a filter bank such as a QMF bank, and using non-uniform grouping of QMF bands to create perceptual parameter bands. frequency scale.
[0042] In the ACPL decoder, the decoded reduced number of mono channels M 'along with the stereo parameters a' and b 'and the decorrelated version M "(decorr (M')) are used as input values to reconstruct the side signal approximation based on the following equation:
S '= a' * M + b '* decorr (M') (equation 3)
[0043] Then L 'and R' are computed as:
L '= M' + S '(equation 4)
R '= M' - S '(equation 5)
[0044] A pair of parameters (a, b) can be viewed as a point in the two-dimensional plane (a, b). The parameters a, b are related to the received stereo image, where the parameter a is mainly related to the position of the received sound source (e.g. left or right), and where the parameter b is mainly related to the size or width of the received sound source (small and well focused or wide and enveloping). Table 1 lists some typical examples of received stereo images and the corresponding values for the parameters a, b.
Table 1.
<td>Point</td><td>Parameter values</td><td>Signal description</td>
<td>Left</td><td>a = l, b = 0</td><td>Signal completely shifted to the left side, ie R = 0</td>
<td>Center</td><td>a = 0, b = 0</td><td>Signal in the phantom center, i.e. L = R.</td>
<td>Right</td><td>a = -l, b = 0</td><td>Signal completely shifted to the right side, ie L = 0</td>
<td>Widely</td><td>a = 0, b = 1</td><td>Broad signal, L and R are not correlated and have the same level.</td>
[0045] Note that b is never negative. It should also be noted that even though b and the absolute value of a are often in the range 0 to 1, they can also have absolute values greater than 1, e.g. for strong phase-shifted components in L and R, i.e. the correlation between L and R is negative.
[0046] The problem at this point is in designing a technique for quantizing the parameters a, b for transmission as side information in a parametric stereo / spatial coding system. A simple approach available in the art is to use uniform quantization and quantize a and b independently, i.e. to use two scalar quantizers. A typical quantization step size is delta = 0.1 for fine quantization or delta = 0.2 for coarse quantization. The lower left and right panels in Fig. 5 show points in the plane (a, b) that can be represented by this kind of quantization scheme for both fine and coarse quantization. Typically, the quantized parameters a and b are independently entropy-coded using a differential-time or difference-frequency coding in conjunction with Huffman coding.
[0047] However, the inventors have found that the performance (in terms of rate disturbance) of the parameter quantization can be improved over scalar quantization considering perceptual aspects. Especially the sensitivity of the human auditory system to small changes in parameter values (such as errors introduced by quantization) depends on the position in the plane (a, b). Repeated experiments to test the audibility of slight changes or "barely perceptible differences" (JND) just noticable difference) indicate that the JNDs in the case of a and b are substantially smaller than for the sources the sounds with a perceivable stereo image which is represented by points (1, 0) and (-1, 0) in the plane (a, b). Thus, the uniform quantization of a and b can be too coarse (with audible artifacts) in the areas near (1, 0) and (-1, 0) and unnecessarily accurate (causing unnecessarily fast side information transfer) in other areas such as nearby (0, 0) and (0, 1). It is also possible to consider a vector quantizer for (a, b) to obtain a common and non-uniform quantization of the stereo parameters a and b. However, the stereo quantizer is computationally more complex and furthermore entropy coding (difference-time or -frequency) should be adapted which would also more complicated.
[0048] Accordingly, in this application a new non-uniform quantization scheme for the parameters a and b is introduced. The non-uniform quantization scheme for a and b uses a location dependent JND (such as a vector quantizer), but may be implemented as a minor modification of the uniform and independent quantization of a and b according to state of the art. Moreover, the time-time or frequency-difference entropy coding may also remain unchanged. Only the Huffman code book will need to be updated to reflect changes in the index ranges and symbol probability.
[0049] The resulting quantization scheme is shown in Figs. 2 and 3, wherein Fig. 2 relates to a means 116 for quantizing the stereo parameters of the encoder part 110, and Fig. 3 relates to a means 126 for dequantizing the stereo parameters of the decoder part 120. The quantization scheme of the stereo parameters starts from applying non-uniform scalar quantization of the parameter a (numbered 104a in Fig. 2) to the quantization means Q (numbered 202 in Fig. 2). The quantized parameter 106a is forwarded to the multiplexer 118. The quantized parameter is also dequantized directly.
-10in the means of dequantization of Q<sub>and</sub>'' (labeled 204 in Fig. 2) to parameters a '. Since quantized parameter 106a is dequantized to a '(denoted by 104a' in FIG. 3) in the decoder portion 120, a 'will be identical in both the encoder portion 110 and the decoder portion 120 of system 100. Then a' is used to calculate the scaling factor s (implemented by scaling means 206), which is used to quantize b regardless of the actual value of a. The parameter b (in Fig. 2 104b) is divided by the scaling factor s (performed by inverting means 208 and multiplier 210) and is then sent to another non-uniform scalar quantizer Qb (numbered 212 in FIG. 2) from which the quantized parameter 106b is transmitted forward. The process is partially reversed at means 126 to dequantize the stereo parameters, as in Fig. 3. The received quantized parameters 106a and 106b are dequantized in a dequantizing means Q<sub>and</sub><sup>1</sup> (numbered 304 in Figure 3) and Qb '(designated 308 in Figure 3) to a' (designated 104a 'in Figure 3) and b', previously divided by the scaling factor s in the encoder portion 110. Scaling Means 306 determine the scaling factor s from the dequantized parameter a '(104a) in the same way as the scaling means 206 in the encoder portion 110. The scaling factor is then multiplied by the dequantization result of the quantized parameter 106b in the multiplier 310 to obtain the dequantized parameter b '(denoted in FIG. 3 by 104b'). Correspondingly, the dequantization of a and the scaling factor calculations are implemented in both the encoder 110 and decoder portion 120, ensuring that exactly the same s value is used for encoding and decoding b.
Non-uniform quantization of a and b is based on a common non-uniform quantizer for values in the range 0 to 1, where the quantization step sizes for a value of about 1 are four times the value for the quantization step sizes for a value of about 0, and the quantization step size increases with the parameter value . For example, the quantization step size may increase practically linearly with an index identifying the corresponding dequantized value. For a quantizer with 8 intervals (i.e. 9 indices), the following values can be obtained, with the quantization step size being the difference between two adjacent dequantized values.
Table 2: Dequantized values from 0 to 1
<td>Table of Contents</td><td>Value</td><td>Table of Contents</td><td>Value</td>
<td> 0</td><td> 0</td><td> 5</td><td> 0,4844</td>
<td> 1</td><td> 0,0594</td><td> 6</td><td> 0,6375</td>
<td> 2</td><td> 0,1375</td><td> 7</td><td> 0,8094</td>
<td> 3</td><td> 0,2344</td><td> 8</td><td> 1,0000</td>
<td> 4</td><td> 0,3500</td><td></td><td></td>
[0051] The above table is an example of a quantization scheme that can be used as a Qb 'dequantization means (reference number 308 in Fig. 3). However, in the case of
-11 parameter and a larger range of values should be handled. Example of a quantization scheme for the Q dequantization means<sub>and</sub>'' (numbered 304 in Figure 3) may simply be constructed by mirroring and combining non-uniform quantization intervals shown in Table 2 above to provide a quantizer that represents values in the range -2 to 2, where the quantization step size for the value about -2, 0 and 2 are about four times larger than the quantization step size for the -li values. The resulting values are shown in the following table 3.
Table 3: Dequantized values in the range -2 to 2
<td>Table of Contents</td><td>Value</td><td>Table of Contents</td><td>Value</td>
<td> 0</td><td> -2,000</td><td> 17</td><td> 0,1906</td>
<td> 1</td><td> -1,8094</td><td> 18</td><td> 0,3625</td>
<td> 2</td><td> -1,6375</td><td> 19</td><td> 0,5156</td>
<td> 3</td><td> -1,4844</td><td> 20</td><td> 0,6500</td>
<td> 4</td><td> -1,3500</td><td> 21</td><td> 0,7656</td>
<td> 5</td><td> -1,2344</td><td> 22</td><td> 0,8625</td>
<td> 6</td><td> -1,1375</td><td> 23</td><td> 0,9406</td>
<td> 7</td><td> -1,0594</td><td> 24</td><td> 1,000</td>
<td> 8</td><td> -1,000</td><td> 25</td><td> 1,0594</td>
<td> 9</td><td> -0,9406</td><td> 26</td><td> 1,1375</td>
<td> 10</td><td> -0,8625</td><td> 27</td><td> 1,2344</td>
<td> 11</td><td> -0,7656</td><td> 28</td><td> 1,3500</td>
<td> 12</td><td> -0,6500</td><td> 29</td><td> 1,4844</td>
<td> 13</td><td> -0,5156</td><td> 30</td><td> 1,6375</td>
<td> 14</td><td> -0,3625</td><td> 31</td><td> 1,8094</td>
<td> 15</td><td> -0,1906</td><td> 32</td><td> 2,000</td>
<td> 16</td><td> 0</td><td></td><td></td>
[0052] Fig. 4 shows the value of the scaling factor s for a. It is a segmented linear function where s = 1 (i.e. no scaling) for a = -1 and a = 1 and s = 4 (4 times coarse quantization b) for a = -2, a = 0 and a = 2. It should be emphasized that the function shown in Fig. 4 is an example and that other functions are also theoretically possible. The same reasoning applies to quantization schemes.
[0053] The resultant non-uniform quantization of a and b is shown in the upper left pane of Fig. 5, where each point in the plane (a, b) that may be represented by such quantizer is marked with a cross. Around the most sensitive points (1, 0) and (-1, 0),
The quantization step size for both a and b is about 0.06, while a and b around (0, 0) is 0.2. Thus, the quantization steps are much more suited to JND than those for unitary scalar quantization a and b.
[0054] If a more coarse quantization is required, it is possible to simply omit every second dequantized value of non-uniform quantizers, thus doubling the quantization step size. Table 4 shows the following rough non-uniform quantizers for the parameter b, and the non-uniform quantizers for the parameter a were obtained analogously to what is shown above.
Table 4: Dequantized values for coarse dequantization from 0 to 1
<td>Table of Contents</td><td>Value</td>
<td> 0</td><td> 0</td>
<td> 1</td><td> 0,1375</td>
<td> 2</td><td> 0,3500</td>
<td> 3</td><td> 0,6375</td>
<td> 4</td><td> 1,000</td>
[0055] The scaling function shown in Fig. 4 remains unchanged for coarse quantization and the resulting coarse quantizer for (a, b) is shown in the upper right pane in Fig. 5. Such coarse quantization may be desirable if the coding system is operated at very low target transfer rates, where it may be advantageous to spend the saved bits by coarse quantization of the stereo parameters to encode the signal of the reduced number of mono M channels (indicated by 103 in FIG. 1).
[0056] The performance difference between non-uniform and uniform quantization of stereo parameters a and b is shown in Fig. 6. The differences are shown for fine and coarse quantization. The average bit consumption per second corresponding to 11 hours of music is shown. From this figure, it can be concluded that the bit consumption for nonuniform quantization is substantially lower than for uniform quantization. Furthermore, it can be concluded that coarse nonuniform quantization reduces bit-per-second consumption to a greater extent than coarse uniform quantization.
[0057] Finally, in Fig. 7, a block diagram is shown, an embodiment of a 5-3-5 parametric multi-channel coding and decoding system 700. A multi-channel signal including the left front channel 701, the left surround channel 702, the front center channel 703, the right front channel 704 and the right surround channel 705 is received by the encoder portion 710 of the system 700. The left front channel 701 signals of the left surround channel 702 are sent as input to the first advanced coupling encoder (ACPL) 712 generating the left reduced number of channels 706 and stereo parameters ai. (designated with number 708a) and bt (designated with number 708b). Similarly, the right front channel signals 704 of the right surround channel 705 are transmitted as input to
The second Advanced Coupling Encoder (ACPL) 713 generating a right-hand reduced number of channels 707 and the stereo parameters aR (indicated by number 709a) and Or (indicated by number 709b). In addition, the encoder portion 710 includes a 3-channel reduction encoder 714 processing the left channel reduction 706, center front channels 703 and right channel reduction signals 707 per bit stream 722, first stereo parameter quantization means 715, generating a first stereo quantized stream 720 , based on stereo parameters 708a and 708b, second stereo parameter quantization means 716, generating a second stream of quantized stereo parameters 724, based on stereo parameters 709a and 709b, and a mux 730 that generates a final bit stream 735 that also includes quantized stereo parameters that are transferred to decoder portion 740. Decoder portion 740 includes a demultiplexer 742 that receives the incoming final data stream 735 and generates a data stream 722, first quantized stereo parameters 720 and second quantized stereo parameters 724. The first stream of quantized stereo parameters 720 is received by first stereo parameter dequantization means 745 that generates dequantized stereo parameters 708a 'and 708b'. A second stream of quantized stereo parameters 724 is received by second stereo parameter dequantization means 746 that generates stereo dequantized parameters 709a 'and 709b';. Data stream 722 is received by a 3-channel decoder 744 that forms a re-generated left reduced number of channels 706 ', a reconstructed center front channel 703' and a re-generated right reduced number of channels 707 '. The first ACPL decoder 747 receives the dequantized stereo parameters 708a 'and 708b' and regenerates the left reduced number of channels 706 'and generates a reconstructed left front channel 701' and a reconstructed left surround channel 702 '. Similarly, the second ACPL decoder 748 receives the dequantized stereo parameters 709a ', 709b' and the regenerated right reduced number of channels 707 'and generates a reconstructed right front channel 704' and a reconstructed right surround channel 705 '.
Equivalents, Extensions, Alternatives and Miscellaneous
[0058] Further preferred embodiments of the invention will become apparent to those skilled in the art from the foregoing description. Although the description and figures of the drawings disclose preferred embodiments, the invention is not limited to these specific examples. Many modifications and variations can be made without departing from the scope of the invention as defined in the appended claims. The reference marks appearing in the claims are not to be construed as limiting.
[0059] Moreover, those skilled in the art will appreciate from the figures, description and appended claims that variations can be made to the illustrated preferred embodiments. In the claims, the term "comprising" does not exclude other elements or steps, and the singular does not exclude plurality. Although some measures are mentioned separately in the dependent claims, this does not mean that they cannot be applied simultaneously.
[0060] The systems and methods disclosed above may be implemented with programs, firmware, hardware, or a combination thereof. In the case of a hardware implementation, the division of tasks among the functional units described in the above description does not necessarily correspond to the division into physical units; on the contrary, one physical component may have multiple functions and one task may be performed by multiple physical components working together. Some or all components may be implemented as a program using a digital signal processor or microprocessor, or implemented in hardware or as an application specific integrate circuit. This type of software can be distributed over computer-readable media, which may include computer storage media (or tangible media) and communication media (or volatile media). As is well known to those skilled in the art, computer storage media includes both volatile and non-volatile removable and non-removable media, implemented by any method or technology for storing information, e.g., computer readable instructions, data structures of program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash or other memory technologies, CD ROM, DVD ( digital versatile disk) or other types of optical memories, magnetic cassettes, magnetic tapes, magnetic disk memories, or other types of magnetic storage devices or other media that can be used to save the required information and that can be accessed from computer. Moreover, those skilled in the art know that communication media typically includes computer readable instructions, data structures, program modules, or other data in a modulated data signal, e.g., a carrier wave or other transport mechanism, and includes an information delivery medium.
20 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20
55 members in 23 offices
Priority claims8
| Document | Office | Kind | Date |
|---|---|---|---|
| 201361877166 | United States of America | P | |
| 201361877166 | United States of America | P | |
| 14761831 | European Patent Office (EPO) | A | |
| 2014069040 | European Patent Office (EPO) | W | |
| 2014069040 | European Patent Office (EPO) | W | |
| EP20140761831 | – | – | – |
| US201361877166P | – | – | – |
| WO2014EP69040 | – | – | – |
Members55
| Document | Office | Kind | |
|---|---|---|---|
| CA2922256A1 | Canada | A1 | |
| WO2015036349A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW201521011A | Taiwan Province of China | A | |
| AU2014320538A1 | Australia | A1 | |
| SG11201601144WA | Singapore | A | |
| AR097618A1 | Argentina | A1 | |
| KR20160042113A | Republic of Korea | A | |
| IL244153A0 | Israel | A0 | |
| IL244153D0 | Israel | D0 | |
| CN105531763A | China | A | |
| MX2016002793A | Mexico | A | |
| EP3044788A1 | European Patent Office (EPO) | A1 | |
| US2016217800A1 | United States of America | A1 | |
| IL244153A | Israel | A | |
| JP2016531327A | Japan | A | |
| CL2016000571A1 | Chile | A1 | |
| HK1220037A | Hong Kong, China | A | |
| HK1220037A1 | Hong Kong, China | A1 | |
| TWI579831B | Taiwan Province of China | B | |
| US9672837B2 | United States of America | B2 | |
| EP3044788B1 | European Patent Office (EPO) | B1 | |
| CN105531763B | China | B | |
| US2017238208A1 | United States of America | A1 | |
| RU2628898C1 | Russian Federation | C1 | |
| KR101777631B1 | Republic of Korea | B1 | |
| JP6201057B2 | Japan | B2 | |
| AU2014320538B2 | Australia | B2 | |
| CA2922256C | Canada | C | |
| DK3044788T3 | Denmark | T3 | |
| ES2645839T3 | Spain | T3 | |
| PL3044788T3This record | Poland | T3 | |
| NO2996227T3 | Norway | T3 | |
| UA116482C2 | Ukraine | C2 | |
| EP3321932A1 | European Patent Office (EPO) | A1 | |
| MX356805B | Mexico | B | |
| US10057808B2 | United States of America | B2 | |
| HK1247432A | Hong Kong, China | A | |
| HK1247432A1 | Hong Kong, China | A1 | |
| US2018352475A1 | United States of America | A1 | |
| US10383003B2 | United States of America | B2 | |
| US2019320348A1 | United States of America | A1 | |
| US10694424B2 | United States of America | B2 | |
| EP3321932B1 | European Patent Office (EPO) | B1 | |
| US2020389815A1 | United States of America | A1 | |
| BR112016005192B1 | Brazil | B1 | |
| AR115819A2 | Argentina | A2 | |
| AR115820A2 | Argentina | A2 | |
| MY187124A | Malaysia | A | |
| US11297533B2 | United States of America | B2 | |
| US2022295347A1 | United States of America | A1 | |
| US11838798B2 | United States of America | B2 | |
| US2024155427A1 | United States of America | A1 | |
| MY204045A | Malaysia | A | |
| US12213004B2 | United States of America | B2 | |
| US2025261039A1 | United States of America | A1 |
Numbers
- Publication, DOCDB
- 3044788
- Publication, EPODOC
- PL3044788T
- Application
- 761831
- Application, DOCDB
- 14761831
- Application, EPODOC
- PL20140761831T
Titles2
- English
- NON-UNIFORM PARAMETER QUANTIZATION FOR ADVANCED COUPLING
- Polish
- Niejednolita kwantyzacja parametrów dla zaawansowanego sprzęgania
Classification
- CPC, 8
- G10L19/008
- H04W28/065
- G10L19/035
- H03M7/30
- H04S1/007
- H04S2420/03
- G10L19/038
- H04W28/0215
- IPC, 2
- G10L19 035
- G10L19 008
