Efficient and scalable parametric stereo coding for low bitrate audio coding applications
Abstract
Method for encoding a power spectral envelope of a stereophonic audio signal or of a multi-channel audio signal having two channels, the two channels having a set of frequency bands, comprising: calculating an audio signal balance parameter Stereophonic or two channels for each frequency band and a level parameter that represents the total power of the two channels for each frequency band.

Term
Term ended
Projected expiry passed 10 July 2022, 4.2 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
24 claims: 12 independent, 12 dependent
- 1ES 2 344 145 T3 ES 2 344 145 T3 CLAIMS REIVINDICACIONES 1. Method for encoding a power spectral envelope of a stereophonic audio signal or of a multichannel audio signal having two channels, the two channels having a set of frequency bands, comprising:1. Método para codificar una envolvente espectral de potencia de una señal de audio estereofónica o de una señal de audio multicanal que tiene dos canales, teniendo los dos canales un conjunto de bandas de frecuencia, que comprende: calcular un parámetro de equilibrio de la señal de audio estereofónica o de los dos canales para cada banda de frecuencia y un parámetro de nivel que representa la potencia total de los dos canales para cada banda de frecuencia. calculating a balance parameter of the stereophonic or two-channel audio signal for each frequency band and a level parameter representing the total power of the two channels for each frequency band.
- 4Método según una cualquiera de las reivindicaciones precedentes, que comprende además la siguiente etapa:Four. Method according to any one of the preceding claims, further comprising the following step: adaptively calculate a first channel level and a second channel level instead of calculating the level parameter and the balance parameter. calcular de manera adaptativa un nivel del primer canal y un nivel del segundo canal en lugar de calcular el parámetro de nivel y el parámetro de equilibrio.
- 5Method according to one of the preceding claims, further comprising the following step:5. Método según una de las reivindicaciones precedentes, que comprende además la siguiente etapa: convert the level parameter to a dB representation using an arbitrary reference power and convert the balance parameter to a dB representation. convertir el parámetro de nivel en una representación en dB usando una potencia de referencia arbitraria y convertir el parámetro de equilibrio en una representación en dB.
- 7Method according to one of the preceding claims, further comprising the following steps:7. Método según una de las reivindicaciones precedentes, que comprende además las siguientes etapas: codificar a modo delta y codificar a modo Huffman los parámetros de equilibrio y los parámetros de nivel;y transmitir y almacenar los parámetros con codificación delta y con codificación Huffman. delta encoding and Huffman encoding equilibrium parameters and level parameters;and transmitting and storing the delta-encoded and Huffman-encoded parameters.
- 8Method according to one of the preceding claims, in which the equilibrium parameter is progressively quantified. 8. Método según una de las reivindicaciones precedentes, en el que el parámetro de equilibrio se cuantifica de manera progresiva.
- 9Method according to one of the preceding claims, in which the level parameter is not progressively quantized. 9. Método según una de las reivindicaciones precedentes, en el que el parámetro de nivel no se cuantifica de manera progresiva.
- 10Method according to one of the preceding claims, further comprising the following step:10. Método según una de las reivindicaciones precedentes, que comprende además la siguiente etapa: codificar a modo delta el parámetro de equilibrio de manera adaptativa en tiempo o en frecuencia, en el que la codificación delta en tiempo se usa cuando una fuente tiene una característica en tiempo más estacionaria y una mayor radiación no uniforme, y en el que la codificación delta en frecuencia se usa cuando la fuente tiene una característica en tiempo menos estacionaria y una menor radiación no uniforme. delta encode the equilibrium parameter in time or frequency adaptive fashion, where delta time encoding is used when a source has a more stationary time characteristic and higher non-uniform radiation, and where encoding Frequency delta is used when the source has a less stationary time characteristic and less non-uniform radiation.
- 11Aparato para codificar una envolvente espectral de potencia de una señal de audio estereofónica o una señal de audio multicanal o la señal que tiene dos canales, teniendo los dos canales un conjunto de bandas de frecuencia, que comprende:eleven. Apparatus for encoding a power spectral envelope of a stereophonic audio signal or a multichannel audio signal or the signal having two channels, the two channels having a set of frequency bands, comprising: a calculator for calculating a balance parameter for each frequency band and a level parameter representing the total power of the stereophonic signal or the two channels for each frequency band. un calculador para calcular un parámetro de equilibrio para cada banda de frecuencia y un parámetro de nivel que representa la potencia total de la señal estereofónica o los dos canales para cada banda de frecuencia. ES 2 344 145 T3 ES 2 344 145 T3
- 12Encoder comprising:12. Codificador que comprende: a low-band encoder for encoding the stereophonic audio signal or a multi-channel audio signal having two channels to obtain a low-band encoded output signal;and a parametric stereo encoder for estimating a high-band power spectral envelope of the signal, the parametric stereo encoder having an apparatus for encoding according to claim 11. un codificador de banda baja para codificar la señal de audio estereofónica o una señal de audio multicanal que tiene dos canales para obtener una señal de salida codificada de banda baja;y un codificador estereofónico paramétrico para estimar una envolvente espectral de potencia de banda alta de la señal, teniendo el codificador estereofónico paramétrico un aparato para codificar según la reivindicación 11.
- 18Encoding method, comprising:18. Método de codificación, que comprende: codificar la señal estereofónica de audio o una señal de audio multicanal que tiene dos canales para obtener una señal de salida codificada de banda baja;y estimar una envolvente espectral de potencia de banda alta de la señal, en el que en la etapa de estimación se realiza un método según una de las reivindicaciones 1 a 10. encoding the stereophonic audio signal or a multichannel audio signal having two channels to obtain a low-band encoded output signal;and estimating a high-band power spectral envelope of the signal, wherein in the estimation step a method according to one of claims 1 to 10 is performed.
- 19Decoder for decoding an encoded audio bitstream, comprising:19. Decodificador para decodificar una corriente de bits de audio codificada, que comprende: a demultiplexer for demultiplexing the encoded bitstream to obtain a low-band center decoder signal and a high-band level parameter, the high-band level parameter representing the total power of two channels of a signal in a frequency band of the high band of the signal that has the two channels;un demultiplexor para demultiplexar la corriente de bits codificada para obtener una señal de decodificador central de banda baja y un parámetro de nivel de banda alta, representando el parámetro de nivel de banda alta la potencia total de dos canales de una señal en una banda de frecuencia de la banda alta de la señal que tiene los dos canales;a central low-band decoder for producing a low-band output signal, the low-band output signal having a low-band mono signal or a low-band stereo signal;and a high frequency reconstruction unit for generating a synthetic high band using the low band output signal and the high band level parameter and for combining the synthetic high band and the low band output signal. un decodificador central de banda baja para producir una señal de salida de banda baja, teniendo la señal de salida de banda baja una señal monofónica de banda baja o una señal estereofónica de banda baja;y una unidad de reconstrucción de alta frecuencia para generar una banda alta sintética usando la señal de salida de banda baja y el parámetro de nivel de banda alta y para combinar la banda alta sintética y la señal de salida de banda baja.
- 24Method for decoding an encoded audio bitstream, comprising:24. Método para decodificar una corriente de bits de audio codificada, que comprende: demultiplexar la corriente de bits codificada para obtener una señal de decodificador central de banda baja y un parámetro de nivel de banda alta, representando el parámetro de nivel de banda alta la potencia total de dos canales de una señal en una banda de frecuencia de la banda alta de la señal que tiene los dos canales;demultiplexing the encoded bitstream to obtain a low-band center decoder signal and a high-band level parameter, the high-band level parameter representing the total power of two channels of a signal in a frequency band of the band high signal that has both channels;decodificar la señal de decodificador central de banda baja para producir una señal de salida de banda baja, teniendo la señal de salida de banda baja una señal monofónica de banda baja o una señal estereofónica de banda baja;y generar mediante reconstrucción de alta frecuencia una banda alta sintética usando la señal de salida de banda baja y el parámetro de nivel de banda alta y combinando la banda alta sintética y la señal de salida de banda baja. decoding the low-band center decoder signal to produce a low-band output signal, the low-band output signal having a low-band mono signal or a low-band stereo signal;and generating by high frequency reconstruction a synthetic high band using the low band output signal and the high band level parameter and combining the synthetic high band and the low band output signal.
Independent claims12
40 paragraphs in 4 sections, as filed
ES 2 344 145 T3
DESCRIPTION
Efficient and expandable parametric stereo coding for low bit rate applications.
The present invention relates to low bit rate audio source coding systems. Various parametric representations of stereophonic properties of an input signal are introduced and the application of them on the decoder side is explained, ranging from pseudo stereophonic encoding to full stereophonic encoding of spectral envelopes, the latter of these being especially suitable for encoders- HFR-based decoders.
Audio source encoding techniques can be divided into two classes: natural audio encoding and speech encoding. At medium to high bit rates, natural audio coding is typically used for voice and music signals, and stereophonic transmission and playback are possible. In applications where only low bit rates are available, for example, in Internet audio transmissions directed at users with slow modem dial-up connections, or in emerging digital AM broadcasting systems, monophonic encoding of the device is unavoidable. audio program material. However, a stereophonic feel can still be desired, particularly when listening with headphones, in which case a pure monophonic signal is perceived as coming from "inside the head", which can be an unpleasant experience.
One approach to dealing with this problem is to synthesize a stereo signal on the decoder side from a received pure mono signal. Several different "pseudo stereo" generators have been proposed over the years. For example, US Patent 5,883,962 describes enhancing mono signals by adding lag / lag versions of a signal to the raw signal, thereby creating a stereo illusion. With this, the processed signal is added to the original signal for each of the two outputs at equal levels but with opposite signs, ensuring that the enhancement signals are canceled if the two channels are added later to the signal path. A similar system is shown in PCT WO 98/57436, albeit without the above monophonic compatibility of the enhanced signal. The prior art methods have in common that they are applied as post-process only. In other words, no information about the degree of stereo width is provided to the decoder, leaving aside the position on the stereo sound stage. In this way, the pseudo-stereo signal may or may not resemble the stereo character of the original signal. A particular situation where prior art systems are deficient is when the original signal is a pure mono signal, which is often the case in voice recordings. This mono signal is blindly converted to a synthetic stereo signal at the decoder, which in the case of speech causes disturbing artifacts and can reduce the clarity and intelligibility of speech. Document J. Herre et al. "Intensity Stereo Coding", preprint # 3799 filed at the AES Convention, February 26, 1994, discloses an audio coding method for a stereo source that involves determining the directional parameter from which factors can be derived. scale of a given frequency band on each channel. This document does not indicate how a band-limited encoded signal can be reconstructed into a wider-band decoded stereo audio signal.
Other prior art systems aimed at true stereophonic transmission at low bit rates typically employ an addition and subtraction coding scheme. In this way, the original left (L) and right (R) signals are converted into an addition signal, S = (L + R) / 2, and a subtraction signal, D = (LR) / 2, and then they are coded and processed. The receiver decodes the S and D signals, recreating the original L / R signal through the operations L = S + D, and R = S - D. The advantage of this is that a redundancy is very often found between L and R, with less information in D to be encoded, requiring fewer bits, than in S. Clearly, the extreme case is a pure monophonic signal, that is, L and R are identical. A conventional L / R codec encodes this mono signal twice, while an S / D codec detects this redundancy, and the D signal does not require (ideally) any bits at all. Another extreme is represented by the situation in which R = -L, corresponding to “out of phase” signals. Now the S signal is zero, while the D signal computes for L. Again, the S / D scheme has a clear advantage over standard L / R encoding. However, consider the situation where, for example, R = 0 during a transition, which was not uncommon in the early days of stereo recordings. S and D are equal to L / 2, and the S / D scheme offers no advantage. In contrast, L / R encoding handles this very well: the R signal does not require any bits. For this reason, prior art codec-decoders employ adaptive switching between these two coding schemes, depending on which method is most beneficial to use at any given time. The above examples are purely theoretical (except in the dual mono case, which is common in voice-only programs). In this way, real-world stereo program material contains significant amounts of stereo information, and even if the above switching is carried out, the resulting bit rate is often still too high for many applications. Furthermore, as can be seen from the resynthesizing relationships above, a very imprecise quantization of the D signal in an attempt to further reduce the bit rate is not feasible since quantization errors translate into level errors that they cannot be neglected in the L and R signals.
It is an object of the present invention to provide an improved concept for encoding a power spectral envelope or decoding an encoded bit stream.
ES 2 344 145 T3
This object is achieved by an encoding method according to claim 1, an encoding apparatus according to claim 11, a decoder according to claim 19 or a method of decoding an encoded bit stream according to claim 24. In embodiments, it is used a detection of stereophonic properties of signals before encoding and transmission. In the simplest form, a detector measures the amount of stereo perspective that is presented in the input stereo signal. This quantity is then transmitted as a stereo width parameter, along with a coded mono sum of the original signal. The receiver decodes the mono signal and applies the appropriate amount of stereo width using a pseudo stereo generator that is controlled by this parameter. As a special case, a mono input signal is signaled as zero stereo amplitude and, correspondingly, no stereo synthesis is applied at the decoder. According to the invention, useful measurements of stereophonic amplitude can be obtained, for example, from the differential signal or from the cross-correlation of the original left and right channel. The value of these calculations can be represented in a small number of states that are transmitted at a suitable fixed rate in time, or on a basis according to need. The invention also teaches how to filter synthesized stereo components to reduce the risk of unmasking coding artifacts that are normally associated with low bit rate coding signals.
Alternatively, the global stereo balance or location in the stereo field is detected at the encoder. This information, optionally together with the amplitude parameter above, is efficiently transmitted as a balance parameter, along with the encoded mono signal. In this way, offsets to either side of the sound stage can be recreated in the decoder by correspondingly altering the gains of the two output channels. This stereophonic-balance parameter can be obtained from the ratio of the left and right signal strengths. Transmission of the two types of parameters requires very few bits compared to full stereo coding, thus keeping the overall bit rate demand low. In a more elaborate version of the invention, offering a more accurate parametric stereo description, several stereo width and balance parameters are used, each representing independent frequency bands.
The balance parameter, generalized to a frequency band operation, together with a corresponding band operation of a level parameter, calculated as the sum of the left and right signal powers, allows a new representation, arbitrarily detailed, of the power spectral density of a stereophonic signal. A particular benefit of this representation, in addition to the benefits of stereo redundancy, which S / D systems also take advantage of, is that the equilibrium signal can be quantized less precisely than the said level given that the quantization error, when converted back to a stereophonic spectral envelope, it causes a "space error", that is, the perceived location in the stereophonic panorama, rather than a level error. Analogous to a traditional switched L / R and S / D system, the level / balance scheme can be adaptively disrupted in favor of an IL level / IR level signal, which is most efficient when the overall signal is strongly out of phase toward any channel. The above spectral envelope coding scheme can be used whenever efficient coding of power spectral envelopes is required, and can be incorporated as a tool in newer stereo source codecs. A particularly interesting application is in HFR systems that are guided by information about the high band envelope of the original signal. In such a system, the low band is encoded and decoded by means of an arbitrary codec, and the high band is regenerated in the decoder using the decoded low-band signal and the transmitted high-band envelope information [PCT WO 98/57436]. In addition, the possibility is offered to build an expandable HFR-based stereo codec by locking the envelope encoding to level / balance operation. Hereby, the level values are fed into the primary bit stream which, depending on the implementation, normally decodes to a mono signal. The balance values are fed into the secondary bit stream that is supplied, in addition to the primary stream of bits, to receivers close to the transmitter, taking as an example an IBOC (In-Band OnChannel) digital AM broadcasting system. When the two bit streams are combined, the decoder produces a stereo output signal. In addition to the level values, the primary bit stream can contain stereo parameters, for example an amplitude parameter. In this way, decoding this bit stream alone already produces stereo output which is enhanced when both bit streams are available.
The present description will now be described by way of illustrative examples, without limiting the scope of the invention defined in the appended claims, with reference to the accompanying drawings, in which:
Figure 1 illustrates a source coding system containing an encoder enhanced by a parametric stereophonic encoder module, and a decoder enhanced by a parametric stereophonic decoder module, Figure 2a is a schematic block of a parametric stereophonic decoder module, Figure 2b is a schematic block of a pseudo stereo generator with control parameter inputs, Figure 2c is a schematic block of a balance adjuster with control parameter inputs, Figure 3 is a schematic block of a parametric stereo decoder module using multi-band pseudo-stereo generation combined with multi-band balance adjustment,
ES 2 344 145 T3 Figure 4a is a schematic block on the encoder side of an expandable HFR-based stereo codec, employing level / balance encoding of the spectral envelope according to one embodiment of the invention, and Figure 4b is a corresponding decoder side schematic block according to an embodiment of the invention.
The embodiments described below are purely illustrative for the principles of the present invention. It is understood that modifications and variations to the arrangements and details described herein will be apparent to others skilled in the art. Therefore, it is intended to be limited only to the scope of the claims set forth below, and not to the specific details presented by way of description and explanation of the embodiments herein. For clarity, all the examples shown below will assume two-channel systems, but as is apparent to others skilled in the art, the methods can be applied to multi-channel systems, such as a 5.1 system.
Figure 1 shows how an arbitrary source coding system comprising an encoder 107 and a decoder 115, with the encoder and decoder operating in monaural mode, can be enhanced by parametric stereo coding. L and R indicate the left and right analog input signals, which are fed to a 101 AD transformer. The output of the AD converter is converted to mono 105 and the mono signal is encoded encoded 107. Additionally, the stereophonic signal is directed to a parametric stereophonic encoder 103 which calculates one or more stereophonic parameters to be described below. These parameters are combined with the monophonic signal encoded by means of a multiplexer 109 which forms a stream 111 of bits. The bit stream is stored or transmitted and subsequently extracted on the decoder side by means of a demultiplexer 113. The mono signal is decoded 115 and converted to a stereo signal by a parametric stereo decoder 119 using stereo parameter (s) 117 as the control signal (s). Finally, the stereo signal is routed to DA converter 121, which feeds the analog outputs L 'and R'. The topology according to figure 1 is common to a set of parametric stereophonic coding methods that will be described in detail, starting with the less complex versions. One method of parameterizing stereo properties is to determine the stereo width of the original signal on the encoder side. A first approximation of the stereophonic amplitude is the differential signal, D = LR, since, so to speak, a high degree of similarity between L and R computes for a small value of D and vice versa. A special case is the dual monophonic case in which L = R and therefore D = 0. Therefore, even this simple algorithm is capable of detecting the type of monophonic input signal commonly associated with news broadcasts, in in which case pseudo stereo is not desired. However, a mono signal fed to L and R at different levels does not produce a zero D signal, even though the perceived amplitude is zero. Thus, in practice, more elaborate detectors may be needed, employing, for example, cross-correlation methods. It should be ensured that the value describing the left-right difference or correlation is somehow normalized to the overall signal level to achieve a level-independent detector. A problem with the detector mentioned above is the case where monophonic speech is mixed with a much weaker stereo signal, for example stereophonic noise or background music during speech-to-music / music-to-speech transitions. During voice pauses, the detector will then indicate a wide stereo signal. This is solved by normalizing the stereo amplitude value with a signal that contains information of the previous global energy level, for example, a signal of reduction of the peak of the total energy. Furthermore, to prevent the stereo amplitude detector from being triggered by high-frequency noise or high-frequency distortion from different channel, the detector signals must be pre-filtered by a low-pass filter, usually with a cutoff frequency to some extent. above a second voice formant and optionally also through a high pass filter to avoid unbalanced signal offsets or hum. Regardless of the detector type, the calculated stereo width is represented by a finite set of values that covers the entire range, from mono to wide stereo.
Figure 2a provides an example of the contents of the parametric stereo decoder presented in Figure 1. The block designated "balance", 211, controlled by parameter B, will be described later, and should be considered as omitted for the time being. The block called "amplitude", 205, takes a monophonic input signal and synthetically recreates the feeling of a stereophonic amplitude, the amount of amplitude being controlled by parameter W. The optional parameters S and D will be described later. Subjectively better sound quality can often be achieved by incorporating a crossover filter comprising a low pass filter 203 and a high pass filter 201 to keep the low frequency range "tight" and unaffected. Here only the output of the high pass filter is directed to the amplitude block. The stereo output from the amplitude block is added to the mono output of the low pass filter via 207 and 209, forming the stereo output signal.
Any prior art pseudo stereo generator can be used for the amplitude block, such as those mentioned in the background section, or a Schroeder (multitap delay) or reverb type early reflection simulation unit. Figure 2b gives an example of a pseudo-stereo generator, powered by a monophonic M signal. The amount of stereo width is determined by the gain of 215, and this gain is a function of the stereo width parameter W. The higher the gain, the wider the stereophonic feel, zero gain corresponds to pure monophonic reproduction. The output of 215 is delayed, 221, and added, 223 and 225, to the two direct signal instances, using opposite signs. In order not to significantly alter the overall playback level when the stereo width is changed, you can
ES 2 344 145 T3 incorporates a direct signal compensation attenuation, 213. For example, if the gain of the delayed signal is G, the gain of the direct signal can be selected as sqrtQ-G<sup>2</sup>). High-frequency roll-off can be incorporated into the delay signal path, 217, which helps prevent pseudo-stereo unmasking of encoding artifacts. Optionally, the crossover filters, the roll-off filter and the delay parameters can be sent in the bit stream, offering more possibilities to mimic the stereophonic properties of the original signal, as also shown in Figures 2a and 2b as the X, S and D signals. If a reverb unit is used to generate a stereo signal, the reverb reduction may sometimes not be desired after the end of a sound. However, these unwanted reverb tails can be easily attenuated or completely eliminated by simply altering the gain of the reverb signal. A detector designed to find endings of sounds can be used for this purpose. If the reverb unit generates artifacts in some specific signals, for example transients, a detector for those signals can also be used to attenuate them.
A method for detecting stereophonic properties according to the invention is described as follows. Again, L and R indicate the left and right input signals. The corresponding signal powers are then given by P<sub>l</sub>-L<sup>2</sup> and Pr ~ R<sup>2</sup>. Now a measure of stereophonic balance can be calculated as the quotient between the two signal powers, or more specifically as B = (PL + e) / (P<sub>R</sub> + e), where e is a very small arbitrary number that eliminates division by zero. The equilibrium parameter, B, can be expressed in dB given by the relation B® == 10 log<sub>10</sub>(B). As an example, the three cases P<sub>L</sub> = 10P<sub>R</sub>, P<sub>L</sub> = P<sub>R</sub> And p<sub>L</sub> = 0.1 P<sub>R</sub> they correspond to balance values of + 10 dB, 0 dB, and -10 dB respectively. Clearly, these values represent the "left", "center" and "right" locations. Tests have shown that the range of the balance parameter can be limited, for example, to +/- 40 dB, since these extreme values are already perceived as if the sound originated entirely from one of the two speakers or headphone drivers. This limitation reduces the signal space to be covered in transmission, thus offering a reduction in bit rate. Furthermore, a progressive quantization scheme can be employed whereby smaller quantization stages are used around zero and larger stages, towards the outer limits, further reducing the bit rate. Equilibrium is often constant over time for extended steps. Thus, a last step can be carried out to significantly reduce the average number of bits required: after transmission of an initial balance value, only the differences between consecutive balance values are transmitted, using entropy coding. Very often this difference is zero, which is therefore signaled by the shortest possible code word. Clearly, in applications where bit errors are possible, this delta encoding must be readjusted to a suitable time interval to eliminate the uncontrolled propagation of errors.
The most rudimentary decoder's use of the balance parameter is simply to offset the mono signal towards one of the two playback channels by feeding the mono signal to the two outputs and adjusting the gains accordingly, as illustrated in figure 2c, blocks 227 and 229, with control signal B. This is analogous to turning the “pan” knob on a mixing console, synthetically “moving” a mono signal between the two stereo speakers.
The balance parameter can be sent in addition to the amplitude parameter described above, offering the possibility of placing and extending the sound image on the sound stage in a controlled way, offering flexibility by simulating the original stereophonic feel. One problem with combining pseudo-stereo generation, as mentioned earlier, and parameter-controlled balancing is the unwanted input of signals from the pseudo-stereo generator at balance positions away from the center position. This is solved by applying a function that favors monophonic character at the stereo width value, resulting in a greater attenuation of the stereo width value at equilibrium positions in the extreme lateral position and less attenuation or no attenuation at the equilibrium positions. close to the center position.
The methods described so far are intended for applications with a very low bit rate. In applications where higher bit rates are available, more elaborate versions of the above span and balance methods can be used. Stereo width detection can be done in various frequency bands, resulting in individual stereo width values for each frequency band. Similarly, the balance calculation can work in a multiband way, which is equivalent to applying different filter curves to two channels that are fed by a mono signal. Figure 3 shows an example of a parametric stereophonic decoder using a set of N pseudo-stereo generators according to figure 2b, represented by blocks 307, 317 and 327, combined with a multiband balance setting, represented by blocks 309, 319 and 329, as described in Figure 2c. The individual transmission bands are obtained by feeding the mono input signal, M, to a set of 305, 315 and 325 band pass filters. The stereo outputs of the transmit band from the balance adjusters, 311, 321, 313, 323 are added, forming the stereo output signal, LyR. The previously scalar amplitude and equilibrium parameters are now replaced by the W (k) and B (k) arrangements. In Figure 3, each pseudo-stereo generator and balance adjuster has unique stereo parameters. However, to reduce the total amount of data to be transmitted or stored, the parameters of various frequency bands in groups can be averaged in the encoder, and this smaller number of parameters applied to the corresponding groups of amplitude blocks. and balance in the decoder. Clearly, different grouping schemes and lengths can be used for the W (k) and B (k) arrangements. S (k) represents the earnings of
ES 2 344 145 T3 the paths of the delay signals in the amplitude blocks, and d (k) represents the delay parameters. Again, S (k) and D (k) are optional in the bit stream.
The parametric balance coding method, especially for lower frequency bands, can give somewhat unstable behavior due to lack of frequency resolution or due to too many sound events happening at the same time in one frequency band but in different equilibrium positions. These equilibrium problems are typically characterized by a deviating equilibrium value for just a short period of time, usually one or a few consecutive calculated values, dependent on the update rate. To avoid disturbing equilibrium problems, a stabilization process can be applied to the equilibrium data. This process can use a number of balance values before and after the current time position to calculate the mean value of these. The mean value can then be used as a limiting value for the current balance value, that is, the current balance value should not be allowed to go beyond the mean value. The current value is then limited to the interval between the last value and the mean value. Optionally, the current equilibrium value can be allowed to exceed the limited values by a certain excess factor. Furthermore, the excess factor, as well as the number of equilibrium values used to calculate the mean, should be viewed as frequency-dependent properties and therefore individual for each frequency band.
At low balance information update ratios, lack of temporal resolution can cause mis-timing between the movements of the stereo image and the actual sound events. To improve this behavior with respect to synchronization, an interpolation scheme based on identifying sound events can be used. Here interpolation refers to the interpolation between two consecutive equilibrium values in time. By studying the mono signal on the receiver side, information about the beginnings and ends of different sound events can be obtained. One way is to detect a sudden increase or decrease in signal energy in a particular frequency band. Interpolation, after guiding from that energy envelope over time, should ensure that changes in equilibrium position should preferably be made during time slots containing little signal energy. Since the human ear is more sensitive to the inputs than the output portions of a sound, the interpolation scheme benefits from finding the beginning of a sound, applying, for example, peak hold to the energy and then letting the equilibrium value increases are a function of the peak holding energy, where a small energy value gives a large increase and vice versa. For time segments containing energy uniformly distributed in time, for example, as for some stationary signals, this interpolation method is equal to the linear interpolation between the two equilibrium values. If the equilibrium values are ratios of left and right energies, logarithmic equilibrium values are preferred, for reasons of left-right symmetry. Another advantage of applying the full interpolation algorithm in the logarithmic domain is the tendency of the human ear to relate levels on a logarithmic scale.
Also, for low update ratios of stereo width gain values, interpolation can be applied to them. A simple way is to linearly interpolate between two consecutive stereo width values in time. More stable stereo width behavior can be achieved by smoothing the stereo width gain values over a longer time segment containing various stereo width parameters. Using smoothing with different attack and broadcast time constants is a very suitable system for program material containing mixed or interleaved voice and music. An appropriate design of this type of smoothing filter is produced using a short attack time constant to achieve a short rise time and therefore immediate response to stereo music inputs, and a long broadcast time for get a long hang time. In order to be able to quickly switch from a wide stereo mode to a mono mode, which may be desirable for sudden voice inputs, there is a possibility to bypass or reset the smoothing filter by signaling this event. In addition, attack time constants, broadcast time constants, and other smoothing filter characteristics can also be signaled by an encoder.
For signals containing masked distortion from a psychoacoustic codec, a common problem when inputting stereo information based on the encoded mono signal is a distortion unmasking effect. This phenomenon commonly referred to as “stereo unmasking” is the result of off-centered sounds that do not meet the masking criteria. The problem with stereophonic unmasking can be solved or partially solved by introducing, on the decoder side, a detector intended for these situations. Known technologies for measuring signal-to-mask ratios can be used to detect possible stereophonic unmasking. Once detected it can be explicitly flagged or the stereo parameters can simply be lowered.
On the encoder side, one option, taught by the invention, is to employ a Hilbert transformer in the input signal, eg a 90 degree phase shift is introduced between the two channels. By subsequently forming the mono signal by adding the two signals, a better balance is achieved between a centered mono signal and "true" stereo signals since the Hilbert transformation introduces a 3 dB attenuation for center information. In practice this improves the monophonic encoding of, for example, contemporary pop music, in which, for example, lead singers and bass are typically recorded using a single monophonic source.
ES 2 344 145 T3
The multiband balance parameter method is not limited to the type of application described in Figure 1. It can be advantageously used as long as the objective is to efficiently encode the power spectral envelope of a stereo signal. Therefore, it can be used as a tool in stereophonic codecs in which, in addition to a stereophonic spectral envelope, a corresponding stereophonic residue is encoded. The global power P is defined by P = P<sub>L</sub> + P<sub>R</sub>, where P<sub>L</sub> And p<sub>R</sub> they are signal strengths, as described above. Note that this definition does not take into account the right-to-left phase relationships. (For example, identical left and right signals, but with opposite sign, do not produce a total power of zero.) Analogously to B, P can be expressed in dB as P® = 10 log<sub>10</sub> (P / P<sub>ref</sub>) where P<sub>ref</sub> is an arbitrary reference power and delta values can be entropy encoded. In contrast to the equilibrium case, non-progressive quantization is used for P. To represent the spectral envelope of a stereophonic signal, P and B are calculated for a set of frequency bands, usually, but not necessarily, with bandwidths that they are related to the critical bands of the human ear. For example, such bands can be formed by grouping channels in a constant bandwidth filter bank, with PL and PR being calculated as the frequency and time averages of the squares of the subband samples that correspond to the respective time band and period. . The sets P<sub>0</sub>, P<sub>1</sub>, P<sub>2</sub>, ..., P<sub>N-1</sub> and B<sub>0</sub>, B<sub>1</sub> B<sub>2</sub>, ... B<sub>N-1</sub>, in which the subscripts indicate the frequency band in a representation of N bands, are encoded in delta and Huffman mode, transmitted or stored, and finally decoded into the quantized values that were calculated in the encoder. The last stage is to convert P and B back to P<sub>L</sub> And p<sub>R</sub>. As can easily be seen from the definitions of P and B, the inverse relations are (ignoring e in the definition of B) P<sub>L</sub> = BP / (B + 1) and P<sub>R</sub> = P / (B + 1).
An especially interesting application of the above envelope encoding method is encoding high-band spectral envelopes for HFR-based codec-decoders. In this case, no residual high-band signal is transmitted. Instead, this residual is obtained from the low band. Therefore, there is no strict relationship between envelope and residual representation, and an envelope quantization is more critical. To study the effects of quantification, Pq and Bq indicate the quantified values of P and B respectively. Pq and Bq are then inserted in the previous relations and the sum is formed: P<sub>L</sub> q + P<sub>R</sub> q = BqPq / (Bq + 1) + Pq / (Bq + 1) = Pq (Bq + 1) / (Bq + 1) = Pq. The interesting feature here is that Bq is removed and the error in the total power is determined solely by the quantization error in P. This implies that even though B is intensely quantized, the perceived level is correct assuming sufficient precision is used in the P quantization. In other words, the distortion at B represents the distortion in space, rather than in level. As long as sound sources are stationary in space over time, this distortion in stereophonic perspective is also stationary and difficult to notice. As already discussed, the quantization of the stereophonic balance can also be less precise towards the outer extremes since an error given in dB corresponds to a smaller error in the perceived angle when the angle with respect to the center line is large, due to the properties of the human ear.
When quantizing frequency-dependent data, for example, multi-band stereo amplitude gain values or multi-band balance values, the resolution and range of the quantization method can be advantageously selected to fit the properties of a scale of perception. If such a scale is made as a function of frequency, different quantization methods, or so-called quantization classes, can be chosen for the different frequency bands. The encoded parameter values representing the different frequency bands should then in some cases, even if they have identical values, be interpreted in different ways, that is, decoded into different values.
Analogously to a switched L / R to S / D coding scheme, the P and B signals can be adaptively replaced by the P signals<sub>L</sub> And p<sub>R</sub>, to better cope with extreme signals. As taught by PCT / SE00 / 00158, the delta encoding of envelope samples can be switched from delta-in-time to delta-in-frequency depending on which address is most efficient with respect to the number of bits at a time. particular. The balance parameter can also take advantage of this scheme: consider, for example, a source that moves in time through the stereophonic field. Clearly, this corresponds to a successive change of equilibrium values with respect to time which, depending on the speed of the source versus the update speed of the parameters, can correspond to large delta-in-time values, corresponding to large words. code when using entropy encoding. However, assuming the source has uniform sound radiation versus frequency, the delta-in-frequency values of the equilibrium parameter are zero at any point in time, again corresponding to small code words. Therefore, in this case a lower bit rate is achieved by using the delta encoding address in frequency. Another example is a source that is stationary in space, but has non-uniform radiation. Now, the delta-frequency values are large and the preferred choice is delta-in-time.
The P / B coding scheme offers the possibility of building an expandable HFR codec, see figure 4. An expandable codec is characterized by the fact that the bit stream is divided into two or more parts, receiving and transmitting optional. decoding of higher order parts. The example involves two bit stream parts, hereinafter referred to as primary 419 and secondary 417, but extension to a larger number of parts is also clearly possible. The encoder side, Figure 4a, comprises an arbitrary low-band stereo encoder 403 operating on the stereo input signal, IN (the trivial steps of the AD, or respectively, DA conversion are not shown in the figure), an encoder Parametric stereo that calculates the high-band spectral envelope and optionally additional stereo parameters 401, which also work on the stereo input signal, and two multiplexers 415 and 413 for the primary and secondary bit streams respectively. In this application, the high-band envelope encoding is locked to P / B operation, and the
ES 2 344 145 T3 signal P, 407, is sent to the primary stream of bits by 415, while signal B, 405, is sent to the secondary stream of bits, by 413.
For the low band codec there are different possibilities: it can operate constantly in the S / D mode, and the S and D signals can be sent to the primary and secondary bit streams respectively. In this case, a decoding of the primary bit stream results in a full band mono signal. Of course, this monophonic signal can be enhanced by parametric stereophonic methods according to the invention, in which case the stereophonic parameter (s) must also be in the primary bit stream. Another possibility is to feed a coded low-band stereo signal to the primary bit stream, optionally together with high-bandwidth and balance parameters. Now decoding the primary bit stream results in true stereo for the low band, and a very realistic pseudo stereo for the high band as the stereo properties of the low band are reflected in the high frequency reconstruction. In other words, even though the available high-band envelope representation or inaccurate spectral structure is in mono mode, the synthesized high-band residual structure or spectral fine structure is not. In this type of implementation, the secondary bitstream may contain more low-band information which, combined with that of the primary bitstream, produces higher quality low-band reproduction. The topology of Figure 4 illustrates both cases since the low band encoder output primary and secondary signals 411 and 409 connected to 415 and 417 respectively may contain some of the types of signals described above.
The bit streams are transmitted or stored and only 419 or both 419 and 417 are fed to the decoder, Figure 4b. The primary stream of bits is demultiplexed by 423 into the primary signal 429 of the central lowband decoder and the signal P, 431. Similarly, the secondary stream of bits is demultiplexed by 421 in the secondary signal 427 of the central band decoder low and signal B, 425. The low-band signal (s) is directed to the low-band decoder 433, which produces an output 435, which again, in case of decoding only the primary stream of bits, can be of any of the types described above (monophonic or stereophonic). Signal 435 feeds the HFR unit 437, generating a synthetic high band and adjusting according to P, which is also connected to the HFR unit. The decoded low band is combined with the high band in the HFR unit, and the low band and / or the high band is optionally enhanced by a pseudo stereo generator (also located in the HFR unit) before finally being fed to the system outputs , forming the output signal, OUT. When the secondary bit stream 417 is present, the HFR unit also gets signal B as an input signal 425, and 435 is in stereo mode, with the system producing a full stereo output signal and the pseudo-stereo generators, if any. , they skip.
In other words, a method for encoding stereophonic properties of an input signal includes, in an encoder, the step of calculating an amplitude parameter that signals a stereophonic amplitude of said input signal, and in a decoder, a step of generating a stereo output signal, using said amplitude parameter to control a stereo width of said output signal. The method further comprises in said encoder, forming a monophonic signal from said input signal, wherein, in said decoder, said generation involves a pseudo-stereophonic method operating on said monophonic signal. The method further involves dividing said mono signal into two signals as well as adding a delayed version (s) of said mono signal to said two signals, at a level (s) controlled by said amplitude parameter. The method further includes that said delayed version (s) are high pass filtered and progressively attenuated to higher frequencies before being added to said two signals. The method further includes that said amplitude parameter is a vector, and the elements of said vector correspond to independent frequency bands. The method further includes that if said input signal is dual mono type, said output signal is also dual mono type.
A method for encoding stereophonic properties of an input signal includes, in an encoder, calculating a balance parameter that signals a stereophonic balance of said input signal, and in a decoder, generating a stereophonic output signal, using said balance parameter. to control a stereophonic balance of that output signal.
In this method, in said encoder, a mono signal is formed from said input signal, and in said decoder, said generation involves dividing said mono signal into two signals, and said control involves adjusting the levels of said two signals. The method further includes calculating a power for each channel of said input signal, and said balance parameter is calculated from a quotient between said powers. The method further includes that said powers and said balance parameter are vectors in which each element corresponds to a specific frequency band. The method further includes that in said decoder it is interpolated between two consecutive values in time of said balance parameters so that the instantaneous value of the corresponding power of said monophonic signal controls the inclination that the instantaneous interpolation must have. The method further includes that said interpolation method is performed on equilibrium values represented as logarithmic values. The method further includes that said equilibrium parameter values are limited to a range between a previous equilibrium value and an equilibrium value extracted from other equilibrium values by means of a mean filter or other filtering process, in which said range can be further extended by moving the edges of that range by a certain factor. The method further includes that said method of extracting boundary edges for the equilibrium values depends, for a multiband system, on the frequency. The method further includes calculating an additional level parameter as a sum of vectors of said powers and sending it to said decoder, said decoder thus providing a representation of a spectral envelope of said input signal. The method
ES 2 344 145 T3 further includes that said level parameter and said balance parameter are adaptively replaced by said powers. The method further includes that said spectral envelope is used to control an HFR process in a decoder. The method further includes that said level parameter is fed to a primary stream of bits from an expandable HFR-based stereo codec, and said balance parameter is fed to a secondary stream of bits from said codec. Said monophonic signal and said amplitude parameter are fed to said primary stream of bits. Furthermore, said amplitude parameters are processed by a function that gives smaller values for an equilibrium value that corresponds to an equilibrium position further from the central position. The method further includes that a quantization of said equilibrium parameter employs smaller quantization stages around a central position and larger stages towards external positions. The method further includes that said amplitude parameters and said equilibrium parameters are quantized using a quantization method with respect to resolution and range which, for a multiband system, is frequency dependent. The method further includes that said equilibrium parameter is adaptively delta encoded in either time or frequency. The method further includes that said input signal is passed through a Hilbert transformer before forming said mono signal.
An apparatus for parametric stereophonic coding includes, in an encoder, means for calculating an amplitude parameter that signals a stereophonic amplitude of an input signal, and means for forming a monophonic signal from said input signal, and, in a decoder, means for generating a stereophonic output signal from said monophonic signal, using said amplitude parameter to control a stereophonic amplitude of said output signal.
Contents4
4 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4
143 members in 13 offices
Priority claims12
| Document | Office | Kind | Date |
|---|---|---|---|
| 0102481 | Sweden | A | |
| 0102481 | Sweden | A | |
| 0200796 | Sweden | A | |
| 0200796 | Sweden | A | |
| 0202159 | Sweden | A | |
| 0202159 | Sweden | A | |
| 010248105017007 | – | – | – |
| 0200796 | – | – | – |
| 0202159 | – | – | – |
| SE20010002481 | – | – | – |
| SE20020000796 | – | – | – |
| SE20020002159 | – | – | – |
Members143
| Document | Office | Kind | |
|---|---|---|---|
| SE0102481D0 | Sweden | D0 | |
| SE0200796D0 | Sweden | D0 | |
| SE0202159D0 | Sweden | D0 | |
| WO03007656A1 | World Intellectual Property Organization (WIPO) | A1 | |
| KR20040019042A | Republic of Korea | A | |
| EP1410687A1 | European Patent Office (EPO) | A1 | |
| CN1524400A | China | A | |
| HK1062624A1 | Hong Kong, China | A1 | |
| JP2004535145A | Japan | A | |
| US2005053242A1 | United States of America | A1 | |
| EP1410687B1 | European Patent Office (EPO) | B1 | |
| KR20050099559A | Republic of Korea | A | |
| KR20050099560A | Republic of Korea | A | |
| AT305715T | Austria | T | |
| ATE305715T1 | Austria | T1 | |
| KR20050100011A | Republic of Korea | A | |
| KR20050100012A | Republic of Korea | A | |
| DE60206390D1 | Germany | D1 | |
| EP1600945A2 | European Patent Office (EPO) | A2 | |
| EP1603117A2 | European Patent Office (EPO) | A2 | |
| EP1603118A2 | European Patent Office (EPO) | A2 | |
| EP1603119A2 | European Patent Office (EPO) | A2 | |
| US2006023888A1 | United States of America | A1 | |
| US2006023891A1 | United States of America | A1 | |
| US2006023895A1 | United States of America | A1 | |
| US2006029231A1 | United States of America | A1 | |
| ES2248570T3 | Spain | T3 | |
| JP2006074818A | Japan | A | |
| JP2006085183A | Japan | A | |
| JP2006087130A | Japan | A | |
| JP2006087131A | Japan | A | |
| CN1758335A | China | A | |
| CN1758336A | China | A | |
| CN1758337A | China | A | |
| CN1758338A | China | A | |
| HK1080206A1 | Hong Kong, China | A1 | |
| HK1080208A1 | Hong Kong, China | A1 | |
| HK1080979A1 | Hong Kong, China | A1 | |
| DE60206390T2 | Germany | T2 | |
| CN1279790C | China | C | |
| KR100649299B1 | Republic of Korea | B1 | |
| KR100666813B1 | Republic of Korea | B1 | |
| KR100666814B1 | Republic of Korea | B1 | |
| KR100666815B1 | Republic of Korea | B1 | |
| KR100679376B1 | Republic of Korea | B1 | |
| EP1603117A3 | European Patent Office (EPO) | A3 | |
| EP1603119A3 | European Patent Office (EPO) | A3 | |
| EP1600945A3 | European Patent Office (EPO) | A3 | |
| EP1603118A3 | European Patent Office (EPO) | A3 | |
| US7382886B2 | United States of America | B2 | |
| EP2015292A1 | European Patent Office (EPO) | A1 | |
| HK1124950A1 | Hong Kong, China | A1 | |
| EP2015292B1 | European Patent Office (EPO) | B1 | |
| JP2009217290A | Japan | A | |
| AT443909T | Austria | T | |
| ATE443909T1 | Austria | T1 | |
| DE60233835D1 | Germany | D1 | |
| US2009316914A1 | United States of America | A1 | |
| DK2015292T3 | Denmark | T3 | |
| EP1603119B1 | European Patent Office (EPO) | B1 | |
| JP2010020342A | Japan | A | |
| AT456124T | Austria | T | |
| ATE456124T1 | Austria | T1 | |
| ES2333278T3 | Spain | T3 | |
| US2010046761A1 | United States of America | A1 | |
| US2010046762A1 | United States of America | A1 | |
| DE60235208D1 | Germany | D1 | |
| JP4447317B2 | Japan | B2 | |
| EP1603117B1 | European Patent Office (EPO) | B1 | |
| AT464636T | Austria | T | |
| ATE464636T1 | Austria | T1 | |
| ES2338891T3 | Spain | T3 | |
| DE60236028D1 | Germany | D1 | |
| JP4474347B2 | Japan | B2 | |
| HK1080206B | Hong Kong, China | B | |
| CN1758336B | China | B | |
| ES2344145T3This record | Spain | T3 | |
| HK1080979B | Hong Kong, China | B | |
| CN1758335B | China | B | |
| EP2249336A1 | European Patent Office (EPO) | A1 | |
| CN101887724A | China | A | |
| CN1758338B | China | B | |
| CN1758337B | China | B | |
| JP2011034102A | Japan | A | |
| EP1600945B1 | European Patent Office (EPO) | B1 | |
| AT499675T | Austria | T | |
| ATE499675T1 | Austria | T1 | |
| CN101996634A | China | A | |
| DE60239299D1 | Germany | D1 | |
| HK1080208B | Hong Kong, China | B | |
| HK1145728A1 | Hong Kong, China | A1 | |
| JP2011101406A | Japan | A | |
| JP4700467B2 | Japan | B2 | |
| US8014534B2 | United States of America | B2 | |
| JP4786987B2 | Japan | B2 | |
| US8059826B2 | United States of America | B2 | |
| US8073144B2 | United States of America | B2 | |
| US8081763B2 | United States of America | B2 | |
| US8116460B2 | United States of America | B2 | |
| JP4878384B2 | Japan | B2 |
Numbers
- Publication, DOCDB
- 2344145
- Publication, EPODOC
- ES2344145T
- Application
- 5017007
- Application, DOCDB
- 05017007
- Application, EPODOC
- ES20050017007T
Titles2
- English
- EFFECTIVE AND EXTENDABLE PARAMETRIC STEREOPHONIC CODING FOR LOW-SPEED BITS TRANSFER APPLICATIONS.
- Spanish
- CODIFICACION ESTEREOFONICA PARAMETRICA EFICAZ Y AMPLIABLE PARA APLICACIONES DE BAJA VELOCIDAD DE TRANSFERENCIA DE BITS.
Classification
- CPC, 6
- G10L19/24
- H04S5/00
- G10L19/008
- G10L19/0204
- H04S1/007
- H04S3/002
- IPC, 8
- G10L19 008
- G10L19 02
- G10L19 14
- G10L19 24
- H04S
- H04S1 00
- H04S3 00
- H04S5 00