Audio decoding
Abstract
Method of decoding an audio stream, the method comprising the steps of: reading an encoded audio stream (AS) that includes sinusoidal codes (r) representing a phase (psi) for each track of linked sinusoidal components, for each hint, generate (56) a monotonously changing value (¿psi) in general from said codes (r) representing said phase; filter (58) said generated value to provide a frequency estimate (¿omega) for a track; and employing (32) said generated values and said frequency estimates to synthesize said sinusoidal components of said audio signal.

Term
Term ended
Projected expiry passed 6 November 2023, 2.9 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
3 claims: 2 independent, 1 dependent
- 1ES 2 298 568 T3 REIVINDICACIONES 1. Procedimiento de descodificación de un flujo de audio, comprendiendo el procedimiento las etapas de:leer un flujo de audio (AS') codificado que incluye códigos (r) sinusoidales que representan una fase (ψ) para cada pista de componentes sinusoidales enlazadas, para cada pista, generar (56) un valor (ψ) monótonamente cambiante en general a partir de dichos códigos (r) que representan dicha fase;filtrar (58) dicho valor generado para proporcionar una estimación de frecuencia (ώ) para una pista;y emplear (32) dichos valores generados y dichas estimaciones de frecuencia para sintetizar dichas componentes sinusoidales de dicha señal de audio.
- 2Reproductor (3) de audio que comprende:medios para leer un flujo de audio (AS') codificado que incluye códigos (r) sinusoidales que representan una fase (ψ) para cada pista de componentes sinusoidales enlazadas, un desempaquetador (56) de fase para generar, para cada pista, un valor (ψ) monótonamente cambiante en general a partir de dichos códigos (r) que representan dicha fase;un filtro (58) para filtrar dicho valor generado para proporcionar una estimación de frecuencia (ώ) para una pista;y un sintetizador (32) dispuesto para emplear dichos valores generados y dichas estimaciones de frecuencia para sintetizar dichas componentes sinusoidales de dicha señal de audio.
- 3Sistema de audio que comprende un codificador (1) de audio y un reproductor (3) de audio según la reivindicación 2.
Independent claims3
66 paragraphs in 5 sections, as filed
ES 2 298 568 T3
DESCRIPTION
Audio decoding.
Field of the invention
The present invention relates to encoding and decoding of audio signals.
Background of the invention
Referring now to Figure 1, a parametric coding scheme, in particular a sinusoidal encoder, is described in PCT patent application No. WO01 / 69593. In this encoder, an input audio signal x (t) is divided into several segments or frames (overlap), typically 20 ms in length. Each segment is broken down into transient, sinusoidal, and noise components. (It is also possible to obtain other components of the input audio signal such as harmonic complexes although these are not very important for the purposes of the present invention).
In the sinusoidal analyzer 130, the x2 signal for each segment is modeled using a number of sinusoids represented by amplitude, frequency, and phase parameters. This information is typically extracted for an analysis interval by performing a Fourier Transform (FT) that provides a spectral representation of the interval that includes: frequencies; amplitudes for each frequency; and phases for each frequency where each phase is in the interval {-π, π}. Once the sinusoidal information for a segment is estimated, a tracking algorithm is started. This algorithm uses a cost function to link sinusoids together segment by segment to obtain so-called tracks. Therefore, the tracking algorithm returns C codes<sub>S</sub> sinusoids comprising sinusoidal tracks that start at a specific point in time, evolve for a certain amount of time over a plurality of time segments, and then stop.
In such sinusoidal coding, frequency information is normally transmitted for the tracks formed in the encoder. This can be done economically, since the tracks are defined as having a slowly varying frequency, and therefore the frequency can be transmitted efficiently by differential time coding. (In general, amplitude over time can also be differentially encoded.)
Contrary to frequency, phase transmission is considered expensive. In principle, if the frequency is (almost) constant, the phase as a function of the track segment index should comply with (almost) linear behavior. However, when transmitted, the phase is limited to the interval {-π, π} as provided by the Fourier transform. Due to this 2π modulo phase representation, the relationship between structural frames of the phase is lost and, at first glance, it appears to be a white stochastic variable.
However, since the phase is the integral of the frequency, the phase does not need to be transmitted, in principle. This is called phase continuation and it reduces the bit rate significantly.
In phase continuation, only the frequency is transmitted and the phase is recovered in the decoder from the frequency data taking advantage of the integral relationship between phase and frequency. However, it is known that the phase can only be roughly recovered using phase continuation. If frequency errors occur, due to measurement errors in frequency or due to quantization noise, the phase, which is reconstructed using the integral relationship, will normally show an error that has the character of an offset. This is because frequency errors have roughly a white noise character. Integration amplifies low frequency errors and therefore the recovered phase will tend to shift away from the actually measured phase. This leads to audible artifacts.
This is illustrated in Figure 2 (a) where ψ and Ω are the actual frequency and phase for a track. In both the encoder and decoder the frequency and phase have an integral relationship represented by I. The quantization process in the encoder is modeled as additive white noise n. In the decoder, the recovered phase ψ therefore includes two components: the real phase ψ and a component ε<sub>2</sub> noise, where both the spectrum of the recovered phase and the power spectral density function of the noise ε<sub>2</sub> they have a pronounced low-frequency character.
Thus, it can be seen that in phase continuation, since the recovered phase is the integral of a low frequency signal, the recovered phase is itself a low frequency signal. However, the noise introduced in the reconstruction process is also predominant in this low frequency range. Therefore, it is difficult to separate these sources with the idea of filtering out the noise n introduced during encoding.
Description of the invention
According to the present invention there is provided a method according to claim 1, and an audio player according to claim 2.
According to the invention, in the decoder, the frequency can be roughly recovered from the quantized phase information using finite differences as an approximation for differentiation. The noise component of the recovered frequency has a pronounced high-frequency behavior under the
ES 2 298 568 T3 assumption that noise introduced by phase quantization is almost spectrally flat. This is illustrated in Figure 2 (b), where within the encoder and decoder, the frequency is represented as the phase differential (D). Again, noise n is introduced into the encoder and thus the decoder, the frequency Ω. recovered includes two components: the actual frequency Ω and a component ε<sub>4</sub> noise where the frequency is almost a DC signal and the noise is mainly in the high frequency range. However, since the underlying frequency has a low-frequency behavior and the added noise has a high-frequency behavior, the component ε<sub>4</sub> Noise from the recovered frequency can be reduced by a low pass filter.
Brief description of the drawings
Figure 1 shows an audio encoder;
Figures 2 (a) and 2 (b) illustrate the relationship between phase and frequency in prior art systems and in audio systems according to the present invention, respectively;
Figures 3 (a) and 3 (b) show a sinusoidal encoder component of the audio encoder of Figure 1;
Figure 4 shows an audio player in which an embodiment of the invention is implemented; and Figures 5 (a) and 5 (b) show a preferred embodiment of a sinusoidal synthesizer component of the audio player of Figure 4; and Figure 6 shows a system comprising an audio encoder and an audio player according to the invention.
Description of the preferred embodiment
Preferred embodiments of the invention will now be described with reference to the accompanying drawings, in which similar components have been given similar reference numerals and, unless otherwise stated, perform a similar function. Encoder 1 is a sinusoidal encoder of the type described in PCT patent application No. WO 01/69593, Figure 1. The operation of this prior art encoder and its corresponding decoder has been well described and the description is provided herein only as important to the present invention.
The audio encoder 1 samples an input audio signal at a certain sampling frequency which results in a digital x (t) representation of the audio signal. Encoder 1 then separates the sampled input signal into three components: transient signal components, continuous deterministic components, and continuous stochastic components. The audio encoder 1 comprises a transient encoder 11, a sinusoidal encoder 13, and a noise encoder 14.
The transient encoder 11 comprises a transient detector 110 (TD), a transient analyzer 111 (TA) and a transient synthesizer 112 (TS). First, the signal x (t) enters the transient detector 110. This detector 110 estimates whether there is a transient signal component and its position. This information is supplied to the transient analyzer 111. If the position of a transient signal component is determined, the transient analyzer 111 attempts to extract (the main part of) the transient signal component. It compares a shape function with a signal segment that preferably starts at an estimated start position, and determines the content under the shape function, using for example a (small) number of sinusoidal components. This information is contained in the C code<sub>T</sub> transients and more detailed information on the generation of the transient CT code is provided in PCT patent application No. WO 01/69593.
The C code<sub>T</sub> transients is provided to transient synthesizer 112. The synthesized transient signal component is subtracted from the input signal x (t) at subtractor 16, resulting in a signal x1. A gain control (GC) mechanism (12) is used to produce x2 from x1.
The x2 signal is provided to the sinusoidal encoder 13 where it is analyzed in a sinusoidal analyzer 130 (SA), which determines the sinusoidal (deterministic) components. Therefore, it will be appreciated that although the presence of the transient analyzer is desirable, it is not necessary and the invention can be implemented without such an analyzer. Alternatively, as mentioned above, the invention can also be implemented with eg a harmonic complex analyzer.
In short, the sinusoidal encoder encodes the input x2 signal as linked sinusoidal component tracks from one frame segment to the next. Referring now to Fig. 3 (a), in the same manner as in the prior art, each segment of the input signal x2 is transformed to the frequency domain in a Fourier transform (FT) unit 40. For each segment, the FT unit provides measured amplitudes A, phases φ and frequencies ω. As previously mentioned, the phase interval provided by the Fourier transform is restricted to -π <φ <π. A tracking algorithm (TA) unit 42 takes the information for each segment and using an appropriate cost function, links sinusoids from one segment to the next, thus producing a sequence of phases φ (φ and frequencies ω (Χ ) measurements for each track.
ES 2 298 568 T3
Contrary to the prior art, according to the present invention the codes C<sub>S</sub> Sinusoids ultimately produced by analyzer 130 include phase information, and the frequency is reconstructed from this information in the decoder.
As mentioned above, however, the measured phase is restricted to a representation of modulo 2π. Thus, in encoder 1 the analyzer comprises a phase unwrapper (PU) 44 where the modulo 2π phase representation is unpacked to expose the phase behavior between structural frames for a track ψ. When the frequency in sinusoidal tracks is nearly constant, it will be seen that the unpacked phase ψ will normally be a linearly increasing (or decreasing) function and this makes economical phase transmission possible. The unpacked phase ψ is provided as input to a phase encoder (PE) 46 which outputs suitable representation levels r to be transmitted.
Referring now to the operation of the phase unpacker 44, as mentioned above, the actual phase ψ and the actual frequency Ω for a track are related by:
ψ (ί) = £ Q (r) t7r + ψ (Τ ^) Equation 1 where T<sub>or</sub> an instant of reference time.
A sinusoidal track in frames k = K, K + 1 ... K + L-1 has frequencies m (k) measured (expressed in radians per second) and phases 0 (k) measured (expressed in radians). The distance between the center of the frames is given by U (update rate expressed in seconds). The measured frequencies are assumed to be samples of the track Ω of continuous frequency in the underlying time assumed with ω (φ = Ω ^ y) and, similarly, the measured phases are samples of the track ψ of continuous phase in time associated with 0 (k) = <XkU) mod (2n). For sinusoidal coding it is assumed that Ω is a nearly constant function.
Assuming that the frequencies are nearly constant within a segment, Equation 1 can be approximated as follows:
| r (W) = Ω (ί) <* + F ((* -1) 0. W) + a (k - l)) tZ / 2 + F ((* - 1) U) ·
Equation 2
Therefore, it will be seen that by knowing the phase and frequency for a given segment and the frequency of the next segment, it is possible to estimate an unpacked phase value for the next segment, and so on for each segment in a track.
In the preferred embodiment, the phase unpacker determines an unpacking factor m (k) at time k:
y (kU) = $ (k) + m (k) 2n Equation 3
The unpacked factor m (k) tells the phase unpacker 44 the number of cycles that have to be added to obtain the unpacked phase.
By combining equations 2 and 3, the phase unpacker determines an incremental unpacking factor e based on the following:
2ae (¿) = - m (k -1)} = + a> (k - \)} U / 2 - {¿(Λ) - <j> (k -1)} where e should be an integer. However, due to model and measurement errors, the incremental unpacking factor will not be exactly an integer, so:
e (k) = network<sub>OR</sub>nd ([{d) (k) + O) (k - \)} U / 2 - {# *) - -1)}] / (2π)) assuming that the model and measurement errors are small.
ES 2 298 568 T3
Taking the incremental unpacking factor e, the m (k) is calculated from equation (3) as the cumulative sum where, without loss of generality, the phase unpacker starts in the first frame K with m (K) = 0, and from m (k) and 0 (k) the phase ^ (kU) (unpacked) is determined.
In practice, the sampled ^ (kU) and O (kU) data are distorted by measurement errors:
ΰ> (*) = Ω (Η7) + /?<sub>2</sub>(λ :), where ε<sub>λ</sub> and ε<sub>2</sub> they are phase and frequency errors, respectively. In order to prevent the determination of the unpacking factor from becoming ambiguous, the measurement data needs to be determined with sufficient precision. Therefore, in encoder 1, tracking is restricted such that:
<5 (t) = «(4) - ({or (4) + or (4 - 1)} U / 2 - {« 4) - ^ (4-1))) / (2π) <<J<sub>0</sub> where δ is the error in the rounding operation. The error δ is determined mainly by the errors in ω due to multiplication with U. Suppose that ω is determined from the maximum of the absolute value of the Fourier transform from a sampled version of the input signal with frequency F<sub>s</sub> sampling and that the resolution of the Fourier transform is 2n / L<sub>to</sub> where L<sub>to</sub> the scan size. In order to be within the considered limit, one has to:
<img file="ES2298568T3_D0001.tif" />
This means that the scan size should be a few times larger than the update size for unpacking to be accurate, for example adjusting δ<sub>0</sub>= 1/4, the scan size should be four times the update size (neglecting the errors ε<sub>λ</sub> in phase measurement).
The second precaution that can be taken to avoid decision errors in the rounding operation is to define tracks appropriately. In the tracking unit 42, the sinusoidal tracks are typically defined by considering differences in amplitude and frequency. Additionally, it is also possible to take phase information into account in the link criterion. For example, the prediction error ε can be defined as the difference between the measured value and the predicted value φ according to ε - - $ (*)} mod2fl · where the predicted value can be taken as
<img file="ES2298568T3_D0002.tif" />
Therefore, preferably the tracking unit 42 prohibits tracks where ε is greater than a certain value (eg ε> π / 2), resulting in an unambiguous definition of e (k).
Additionally, the encoder can calculate the phases and frequencies as they will be available at the decoder. If the phases or frequencies that will become available in the decoder differ too much from the phases and / or frequencies as they are present in the encoder, it may be decided to interrupt a track, that is, to signal the end of a track and start a new one. using the current frequency and phase and its linked sinusoidal data.
The sampled unpacked phase ^ (kU) produced by phase unpacker 44 (PU) is provided as input to phase encoder (PE) 46 to produce a set of representation levels r. Techniques are known for the efficient transmission of a monotonically changing characteristic in general such as the unpacked phase. In Figure 3 (b), Adaptive Differential Pulse Code Modulation (ADPCM) is employed. In this case, a predictor 48 (PF) is used to estimate the phase of the next track segment and encode the difference only in a quantizer 50 (Q). Since ψ is expected to be a nearly linear function and for simplicity, the predictor 48 is chosen as a second-order filter of the form:
ES 2 298 568 T3
<img file="ES2298568T3_D0003.tif" />
where x is the input and y is the output. However, it will be appreciated that it is also possible to take other functional relationships (including higher order relationships) and include adaptive (backward or forward) adaptation of the filter coefficients. In phase encoder 46, a backward adaptive control (QC) mechanism 52 is used for simplicity to control quantizer 50. Likewise, forward adaptive control is also possible but would require additional bitrate overhead.
As can be seen, the initialization of the encoder (and decoder) for a track begins with the knowledge of the phase φ (0) and the starting frequency ω (0). These are quantized and transmitted through a separate mechanism. Additionally, the initial quantization step used in the encoder quantization controller 52 and the corresponding controller 62 in the decoder, Figure 5 (b), is either transmitted or set to a certain value in both the encoder and the decoder. Finally, the end of a track can be signaled either in a separate lateral stream or as a single symbol in the bit stream of the phases.
From code C<sub>S</sub> sinusoidal generated with the sinusoidal encoder, the sinusoidal signal component is reconstructed by a sinusoidal synthesizer 131 (SS) in the same manner as will be described for the sinusoidal (SS) synthesizer 32 of the decoder. This signal is subtracted by subtractor 17 from input x2 to sinusoidal encoder 13, resulting in a remaining signal x3. The residual signal x3 produced by the sinusoidal encoder 13 is passed to the noise analyzer 14 of the encoder 1 which produces a C code<sub>N</sub> representative of this noise, as described in, for example, PCT patent application No. PCT / EP00 / 04599.
Finally, in a multiplexer 15, an audio stream AS is constituted (audio stream) that includes the C codes<sub>T</sub>, C<sub>S</sub> and C<sub>N</sub>. The audio stream AS is provided to, for example, a data bus, an antenna system, a storage medium, etc.
Figure 4 shows an audio player 3 suitable for decoding an audio stream AS ', for example, generated by an encoder 1 of Figure 1, obtained from a data bus, antenna system, storage medium, etc. . The audio stream AS 'is demultiplexed in a demultiplexer 30 to obtain the codes CT, CS and C<sub>N</sub>. These codes are provided to a transient synthesizer 31, a sinusoidal synthesizer 32, and a noise synthesizer 33 respectively. From the transient CT code, the transient signal components are calculated on the transient synthesizer 31. In case the transient code indicates a shape function, the shape is calculated based on the received parameters. Furthermore, the shape content is calculated based on the frequencies and amplitudes of the sinusoidal components. If the C code<sub>T</sub> of transients indicates a step, then no transients are calculated. The total transient signal yT is a sum of all transients.
The C code<sub>S</sub> sinusoidal that includes the information encoded by the analyzer 130 is used by the sinusoidal synthesizer 32 to generate the signal and<sub>S</sub>. Referring now to Figures 5 (a) and (b), sinusoidal synthesizer 32 comprises a phase decoder 56 (PD) compatible with phase encoder 46. In this case, the dequantizer 60 (DQ, dequantiser) together with a second order prediction filter (PF) 64 produces (an estimate of) the unpacked phase ψ from: the representation levels r, the information initial φ (0), ω (0) provided to prediction filter 64 (PF) and initial quantization step for quantization controller 62 (QC).
As illustrated in Figure 2 (b), the frequency can be recovered from the unpacked phase ψ by differentiation. Assuming that the phase error in the decoder is approximately white and since the differentiation amplifies the high frequencies, the differentiation can be combined with a low-pass filter to reduce noise and thus to obtain an accurate estimate of the frequency in the decoder.
In the preferred embodiment, a filtering unit (FR) 58 approximates the differentiation that is necessary to obtain the frequency ω from the unpacked phase by methods such as forward, backward, or center differences. This enables the decoder to output phases ψ and frequencies ω which can be used in a conventional way to synthesize the sinusoidal component of the encoded signal.
At the same time, when the sinusoidal components of the signal are being synthesized, the noise CN code is supplied to a noise synthesizer 33, which is primarily a filter, which has a frequency response approaching the spectrum noise. The NS 33 generates reconstructed yN noise by filtering a white noise signal with the noise code CN. The total signal y (t) comprises the sum of the transient signal yT and the product of any amplitude decompression (g) and the sum of the sinusoidal signal yS and the noise signal yN. The audio player comprises two adders 36 and 37 to add the respective signals. The total signal is provided to an output unit 35, which is for example a loudspeaker.
Figure 6 shows an audio system according to the invention comprising an audio encoder 1 as shown in Figure 1 and an audio player 3 as shown in Figure 4. Such a system offers playback and recording features. The audio stream AS is provided from the audio encoder to the audio player over a communication channel 2, which can be a wireless connection, a data bus 20 or a storage medium.
ES 2 298 568 T3 nation. In case the communication channel 2 is a storage medium, the storage medium may be fixed in the system or it may be a removable disk, memory card, etc. Channel 2 communication may be part of the audio system, but will often be outside of the audio system, however.
Contents5
7 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7
23 members in 14 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 02080002 | European Patent Office (EPO) | A | |
| 02080002 | European Patent Office (EPO) | A | |
| 20020080002 | European Patent Office (EPO) | – | |
| 0208000203758591 | – | – | – |
| EP20020080002 | – | – | – |
Members23
| Document | Office | Kind | |
|---|---|---|---|
| WO2004051627A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU2003274617A1 | Australia | A1 | |
| AU2003274617A8 | Australia | A8 | |
| MXPA05005601A | Mexico | A | |
| KR20050086871A | Republic of Korea | A | |
| EP1568012A1 | European Patent Office (EPO) | A1 | |
| BR0316663A | Brazil | A | |
| CN1717719A | China | A | |
| PL376861A1 | Poland | A1 | |
| RU2005120380A | Russian Federation | A | |
| US2006036431A1 | United States of America | A1 | |
| JP2006508394A | Japan | A | |
| EP1568012B1 | European Patent Office (EPO) | B1 | |
| AT381092T | Austria | T | |
| ATE381092T1 | Austria | T1 | |
| DE60318102D1 | Germany | D1 | |
| ES2298568T3This record | Spain | T3 | |
| DE60318102T2 | Germany | T2 | |
| RU2353980C2 | Russian Federation | C2 | |
| CN100559467C | China | C | |
| US7664633B2 | United States of America | B2 | |
| JP4606171B2 | Japan | B2 | |
| KR101016995B1 | Republic of Korea | B1 |
Numbers
- Publication
- 2298568
- Publication, DOCDB
- 2298568
- Publication, EPODOC
- ES2298568T
- Application
- 3758591
- Application, DOCDB
- 03758591
- Application, EPODOC
- ES20030758591T
Titles2
- Spanish
- DESCODIFICACION DE AUDIO.
- English
- AUDIO DECODING.
Classification
- CPC, 2
- G10L19/093
- G10L19/02
- IPC, 3
- G10L19 08
- G10L19 093
- G10L25 90