Parametric audio coding
Abstract
Method for encoding (11) an audio signal (I, D) of at least two channels, said method comprising: - determine (110) common frequencies (fcom) in the at least two channels (I, D) of the audio signal, common frequencies that occur in at least two of the at least two channels of the audio signal, and - represent (111) respective sinusoidal elements in the respective channels on a given common frequency by a representation of the given common frequency (fcom) and a representation of the amplitudes (A, A) respective of the respective sinusoidal elements at the given common frequency.

Term
Term ended
Projected expiry passed 17 January 2023, 3.7 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
25 claims: 5 independent, 20 dependent
- 1ES 2 255 678 T3 REIVINDICACIONES 1. Método para codificar (11) una señal (I, D) de audio de al menos dos canales, comprendiendo dicho método:determinar (110) frecuencias comunes (f com ) en los al menos dos canales (I, D) de la señal de audio, frecuencias comunes que ocurren en al menos dos de los al menos dos canales de la señal de audio, y representar (111) elementos sinusoides respectivos en los respectivos canales en una frecuencia común dada mediante una representación de la frecuencia (f com ) común dada y una representación de las amplitudes (A, AA) respectivas de los elementos sinusoidales respectivos en la frecuencia común dada.
- 2Método de codificación según la reivindicación 1, en el que la representación de las amplitudes (A, AA) respectivas comprende una amplitud (A) promedio y una amplitud (AA) de diferencia.
- 3Método de codificación según la reivindicación 1, en el que la representación de las amplitudes (A, AA) respectivas comprende una amplitud (A) máxima y una amplitud (AA) de diferencia.
- 4Método de codificación según la reivindicación 1, en el que las frecuencias no comunes se codifican como frecuencias comunes, en las que la representación de la amplitud incluye una indicación para indicar el al menos un canal en el que no ocurre la frecuencia.
- 5Método de codificación según la reivindicación 1, en el que además de las frecuencias comunes, se codifican independientemente las frecuencias no comunes.
- 6Método según la reivindicación 5, en el que las frecuencias no comunes se agrupan en el flujo de audio codificado en un bloque separado.
- 7Método según la reivindicación 6, en el que las frecuencias comunes se agrupan y se incluyen en la señal de audio codificada antes del bloque de frecuencias no comunes.
- 8Método según la reivindicación 6, en el que los parámetros de los elementos sinusoidales en las frecuencias comunes se incluyen en una capa base y los parámetros de las sinusoides en las frecuencias no comunes se incluyen en una capa de refuerzo.
- 9Método según la reivindicación 1, en el que el método comprende la etapa de combinar representaciones de potencia o de energía respectivas de los al menos dos canales para obtener una representación común y en el que la etapa de determinar las frecuencias comunes se realiza basándose en la representación común.
- 10Método según la reivindicación 9, en el que la etapa de combinación incluye añadir espectros de potencia de los al menos dos canales y en el que la representación común es un espectro de potencia común.
- 11Método según la reivindicación 1, en el que los parámetros de frecuencia y amplitud se incluyen en una capa base y la amplitud delta se incluye en una capa de refuerzo.
- 12Método según la reivindicación 1, en el que se determinan respectivas fases de los sinusoides respectivos en la frecuencia común dada y en el que se incluye una representación de las fases respectivas en la señal de audio codificada.
- 13Método según la reivindicación 12, en el que la representación de las fases respectivas incluye una fase promedio y una fase de diferencia.
- 14Método según la reivindicación 12, en el que la representación de las fases respectivas incluye una fase del canal con una amplitud máxima, y una fase de diferencia.
- 15Método según la reivindicación 12, en el que la representación de las fases respectivas sólo se incluye en la señal para los sinusoides que tienen una frecuencia hasta cierta frecuencia umbral.
- 16Método según la reivindicación 15, en el que la frecuencia umbral dada es alrededor de 2 kHz.
- 17Método según la reivindicación 12, en el que la representación de las fases respectivas sólo se incluye en la señal para los sinusoides que tengan una diferencia de amplitud con al menos uno de los otros canales hasta cierto umbral de amplitud.
- 18Método según la reivindicación 17, en el que el umbral de amplitud dado es de 10 dB.
- 19Codificador (11) para codificar una señal (I, D) de audio de al menos dos canales, comprendiendo dicho codificador:ES 2 255 678 T3 medios (110) para determinar frecuencias (f com )comunes en los al menos dos canales (I, D) de la señal de audio, frecuencias comunes que ocurren en al menos dos de los al menos dos canales de la señal de audio medios (111) para representar elementos sinusoidales respectivos en canales respectivos en una frecuencia común dada mediante una representación de la frecuencia (f com ) común dada y una representación de las amplitudes (A, AA) respectivas de los elementos sinusoidales respectivos en la frecuencia común dada.
- 20Aparato (1) para transmitir o grabar, comprendiendo dicho aparato una unidad (10) de entrada para recibir una señal (S) de audio de al menos dos canales (I, D), un codificador (11) según la reivindicación 19 para codificar la señal (S) de audio para obtener una señal ([S]) de audio codificada, y una unidad de salida para proporcionar la señal ([S]) de audio codificada.
- 21Señal ([S]) de audio codificada que representa una señal (I, D) de audio de al menos dos canales que comprende:representaciones de frecuencias (fcom) comunes, frecuencias comunes que representan frecuencias que ocurren en al menos dos de los al menos dos canales de la señal [S] de audio, y para una frecuencia (f com ) común dada, una representación de amplitudes (A, AA) respectivas que representa elementos sinusoidales respectivos en canales respectivos en la frecuencia común dada.
- 22Medio (2) de almacenamiento que tiene almacenado en el mismo una señal según la reivindicación 21.
- 23Método para decodificar (31) una señal ([S]) de audio codificada, comprendiendo dicho método:recibir (31) la señal ([S]) de audio codificada que representa una señal (I, D) de audio de al menos dos canales, comprendiendo la señal de audio codificada representaciones de frecuencias (f com ) comunes, frecuencias comunes que representan frecuencias que ocurren en al menos dos de los al menos dos canales de la señal [S] de audio, y para una frecuencia (f com ) común dada, una representación de amplitudes (A, AA) respectivas que representan elementos sinusoidales respectivos en canales respectivos en la frecuencia común dada, y generar (31) las frecuencias comunes en las amplitudes respectivas en los al menos dos canales (I, D) para obtener una señal (S’) de audio decodificada.
- 24Decodificador (31) para decodificar una señal ([S]) de audio codificada, comprendiendo dicho decodificador:medios (31) para recibir la señal ([S]) de audio codificada que representan una señal (I, D) de audio de al menos dos canales, comprendiendo la señal de audio codificada representaciones de frecuencias (fcom) comunes, frecuencias comunes que representan frecuencias que ocurren en al menos dos de los al menos dos canales de la señal [S] de audio, y para una frecuencia (f com ) común dada, una representación de amplitudes (A, AA) respectivas que representan elementos sinusoidales respectivos en canales respectivos en la frecuencia común dada, y medios (31) para generar las frecuencias comunes en las amplitudes respectivas en los al menos dos canales (I, D) para obtener una señal (S’) de audio decodificada.
- 25Receptor o aparato (3) reproductor, comprendiendo el aparato:una unidad (30) de entrada para recibir una señal ([S]) de audio codificada, un decodificador (31) según la reivindicación 24 para decodificar la señal ([S]) de audio codificada para obtener una señal (S’) de audio decodificada, y una unidad (32) de salida para proporcionar la señal (S’) de audio decodificada.
Independent claims25
79 paragraphs in 4 sections, as filed
IS 2 255 678 T3
DESCRIPTION
Parametric audio coding.
The present invention relates to parametric audio coding.
Heiko Purnhagen, "Advances in parametric audio coding", Proc. 1999 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics, New Paltz, New York, Oct. 17-20, 1999 reports that parametric modeling provides a representation efficient of general audio signals and is used in very low bit rate audio encoding. It is based on the decomposition of an audio signal into elements that are described by suitable source models and represented by parameters of the models (such as frequency and amplitude of a pure tone). The perception models are used in the decomposition of the signal and in the coding of the parameters of the models.
An object of the invention is to provide an advantageous parameterization of a multi-channel audio signal (for example stereo). To this end, the invention provides a coding method, an encoder, an encoded audio signal, a storage medium, a decoding method, and a decoder, as defined in the independent claims. Advantageous embodiments are defined in the dependent claims.
It is noted that stereo audio coding as known in the prior art. For example, the two channels left (L) and right (R) can be encoded independently. This can be done by two independent encoders arranged in parallel or by time multiplexing in one encoder. Typically the two channels can be encoded more efficiently using cross channel correlation (and irrelevancies) in the signal. Reference is made to the MPEG-2 audio standard (ISO / IEC 13818-3, pages 5,6) which discloses a “joint” (dual channel) stereo encoding. Stereo joint encoding takes advantage of redundancy between the left and right channels to reduce the audio bit rate. Two forms of joint stereo coding are possible: MS stereo and intensity stereo. MS stereo is based on the coding of the sum (L + D) and difference (ID) signal rather than the left (I) and right (D) channels. Intensity coding is based on retaining at high frequencies only the energy envelope of the right (D) and left (I) channels. Direct application of the MS stereo coding principle in parametric coding rather than subband coding would result in a parameterized sum signal and a parameterized difference signal. The formation of the sum signal and the difference signal before encoding can result in the generation of additional frequency elements in the audio signal to be encoded, which reduces the efficiency of parametric encoding. Direct application of the intensity stereo coding principle in a parametric coding scheme would result in a low frequency part with independently coded channels and a high frequency part that includes only the energy envelope of the left and right channels.
According to a first aspect of the invention, common frequencies are determined in the at least two channels of the audio signal, common frequencies occurring in at least two of the at least two channels, and respective sinusoidal elements in respective channels at a common frequency. Given are represented by a representation of the given common frequency, and a representation of the respective amplitudes of the respective sinusoidal elements at the given common frequency. This aspect is based on the observation that a given frequency generated by a given source has a high probability of having an element in each of the channels. These signal elements will have their frequency in common. This is true because the signal transformations that can occur in transmission from the sound source by recording equipment to the listener will not normally affect the frequency elements differentially on several or all channels. Thus, common elements in the various signal channels can be represented by a single, common frequency. The respective amplitudes (and phases) of the respective elements in the respective channels may be different. Thus, by encoding the sinusoids with a common frequency and a representation of the respective amplitudes, efficient compression encoding of the audio signal is achieved; only one parameter is needed to encode a given common frequency (which occurs on multiple channels). Furthermore, such parameterization is advantageously applied with a suitable psychoacoustic model.
Once a common frequency has been found, the other parameters describing the elements on each respective channel can be plotted. For example, for a stereo signal that is represented with sinusoidal elements, the mean and difference of the amplitudes (and optionally the respective phases) can be encoded. In a further embodiment the largest amplitude in the encoded audio stream is encoded along with a difference amplitude, where the sign of the difference amplitude can determine the dominant channel for this frequency.
Since there is likely to be some degree of correlation between the left and right channels, an entropy coding of the sinusoidal parameters can be used which would result in more efficient coding of the stereo signal. Furthermore, irrelevant information within the representation of common elements can be eliminated, for example interaural phase differences at high frequencies are inaudible and can be set to zero.
Any frequency that occurs in the channels can be encoded as a common frequency. If a frequency that occurs on one channel does not occur on another channel, then the amplitude representation must be coded so that it results in zero amplitude for the channel on which the frequency does not occur.
IS 2 255 678 T3
Uncommon frequencies can also be represented as independent sinusoids on the respective channels. Uncommon frequencies can be encoded in a separate parameter block. It is also possible to produce a first block of parameters that includes common frequencies that are common to all channels, a second block of parameters that includes frequencies that are common to a subset (predetermined) of all channels, a third block of parameters that includes frequencies that are common to an additional (default) subset of all channels, and so on until a last block of parameters that includes the frequencies that occur in a single channel and that are encoded independently.
A common frequency can be represented as an absolute frequency value, but also as a frequency that changes with time, for example, a first derivative df / dt. Furthermore, common frequencies can be differentially encoded relative to other common frequencies.
Common frequencies can be found by estimating the frequencies considering two or more channels at the same time.
In a first embodiment, the frequencies are determined independently for the respective channels, followed by a comparison step to determine the common frequencies. Determination of the frequencies occurring on the respective channels can be done by a conventional matching-pursuit logarithm (see for example SG Mallat and Z. Zhang, "Matching pursuits with time-frequency dictionaries", IEEE trans. On Signal Processing, Vol. 41, No. 12, pp. 3397-3415) or peak width adjustment (see for example R. McAulay and T. Quatieri, “Speech Analysis / Synthesis Based on a Sinusoidal Representation”, IEEE Trans. ASSP, Vol. 34, No. 4, pp. 744-754, August 1986).
In a second embodiment, a combined matching pursuit algorithm is used to determine common frequencies. For example, respective power or energy representations of the at least two channels are combined to obtain a common representation. The common frequencies are then determined based on the common representation. Preferably, the power spectra of the at least two channels are added to obtain a common power spectrum. A conventional logarithm matching pursuit is used to determine the frequencies in this addition spectrum. Frequencies found in this added power spectrum are determined as common frequencies.
In a third embodiment, to determine common frequencies, peak width adjustment is used in the addition power spectra. The frequencies of the maxima found in this common power spectrum can be used as the common frequencies. Logarithmic power spectra could also be added instead of linear power spectra.
Preferably, the phase of the respective elements of the common frequency is also encoded. A common phase can be included in the encoded audio signal, which can be the average phase of the phases in the channels or the phase of the channel with the highest amplitude and a difference phase (interchannel). Advantageously, the difference phase is only encoded up to a given threshold frequency (eg 1.5 kHz or 2 kHz). For frequencies above this threshold, no difference phase is encoded. This is possible without significantly reducing quality, because human sensitivity for interaural phase differences is low for frequencies above this threshold. Therefore, a difference phase parameter is not necessary for frequencies above the given threshold. When decoding it can be assumed that the delta phase parameter is zero for frequencies above the threshold. The decoder is arranged to receive such signals. Above the threshold frequency the decoder does not wait for any code for the difference phases. Since the difference phases in the practical embodiment are not provided with an identifier, it is important that the decoder knows when to expect difference phases and when not. Furthermore, since the human ear is less sensitive to large differences in interaural intensity, delta amplitudes that are greater than a certain threshold, for example 10 dB, can be assumed to be infinite. Consequently, the interaural phase differences do not have to be encoded in this case either.
Frequencies on different channels that differ less than a given threshold can be represented by a common frequency. In this case it is assumed that the differing frequencies originate from the same source frequency. In practical embodiments, the threshold is related to the accuracy of the "matching pursuit" algorithm or peak amplitude adjustment.
In practical embodiments, the parameterization according to the invention is used on a frame basis.
The invention can be applied to any audio signal, including voice signals.
These and other aspects of the invention will be obvious from what will be understood with reference to the accompanying drawings.
In the drawings:
Figure 1 shows an encoder according to an embodiment of the invention; Figure 2 shows a possible implementation of the encoder of Figure 1;
Figure 3 shows an alternative implementation of the encoder of Figure 1, and Figure 4 shows a system according to an embodiment of the invention.
The drawings only show those elements that are necessary to understand the embodiments of the invention.
Figure 1 shows an encoder 11 according to an embodiment of the invention. A multichannel audio signal is fed into the encoder. In this embodiment the multichannel audio signal is a stereo audio signal having a left channel I and a right channel D. The encoder 11 has two inputs: one input for the left channel signal I and another input for the channel signal. right D. Alternatively, the encoder has an input for both I and R channels which are then provided in multiplexed form to encoder 11. The encoder 11 extracts sinusoids from both channels and determines the common frequencies f<sub>com</sub>. The result of the encoding process performed in encoder 11 is an encoded audio signal. The encoded audio signal includes the common frequencies fcom and for each common frequency fcom a representation of the respective amplitudes in the respective channels, for example in the form of a maximum or average amplitude A and a difference amplitude AA (delta).
The following describes how common frequencies can be determined, a first embodiment using a matching pursuit and a second embodiment using a peak width setting.
A realization that uses "matching pursuit"
This method is an extension of the existing matching pursuit algorithms. Matching pursuits are well known in the art. A matching pursuit is an iterative algorithm. Projects the signal onto an element of a correspondence dictionary chosen from a redundant dictionary of time-frequency waveforms. The projection is subtracted from the signal to be approximated in the next iteration. In this way, in the existing matching pursuit algorithms, the parameterization is carried out by determining by iterations a peak of the “projected” power spectrum of a frame of the audio signal, obtaining the optimal amplitude and phase that correspond to the frequency of the peak. and extracting the corresponding sinusoid from the frame being analyzed. This process is repeated iteratively until a satisfactory parameterization of the audio signal is obtained. To obtain common frequencies in a multichannel audio signal, the power spectra of the left and right channels are summed and the peaks of this addition power spectrum are determined. These peak frequencies are used to determine the optimal amplitudes and optionally the phases of the left and right channels (or more).
The multichannel matching pursuit algorithm according to a practical embodiment of the invention comprises the step of separating the multichannel signal into overlapping frames of short duration (for example 10 ms) and iteratively applying the following steps on each of the frames until reach a stopping criterion:
1. The power spectra of each of the channels of the multichannel frame are calculated
2. The power spectra are added to obtain a common power spectrum
3. Determine the frequency at which the "projected" common power spectrum is maximum
Four. For the frequency determined in step 3, the amplitude and phase of the sinusoids that best fit are determined and all these parameters are stored. These parameters are encoded using the common frequencies in combination with a representation of the respective amplitudes, thereby taking advantage of cross-channel correlations and irrelevancies.
5. The sinusoids are subtracted from the corresponding current multichannel frames to obtain an updated residual signal that serves as the next multichannel frame in stage 1.
Embodiment Using "Peak Width Adjustment"
Alternatively, peak width adjustment may be used, including for example the following steps:
1. The power spectra of each of the channels of the multichannel frame are calculated
2. The power spectra are added to obtain a common power spectrum
3. The frequencies corresponding to all the peaks that fall within the power spectrum are determined
Four. The best amplitudes and the best phases are obtained for these determined frequencies.
Figure 2 shows a possible implementation of the encoder of Figure 1, which uses a common power (addition) spectrum of the channels to determine the common frequencies. In the calculation unit 110 a matching pursuit process or a peak amplitude adjustment process is performed as described above using a common power spectrum obtained from the I and D channels. The determined common fcom frequencies are provided.
ES 2 255 678 T3 to the coding unit 111. This coding unit determines the respective amplitudes of the sinusoids (and preferably the phases) in the different channels at a given common frequency.
Alternatively, the respective channels are independently encoded to obtain a set of parameterized sinusoids for each channel. These parameters are subsequently verified for common frequencies. Such an embodiment is shown in Figure 3. Figure 3 shows an alternative implementation of the encoder 11 of Figure 1. In this implementation the encoder 11 comprises two independent parametric encoders 112 and 113. The parameters fi, A<sub>L</sub> and f<sub>D</sub>, TO<sub>D</sub> obtained in these independent encoders are provided to an additional encoding unit 114 which determines the frequencies f<sub>com</sub> common in these two parameterized signals.
Example of encoding a stereo audio signal
Assuming that a stereo audio signal is given with the following characteristics:
<td>channel</td><td>f (Hz)</td><td>A (dB)</td><td>f (Hz)</td><td>A (dB)</td><td>f (Hz)</td><td>A (dB)</td><td>f (Hz)</td><td>A (dB)</td><td>f (Hz)</td><td>A (dB)</td>
<td>I</td><td> 50</td><td> 30</td><td> 100</td><td> 50</td><td> 250</td><td> 40</td><td> -</td><td> -</td><td> 500</td><td> 40</td>
<td>D</td><td> 50</td><td> 20</td><td> 100</td><td> 60</td><td> -</td><td> -</td><td> 200</td><td> 30</td><td> 500</td><td> 35</td>
In practice, in this case the difference in amplitude between the channels is +15 dB or -15 dB at a given frequency, this frequency is considered to occur only in the dominant channel.
Coded independently
The following parameterization can be used to encode the exemplary stereo signal independently.
I (f, A) = (50, 30), (100, 50), (250,40), (500, 40)
D (f, A) = (50, 20), (100, 60), (200, 30), (500, 35)
This parameterization requires 16 parameters.
Using common frequencies and uncommon frequencies
Common frequencies are 50 Hz, 100 Hz, and 500 Hz. To encode this signal:
(fcom, Amax, AA) = (50, 30, 10), (100, 60, -10), (500, 40, 5) (fno-com, A) = (200, -30), (250, 40)
Encoding the stereo audio signal using common and uncommon frequencies requires 13 parameters in this example. Compared to the independently encoded multichannel signal, the use of common frequencies reduces the number of encoding parameters. Furthermore, the values for the delta amplitude are smaller than for the absolute amplitudes as given in the independently encoded multichannel signal. This further reduces the bit rate.
The signal in the AA delta amplitude determines the dominant channel (between two signals). In the example above, a positive amplitude means that the left channel is dominant. The sign can also be used in the representation of the uncommon frequency to indicate for which signal the frequency is valid. The same convention is used here: positive is left (dominant). Alternatively it is possible to provide an average amplitude in combination with a difference amplitude, or consistently the amplitude of a given channel with a difference amplitude relative to the other channel.
Instead of using the sign in the AA delta amplitude to determine the dominant channel, it is also possible to use a bit in the bit stream to indicate the dominant channel. This requires 1 bit, as may also be the case for the sign bit. This bit is included in the bit stream and is used in the decoder. In the case that an audio signal is encoded with more than two channels, more than 1 bit is needed to indicate the dominant channel. This implementation is straightforward.
Use only common frequencies
When using only a representation based on common frequencies, the uncommon frequencies are encoded so that the amplitude of the common frequency in the channel in which no sinusoids occur at that frequency is zero. In practice, a value of for example +15 dB or -15 dB can be used for the delta amplitude to indicate that there is no sinusoid of the current frequency on the given channel. The sign in the amplitude delta AA
ES 2 255 678 T3 determines the dominant channel (between two signals). In this example, a positive amplitude means that the left channel is dominant.
(fcom, A, AA) = (50, 30, 10), (100, 60, -10), (200, 30, -15), (250, 40, 15), (500, 40, 5)
This parameterization requires 15 parameters. For this example, using only common frequencies is less advantageous than using common and uncommon frequencies.
Average frequencies and differences (Fav, AF, A<sub>av</sub>, AA) = (50, 0, 25, 5), (100, 0, 55, -5), (225, 25, 35, 5), (500, 0, 30, 10)
This parameterization requires 16 parameters.
This is an alternative encoding in which the sinusoidal elements in the signal are represented by average frequencies and average amplitudes. It is clear that also compared to this coding strategy, the use of common frequencies is advantageous. It is noted that the use of average frequencies and average amplitudes can be viewed as a separate invention outside the scope of the present application.
Note that it is not strictly the number of parameters but rather the sum of the number of bits per parameter that is important for the bit rate of the resulting encoded audio stream. In this regard, differential encoding typically provides a bit stream reduction for correlated signal elements.
The representation with a common frequency parameter and respective amplitudes (and optionally respective phases) can be viewed as a mono representation, captured at the common frequency, the maximum or average amplitude, the phase of the maximum or average amplitude (optional) of the parameters and a multi-channel extension captured in the delta amplitude and delta phase parameters (optional). Mono parameters can be treated as standard parameters that can be obtained from a sinusoidal mono encoder. Thus, these mono parameters can be used to create links between sinusoids in subsequent frames, to differentially encode parameters along these links, and to perform phase continuation. Additional multichannel parameters can be encoded according to the strategies mentioned above that further take advantage of the stereophonic listening properties. The delta parameters (delta amplitude and delta phase) can also be differentially encoded based on the links that have been made based on the mono parameters. Furthermore, to provide a scalable bit stream, the mono parameters can be included in a base layer, while the multichannel parameters are included in a backing layer.
In tuning the mono components, the cost function (or measure of similarity) is a combination of the cost for the frequency, the cost for the amplitude, and (optionally) the cost for the phase. For stereo elements, the cost function can be a combination of the cost for the common frequency, the cost for the average or maximum amplitude, the cost for the phase, the cost for the delta amplitude, and the cost for the delta phase. Alternatively, it can be used for the cost function for the stereo elements: the common frequency, the respective amplitudes and the respective phases.
Advantageously, sinusoidal parameterization using a common frequency and a representation of the respective amplitudes of that frequency in the respective channels is combined with a transient mono parameterization as disclosed in WO 10/69593-A1. This can be further combined with a mono representation for noise such as that described in WO 01/88904.
Although most of the embodiments described above are related to two-channel audio signals, extension to three or more channels is straightforward.
The addition of an additional channel to an already encoded audio signal can be advantageously carried out as follows: it is sufficient to identify in the encoded audio signal a representation of the amplitudes of the common frequencies present in the extra channel and a representation of the non-frequencies. common. Phase information may also optionally be included in the encoded audio signal.
In a practical embodiment, the average or maximum amplitude and the average phase of the highest amplitude at a common frequency are quantized similarly to the respective quantization of the delta amplitude and the delta phase at the common frequency for the other ( s) channel (s). The practical values for quantification are:
common frequency 0.5% resolution amplitude, delta amplitude 1 dB resolution phase, delta phase 0.25 rad resolution
The proposed multichannel audio coding provides a bit stream reduction when compared to coding the channels separately.
IS 2 255 678 T3
Figure 4 shows a system according to an embodiment of the invention. The system comprises an apparatus 1 for transmitting or storing an encoded audio signal [S]. The apparatus 1 comprises an input unit 10 for receiving an audio signal S of at least two channels. The input unit 10 can be an antenna, microphone, network connection, etc. The apparatus 1 further comprises the encoder 11, as shown in figure 1 for encoding the audio signal S to obtain an audio signal encoded with a parameterization according to the present invention, for example (f<sub>com</sub>, TO<sub>av</sub>, AA) or (f<sub>com</sub>, TO<sub>max</sub>, AA). Parameterization of the encoded audio signal is provided to an output unit 12 which transforms the encoded audio signal into a format [S] suitable for transmission or storage by a transmission medium or a storage medium 2. The system comprises additionally a receiver or reproducing apparatus 3 receiving the encoded audio signal [S] in an input unit 30. The input unit 30 extracts from the encoded audio signal [S] the parameters (f<sub>com</sub>, TO<sub>av</sub>, AA) or (f<sub>com</sub>, TO<sub>max</sub>, AA). These parameters are provided to a decoder 31 which synthesizes a decoded audio signal based on the received parameters generating the common frequencies having the respective amplitudes to obtain the two L and R channels of the decoded audio signal S '. The two L and R channels are provided to an output unit 32 which provides the decoded audio signal S '. The output unit 32 may be a reproduction unit such as a speaker for reproducing the decoded audio signal S '. The output unit 32 may also be a transmitter for further transmitting the decoded audio signal S ', for example, through a home network, etc.
It should be noted that the above-mentioned embodiments illustrate rather than limit the invention, and that those skilled in the art will be able to design many alternative embodiments without departing from the scope of the appended claims. In the claims, any reference sign placed in parentheses shall not be construed as limiting the claim. The word "comprises" does not exclude the presence of other elements or steps than those listed in a claim. The invention can be implemented by physical equipment comprising several defined elements, and by a suitably programmed computer. In a device claim listing multiple media, multiple of these media may be realized in a single piece of hardware. The mere fact that certain measures are recited in different dependent claims does not indicate that a combination of these measures cannot be used to advantage.
Contents4
2 sheets
Sheet 1 Sheet 2
16 members in 10 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 02075639 | European Patent Office (EPO) | A | |
| 02075639 | European Patent Office (EPO) | A | |
| 20020075639 | European Patent Office (EPO) | – | |
| 0373958602075639 | – | – | – |
| EP20020075639 | – | – | – |
Members16
| Document | Office | Kind | |
|---|---|---|---|
| WO03069954A2 | World Intellectual Property Organization (WIPO) | A2 | |
| AU2003201097A1 | Australia | A1 | |
| AU2003201097A8 | Australia | A8 | |
| WO03069954A3 | World Intellectual Property Organization (WIPO) | A3 | |
| KR20040080003A | Republic of Korea | A | |
| EP1479071A2 | European Patent Office (EPO) | A2 | |
| US2005078832A1 | United States of America | A1 | |
| JP2005517987A | Japan | A | |
| CN1705980A | China | A | |
| EP1479071B1 | European Patent Office (EPO) | B1 | |
| AT315823T | Austria | T | |
| ATE315823T1 | Austria | T1 | |
| DE60303209D1 | Germany | D1 | |
| ES2255678T3This record | Spain | T3 | |
| DE60303209T2 | Germany | T2 | |
| JP4347698B2 | Japan | B2 |
Numbers
- Publication
- 2255678
- Publication, DOCDB
- 2255678
- Publication, EPODOC
- ES2255678T
- Application
- 3739586
- Application, DOCDB
- 03739586
- Application, EPODOC
- ES20030739586T
Titles2
- Spanish
- CODIFICACION DE AUDIO PARAMETRICA.
- English
- PARAMETRIC AUDIO CODING.
Classification
- CPC, 2
- G10L19/08
- G10L19/008
- IPC, 3
- G10L19 02
- G10L19 00
- G10L19 08