Parametric encoder for encoding a multi-channel audio signal
Abstract
A parametric audio encoder (100) for generating an encoding parameter (ICC) for an audio channel signal (X1 [b]) of a plurality of audio channel signals (X1 [b], X2 [b]) of a multichannel audio signal, each audio channel signal having (X 1 [b], X2 [b]) audio channel signal values (X1 [k], X2 [k]), the coding parameter being an interchannel coherence parameter (ICC), the parametric audio encoder (100) comprising a parameter generator (105), the parameter generator (105) being for - to determine for the audio channel signal (X1 [b]) of the plurality of channel signals of audio a first set of encoding parameters (IPD [b]) from the audio channel signal values (X1 [k]) of the audio channel signal (X1 [b]) and audio signal values reference (X2 [k]) of a reference audio signal (X2 [b]), wherein the reference audio signal is another audio channel signal (X2 [b]) of the plurality of audio channel signals or a reduced mixed audio signal derived from at least two audio channel signals of the plurality of multichannel audio signals, wherein the first set of coding parameters (IPD [b]) are inter-channel phase difference parameters or sub-band inter-channel phase difference parameters, - to determine for the audio channel signal (X1 [b]) a first average value of coding parameters (IPDmean [i]) on the basis of the first set of coding parameters (IPD [b]) of the first channel of audio signal (X1 [b]), the first average value of coding parameters referring to a current frame of the audio channel signal, wherein the parameter generator (105) is configured to determine the first average value of encoding parameters (IPDmean [i]) of the audio channel signal (X1 [b]) as an average of the first set of encoding parameters (IPD [b]) of the audio channel signal (X1 [b]) through frequency bands [k] or frequency subbands [b], - to determine for the audio channel signal (X1 [b]) a second average of encoding parameters (IPDmean_long_term) based on the first average of encoding parameters (IPDmean [i]) of the audio channel signal ( X1 [b]) and at least one other first half of the encoding parameters (IPDmean [i-1]) of the audio channel signal (X1 [b]), referring to the at least one other first half of the parameters of encoding to an earlier frame of the audio channel signal, and - to determine the coding parameter (ICC) on the basis of the first encoding parameter average (IPDmean [i]) of the audio channel signal (X1 [b]) and the second encoding parameter average ( IPDmean_long_term) of the audio channel signal (X1 [b]); wherein the parameter generator (105) is further configured - to determine an absolute value (IPDdist) of a difference between the second average of encoding parameters (IPDmean_long_term) and the first average of encoding parameters (IPDmean [i]) , and - to determine the coding parameter (ICC) as a determined absolute value function (IPDdist).
Term
5.4 yearsto projected expiry
Projected expiry 17 February 2032, counted from filing; an application has no term until it is granted.
- Priority and filed
- Published
- Today
- Projected expiry
11 claims: 8 independent, 3 dependent
- 1ES 2 555 136 T3 REIVINDICACIONES 1. Un codificador de audio paramétrico (100) para generar un parámetro de codificación (ICC) para una señal de canal de audio (X-i[b]) de una pluralidad de señales de canales de audio (X-i[b], X2[b]) de una señal de audio multicanal, teniendo cada señal de canal de audio (X-i[b], X2[b]) valores de señal de canal de audio (X-i[k], X2[k]), siendo el parámetro de codificación un parámetro de coherencia intercanales (ICC), comprendiendo el codificador de audio paramétrico (100) un generador de parámetros (105), estando el generador de parámetros (105) para - para determinar para la señal de canal de audio (X-i[b]) de la pluralidad de señales de canales de audio un primer conjunto de parámetros de codificación (IPD[b]) a partir de los valores de señal de canal de audio (X-i[k]) de la señal de canal de audio (X-i[b]) y valores de señal de audio de referencia (X2[k]) de una señal de audio de referencia (X2[b]), en donde la señal de audio de referencia es otra señal de canal de audio (X2[b]) de la pluralidad de señales de canales de audio o una señal de audio mezclada reducida derivada de al menos dos señales de canales de audio de la pluralidad de señales de audio multicanales, en donde el primer conjunto de parámetros de codificación (IPD[b]) son parámetros de diferencia de fase intercanales o parámetros de diferencia de fase intercanales de sub-banda, - para determinar para la señal de canal de audio (X-i[b]) un primer valor medio de parámetros de codificación (IPDmean[i]) sobre la base del primer conjunto de parámetros de codificación (IPD[b]) del primer canal de señal de audio (X-i[b]), refiriéndose el primer valor medio de parámetros de codificación a una trama corriente de la señal de canal de audio, en donde el generador de parámetros (105) está configurado para determinar el primer valor medio de parámetros de codificación (IPDmean[i]) de la señal de canal de audio (X-i[b]) como una media del primer conjunto de parámetros de codificación (IPD[b]) de la señal de canal de audio (X-i[b]) a través de bandas de frecuencias [k] o de sub-bandas de frecuencia [b], - para determinar para la señal de canal de audio (X-i[b]) una segunda media de parámetros de codificación (IPDmean_long_term) en función de la primera media de parámetros de codificación (IPDmean[i]) de la señal de canal de audio (X-i[b]) y al menos una otra primera media de parámetros de codificación ( IPDmean[i-1]) de la señal de canal de audio (X-i[b]), refiriéndose a la al menos una otra primera media de parámetros de codificación a una trama anterior de la señal de canal de audio, y - para determinar el parámetro de codificación (ICC) sobre la base de la primera media de parámetros de codificación (IPDmean[i]) de la señal de canal de audio (X-i[b]) y la segunda media de parámetros de codificación (IPDmean_long_term) de la señal de canal de audio (X-i[b]);en donde el generador de parámetros (105) está configurado, además - para determinar un valor absoluto (IPDdist) de una diferencia entre la segunda media de parámetros de codificación (IPDmean_long_term) y la primera media de parámetros de codificación (IPDmean[i]), y - para determinar el parámetro de codificación (ICC) como una función de valor absoluto determinado (IPDdist).
- 2El codificador de audio paramétrico (100) según la reivindicación 1, en donde el generador de parámetros (105) está configurado para determinar diferentes de fase de valores de señal de canal de audio siguientes (X-i[k]) con el fin de obtener el primer conjunto de parámetros de codificación (IPD[b]).
- 3El codificador de audio paramétrico (100) según una de las reivindicaciones precedentes, en donde la señal de canal de audio (X-i[b]) y la señal de audio de referencia (X2[b]) son señales del dominio de frecuencia y en donde los valores de la señal de canal de audio (X-i[k]) y los valores de la señal de audio de referencia (X2[k]) están asociados con las bandas de frecuencia (k) o las sub-bandas de frecuencia (b).
- 4El codificador de audio paramétrico (100) según una de las reivindicaciones precedentes, que comprende, además, un transformador (FFT) para transformar una pluralidad de señales de canales de audio en el dominio temporal (x-i[n], X2[n]) en el dominio de la frecuencia para obtener la pluralidad de señales de canales de audio (X-i[b], X2[b]).
- 5El codificador de audio paramétrico (100) según una de las reivindicaciones precedentes, en donde el generador de parámetros (105) está configurado para determinar el primer conjunto de parámetros de codificación (IPD[b]) para cada contenedor de frecuencia ([k]) o para cada sub-banda de frecuencia ([b]) de las señales de canales de audio (X-i[b], X2[b]).
- 6El codificador de audio paramétrico (100) según una de las reivindicaciones precedentes, en donde el generador de parámetros (105) está configurado para determinar la segunda media de parámetros de codificación (IPDmean_long_term) de la señal de canal de audio (X-i[b]) como una media de una pluralidad de primeras medias de parámetros de codificación (IPDmean[i]) en una pluralidad de tramas de la señal de canal de audio (X-i[b]), en donde cada primera media de parámetros de codificación (IPDmean[i]) está asociada a una primera trama (i) de la señal de ES 2 555 136 T3 audio multicanal.
- 7El codificador paramétrico (100) según una de las reivindicaciones precedentes, en donde el generador de parámetros (105) está configurado para determinar el parámetro de codificación (ICC) a partir de una diferencia entre un primer valor paramétrico (d) y el valor absoluto determinado (IPDdist) multiplicado por un segundo valor paramétrico (e).
- 8El codificador de audio de paramétrico (100) según la reivindicación 7, en donde el generador de parámetros (105) está configurado para establecer el primer valor paramétrico (d) a uno y para establecer el segundo valor paramétrico (e) a uno.
- 9El codificador de audio paramétrico (100) según una de las reivindicaciones precedentes, que comprende, además, un generador de señales mezcladas reducidas para superponer al menos dos de las señales de canales de audio de la señal de audio multicanal con el fin de obtener una señal mezclada reducida, un codificador de audio, en particular, un codificador mono, para codificar la señal mezclada reducida para obtener una señal de audio codificada y un combinador para combinar la señal de audio codificada con un parámetro de codificación correspondiente.
- 10Un método (400) para generar un parámetro de codificación (ICC) para una señal de canal de audio (X-i[b]) de una pluralidad de señales de canales de audio (X-i[b], X2[b]) de una señal de audio multicanal, teniendo cada señal de canal de audio (X-i[b], X2[b]) valores de señales de canales de audio (X-i[k], X2[k]), siendo el parámetro de codificación un parámetro de coherencia intercanales (ICC), comprendiendo el método (400):- la determinación (407) para la señal de canal de audio (X-i[b]) de la pluralidad de señales de canales de audio un primer conjunto de parámetros de codificación (IPD[b]) a partir de los valores de señales de canales de audio (X-i[k]) de la señal de canal de audio (X-i[b]) y los valores de la señal de audio de referencia (X2[k]) de una señal de audio de referencia (X2[b]), en donde la señal de audio de referencia es otra señal de canal de audio (X2[b]) de la pluralidad de señales de canales de audio o una señal de audio mezclada reducida derivada de al menos dos señales de canales de audio de la pluralidad de señales de audio multicanales, en donde el primer conjunto de parámetros de codificación (IPD[b]) son parámetros de diferencia de fase intercanal o parámetros de diferencia de fase intercanal de sub-banda, - la determinación (409) para la señal de canal de audio (X-i[b]) de una primera media de parámetros de codificación (IPDmean[i]) sobre la base del primer conjunto de parámetros de codificación (IPD[b]) de la señal de canal de audio (X-i[b]), refiriéndose la primera media de parámetros de codificación a una trama corriente de la señal de canal de audio, en donde la primera media de parámetros de codificación (IPDmean[i]) se determina como una media del primer conjunto de parámetros de codificación (IPD[b]) de la señal de canal de audio (X-i[b]) a través de bandas de frecuencia [k] o sub-bandas de frecuencia [b], - la determinación (411) para la señal de canal de audio (X-i[b]) de una segunda media de parámetros de codificación (IPDmean_long_term) sobre la base de la primera media de parámetros de codificación (IPDmean[i]) de la señal de canal de audio (X-i[b]) y al menos una otra primera media de parámetros de codificación ( IPDmean[i-1]) de la señal de canal de audio (X-i[b]), refiriéndose a por lo menos otra primera media de parámetros de codificación a una trama anterior de la señal de canal de audio, y - la determinación (413) del parámetro de codificación (ICC) sobre la base de la primera media de parámetros de codificación (IPDmean[i]) de la señal de canal de audio (X-i[b]) y la segunda media de parámetros de codificación (IPDmean_long_term) de la señal de canal de audio (X-i[b]);en donde la determinación (413) del parámetro de codificación (ICC) sobre la base de la primera media de parámetros de codificación (IPDmean[i]) de la señal de canal de audio (X-i[b]) y la segunda media de parámetros de codificación (IPDmean_long_term) de la señal de canal de audio comprende: - la determinación de un valor absoluto (IPDdist) de una diferencia entre la segunda media de parámetros de codificación (IPDmean_long_term) y la primera media de parámetros de codificación (IPDmean[i]), y - la determinación del parámetro de codificación (ICC) como una función del valor absoluto determinado (IPDdist).
- 11Un programa informático que está configurado para poner en práctica el método según la reivindicación 10 cuando se ejecuta en un ordenador.
Independent claims11
231 paragraphs in 15 sections, as filed
ES 2 555 136 T3
DESCRIPTION
Parametric encoder to encode a multichannel audio signal
FIELD OF THE INVENTION
The present invention relates to audio coding.
BACKGROUND OF THE INVENTION
Parametric audio coding in stereo or multichannel as described, eg, in C. Faller and F. Baumgarte, "Efficient representation of spatial audio signal using perceptual parameterization", in Proc. IEEE Workshop on Appl. of Sig. Proc. for audio and acoustics, October 2001, pages 199-202, uses spatial tracks to synthesize multichannel audio signals from down-mixed audio signals - typically mono or stereo - with multichannel audio signals having more channels than mixed audio signals reduced. Under normal conditions, down-mixed audio signals result from a superposition of a plurality of audio channel signals from a multi-channel audio signal, eg, from a stereo audio signal. These channels are encoded in waveform and secondary information, that is, the spatial tracks, referred to the original channel signal ratios are added as encoding parameters to the encoded audio channels. The decoder uses this secondary information to regenerate the original number of audio channels based on the decoded waveform encoded channels.
A basic parametric stereo encoder can use Interchannel Level Differences (ILD) as a necessary track to generate the stereo signal from the monaural down-mixed audio signal. More sophisticated encoders can also use inter-channel coherence (ICC), which can represent a degree of similarity between the signals of audio channels, that is, audio channels. Furthermore, when encoding binaural stereo signals, eg, for 3D audio signals or the surrounding headphone-based rendering, also an inter-channel phase difference (IPD) can play an important role in reproducing phase / delay differences between channels.
ICC track synthesis can be of importance for most audio and musical content to regenerate ambient environment, stereo reverb, font width, and other perceptions related to special printing as described in J. Blauert, Spatial Hearing: The Psychophysics of Human Sound Localization, The MIT Press, Cambridge, Massachusetts, United States, 1997.
Coherence synthesis can be performed using so-called frequency domain de-correlators as described in E. Schuijers, W. Oomen, B. den Brinker and J. Breebaart, "Advances in Parametric Coding for Audio of high quality ”, in Preprint 114th Conv. Aud. Eng. Soc., March 2003. However, known synthesis methods for estimating spatial tracks and synthesizing multichannel audio signals may suffer from increased complexity, see EP 1565036. Furthermore, the use of ICC parameters, eg, in addition to other parameters, such as inter-channel level differences (ICLDs) and inter-channel phase differences (ICPDs) can increase bit rate overhead.
SUMMARY OF THE INVENTION
It is the object of the invention to provide a concept for estimating coding parameters representing interchannel relationships between channels of a multichannel audio signal for efficient coding of audio signals.
This objective is achieved by the characteristics set out in the independent claims. Other embodiments are apparent from the dependent claims, the description, and the figures.
In order to be able to describe the invention in detail, the following terms, abbreviations and notations will be used:
BCC: Binaural track encoding, stereo or multichannel signal encoding using reduced mix and binaural tracks (or spatial parameters) to describe the relationships between channels.
Binaural tracks: Interchannel tracks between the input signals in the left ear and in the right hate (see also ITD, ILD, and IC).
CLD: Channel level difference, same as ICLD.
FFT: Fast realization of the DFT, indicated as Fast Fourier Transform.
STFT: Short duration Fourier transform.
ES 2 555 136 T3
HRTF: Headphone related transfer function, sound modeling transduction from a source to the left and right ear inputs in a free field.
IC: Interaural coherence, that is, degree of similarity between the input signals from the left ear and the right ear. This term is also sometimes referred to as IAC or interaural cross correlation (IACC).
ICC: Interchannel coherence, interchannel correlation.
ICPD: Interchannel Phase Difference. Average phase difference between a pair of signals.
ICLD: Interchannel level difference.
ICTD: Interchannel time difference.
ILD: Interaural level difference, that is, difference in level between the input signals from the left ear and the right ear. Also sometimes referred to as interaural intensity difference (IID)
IPD: Interaural phase difference, that is, the phase difference between the input signals from the left ear and the right ear.
ITD: Interaural time difference, that is, time difference between the input signals from the left ear and the right ear. It is also sometimes referred to as an interaural delay.
Mixing: Given a number of source signals (eg, separately recorded instruments, multitrack recording), the process of generating stereo or multichannel audio signals intended for spatial audio reproduction indicated by mixing.
Spatial Audio: Audio signals that, when reproduced through a suitable reproduction system, operationally recall a spatial image of the auditorium.
Spatial clues: Important clues for spatial perception. This term is used for tracks between channel pairs of a stereo or multichannel audio signal (see also ICTD, ICLD, and ICC), also referred to as spatial parameters or binaural tracks.
In accordance with the first aspect of the inventive idea, the invention relates to a parametric audio encoder for generating a coding parameter for an audio channel signal for a plurality of audio channel signals of a multi-channel audio signal, each audio channel signal having channel signal values, the encoding parameter being an inter-channel coherence parameter, the parametric audio encoder comprising a parameter generator, the parameter generator being configured to
- determining for the audio channel signal of the plurality of audio channel signals a first set of coding parameters from the values of the audio channel signals of the values of the reference audio signal and of the signal of audio channel of a reference audio signal, wherein the reference audio signal is another audio channel signal of the plurality of audio channel signals, wherein the first set of encoding parameters are inter-channel phase difference parameters or sub-band inter-channel phase difference parameters,
- to determine for the audio channel signal a first average of coding parameters on the basis of the first set of coding parameters of the audio channel signal, the first average of coding parameters referring to a current frame of the signal audio channel, wherein the parameter generator is configured to determine the first average of audio channel signal encoding parameters as an average of the first set of audio channel signal encoding parameters across the frequency or frequency bands. the frequency sub-bands,
- to determine for the audio channel signal, a second mean of coding parameters based on the first mean of coding parameters of the audio channel signal and at least one other first mean of coding parameters of the signal audio channel, referring to the at least one other first of encoding parameters to a previous frame of the audio channel signal, and
- to determine the encoding parameter on the basis of the first average of encoding parameters of the audio channel signal and the second average of encoding parameters of the audio channel signal;
ES 2 555 136 T3 where the parameter generator is configured, in addition,
- to determine an absolute value of a difference between the second mean of encoding parameters and the first mean of encoding parameters, and
- to determine the encoding parameter as a function of the determined absolute value.
Using the current and previous frames of the audio channel signal, the long-term averaging operation can be efficiently performed.
The reference audio signal may be one of the audio channel signals of the multichannel audio signal. In particular, the reference audio signal may be a left ear or right ear audio channel signal of a stereo signal constituting an embodiment of the two-channel multi-channel signal. However, the reference audio signal can be any signal that constitutes a reference for determining the encoding parameters. Said reference signal may be constituted by a monaural reduced mixed audio signal after the reduced mixing of the channels of the multichannel audio signal or one of the channels of a reduced mixed audio signal after the reduced mixing of the channels. of the multichannel audio signal.
Said parameters such as inter-channel phase difference or sub-band inter-channel phase difference represent a degree of similarity between the audio signals and therefore can be used by the encoder to reduce the information to be transmitted and thus reduce complexity. of the calculation.
By averaging the first set of encoding parameters of the audio channel signal across the frequency bands or frequency sub-bands, the parametric audio encoder provides a short duration average of the audio signal in where all frequency components are considered.
By using that difference between the second mean of encoding parameters and the first mean of encoding parameters, the parametric audio encoder provides a measure for the difference between the long-term mean and the short-run mean and is therefore able to to predict the behavior of the voice or music.
When the encoding parameter is provided as a function of the determined absolute value, there is a relationship between the encoding parameter and the determined absolute value, which can be used to efficiently calculate the encoding parameter. The computational complexity is reduced in this way.
The parametric audio encoder can have low complexity since it does not require coherence or correlation computation. It even provides an accurate estimate of the relationship between the audio channels when the ICC value is quantized with an approximate quantizer that requires only a few steps. In particular, for music signals, but also for voice signals, the use of the encoding parameter for encoding the audio signals is important because the output music sounds more natural with the correct acoustic scene width and not " dry". For a low bit rate parametric stereo audio coding system, the bit budget is limited and only a full band ICC is transmitted, the coding parameter being capable of representing the overall correlation between the channels.
In a first possible way of implementing the parametric audio encoder in accordance with the first aspect of the inventive idea, the parameter generator is configured to determine the phase differences of the following audio channel signal values to obtain the first set of encoding parameters.
The phase differences of the following audio channel signals are required to reproduce the phase differences and / or delay between the channels. When phase differences are reproduced, voice and music sound more natural.
In a second possible form of implementation of the parametric audio encoder in accordance with the first aspect of the inventive idea or in accordance with the preceding form of implementation of the first aspect of the inventive idea, the audio channel signal and the reference audio signal are frequency domain signals and the values of the audio channel signals and the reference audio signal values are associated with the frequency bands or the sub -frequency bands.
The frequency resolution used is largely motivated by the frequency resolution of the auditorium system. Psychoacoustics suggests that spatial perception is most likely based on a critical band representation of the acoustic input signal. This frequency resolution is considered using an invertible filter bank with sub-bands with bandwidths equal or proportional to the critical bandwidth of the auditorium system. In this way, the parametric audio encoder can be well adapted to human perception.
ES 2 555 136 T3
In a third possible form of implementation of the parametric audio encoder in accordance with the first aspect of the inventive idea or in accordance with any of the preceding forms of implementation of the first aspect, the parametric audio encoder further comprises: a transformer for transforming a plurality of audio channel signals in the time domain into the frequency domain to obtain the plurality of audio channel signals.
Equalization of the channel impulse response can be performed efficiently in the frequency domain since the convolution in the time domain is a multiplication in the frequency domain. Thus, performing the parametric audio encoder calculations in the frequency domain can result in higher efficiency with respect to computational complexity or higher accuracy.
In a fourth possible embodiment of the parametric audio encoder, in accordance with the first aspect of the inventive idea or in accordance with any of the preceding embodiments of the first aspect, the parameter generator is configured to determine the first set of encoding parameters for each frequency container or for each frequency sub-band of the audio channel signals.
The parametric audio encoder can limit the determination of the first set of coding parameters to frequency bands or frequency sub-bands that are perceivable by the human ear and thus, complexity is reduced.
In a fifth possible embodiment of the parametric audio encoder in accordance with the first aspect of the inventive idea or in accordance with any of the preceding embodiments of the first aspect, The parameter generator is configured to determine the second mean of coding parameters of the audio channel signal as an average of a plurality of the first means of coding parameters over a plurality of frames of the audio channel signal , wherein each first half of the coding parameters is associated with a frame and of the multichannel audio signal.
Through this averaging, the parametric audio encoder provides a long-term average of the audio signal where the characteristic properties of the voice signal or the music signal are considered.
In a sixth possible embodiment of the parametric audio encoder in accordance with the first aspect of the inventive idea or in accordance with any of the preceding embodiments of the first aspect, the parameter generator is configured to determine the encoding parameter to starting from a difference between a first parametric value and the determined absolute value multiplied by a second parametric value.
When the encoding parameter is provided as a difference between the first parametric value and the determined absolute value, there is a relationship between the encoding parameter and the determined absolute value, which can be used to efficiently calculate the encoding parameter. This reduces the complexity of the calculation.
In a seventh possible embodiment of the parametric audio encoder in accordance with the first aspect of the inventive idea or in accordance with any of the preceding embodiments of the first aspect, the parameter generator is configured to set the first parametric value to one and to set the second parametric value to one.
Through this relationship, the parametric audio encoder is able to efficiently calculate the encoding parameter. This reduces the complexity of the calculation.
In an eighth possible embodiment of the parametric audio encoder in accordance with the first aspect of the inventive idea or in accordance with any of the preceding embodiments of the first aspect, the parametric audio encoder further comprises a signal generator downmixed for overlapping at least two of the audio channel signals of the multichannel audio signal to obtain a downmixed signal, an audio encoder, in particular a monaural encoder, for encoding the reduced mixed signal to obtain an encoded audio signal and a combiner for combining the encoded audio signal with a corresponding encoding parameter.
The reduced mixed signal and the encoded audio signal can be used as a reference signal for the parameter generator. Both signals include the plurality of audio channel signals and thus provide greater accuracy than a single channel signal taken as a reference signal.
In a ninth possible embodiment of the parametric audio encoder in accordance with the first aspect of the inventive idea, the current frame of the audio channel signal is contiguous with the previous frame or the audio channel signal.
ES 2 555 136 T3
When both frames are contiguous, so-called voltage spikes are detected in the audio channel signals in the middle and can be seen in the parametric audio encoder. In this way, the coding is more accurate than coding where such voltage peaks cannot be detected.
In accordance with a second aspect of the inventive idea, the invention relates to a parametric audio encoder for generating a coding parameter for an audio channel signal from a plurality of audio channel signals from a multi-channel audio signal, each audio channel signal having values of audio channel signals, the encoding parameter being an interchannel coherence parameter, the parametric audio encoder comprising a parameter generator, the parameter generator being configured
- to determine for the audio channel signal of the plurality of audio channel signals a first set of encoding parameters from the audio channel signal values of the audio channel signal and the signal values reference audio signal from a reference audio signal, wherein the reference audio signal is a reduced mixed audio signal derived from at least two audio channel signals from among the plurality of multichannel audio signals, wherein the first set of encoding parameters are phase difference parameters interchannel or phase difference parameters between sub-band channels,
- to determine for the audio channel signal a first average of coding parameters based on the first set of coding parameters of the audio channel signal, the first average of coding parameters referring to a current frame of the audio channel signal audio channel, wherein the parameter generator is configured to determine the first average of audio channel signal encoding parameters as an average of the first set of audio channel signal encoding parameters across the frequency or frequency bands. the frequency sub-bands,
- to determine for the audio channel signal a second average of coding parameters based on the first average of coding parameters of the audio channel signal and at least one other first average of coding parameters of the channel signal of audio, the at least one other encoding parameter mean referring to a previous frame of the audio channel signal and
- to determine the encoding parameter on the basis of the first average of encoding parameters of the audio channel signal and the second average of encoding parameters of the audio channel signal;
where the parameter generator is configured in addition
- to determine an absolute value of a difference between the second mean of encoding parameters and the first mean of encoding parameters, and
- to determine the encoding parameter as a function of the determined absolute value.
Using the current and previous frames of the audio channel signal, the long-term averaging operation can be efficiently performed.
The reference audio signal may be one of the audio channel signals of the multichannel audio signal. In particular, the reference audio signal may be a left or right ear channel signal of a stereo signal constituting an embodiment of a two-channel multi-channel signal. However, the reference audio signal can be any signal that constitutes a reference for determining the encoding parameters. Said reference signal can be formed by a reduced mixed audio signal after the reduced mixing of the channels of the multichannel audio signal or an output from a monaural encoder.
Parameters such as the interchannel phase difference or the phase difference between sub-band channels represent a degree of similarity between the audio signals and therefore can be used by the encoder to reduce the information to be transmitted and thus, reduce the complexity of the calculation.
By averaging the first set of encoding parameters of the audio channel signal across frequency bands or frequency sub-bands, the parametric audio encoder provides a short-term average of the audio signal where it is they consider all components of frequency.
By using that difference between the second mean of encoding parameters and the first mean of encoding parameters, the parametric audio identifier provides a measure of the difference between the long-term mean and the short-run mean and is therefore capable of to predict the behavior of the voice or music.
When the encoding parameter is given as a function of the given absolute value, there is a
ES 2 555 136 T3 relationship between the encoding parameter and the determined absolute value, which can be used to efficiently calculate the encoding parameter. This reduces the complexity of the calculation.
The parametric audio encoder can have low complexity since it does not require a correlation or coherence calculation. It even provides an accurate estimate of the relationship between the audio channels when the ICC value is quantized with a coarse quantizer that requires only a few steps. In particular, for music signals, but also for voice signals, the use of the encoding parameter for encoding the audio signals is important because the output music sounds more natural with the correct acoustic scene width and not " dry". For a very low bit rate parametric stereo audio coding system, the bit budget is limited and only a full band ICC is transmitted, the coding parameter being capable of representing the overall correlation between the channels.
In a first possible embodiment of the parametric audio encoder in accordance with the second aspect of the inventive idea, the parameter generator is configured to determine phase differences of signal values of subsequent audio channels in order to obtain the first encoding parameter set. The phase differences of the following audio channel signals are required to reproduce the phase differences and / or delay between the channels. When phase differences are reproduced, voice and music sound more natural.
In a second possible embodiment of the parametric audio encoder in accordance with the second aspect of the inventive idea or in accordance with the preceding embodiment of the second aspect, the audio channel signal and the reference audio signal are signals Frequency domain and audio channel signal values and reference audio signal values are associated with frequency bands or frequency sub-bands.
The frequency resolution used is largely motivated by the frequency resolution of the auditorium system. Psychoacoustics suggests that spatial perception is most likely based on a critical band representation of the acoustic input signal. This frequency resolution is considered using an invertible filter bank with sub-bands and bandwidths equal to or proportional to the critical bandwidth of the auditorium system. In this way, the parametric audio encoder can be suitably adapted for human perception.
In a third possible embodiment of the parametric audio encoder in accordance with the second aspect of the inventive idea or in accordance with any of the preceding embodiments of the second aspect, the parametric audio encoder further comprises a transformer for transforming a plurality of time domain audio channel signals in the frequency domain in order to obtain the plurality of audio channel signals.
Equalization of the channel impulse response can be performed efficiently in the frequency domain since the convolution in the time domain is a multiplication in the frequency domain. Thus, performing the parametric audio encoder calculations in the frequency domain can result in higher efficiency with respect to computational complexity or in higher accuracy.
In a fourth possible embodiment of the parametric audio encoder in accordance with the second aspect of the inventive idea or in accordance with any of the preceding embodiments of the second aspect, the parameter generator is configured to determine the first set of parameters encoding for each frequency container or for each frequency sub-bands of the audio channel signals.
The parametric audio encoder can limit the detection of the first set of coding parameters to frequency bands or frequency sub-bands that are perceivable by the human ear and thus, complexity is reduced.
In a fifth possible embodiment of the parametric audio encoder in accordance with the second aspect of the inventive idea or in accordance with any of the preceding embodiments of the second aspect, the parameter generator is configured to determine the second mean of coding parameters of the audio channel signal as an average of a plurality of first means of coding parameters during a plurality of frames of the audio channel signal, wherein Each first half of the encoding parameters is associated with one frame of the multichannel audio signal.
Through this averaging, the parametric audio encoder provides a long-term average of the audio signal where the characteristic properties of the voice signal or the music signal are considered.
In a sixth possible embodiment of the parametric audio encoder in accordance with the second aspect of the inventive idea or in accordance with any of the preceding embodiments of the second aspect, the parameter generator is configured to determine the encoding parameter to start from a
ES 2 555 136 T3 difference between a first parametric value and the absolute value determined multiplied by a second parametric value.
When the encoding parameter is provided as a difference between the first parametric value and the determined absolute value, there is a relationship between the encoding parameter and the determined absolute value, which can be used to efficiently calculate the encoding parameter. This reduces the complexity of the calculation.
In a seventh possible embodiment of the parametric audio encoder in accordance with the second aspect of the inventive idea or in accordance with any of the preceding embodiments of the second aspect, the parameter generator is configured to set the first parametric value to one and to set the second parametric value to one.
Through this relationship, the parametric audio encoder is able to efficiently calculate the encoding parameter. This reduces the complexity of the calculation.
In an eighth possible embodiment of the parametric audio encoder in accordance with the second aspect of the inventive idea or in accordance with any of the preceding embodiments of the second aspect, the parametric audio encoder further comprises a signal generator downmixed for overlapping at least two of the audio channel signals of the multichannel audio signal to obtain a downmixed signal, an audio encoder, in particular a monaural encoder, for encoding the reduced mixed signal in order to obtain an encoded audio signal and a combiner for combining the encoded audio signal with a corresponding encoding parameter.
The reduced mixed signal and the encoded audio signal can be used as a reference signal for the parameter generator. Both signals include the plurality of audio channel signals and thus provide greater accuracy than a single channel signal taken as a reference signal.
In a ninth embodiment of the parametric audio encoder in accordance with the second aspect of the inventive idea, the current frame of the audio channel signal is contiguous with the previous frame of the audio channel signal.
When both frames are contiguous, voltage spikes are detected in the audio channel signals in the averaging and can be seen in the parametric audio encoder. In this way, the coding is more accurate than coding where no voltage spikes can be detected.
In accordance with a third aspect of the inventive idea, the invention relates to a method for generating a coding parameter for an audio channel signal from a plurality of audio channel signals of a multi-channel audio signal, each signal having of audio channel audio channel signal values, the encoding parameter being an interchannel coherence parameter, the method comprising:
determining for the audio channel signal of the plurality of audio channel signals a first set of encoding parameters from the audio channel signal values of the audio channel signal and the signal values of reference audio of a reference audio signal, wherein the reference audio signal is another audio channel signal of the plurality of audio channel signals, wherein the first set of encoding parameters are inter-channel phase difference parameters or sub-band inter-channel phase difference parameters,
determining for the audio channel signal a first average of coding parameters on the basis of the first set of coding parameters of the audio channel signal, the first average of coding parameters referring to a current frame of the audio channel signal. audio channel, wherein the first mean of coding parameters is determined as an average of the first set of coding parameters of the audio channel signal across frequency bands or frequency sub-bands,
determining for the audio channel signal a second average of coding parameters on the basis of the first average of coding parameters of the audio channel signal and at least another first average of coding parameters of the audio channel signal. audio, referring to at least one other first half of encoding parameters to a previous frame of the audio channel signal, and
- determining the encoding parameter on the basis of the first average of encoding parameters of the audio channel signal and the second average of encoding parameters of the audio channel signal;
wherein the determination of the encoding parameter based on the first mean of encoding parameters of the audio channel signal and the second mean of encoding parameters of the audio channel signal comprises:
ES 2 555 136 T3
- determining an absolute value of a difference between the second mean of encoding parameters and the first mean of encoding parameters, and
- determining the encoding parameter as a function of the determined absolute value.
The method can be efficiently performed on a processor.
The reference audio signal may be one of the audio channel signals of the multichannel audio signal. In particular, the reference audio signal may be a left or right audio channel signal of a stereo signal constituting an embodiment of a two-channel multi-channel signal. However, the radio frequency audio signal can be any signal that constitutes a reference for determining the encoding parameters. Said reference signal can be constituted by a monaural reduced mixed audio signal after the reduced mixing of the channels of the multichannel audio signal or one of the channel of a reduced mixed audio signal after the reduced mixing of the channels of the multichannel audio signal.
In accordance with a fourth aspect of the inventive idea, the invention relates to a method for generating a coding parameter for an audio channel signal from a plurality of audio channel signals of a multi-channel audio signal, each signal having of audio channel audio channel signal values, the encoding parameter being an interchannel coherence parameter, the method comprising:
- determining for the audio channel signal of the plurality of audio channel signals a first set of encoding parameters from the audio channel signal values of the audio channel signal and the audio channel signal values reference audio of a reference audio signal, wherein the reference audio signal is a reduced mixed audio signal derived from at least two audio channel signals from among the plurality of multichannel audio signals, wherein the first set of encoding parameters are phase difference parameters interchannel or sub-band interchannel phase difference parameters,
determining for the audio channel signal a first average of coding parameters on the basis of the first set of coding parameters of the audio channel signal, the first average of coding parameters referring to a current frame of the audio channel signal. audio channel, wherein the first mean of encoding parameters is determined as the mean of the first set of encoding parameters of the audio channel signal across frequency bands or frequency sub-bands,
determining for the audio channel signal a second mean of coding parameters on the basis of the first mean of coding parameters of the audio channel signal and at least one other first mean of coding parameter of the channel signal of audio, the at least one other first half of encoding parameters referring to a previous frame of the audio channel signal and
- determining the encoding parameter on the basis of the first average of encoding parameters of the audio channel signal and the second average of encoding parameters of the audio channel signal;
wherein the determination of the encoding parameter based on the first mean of encoding parameters of the audio channel signal and the second mean of encoding parameters of the audio channel signal comprises:
- determining an absolute value of a difference between the second mean of encoding parameters and the first mean of encoding parameters, and
- determining the encoding parameter as a function of the determined absolute value.
The method can be performed efficiently on a processor.
The reference audio signal may be one of the audio channel signals of the multi-channel audio signal. In particular, the reference audio signal may be a left or right audio channel signal of a stereo signal constituting an embodiment of a two-channel multi-channel signal. However, the reference audio signal can be any signal that constitutes a reference for determining the encoding parameters. Said reference signal may be constituted by a monaural reduced mixed audio signal after the reduced mixing of the channels of the multichannel audio signal or one of the channels of a reduced mixed audio signal after the reduced mixing of the channels of the multichannel audio signal.
In accordance with a fifth aspect of the inventive idea, the invention relates to a computer program that is configured to implement the method in accordance with one of the third and fourth aspects of the inventive idea when run on a computer.
ES 2 555 136 T3
The computer program has a reduced complexity and thus can be efficiently implemented in a mobile terminal where battery life must be economized. The duration of the battery life is increased when the computer program is run on a mobile terminal .
The methods described herein can be implemented as software in a Digital Signal Processor (DSP), in a microcontroller or any other secondary processor or as a hardware circuit within an application specific integrated circuit (ASIC).
The invention can be practiced in digital electronic circuits or in computer hardware, firmware, software or one of their combinations.
BRIEF DESCRIPTION OF THE DRAWINGS
Additional embodiments of the invention will now be described with respect to the following Figures, in which:
Figure 1 illustrates a block diagram of a parametric audio encoder in accordance with one embodiment;
Figure 2 illustrates a block diagram of a parametric audio decoder in accordance with one embodiment;
Figure 3 illustrates a block diagram of a parametric stereo audio encoder and decoder in accordance with one embodiment; Y
Figure 4 illustrates a schematic diagram of a method for generating a coding parameter for an audio channel signal in accordance with one embodiment.
DETAILED DESCRIPTION OF THE FORMS OF EMBODIMENT OF THE INVENTION
Figure 1 illustrates a block diagram of a parametric audio encoder 100 in accordance with one embodiment. The parametric audio encoder 100 receives a multichannel audio signal 101 as an input signal and provides a bit stream as an output signal 103. The parametric audio identifier 100 comprises a parameter generator 105 coupled to the multichannel audio signal 101 to generate a coding parameter 115, a reduced mixed signal generator 107 coupled to the multichannel audio signal 101 to generate a reduced mixed signal 111 or a sum sign, an audio encoder 109 coupled to the reduced mixed signal generator 107 to encode the reduced mixed signal 111 to provide an encoded audio signal 113 and a combiner 117, eg, a bitstream forming device coupled to the parameter generator 105 and to the audio encoder 109 to form a bit stream 103 from encoding parameter 115 and encoded signal 113.
The parametric audio encoder 100 implements an audio coding system for multichannel and stereo audio signals, which only transmits a single audio channel, eg, the down-mix audio channel along with additional parameters that describe "perceptual differences relevant ”between audio channels Xi [b], X2 [b], ..., XM [b]. The coding system is in compliance with binaural track coding (BCC) because binaural tracks play an important role in this regard. As indicated in the Figure, the plurality M of input audio channels X¿b], X2 [b], ..., XM [b] of the multi-channel audio signal 101 are mixed down to one channel of single audio 111, also indicated as the sum signal. For a stereo audio signal M equal to 2, taking into account the “perceptually relevant differences” between the audio channels X¿b], X2 [b], ..., XM [b], the encoding parameter 115, eg, an interchannel time difference (ICTD), an interchannel level difference (ICLD) and / or an interchannel coherence (ICC) is estimated as a function of frequency and time and transmitted as secondary information to decoder 200 described in Figure 2.
The parameter generator 105 implemented by BCC processes the multichannel audio signal 101 with a certain period of time and frequency resolution. The frequency resolution used is largely driven by the frequency resolution of the auditorium system. Psychoacoustics suggests that spatial perception is most likely based on a critical band representation of the acoustic input signal. This frequency resolution is considered using an invertible filter bank with sub-bands with bandwidths equal or proportional to the critical bandwidth of the auditorium system. It is important that the transmitted sum 111 contains all the signal components of the multichannel audio signal 101. The goal is for each signal component to be fully maintained. The simple summation of the audio input channels X-jb], X2 [b], .., XM [b] of the multi-channel audio signal 101 usually results in the amplification or attenuation of signal components. In other words, the power of the signal components in the “simple” sum is usually greater or less than the sum of the power of the corresponding signal components of each channel X¿b], X2 [b], .. ., XM [b]. Therefore, a reduced mixing technique is used by applying the reduced mixing device 107 which equalizes the sum signal 111 so that the power of the signal components in the sum signal 111 is
ES 2 555 136 T3 approximately the same as the corresponding power in all input audio channels Xi [b], X2 [b], ..., XM [b] of the multi-channel audio signal 101. The audio channels input X¿b], X2 [b], ..., XM [b] represent the channel signals for sub-band b. The frequency domain input audio channel is indicated by Xi [k], X2 [k] ,. XM [k], where k represents the frequency index (frequency container), a sub-band b being normally constituted by several frequency bands k.
Given the sum signal 111, the parameter generator 105 synthesizes a multichannel or stereo audio signal 115 so that ICTD, ICLD and / or ICC approximates the corresponding tracks of the original multichannel audio signal 101.
When considering single source binaural spatial impulse responses (BRIRs), there is a relationship between the width of the auditorium operating event and the listening envelope and the estimated IC value for the beginning and ending parts of the BRIR responses. However, the relationship between IC (or ICC) and these properties for general signals (and not just BRIRs) is not simple. Stereo and multichannel audio signals often contain a complex mix of simultaneously active source signals overlaid by reflected signal components that are derived from recording in closed spaces or added by the recording technician to artificially create a spatial impression. Different source signals and their reflections occupy different areas in the time-frequency plane. The foregoing is reflected by ICTD, ICLD and ICC which vary as a function of time and frequency. In this case, the relationship between the instantaneous values of ICTD, ICLD, and ICC and the auditorium event addresses and spatial impression is not obvious. The strategy of the parameter generator 105 is to blindly synthesize these tracks so that they approximate the corresponding tracks of the original audio signal.
In one embodiment, the parametric audio encoder 100 uses filter banks with sub-bands of bandwidths equal to twice the equivalent rectangular bandwidth. An informal listening revealed that BCC's audio quality did not noticeably improve when a higher frequency resolution was chosen. Lower frequency resolution is favorable since it results in fewer ICTD, ICLD and ICC values that need to be transmitted to the decoder and thus at a lower bit rate. With regard to temporal resolution, ICTD, ICLD and ICC are considered at periodic time intervals. In one ICTD embodiment, ICLD and ICC are considered approximately every 4-16 ms. It should be noted that unless the tracks are considered at very short intervals, the priority effect is not directly considered.
The often perceptually small difference achieved between the reference signal and the synthesized signal implies that tracks related to a wide range of auditorium spatial image attributes are implicitly considered by synthesizing ICTD, ICLd, and ICC at periodic time intervals. The bit rate required for the transmission of these spatial tracks is only a few kb / s and thus the parametric audio encoder 100 is capable of transmitting multichannel and stereo audio signals at bit rates close to those required for a single audio channel. Figure 4 illustrates a method in which ICC is estimated as encoding parameter 115.
The parametric audio encoder 100 comprises a reduced mixed signal generator 107 for superimposing at least two of the audio channel signals of the multichannel audio signal 101 to obtain the reduced mixed signal 111, the audio encoder 109, in particular a monaural encoder, to encode the reduced mixed signal 111 in order to obtain the encoded audio signal 113 and the combiner 117 to combine the encoded audio signal 113 with a corresponding encoding parameter 115.
The parametric audio encoder 100 generates the encoding parameter 115 for an audio channel signal from among the plurality of audio channel signals indicated as X¿b], X2 [b], ..., XM [b] of the multi-channel audio signal 101. Each of the audio channel signals X¿b], X2 [b], ..., XM [b] may be a digital signal comprising values of digital audio channel signals in the frequency domain indicated as X¿k], X2 [k],., XM [k].
An exemplary audio channel signal for which the parametric audio encoder 100 generates the encoding parameter 115 is the first audio channel signal X¿b] with signal values X¿k]. The parameter generator 105 determines for the audio channel signal X¿b] a first set of encoding parameters indicated as IPD [b] from the audio channel signal values X¿k] of the channel signal X¿b Audio Signal] and from the reference audio signal values of a reference audio signal.
An audio channel signal that is used as a reference audio signal is the second audio channel signal X2 [b], as an example. Similarly, any of the other audio channel signals X¿b], X2 [b], ..., XM [b] can serve as a reference audio signal. In accordance with a first aspect of the inventive idea, the reference audio signal is another audio channel signal of the audio channel signals that is not equal to the audio channel signal X¿b] for which it is generates encoding parameter 115.
In accordance with a second aspect of the inventive idea, the reference audio signal is a reduced mixed audio signal derived from at least two audio channel signals from among the plurality of signals.
ES 2 555 136 T3 multi-channel audio 101, eg, derived from the first audio channel signal Xi [b] and the second audio channel signal X2 [b]. In one embodiment, the reference audio signal is the downmixed signal 111, also called the sum signal generated by the downmix device 107. In another embodiment, the reference audio signal is the encoded signal 113 provided by encoder 109.
An exemplary reference audio signal used by parameter generator 105 is the second audio channel signal X2 [b] with signal values X2 [k].
The parameter generator 105 determines for the audio channel signal X¿b] a first mean of coding parameters, indicated as IPDmean [i] on the basis of the first set of coding parameters IPD [b] of the channel signal audio X¿b].
The parameter generator 105 determines for the audio channel signal X¿b] a second mean of encoding parameters, indicated by IPDmean_long_term, based on the first mean of encoding parameters IPDmean [i] of the channel signal of audio X¿b] and at least one other first mean of encoding parameters indicated as IPDmean [i-1] of the audio channel signal X¿b]. In one embodiment, the first average of IPDmean coding parameters [i] refers to a current frame i of the audio channel signal X¿b] and the other first average of IPDmean coding parameters [i-1] refers to a previous frame i-1 of the audio channel signal X¿b]. In one embodiment, the previous frame i-1 of the audio channel signal X¿b] is the frame i-1 received before the current frame i without any other intermediate frames. In one embodiment, the previous frame iN of the audio channel signal X¿b] is an iN frame received before the current frame i but multiple frames arrived in the intervening period.
The parameter generator 105 determines the encoding parameter 115, denoted ICC, based on the first mean of encoding parameters IPDmean [i] of the audio channel signal X¿b] and on the basis of the second mean of encoding parameters IPDmean_long_term of the X¿b audio channel signal].
The first set of IPD coding parameters [b] are phase differences between channels, level differences between channels, coherence between channels, intensity differences between channels, level differences between sub-band channels, phase differences between channels of sub-band, sub-band inter-channel coherence, intensity differences between sub-band channels, or one of their combinations. An inter-channel phase difference (ICPD) is a mean phase difference between a pair of signals. An interchannel level difference (ICLD) is the same as an interaural level difference (ILD), that is, a level difference between input signals from the left ear and the right ear, but more generally defined between any pair of signals. , eg, a pair of speaker signals, a pair of input signals to the ear, etc. An interchannel coherence or interchannel correlation is the same as an interaural coherence (IC), that is, the degree of similarity between the input signals from the left ear and the right ear, but more generally defined between any pair of signals, eg , speaker signal pair, ear input signal pair, etc. An intercanal temporal difference (ICTD) is the same as an interaural temporal difference (ITD), sometimes also referred to as an interaural delay, that is, a temporal difference between the input signals from the left and right ears but more generally defined between any pair of signals, eg, pairs of speaker signals, pair of input signals to the ears, etc. Level differences between sub-band channels, phase differences between sub-band channels, coherence between sub-band channels, and intensity differences between sub-band channels are related to the previously specified parameters with respect to to the sub-band bandwidth.
The parameter generator 101 determines the following audio channel signal value phase differences X¿k] to obtain the first set of IPD encoding parameters [b]. In one embodiment, the audio channel signal X¿b] and the reference audio signal X2 [b] are frequency domain signals and the audio channel signal values X¿k] and the values X2 reference audio signals [k] are associated with frequency bands indicated as [k] or frequency sub-bands, indicated as [b]. In one embodiment, the parametric audio encoder 100 comprises a transformer, eg, an FFT device for transforming a plurality of audio channel signals of the time domain X¿n], X2 [n] in the frequency domain in order to obtain the plurality of audio channel signals X¿b], X2 [b]. In one embodiment, the parameter generator 101 determines the first set of IPD coding parameters [b] for each frequency container [k] or for each frequency sub-band [b] of the audio channel signals X ¿B], X2 [b].
In a first stage, the parameter generator 105 applies a time-frequency transform on the input channel of the time domain, eg, the first input channel x¿n] and the reference channel of the time domain, eg, the second input channel X2 [n]. In the case of stereo playback, these are the left and right channels. In a preferred embodiment, the time-frequency transform is a Fast Fourier Transform (FFT). In an alternative embodiment, the frequency-time transform is a cosine modulation filter bank or a complete filter bank.
In a second stage, the parameter generator 105 calculates a cross spectrum for each frequency container [b] of the fFt as:
ES 2 555 136 T3 c [b] = X ^ bJX ^ b] where c [b] is the crossed spectrum of the frequency container [b] and X¿b] and X2 [b] are the FFT coefficients of the two channels. The asterisk * indicates a complete conjugation. For this case, a sub-band [b] corresponds directly to a frequency container [k], while the frequency container [b] and [k] represent exactly the same frequency container.
alternatively, the parameter generator 105 calculates the subband crossover spectrum [b] as:
c [b] = ΣΐΧ'ΧιΜΧίΜ.
where c [b] is the crossover spectrum of sub-band [b] and Xi [k] and X2 [k] are the FFT coefficients of the two channels. The asterisk * indicates a complete conjugation, kb is the initial band or sub-band b and kb + i is the initial band of the adjacent sub-band b + 1. Therefore, the FFT frequency bands [k] between kb and kb + i-1 represent the subbands [bj.
For phase differences between channels (IPDs) they are calculated by sub-band on the basis of the crossed spectrum as:
JPD (b] = ¿c [b] where operation A is the argument operator to calculate the angle of c [bj.
In one embodiment, the parameter generator 101 determines the first mean of IPD encoding parameters<sub>I</sub>an [i] of the audio channel signal Xi [b] as an average of the first set of IPD encoding parameters [b] of the audio channel signal Xi [b] through the frequency bins [b] or frequency subbands [bj.
The averaged IPD (IPD<sub>I</sub>an) through the frequency bands [b] or the frequency sub-bands [b] is calculated as defined in the following equation:
IPD mean ~ ££<sub>=1</sub>lPD [k] where K is the number of frequency bands or frequency sub-bands that are taken into account for the calculation of the mean.
In one embodiment, the parameter generator 101 determines the second average of IPD encoding parameters<sub>m</sub>ean_iong_term of the audio channel signal Xi [b] as an average of a plurality of first average IPD encoding parameters<sub>I</sub>an [i] through a plurality of frames of the audio channel signal Xi [b], where each first average of IPD encoding parameters<sub>I</sub>an [i] is associated with a frame [i] of the multichannel audio signal.
Based on the mean IPD<sub>I</sub>Still previously calculated, the parameter generator 105 calculates a long-term average of IPD. The mean IPD<sub>m</sub>ean_iong_term is calculated as the average over the last N frames (for example N can be set to 10).
<img file="ES2555136T3_D0001.tif" />
nwanjong term
In one embodiment, the parameter generator 101 determines an absolute value IPDd¡<sub>s</sub>t of a difference between the second mean of IPD encoding parameters<sub>m</sub>ean_iong_term and the first mean of IPD encoding parameters<sub>I</sub>an [i]
In order to evaluate the stability of the IPD parameter, the distance between IPDmean and IPD<sub>I</sub>an_iong_term (IPDdist) is calculated in this regard, indicating the evolution of IPD during the last N frames. In a preferred embodiment, the distance between the local and long-term IPD values is calculated as the absolute value of the difference between the local and long-term mean.
ES 2 555 136 T3
IPDdjst - abs (IPD<sub>mean</sub> IPDmean_long_term)
It can be deduced that if the parameter of IPDmean is stable through the previous frames, the distance IPDd¡<sub>s</sub>t becomes close to 0. The distance is then equal to zero when the phase difference is stable over time. This distance provides a good estimate of channel similarity.
In one embodiment, the parameter generator 101 determines the ICC encoding parameter as a function of the determined absolute value IPDdist. In one embodiment, the parameter generator 101 determines the ICC encoding parameter from a difference between a first parametric value d and the determined absolute value IPDd¡<sub>s</sub>t multiplied by a second parametric value e. In one embodiment, the parameter generator 101 sets the first parametric value to one and sets the second parametric value to one.
The coherence or ICC parameter is calculated as ICC = 1-IPDdist, since ICC and IPDd¡<sub>s</sub>They have an indirect inverse relationship. The ICC value is close to 1 when the channels are similar and IPDd¡<sub>s</sub>t is set equal to zero in that case.
Alternatively, the equation to define the relationship between ICC and IPDd¡<sub>s</sub>t is defined as ICC = d - e. IPDd¡<sub>s</sub>tWith the values of d and e being chosen to best represent the inverse relationship between the two parameters. In another embodiment, the relationship between ICC and IPDd¡<sub>s</sub>t is obtained from a large database and then generalized as ICC = f (IPD<sub>d</sub>St).
During the correlated segment of the audio signal (for example, for the voice signal), the value of IPDdistes small, and during the diffuse parts of the audio input (for example, for the music signal), this IPDd parameter<sub>s</sub>t becomes much larger and will have a value close to 1 if the input channels are not correlated. Thus, ICC and IPDd¡<sub>s</sub>They have an indirect inverse relationship.
Figure 2 illustrates a block diagram of a parametric audio decoder 200 in accordance with one embodiment. Parametric audio decoder 200 receives a bit stream 203 transmitted over a communication channel as an input signal and provides a decoded multi-channel audio signal 201 as an output signal. Parametric audio decoder 200 comprises a bit stream decoder 217 coupled to bit stream 203 to decode bit stream 203 into an encoding parameter 215 and an encoded signal 213, a decoder 209 coupled to bit stream decoder 217 to generate a sum signal 211 from the coded signal 213, a parametric decoder 205 coupled to the bit stream decoder 217 to decode a parameter 221 from the encoding parameter 215 and a synthesizer 205 coupled to the parametric decoder 205 and the decoder 209 to synthesize the decoded multichannel audio signal 201 from the parameter 221 and the signal adds 211.
The parametric audio decoder 200 generates the output channels of its multichannel audio signal 201 such as ICTD, ICLD and / or ICC between the channels in proximity to those of the original multichannel audio signal. The described system is capable of representing multichannel audio signals at a bit rate only slightly higher than that required to represent a monaural audio signal. This is because the ICTD, ICLD, and ICC values estimated between a pair of channels contain approximately two orders of magnitude less information than an audio waveform. Not only the low binary rate but also the backward compatibility aspect is of interest. The sum signal transmitted corresponds to a reduced monaural mix of the multichannel or stereo signal.
Figure 3 illustrates a block diagram of a parametric stereo audio encoder 301 and a decoder 303 in accordance with one embodiment. The parametric stereo audio decoder 301 corresponds to the parametric audio decoder 100 as described with respect to Figure 1 but the multi-channel audio signal 101 is a stereo audio signal with left 305 and right 307 audio channels.
The parametric stereo audio encoder 301 receives the stereo audio signal 305, 307 comprising a left channel audio signal 305 and an audio channel audio signal 307, as an input signal and provides a bit stream as a signal signal. exit 309. The parametric stereo audio encoder 301 comprises a parameter generator 311 coupled to the stereo audio signal 305, 307 to generate spatial parameters 313, a reduced mixed signal generator 315 coupled to the stereo audio signal 305, 307 to generate a signal reduced mixed 317 or sum signal, a monaural encoder 319 coupled to the reduced mixed signal generator 315 to encode the reduced mixed signal 317 to provide an encoded audio signal 321 and a bitstream combiner 323 coupled to the parameter generator 311 and the monaural encoder 319 to combine the parameter encoding 313 and the encoded audio signal 321 for a bit stream to provide the output signal 309. In parameter generator 311, spatial parameters 313 are extracted and quantized before they are multiplexed in the bit stream.
The parametric stereo audio decoder 303 receives the bit stream, that is, the output signal 309 of the parametric stereo audio encoder 301 transmitted through a communications channel, as a signal
ES 2 555 136 T3 input and provides a decoded stereo audio signal with left channel 325 and right channel 327 as the output signal. The parametric stereo audio decoder 303 comprises a bitstream decoder 329 coupled to the received bitstream 309 to decode the bitstream 309 into encoding parameters 331 and an encoded signal 333, a monaural decoder 335 coupled to the streamer decoder bits 329 to generate a sum signal 337 from the encoded signal 333, a spatial parametric decoder 339 coupled to the bit stream decoder 329 to decode spatial parameters 341 from the encoding parameters 331 and a synthesizer 343 coupled to the spatial parametric decoder or resolution system 339 and the monaural decoder 335 to synthesize the signal from decoded stereo audio 325, 327 from spatial parameters 341 and sum signal 337.
The processing in the parametric stereo audio encoder 301 is capable of extracting delays and calculating the level of the audio signals adaptively in time and frequency to generate the spatial parameters 313, eg, temporal differences between channels (ICTDs) and differences inter-channel levels (ICLDs). In addition, the parametric stereo audio encoder 301 performs time adaptive filtering efficiently for inter-channel coherence synthesis (ICC). In one embodiment, the parametric stereo encoder uses a Short Time Fourier Transform (STFT) which serves as the basis for a filter bank for efficient implementation of low complexity binaural track coding (BCC) systems. calculation. The processing in the parametric stereo audio encoder 301 has low computational complexity and low delay. Making parametric stereo audio coding suitable for affordable implementation in microprocessors or digital signal processors for real-time applications.
The parameter generator 311 illustrated in Figure 3 is functionally the same as the corresponding parameter generator 105 described with respect to Figure 1, except that the quantization and encoding of the spatial tracks has been added for illustrative purposes. The sum signal 317 is encoded with a conventional monaural audio encoder 319. In one embodiment, the parametric stereo audio encoder 301 uses a time-frequency transform based on the STFT transform to transform the stereo audio channel signal 305, 307 in the frequency domain. The STFT applies a Discrete Fourier Transform (DFT) to the time window portions of an input signal x (n). A signal frame of N samples is multiplied with a window of length W before a DFT of N points is applied. Adjacent windows are overlapped and shifted by W / 2 samples. The window is chosen so that overlapping windows are added up to a constant value of 1. Therefore, for the inverse transformation, there is no need to set additional windows. A simple inverse DFT of size N with W / 2 sample successive frame timing is used in decoder 303. If the spectrum is not modified, the perfect reconstruction is achieved by overlapping / addition.
Since the uniform spectral resolution of the STFT is not well adapted to human perception, the uniformly spaced spectral coefficients, the output object of the STFT, are grouped into B non-overlapping partitions with bandwidths better adapted for perception. A partition conceptually corresponds to a "sub-band" in accordance with the description with respect to Figure 1. In an alternative embodiment, the parametric stereo audio encoder 301 uses a non-uniform filter bank to transform the stereo audio channel signal 305, 307 in the frequency domain.
In one embodiment, the mixer-reducer 315 determines the spectral coefficients of a partition bo of a sub-band b of the equalized sum signal S<sub>m</sub>(k) 317 through
S<sub>m</sub>{k) - <sup>1</sup> o = l where X<sub>c</sub>, m (k) are the spectra of the input audio channels 305, 307 and eb (k) is a gain factor calculated as
<img file="ES2555136T3_D0002.tif" />
<img file="ES2555136T3_D0003.tif" />
with estimates of partition power,
ES 2 555 136 T3
TO<sub>b</sub> —1 = Σ IAWI<sup>2</sup> .
m. = Afa_i
TO<sub>b</sub>-1 C = Σ I52x<sub>c</sub>,<sub>m</sub>(fc) |<sup>2</sup>.
m = A<sub>b</sub>_j c — 1
To prevent the presence of aids resulting from large gain factors when the attenuation of the sum of the sub-band signals is important, the gain factors eb (k) can be limited to 6 dB, that is, eb (k) <2.
In one embodiment, the parameter generator 311 applies a time-frequency transformation, eg, the STFT as described above or an FFT on the input channels, that is, on the left channel 305 and on the right channel 307. In one embodiment, the time-frequency transformation is a Fast Fourier Transform (FFT). In an alternative embodiment, the time-frequency transformation is performed on a cosine modulation filter bank or on a complex filter bank.
The parameter generator 311 calculates a crossover spectrum for each frequency container [b] of the FFT or STFT as c [b] = XxMxpb]
For this case, a sub-band [b] corresponds directly to a single frequency container [k], representing a frequency container [b] and [k] exactly the same frequency container.
Alternatively, the parameter generator 311 calculates the crossover spectrum for the subband [k] as kb + i-1 c [b] = £ <sup>x</sup>iMXj [k] k = k<sub>b</sub> where c [b] is the crossover spectrum from band b to sub-band k. Xi [k] and X2 [k] are the FFT coefficients of the left channel 305 and of the right channel 307. The asterisk operator * indicates a complex conjugation, kb is the initial band of sub-band k and kb + i is the initial band of the adjacent b + 1 sub-band. Consequently, the frequency bands [k] of an FFT or STFT between kb and k<sub>b</sub>+ i-1 represent the subbands [bj.
The interchannel phase differences (IPDs) are calculated per sub-band based on the crossover spectrum as:
IPD [b] = zc [¿] where the operation Z is the argument operator for calculating the angle of c [bj.
Next, the parameter generator 311 calculates the average value of IPD (IPD<sub>I</sub>an) across the frequency bands or frequency sub-bands as defined in the following equation.
XÍ<sub>=1</sub>IPD [k] mean where K is the number of frequency bands or frequency sub-bands that are taken into account for the calculation of the mean.
Then based on the IPD value<sub>I</sub>Still previously calculated, the parameter generator 311 calculates a long-term average of the IPD. The IPD<sub>m</sub>ean_iong_term is calculated as the average over the last N frames, in one embodiment, with N being set to a value of 10.
ES 2 555 136 T3 and<sup>, v</sup> ipd μ _ / mean LJ
TPD ------- mean_Iong_term
In order to evaluate the stability of the IPD parameter, the parameter generator 311 calculates the distance IPDd!<sub>s</sub>t between IPD<sub>I</sub>an and IPD<sub>m</sub>ean_iong_term that illustrates the evolution of the IPD during the last N frames. In one embodiment, the distance between the local and long-term IPD is calculated as the absolute value of the difference between the local mean and the long-term mean:
IPDdist ~ <sup>to</sup>b<sup>s</sup>(lPDmea<sub>n</sub> "LPD<sub>mean</sub> Jong_term)
It can be deduced that if the parameter of IPD<sub>I</sub>an is stable through the previous frames, the value of the IPDdist distance becomes close to 0. The distance is then equal to zero when the phase difference is stable in the course of time. This distance provides a good estimate of channel similarity.
In one embodiment, the parameter generator 311 calculates the coherence or ICC parameter as ICC = 1IPDdist since ICC and IPDdist have an indirect inverse relationship. The ICC value is close to 1 when the channels are similar and the IPDdist value is equal to 0 in that case.
Alternatively, the parameter generator 311 uses the relationship between ICC and IPDdist defined as ICC = de. IPDdist with d and e being parameters chosen to better represent the inverse relationship between the two parameters ICC and IPDdist In an alternative embodiment, the parameter generator 311 obtains the relationship between ICC and IPDdist through a large database that is generalized as ICC = f (IPDd¡st).
During a correlated segment of an audio signal, for example, for voice signal, the value of IPDdist is small and during diffuse parts of the audio output, for example, for music signal, this parameter IPDdist becomes much larger and will be close to 1 if the input channels are not in correlation. Thus, ICC and IPDdist have an indirect inverse relationship.
Parameter generator 311 uses IPDdist for a rough estimate of the ICC value. The cross spectrum requires a lower complexity than the correlation calculation. Furthermore, in case of calculation of the IPD in the parametric spatial audio encoder, this cross spectrum is already calculated and the total complexity is consequently deduced.
Figure 4 illustrates a schematic diagram of a method 400 for generating an encoding parameter in accordance with one embodiment. The method 400 for generating the ICC coding parameter for an audio channel signal xi [n] from a plurality of audio channel signals xi [n], X2 [n] of a multi-channel audio signal. Each audio channel signal xi [n], X2 [n] has audio channel signal values. Figure 4 illustrates the stereo case where the plurality of audio channel signals comprise a left audio channel xi [n] and a right audio channel X2 [nj. Method 400 comprises:
apply an FFT 401 transform to the left audio channel signal xi [n] and apply an FFT 403 transform to the right audio channel signal X2 [n] to obtain audio channel signals in the frequency domain Xi [b] and X2 [b], where Xi [b] is the left audio channel signal and X2 [b] is the right audio channel signal with respect to the frequency container [b] in the the frequency, alternatively, a filter bank transform is applied to the left audio channel signal xi [n] and to the right audio channel signal X2 [n] to obtain audio channel signals Xi [b], X2 [b] in sub - frequency bands, where [b] indicates the frequency sub-band;
determine 405 a cross-correlation c [b] of each frequency container [b] of the left audio channel signal Xi [b] and the right audio channel signal X2 [b] or alternatively, determine 405 a correlation crossover c [b] of each frequency sub-band [b] of the left audio channel signal Xi [b] and of the right audio channel signal X2 [bj;
determining 407 for the audio channel signal Xi [b] of the plurality of audio channel signals a first set of encoding parameters IPD [b] of the audio channel signal values of the audio channel signal Xi [b] and the reference audio signal values of a reference audio signal X2 [b], wherein the reference audio signal is another audio channel signal X2 [b] of the plurality of audio channel signals or a reduced mixed audio signal derived from at least two audio channel signals of the plurality of signals multichannel audio. Figure 4 illustrates the stereo case, where the determining operation 407 determines for the left audio channel signal Xi [b] the first set of IPD coding parameters [b] and where the reference audio signal is the right audio channel signal X2 [bj;
determine 409 for the audio channel signal Xi [b] a first mean of IPD encoding parameters<sub>I</sub>an [i]
ES 2 555 136 T3 based on the first set of IPD coding parameters [b] of the audio channel signal X¿b];
determine 411 for the audio channel signal Xi [b] a second mean of IPD encoding parameters<sub>m</sub>ean_iong_term based on the first mean of IPD encoding parameters<sub>m</sub>ean [¡] of the audio channel signal Xi [b] and at least one other mean of IPD encoding parameters<sub>m</sub>ean [¡-1] of the audio channel signal Xi [b], the other first half of IPD encoding parameters<sub>m</sub>ean [¡-1] is calculated from the previous N-1 frames of the audio channel signal X¿b]; and determine 413 or calculate the ICC encoding parameter based on the first mean of IPD encoding parameters<sub>m</sub>ean [¡] of the audio channel signal X¿b] and the second mean of IPD encoding parameters<sub>m</sub>ean_iong_term of the audio channel signal Xi [b],
In one embodiment, the first set of IPD coding parameters [b] of the audio channel signal Xi [b] is already available and the method 400 starts with steps 409, 411 and 413 as described above. .
Although not illustrated in Figure 4, method 400 is applicable to the general case of multichannel audio signals, the reference signal then being another audio channel signal or a reduced mixed audio signal as described above with respect to the Figure 1.
In one embodiment, method 400 is processed as follows:
in a first step 401,403, a time-frequency transformation is applied to the input channels (left channel and right channel in the case of stereo). In a preferred embodiment, the time-frequency transformation is performed with a fast Fourier transform (FFT). In an alternative embodiment, the time-frequency transformation can be performed with a cosine modulation filter bank or a complex filter bank.
In a second stage 405, a crossover spectrum for each frequency container of the FFT is calculated by c [b] = Xi [b] X<sub>2</sub>[b] where a sub-band [b] corresponds directly to a single frequency container [k], with the frequency container [b] and [k] representing exactly the same frequency container.
Alternatively, the crossover spectrum can be calculated per sub-band as kp + i-1 c [b] = £ X.MXIM k = k<sub>b</sub> where c [b] is the crossover spectrum of band b or sub-band b. X¿k] and X2 [k] are the FFT coefficients of the two channels (for example left and right channels in case of stereo). The asterisk * indicates a complete conjugation, kb is the initial band of sub-band b and kb + i is the initial band of the adjacent sub-band b + 1. Therefore, the frequency bands [k] of the FFT between kb and kb + i-1 represent the sub-bands [b].
In a third stage 407, the phase differences between channels (IPDs) are calculated per sub-band, based on the crossed spectrum as! PD [b] - ¿c [¿] where the operation Z. is the operator of argument to calculate the angle of c [b].
In a fourth stage 409, the averaged IPD (IPDmean) across the frequency bands (or frequency sub-bands) is also calculated as defined by the following equation:
_ ΣΕ-, ΙΡΡΜ " <sup>v</sup>mean where K is the number of frequency bands or frequency sub-bands that are taken into account for the calculation of the mean.
ES 2 555 136 T3
In a fifth step 411, based on the previously calculated IPDmean value, a long-term average of IPD is determined. The IPD<sub>m</sub>ean_iong_term is calculated as the average over the last N frames (for example, N can be set to 10).
, ρη mcanlongjcrm
In order to evaluate the stability of the IPD parameter, the distance between IPD<sub>I</sub>an and IPD<sub>m</sub>ean_iong_term (IPDdist) is object of calculation, which shows the evolution of IPD during the last N frames. In a preferred embodiment, the distance between the local and long-term IPD is calculated as the absolute value of the difference between the local mean and the long-term mean:
lPD¿j<sub>st</sub> = abs (IPD<sub>raean</sub> IPDmean Jong_term)
It can be deduced that if the parameter of IPDmean is stable through the previous frames, the distance IPDd¡<sub>s</sub>t becomes close to 0. The distance is then equal to zero when the phase difference is stable over time. This distance provides a good estimate of channel similarity.
In a sixth step 413, the ICC parameter or coherence is calculated by ICC = 1-IPDdist, since ICC and IPDdist have an indirect inverse relationship. The ICC value is close to 1 when the channels are similar and IPDdist is equal to zero in that case.
In an alternative embodiment of the sixth step 413, the equation for defining the relationship between ICC and IPDdist is defined as ICC = de. IPDdist with the parameters d and e being chosen to better represent the inverse relationship between the two parameters ICC and IPDdist In another embodiment of the sixth step 413, the relationship between ICC and IPDdist is obtained from a large database and can then be generated as ICC = f (IPDdist).
During a correlated segment of an audio signal (for example, for voice signal), the value of IPDdist is small and during diffuse parts of the audio input (for example, for music signal), this IPDdist parameter becomes much larger and its value will be close to 1 if the input channels are not in correlation. Thus, ICC and IPDdist have an indirect inverse relationship.
From the foregoing one skilled in this art will deduce that a variety of methods, systems, computer programs are disclosed on record carriers and the like.
The present invention also supports a computer program product that includes computer-executable code or computer-executable instructions that, when executed, cause at least one computer to perform the performing and calculating steps described herein.
The present invention also supports a system configured to carry out the implementation and calculation steps described here.
Numerous alternatives, modifications, and variants will be apparent to those skilled in this art in light of the above teachings. Of course, those skilled in the art readily recognize that there are numerous applications of the invention beyond those described herein. Although the present invention has been described with reference to one or more particular embodiments, those skilled in the art recognize that numerous changes can be made without departing from the scope of protection of the present invention. Therefore, it is to be understood that within the scope of the appended claims and their equivalents, the invention may be carried out in a manner other than that specifically described herein.
A corresponding embodiment of the present invention can be applied in the stereo extension encoder of the ITU-T G.722, G.722 Annex B, G.711.1 and / or G.711.1 Annex D standards. In addition, the method described can also be applied for a voice and audio encoder for mobile application as defined in the 3GPP EVS (Enhanced Voice Service) codec.
Contents15
1 priority claim, no other members on record
Priority claims1
| Document | Office | Kind | Date |
|---|---|---|---|
| 2012052734 | European Patent Office (EPO) | W |
Numbers
- Publication
- 2555136
- Application
- 12707055
Titles2
- Spanish
- Codificador paramétrico para codificar una señal de audio multicanal
- English
- Parametric encoder to encode a multichannel audio signal
Classification
- CPC, 4
- G10L19/008
- H04S3/008
- H04S2400/03
- H04S2420/03
- IPC, 2
- H04S3 00
- G10L19 00