Decoding apparatus with decorrelator unit
Abstract
Procedure for encoding an audio signal, the procedure comprising: - generating (S8) a monaural signal comprising a combination of at least two channels (L, R) of input audio, - determining (S2, S3, S4) a set of spatial parameters (ILD, ITD, C) indicative of spatial properties of the at least two input audio channels, including the set of spatial parameters a parameter (C) representing a measure of waveform similarity of the at least two input audio channels, - generate (S5, S6, S7, S9) an encoded signal comprising the monaural signal and the set of spatial parameters characterized in that the similarity measure corresponds to a value of a cross correlation function to a maximum value of said cross correlation function.

Term
Term ended
Projected expiry passed 22 April 2023, 3.4 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
14 claims: 5 independent, 9 dependent
- 1ES 2 300 567 T3 REIVINDICACIONES 1. Procedimiento para codificar una señal de audio, comprendiendo el procedimiento:- generar (S8) una señal monoaural que comprende una combinación de al menos dos canales (L, R) de audio de entrada, - determinar (S2, S3, S4) un conjunto de parámetros (ILD, ITD, C) espaciales indicativos de propiedades espaciales de los al menos dos canales de audio de entrada, incluyendo el conjunto de parámetros espaciales un parámetro (C) que representa una medida de similitud de formas de onda de los al menos dos canales de audio de entrada, - generar (S5, S6, S7, S9) una señal codificada que comprende la señal monoaural y el conjunto de parámetros espaciales caracterizado porque la medida de similitud corresponde a un valor de una función de correlación cruzada a un valor máximo de dicha función de correlación cruzada.
- 2Procedimiento según la reivindicación 1, en el que la etapa de determinar un conjunto de parámetros espaciales indicativos de propiedades espaciales comprende determinar un conjunto de parámetros espaciales en función del tiempo y la frecuencia.
- 3Procedimiento según la reivindicación 2, en el que la etapa de determinar un conjunto de parámetros espaciales indicativos de propiedades espaciales comprende - dividir cada uno de los al menos dos canales de audio de entrada en pluralidades correspondientes de bandas de frecuencia;- para cada una de la pluralidad de bandas de frecuencia determinar el conjunto de parámetros espaciales indicativos de propiedades espaciales de los al menos dos canales de audio de entrada en la banda de frecuencia correspondiente.
- 4Procedimiento según una cualquiera de las reivindicaciones 1 a 3, en el que el conjunto de parámetros espaciales incluye al menos una indicación de posición.
- 5Procedimiento según la reivindicación 4, en el que el conjunto de parámetros espaciales incluye al menos dos indicaciones de posición que comprenden una diferencia de nivel entre canales y una seleccionada de entre una diferencia de tiempo entre canales y una diferencia de fase entre canales.
- 6Procedimiento según la reivindicación 4 ó 5, en el que la medida de similitud comprende información que no puede tenerse en cuenta por las indicaciones de posición.
- 7Procedimiento según una cualquiera de las reivindicaciones 1 a 6, en el que la etapa de generar una señal codificada que comprende la señal monoaural y el conjunto de parámetros espaciales comprende generar un conjunto de parámetros espaciales cuantificados, introduciendo cada uno un error de cuantificación correspondiente relativo al parámetro espacial determinado correspondiente, en el que al menos uno de los errores de cuantificación introducidos se controla para que dependa de un valor de al menos uno de los parámetros espaciales determinados.
- 8Codificador para codificar una señal de audio, comprendiendo el codificador:- medios para generar una señal monoaural que comprende una combinación de al menos dos canales de audio de entrada, - medios para determinar un conjunto de parámetros espaciales indicativos de propiedades espaciales de los al menos dos canales de audio de entrada, incluyendo el conjunto de parámetros espaciales un parámetro que representa una medida de similitud de formas de onda de los al menos dos canales de audio de entrada, y - medios para generar una señal codificada que comprende la señal monoaural y el conjunto de parámetros espaciales, caracterizado porque la medida de similitud corresponde a un valor de una función de correlación cruzada a un valor máximo de dicha función de correlación cruzada.
- 9Aparato para suministrar una señal de audio, comprendiendo el aparato:una entrada para recibir una señal de audio, un codificador según la reivindicación 8 para codificar la señal de audio para obtener una señal de audio codificada, y ES 2 300 567 T3 una salida para suministrar la señal de audio codificada.
- 10Señal de audio codificada, comprendiendo la señal:una señal monoaural que comprende una combinación de al menos dos canales de audio, y un conjunto de parámetros espaciales indicativos de propiedades espaciales de los al menos dos canales de audio de entrada, incluyendo el conjunto de parámetros espaciales un parámetro que representa una medida de similitud de formas de onda de los al menos dos canales de audio de entrada, caracterizado porque la medida de similitud corresponde a un valor de una función de correlación cruzada a un valor máximo de dicha función de correlación cruzada.
- 11Medio de almacenamiento que tiene almacenada en el mismo una señal codificada según la reivindicación 10.
- 12Procedimiento para descodificar una señal de audio codificada, comprendiendo el procedimiento:obtener una señal monoaural a partir de la señal de audio codificada, comprendiendo la señal monoaural una combinación de al menos dos canales de audio, obtener un conjunto de parámetros espaciales a partir de la señal de audio codificada, incluyendo el conjunto de parámetros espaciales un parámetro que representa una medida de similitud de formas de onda de los al menos dos canales de audio, y generar una señal de salida multicanal a partir de la señal monoaural y los parámetros espaciales, caracterizado porque la medida de similitud corresponde a un valor de una función de correlación cruzada a un valor máximo de dicha función de correlación cruzada.
- 13Descodificador para descodificar una señal de audio codificada, comprendiendo el descodificador medios para obtener una señal monoaural a partir de la señal de audio codificada, comprendiendo la señal monoaural una combinación de al menos dos canales de audio, y medios para obtener un conjunto de parámetros espaciales a partir de la señal de audio codificada, incluyendo el conjunto de parámetros espaciales un parámetro que representa una medida de similitud de formas de onda de los al menos dos canales de audio, y medios para generar una señal de salida multicanal a partir de la señal monoaural y los parámetros espaciales, caracterizado porque la medida de similitud corresponde a un valor de una función de correlación cruzada a un valor máximo de dicha función de correlación cruzada.
- 14Aparato para suministrar una señal de audio descodificada, comprendiendo el aparato:una entrada para recibir una señal de audio codificada, un descodificador según la reivindicación 13 para descodificar la señal de audio codificada para obtener una señal de salida multicanal, y una salida para suministrar o reproducir la señal de salida multicanal.
Independent claims14
166 paragraphs in 10 sections, as filed
IS 2 300 567 T3
DESCRIPTION
Parametric representation of spatial audio.
This invention relates to the coding of audio signals and, more particularly, to the coding of multi-channel audio signals.
Within the field of audio coding, it is generally desired to encode an audio signal, for example in order to reduce the bit rate to communicate the signal or the storage requirement to store the signal, without compromising the quality of the audio too much. perception of the audio signal. This is an important issue when the audio signals are to be transmitted via communication channels of limited capacity or when they are to be stored on a storage medium having limited capacity.
Previous solutions in audio encoders that have been suggested to reduce the bitrate of stereo program material include:
“Intensity stereo”. In this algorithm, high frequencies (typically greater than 5 kHz) are represented by a single audio signal (ie, mono), combined with time-varying, frequency-dependent scale factors.
“Stereo M / S” (M / S stereo). In this algorithm, the signal is decomposed into a sum (or central (mid), or common) signal and a difference (or lateral (side), or not common) signal. This decomposition is sometimes combined with time-varying scale factors or principal component analysis. These signals are then independently encoded, either by a transform encoder or waveform encoder. The amount of information reduction achieved by this algorithm is highly dependent on the spatial properties of the original signal. For example, if the original signal is monaural, the difference signal is zero and can be discarded. However, if the correlation of the left and right audio signals is low (which is often the case), this scheme offers only a small advantage.
Parametric descriptions of audio signals have gained interest in recent years, especially in the field of audio coding. It has been shown that transmitting (quantized) parameters that describe audio signals requires only a small transmission capacity to re-synthesize a signal of equal perception at the receiving end. However, today's parametric audio encoders focus on encoding monaural signals, and stereo signals are often processed as dual monos.
European patent application 1 107 232 discloses a method for encoding a stereo signal having an L and R component, in which the stereo signal is represented by one of the following: level and phase differences in the acquisition of parametric information and stereo components of the audio signal. In the decoder, the other stereo component is recovered based on the encoded stereo component and the parametric information. The article "Efficient representation of spatial audio using perceptual parametrization" (Faller C et al, Proceedings of the 2001 IEEE Workshop on the Applications of Signal Processing to Audio and Acoustics) discloses the generation of a binaural signal spatially locating the sources contained in a monophonic sum signal, basing the situation on a set of spatial parameters in critical bands. The article "Subband coding of stereophonic digital audio signals" (Van der Waal RG et al, IEEEICASSP 1991) discloses the use of left-right correlation in a subband codec.
It is an object of the present invention to solve the problem of providing an improved audio coding that achieves a high quality of perception of the recovered signal.
The above and other problems are solved by a method for encoding an audio signal as set forth in claim 1.
The inventor has realized that by encoding a multichannel audio signal as a monaural audio signal and a number of spatial attributes comprising a similarity measure of the corresponding waveforms, the multichannel signal can be recovered with a high perceptual quality. . Another advantage of the invention is the fact that it provides efficient coding of a multichannel signal, that is to say a signal comprising at least a first and a second channel, for example a stereo signal, a quadraphonic signal, etc.
Thus, according to one aspect of the invention, spatial attributes of multi-channel audio signals are parameterized. For general audio coding applications, the transmission of these parameters combined with just a monaural audio signal greatly reduces the transmission capacity required to transmit the stereo signal compared to audio encoders that process channels independently, while maintains the original spatial impression. An important point is that although people receive waveforms from an auditory object twice (once through the left ear and once through the right ear), only a single hearing object is perceived in a certain position and with a certain size. (or spatial ability to spread).
Therefore, it seems unnecessary to describe audio signals as two or more (independent) waveforms, and it would be better to describe multichannel audio as a set of auditory objects, each with its own spatial properties. A
The difficulty that immediately arises is the fact that it is almost impossible to automatically separate individual auditory objects from a given set of auditory objects, for example a musical recording. This problem can be overcome by not dividing the program material into individual auditory objects, but rather by describing the spatial parameters in a way that resembles the efficient (peripheral) processing of the auditory system. When the spatial attributes comprise a measure of (di) similarity of the corresponding waveforms, efficient coding is achieved while maintaining a high level of perception quality.
In particular, the parametric description of multichannel audio presented herein refers to the binaural processing model presented by Breebaart et al. This model is intended to describe the efficient signal processing of the binaural hearing system. For a description of the binaural processing model by Breebaart et al., See Breebaart, J., van de Par, and Kohlrausch, A. (2001a). Binaural processing model based on contralateral inhibition. I. Model setup. J. Acoust. Soc. Am., 110, 1074-1088; Breebaart, J., van de Par, S. and Kohlrausch, A. (2001b). Binaural processing model based on contralateral inhibition. II. Dependence on spectral parameters. J. Acoust. Soc. Am., 110, 1089-1104; and Breebaart, J., van de Par, S. and Kohlrausch, A. (2001c). Binaural processing model based on contralateral inhibition. III. Dependence on temporal parameters .. J. Acoust. Soc. Am., 110, 1105-1117. A brief interpretation is provided below to aid in understanding the invention.
In a preferred embodiment, the set of spatial parameters includes at least one position indication. When the spatial attributes comprise one or more, preferably two, position indications as well as a measure of (di) similarity of the corresponding waveforms, particularly efficient coding is achieved while maintaining a particularly high level of perception quality.
The term position indication encompasses any suitable parameter that conveys information about the position of auditory objects that contribute to the audio signal, for example the orientation of and / or distance from an auditory object.
In a preferred embodiment of the invention, the set of spatial parameters includes at least two position indications comprising an interchannel level difference (ILD) and one selected from an interchannel time difference (ITD). time difference) and an interchannel phase difference (IPD). It is interesting to mention that the difference in level between channels and the time difference between channels are considered as the most important position indications in the horizontal plane.
The similarity measure of the waveforms corresponding to the first and second audio channels corresponds to a value of a cross-correlation function to a maximum value of said cross-correlation function (also known as coherence). The maximum cross-channel correlation is strongly related to the perceptual spatial diffusion capacity (or compactness) of a sound source, that is, it provides additional information that is not taken into account by the position indications above, thus providing a set of parameters with a low degree of redundancy of the information transmitted by them and, therefore, providing efficient coding.
According to a preferred embodiment of the invention, the step of determining a set of spatial parameters indicative of spatial properties comprises determining a set of spatial parameters as a function of time and frequency.
It is an idea of the inventors that it is sufficient to describe spatial attributes of any multichannel audio signal by specifying the ILD, ITD (or IPD) and the maximum correlation as a function of time and frequency.
In another preferred embodiment of the invention, the step of determining a set of spatial parameters indicative of spatial properties comprises
- dividing each of the at least two input audio channels into corresponding pluralities of frequency bands;
- for each of the plurality of frequency bands determining the set of spatial parameters indicative of spatial properties of the at least two input audio channels in the corresponding frequency band.
Thus, the incoming audio signal is divided into several band-limited signals, which (preferably) are linearly spaced on an ERB rate scale. Preferably, the analysis filters show partial overlap in the frequency and / or time domain. The bandwidth of these signals depends on the center frequency, following the ERB rate. Subsequently, preferably for each frequency band, the following properties of the incoming signals are analyzed:
- the level difference between channels, or ILD, defined by the relative levels of the band-limited signal from the left and right signals,
- the time difference (or phase) between channels (ITD or IPD), defined by the delay (or phase shift) between channels corresponding to the position of the peak in the cross-correlation function between channels, and
IS 2 300 567 T3
- the (di) similarity of the waveforms that cannot be taken into account by the ITD or ILD, which can be parameterized by the maximum cross-correlation between channels (that is, the value of the normalized cross-correlation function at the position of the maximum peak, also known as coherence).
The three parameters described above vary over time; however, since the binaural hearing system is very slow to process, the update rate for these properties is quite low (typically tens of milliseconds).
In this case it can be assumed that the (slowly) time-varying properties mentioned above are the only spatial signal properties available to the binaural hearing system, and that from these time and frequency-dependent parameters, the Perceived auditory environment is reconstructed by higher levels of the auditory system.
An important issue in parameter transmission is the precision of the parameter representation (ie the size of the quantization errors), which is directly related to the required transmission capacity.
According to yet another preferred embodiment of the invention, the step of generating a coded signal comprising the monaural signal and the set of spatial parameters comprises generating a set of quantized spatial parameters, each one introducing a corresponding quantization error relative to the corresponding determined spatial parameter , wherein at least one of the introduced quantization errors is controlled to depend on a value of at least one of the determined spatial parameters.
Therefore, the quantization error introduced by the quantization of the parameters is controlled according to the sensitivity of the human auditory system to changes in these parameters. This sensitivity is highly dependent on the values of the parameters themselves. Thus, by controlling the quantization error to depend on the parameter values, improved encoding is achieved.
It is an advantage of the invention that it provides a decoupling of monaural and binaural signal parameters in audio encoders. Thus, difficulties related to stereo audio encoders (such as audibility of interaurally uncorrelated quantization noise compared to interaurally correlated quantization noise, or interaural phase inconsistencies in parametric encoders that encode in dual mono mode).
It is another advantage of the invention that a considerable reduction in the bit rate is achieved in audio encoders due to a low refresh rate and low frequency resolution, required for spatial parameters. The associated bit rate for encoding the spatial parameters is typically 10 kbits / s or less (see the embodiment described below).
Another advantage of the invention is the fact that it can be easily combined with existing audio encoders. The proposed scheme produces a mono signal that can be encoded and decoded with any existing encoding strategy. After monaural decoding, the system described herein regenerates a stereo multichannel signal with the appropriate spatial attributes.
The set of spatial parameters can be used as an enhancement layer in audio encoders. For example, a mono signal is transmitted if only a low bit rate is allowed, while including the spatial enhancement layer the decoder can reproduce stereo sound.
It is stated that the invention is not limited to stereo signals but can be applied to any multi-channel signal comprising n channels (n> 1). In particular, the invention can be used to generate n channels from a mono signal, if they are transmitted (n-1) sets of spatial parameters. In this case, the spatial parameters describe how to form the n different audio channels from the single mono signal.
It is noted that the features of the procedure described above and below can be implemented in software and carried out in a data processing system or other processing means by executing computer-executable instructions. The instructions can be program code means loaded into a memory, such as RAM, from a storage medium or from another computer over a computer network. Alternatively, the features described can be implemented using hardwired circuitry instead of software or in combination with software.
The invention further relates to an encoder for encoding an audio signal as set forth in claim 8.
It is indicated that the above means for generating a monaural signal, the means for determining a set of spatial parameters as well as the means for generating a coded signal can be implemented by any suitable device or circuit, for example as general-purpose or special programmable microprocessors. , digital signal processors (DSP), application-specific integrated circuits (ASIC), logic arrays
ES 2 300 567 T3 programmable (PLA), field programmable gate arrays (FPGA), special purpose electronic circuits, etc. or a combination thereof.
The invention further relates to an apparatus for supplying an audio signal, the apparatus comprising:
- an input to receive an audio signal,
- an encoder as described above and below for encoding the audio signal to obtain an encoded audio signal, and
- an output to supply the encoded audio signal.
The apparatus can be any electronic equipment or part of such equipment, such as fixed or portable computers, portable or fixed radio communication equipment, or other portable or handheld devices, such as multimedia players, recording devices, etc. The term "portable radio communication equipment" includes all equipment such as mobile phones, personal pagers, communicators, ie electronic organizers, smart phones, personal digital assistants (PDAs), pocket computers, or the like.
The input may comprise any device or circuitry suitable for receiving a multi-channel audio signal in digital or analog format, for example via a wired connection, such as a jack line, via a wireless connection, for example a radio signal, or in any other suitable way.
Similarly, the output may comprise any suitable device or circuitry to supply the encoded signal. Examples of such outputs include a network interface to provide the signal to a computer network, such as a LAN, the Internet, or the like, a communications circuitry to communicate the signal through a communications channel, for example a wireless communication channel, etc. In other embodiments, the output may comprise a device for storing a signal on a storage medium.
The invention further relates to an encoded audio signal, as set forth in claim 10.
The invention further relates to a storage medium having such a coded signal stored therein. The term "storage medium" herein comprises, but is not limited to, a magnetic tape, an optical disk, a digital video disk (DVD), a compact disk (CD or CD-ROM), a mini-disk, a hard disk, floppy disk, ferroelectric memory, erasable electrically programmable read-only memory (EEPROM), flash memory, EPROM memory, read-only memory (ROM), static random access memory (SRAM) , a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a ferromagnetic memory, optical storage, charge coupled devices, smart cards, a PCMCIA card, etc.
The invention further relates to a method for decoding an encoded audio signal as set forth in claim 12.
The invention further relates to a decoder for decoding an encoded audio signal as set forth in claim 13.
It is indicated that the above means can be implemented by any suitable device or circuit, such as for example programmable microprocessors of general or special use, digital signal processors (DSP), integrated circuits for specific applications (ASIC), programmable logic arrays (PLA ), field programmable gate arrays (FPGA), special purpose electronic circuits, etc. or a combination thereof.
The invention further relates to an apparatus for supplying a decoded audio signal, the apparatus comprising:
- an input to receive an encoded audio signal,
- a decoder as described above and below to decode the encoded audio signal to obtain a multi-channel output signal,
- an output to supply or reproduce the multichannel output signal.
The apparatus can be any electronic equipment or part of such equipment, as described above.
The input may comprise any device or circuitry suitable for receiving an encoded audio signal. Examples of such inputs include a network interface to receive the signal through a computer network, such as a LAN, the Internet, or the like, a set of communications circuits to receive the signal through a communications channel, for example a wireless communication channel, etc. In other embodiments, the input may comprise a device for reading a signal from a storage medium.
IS 2 300 567 T3
Similarly, the output can comprise any device or circuitry suitable for supplying a multichannel signal in an analog or digital format.
These and other aspects of the invention will become apparent and will be clarified from the embodiments described below with reference to the drawings in which:
Figure 1 shows a flow chart of a method for encoding an audio signal according to an embodiment of the invention;
Figure 2 shows a schematic block diagram of a coding system according to an embodiment of the invention;
Figure 3 illustrates a filtering procedure for use to synthesize the audio signal;
and Figure 4 illustrates a de-correlator for use to synthesize the audio signal.
Figure 1 shows a flow chart of a method for encoding an audio signal according to an embodiment of the invention.
In an initial step S1, the incoming L and R signals are divided into bandpass signals (preferably with a bandwidth that increases with frequency), indicated with the reference number 101, so that their parameters can be analyzed as a function of time. . One possible procedure for time / frequency division is to use the application of a time window function followed by a transform operation, although time-continuous procedures (eg filter banks) could also be used. The time and frequency resolution of this process is preferably matched to the signal; for transient signals a precise time resolution (of the order of a few milliseconds) and an approximate frequency resolution are preferred, while for non-transient signals a more precise frequency resolution and a more approximate time resolution (of the order of tenths milliseconds). Subsequently, in step S2, the level difference (ILD) of corresponding subband signals is determined; in step S3 the time difference (ITD or IPD) of corresponding subband signals is determined; and in step S4 the magnitude of similarity or dissimilarity of the waveforms that cannot be taken into account by the ILD or ITD is described. The analysis of these parameters is explained below.
Stage S2
ILD analysis
The ILD is determined by the difference in level of the signals in a certain time instance for a given frequency band. One method of determining ILD is to measure the root mean square (rms) value of the corresponding frequency band of both input channels and calculate the ratio of these rms values (preferably expressed in dB).
Stage S3
ITD analysis
ITDs are determined by the time or phase alignment that provides the best match between the waveforms of both channels. One procedure to obtain the ITD is to calculate the cross-correlation function between two corresponding subband signals and find the maximum value. The delay that corresponds to this maximum value in the cross-correlation function can be used as the ITD value. A second procedure is to calculate the left and right subband analytical signals (ie, calculate the envelope and phase values) and use the phase difference (mean) between the channels as the IPD parameter.
Stage S4
Correlation analysis
The correlation is obtained by first finding the ILD and ITD that provides the best match between the corresponding subband signals and then measuring the similarity of the waveforms after compensation for the ITD and / or ILD. Therefore, in this context, correlation is defined as the similarity or dissimilarity of corresponding subband signals that cannot be attributed to ILD and / or ITD. A suitable measure for this parameter is the maximum value of the cross-correlation function (that is, the maximum value over a set of delays).
However, not according to the invention, other measurements could also be used, such as the relative energy of the difference signal after ILD and / or ITD compensation compared to the corresponding subband sum signal (preferably also compensated with respect to ILD and / or ITD). This difference parameter is basically a linear transformation of the (maximum) correlation.
IS 2 300 567 T3
In subsequent steps S5, S6 and S7, the determined parameters are quantized. An important issue for parameter transmission is the precision of the parameter representation (ie the size of the quantization errors), which is directly related to the required transmission capacity. In this section, various questions regarding the quantification of spatial parameters will be addressed. The basic idea is to base quantization errors on so-called just-noticeable differences (JND) of spatial identifications. To be more specific, the quantization error is determined by the sensitivity of the human auditory system to changes in parameters. Because the sensitivity to changes in the parameters is highly dependent on the values of the parameters themselves, the following procedures are applied to determine the discrete quantization steps.
Stage S5
Quantification of ILD
From psychoacoustic research it is known that sensitivity to changes in ILD depends on the ILD itself. If the ILD is expressed in dB, deviations of approximately 1 dB from a 0 dB reference can be detected, while changes of the order of 3 dB are required if the difference from the reference level amounts to 20 dB. Therefore, quantization errors can be larger if the left and right channel signals have a larger level difference. For example, this can be applied by first measuring the level difference between the channels, followed by a non-linear (compressive) transformation of the obtained level difference and then a linear quantization process, or by using a look-up table of the values. of available ILDs that have a non-linear distribution. The subsequent embodiment provides an example of such a look-up table.
Stage S6
Quantification of ITD
Sensitivity to changes in ITDs in human subjects can be characterized by having a constant phase threshold. This means that, in terms of delay times, the quantization steps for ITD should decrease with frequency. Alternatively, if ITD is represented in the form of phase differences, the quantization steps should be independent of frequency. One method of implementing this is to take a fixed phase difference as the quantization step and determine the corresponding time delay for each frequency band. This ITD value is then used as a quantization step. Another method is to transmit phase differences that follow a frequency independent quantization scheme. It is also known that, above a certain frequency, the human auditory system is not sensitive to ITDs in fine-structure waveforms. This phenomenon can be exploited by transmitting only ITD parameters up to a certain frequency (usually 2 kHz).
A third bit stream reduction method is to incorporate ITD quantization steps that depend on the correlation and / or ILD parameters of the same subband. For large ILDs, ITDs can be coded less precisely. Furthermore, if the correlation is very low, human sensitivity to changes in ITD is known to be low. Therefore, if the correlation is small, larger ITD quantization errors can be applied. An extreme example of this idea is not transmitting ITD if the correlation is below a certain threshold and / or if the ILD is large enough for the same subband (typically about 20 dB).
Stage S7
Quantification of correlation
The correlation quantization error depends on (1) the correlation value itself and possibly (2) on the ILD. Correlation values close to +1 are encoded with a high precision (that is, a small quantization step), while correlation values close to 0 are encoded with a low precision (a large quantization step). In the embodiment, an example of a set of correlation values distributed non-linearly is given. A second possibility is to use quantization steps for the correlation that depend on the measured ILD of the same subband: for large ILDs (i.e. one channel is energy-dominant), the quantization errors in the correlation become more large. An extreme example of this principle would be to transmit no correlation value for a certain subband if the absolute value of the ILD for that subband is beyond a certain threshold.
In step S8, a monaural S signal is generated from the incoming audio signals, for example as a sum signal of the incoming signal components, determining a dominant signal, generating a main component signal from the components incoming signal, or similar. This process preferably uses the extracted spatial parameters to generate the mono signal, ie, first aligning the subband waveforms using the ITD or IPD before combining.
Finally, in step S9, a coded signal 102 is generated from the monaural signal and the determined parameters. Alternatively, the sum signal and spatial parameters can be communicated as separate signals through the same channel or different channels.
IS 2 300 567 T3
It is indicated that the above procedure can be implemented by a corresponding arrangement, for example implemented as general or special-purpose programmable microprocessors, digital signal processors (DSP), application-specific integrated circuits (ASIC), programmable logic arrays (PLA), field programmable gate arrays (FPGA), special purpose electronic circuits, etc. or a combination thereof.
Figure 2 shows a schematic block diagram of a coding system according to an embodiment of the invention. The system comprises an encoder 201 and a corresponding decoder 202. Decoder 201 receives a stereo signal with two components L and R and generates a coded signal 203 comprising a sum signal S and spatial parameters P which are communicated to decoder 202. Signal 203 can be communicated through any communication channel 204 . Alternatively or additionally, the signal can be stored on removable storage medium 214, for example a memory card, which can be transferred from the encoder to the decoder.
Encoder 201 comprises analysis modules 205 and 206 for analyzing spatial parameters of incoming L and R signals, preferably for each time / frequency slot. The encoder further comprises a parameter extraction module 207 that generates quantized spatial parameters; and a combination module 208 that generates a sum (or dominant) signal consisting of a certain combination of the at least two input signals. The encoder further comprises an encoding module 209 that generates a resulting encoded signal 203 comprising the monaural signal and spatial parameters. In one embodiment, module 209 further performs one or more of the following functions: bit rate allocation, frame synchronization, lossless encoding, and so on.
Synthesis (in decoder 202) is performed by applying the spatial parameters to the summation signal to generate left and right output signals. Therefore, the decoder 202 comprises a decoding module 210 that performs the inverse operation of the module 209 and extracts the sum signal S and the parameters P from the encoded signal 203. The decoder further comprises a synthesis module 211 that recovers the stereo L and R components from the sum (or dominant) signal and the spatial parameters.
In this embodiment, the description of the spatial parameters is combined with a monaural (single channel) audio encoder to encode a stereo audio signal. It should be noted that although the described embodiment works on stereo signals, the general idea can be applied to n-channel audio signals, with n> 1.
In analysis modules 205 and 206, the left and right incoming L and R signals, respectively, are divided into several time frames (for example, each comprising 2048 samples at a sample rate of 44.1 kHz) and are applies a window function to them with a square root Hanning window. Subsequently, the FFTs are calculated. Negative FFT frequencies are discarded and the resulting FFTs are subdivided into groups (subbands) of FFT bins. The number of FFT intervals that are combined into a subband g depends on the frequency: at higher frequencies more intervals are combined than at lower frequencies. In one embodiment, FFT ranges corresponding to approximately 1.8 ERB (Equivalent Rectangular Bandwidth) are grouped together, resulting in 20 subbands to represent the entire audible frequency range. The resulting number of FFT intervals S [g] of each subsequent subband (starting at the lowest frequency) is
S = [4 4 4 5 6 8 9 12 13 17 21 25 30 38 45 55 68 82 100 477]
Therefore, the first three subbands contain 4 FFT intervals, the fourth subband contains 5 FFT intervals, and so on. For each subband, the corresponding ILD, ITD and correlation (r) are calculated. ITD and correlation are calculated simply by zeroing all FFT intervals belonging to other groups, multiplying the resulting FFTs (band limited) from the left and right channels, followed by an inverse FFT transform. The resulting cross-correlation function is scanned for a peak within an inter-channel delay between -64 and +63 samples. The internal delay corresponding to the peak is used as the ITD value, and the value of the cross-correlation function at this peak is used as the correlation between channels of this subband. Finally, the ILD is calculated simply by taking the power ratio of the left and right channels for each subband.
In the combining module 208, the left and right subbands are added after a phase correction (time offset). This phase correction is derived from the ITD calculated for that subband and consists of delaying the left channel subband with ITD / 2 and the right channel subband with -ITD / 2. The delay is realized in the frequency domain by an appropriate modification of the phase angles of each FFT interval. Subsequently, the sum signal is calculated by adding the phase-shifted versions of the left and right subband signals. Finally, to compensate for the uncorrelated or correlated addition, each subband of the sum signal is multiplied by sqrt (2 / (1 + r)), where r is the correlation of the corresponding subband. If necessary, the sum signal can be converted to the time domain (1) by inserting complex conjugates at negative frequencies, (2) inverse FFT, (3) window function mapping, and (4) overlap-add (overlap and sum) .
In the parameter extraction module 207, the spatial parameters are quantized. ILDs (in dB) are quantized to the nearest value of the following I set:
I = [-19 -16 -13 -10 -8 -6 -4 -2 0 2 4 6 8 10 13 16 19]
IS 2 300 567 T3
The ITD quantization steps are determined by a constant phase difference in each subband of 0.1 rad. Therefore, for each subband, the time difference corresponding to 0.1 rad of the subband center frequency is used as the quantization step. For frequencies above 2 kHz, ITD information is not transmitted.
The inter-channel correlation r values are quantized to the closest value of the following set R:
R = [1 0.95 0.9 0.82 0.75 0.6 0.3 0]
This will cost another 3 bits for each correlation value.
If the absolute value of the ILD (quantized) of the current subband is 19 dB, no correlation or ITD values are transmitted for this subband. If the correlation (quantized) value of a certain subband amounts to zero, no ITD value is transmitted for that subband.
Thus, each frame requires a maximum of 233 bits to transmit the spatial parameters. With a frame length of 1024 frames, the maximum bit rate for transmission is 10.25 kbit / s. It should be noted that by using entropy coding or differential coding, this bit rate can be further reduced.
The decoder comprises a synthesis module 211 in which the stereo signal is synthesized from the received sum signal and the spatial parameters. Thus, for this description it is assumed that the synthesis module receives a frequency domain representation of the summation signal as described above. This representation can be obtained by window function and FFT operations of the waveform in the time domain. First, the sum signal is copied to the left and right output signals. Subsequently, the correlation between the left and right signals is modified with a de-correlator. In a preferred embodiment, a de-correlator is used as described below. Subsequently, each subband of the left signal is delayed by -ITD / 2, and the right signal is delayed by ITD / 2, given the ITD (quantized) corresponding to that subband. Finally the left and right subbands are scaled according to the ILD for that subband. In one embodiment, the above modification is done by a filter as described below. To convert the output signals in the time domain, the following steps are performed: (1) insert complex conjugates at negative frequencies, (2) inverse FFT, (3) window function mapping, and (4) overlapadd.
Figure 3 illustrates a filtering procedure for use to synthesize the audio signal. In an initial step 301, the incoming audio signal x (t) is segmented into a number of frames. Segmentation step 301 divides the signal into frames x<sub>n</sub>(t) of a suitable length, for example in the range of 500-5000 samples, for example 1024 or 2048 samples.
Preferably, the segmentation is performed using overlap analysis and synthesis window functions, thus suppressing artifacts that can get into the frame boundaries (see for example Princen, JP, and Bradley, AB: "Analysis / synthesis filterbank design based on time domain aliasing cancellation ”, IEEE transactions on Acoustics, Speech and Signal processing, Vol. ASSP 34, 1986).
In step 302, each of the xn (t) frames is transformed to the frequency domain by applying a Fourier transform, preferably implemented as a fast Fourier transform (FFT). The frequency representation resulting from the nth frame xn (t) comprises a number of frequency components X (k, n), where the parameter n indicates the number of frames and the parameter k indicates the corresponding frequency component or frequency interval at a frequency 0 <k <K. In general, the components X (k, n) in the frequency domain are complex numbers.
In step 303, the desired filter for the current frame is determined according to the received time-varying spatial parameters. The desired filter is expressed as a desired filter response comprising a set of K factors F (k, n), complex weights, 0 <k <K, for the nth frame. The filter response F (k, n) can be represented by two real numbers, that is, its amplitude a (k, n) and its phase ^ (k, n) according to F (k, n) = a (k, n) -exp | jp (k, n) |.
In the frequency domain, the filtered frequency components are Y (k, n) = F (k, n) -X (k, n), that is, they result from a multiplication of the X (k, n) components of frequency of the input signal with the filter response F (k, n). As will be apparent to one skilled in the art, this frequency domain multiplication corresponds to a convolution of the input signal frame xn (t) with a corresponding filter fn (t).
In step 304, the desired filter response F (k, n) is modified before applying it to the current frame X (k, n). In particular, the actual filter response F '(k, n) to be applied is determined as a function of the desired filter response F (k, n) and information 308 about previous frames. Preferably, this information comprises the desired and / or actual filter response of one or more previous frames, depending on
IS 2 300 567 T3
F '(k, n) = a' (k, n) expQ φ'Οςη)] = O [F (W> F (k ^ -1), F (k¿i-2), ... F * (k, n-2), ...].
Thus, by making the actual filter response dependent on the history of previous filter responses, artifacts introduced by changes in the filter response between consecutive frames can be effectively suppressed. Preferably, the actual form of the transform function Φ is selected to reduce overlap-add artifacts that result from dynamically varying filter responses.
For example, the transform function Φ can be a function of a single previous response function, for example F '(k, n) = Φι [F (k, n), F (k, n-1)] or F '(k, n) = Φ<sub>2</sub> [F (k, n), F '(k, n-1)]. In another embodiment, the transform function may comprise a floating average over a number of previous response functions, for example a filtered version of previous response functions, or the like. Preferred embodiments of the transform function Φ will be described in more detail below.
In step 305, the actual filter response F '(k, n) is applied to the current frame by multiplying the frequency components X (k, n) of the current frame of the input signal by the factors F' (k , n) corresponding filter response according to Y (k, n) = F '(k, n) -X (k, n).
In step 306, the resulting processed frequency components Y (k, n) are transformed back into the time domain resulting in frames and<sub>n</sub>(t) filtered. Preferably, the inverse transform is implemented as an inverse fast Fourier transform (IFFT).
Finally, in step 307, the filtered frames are recombined to obtain a signal y (t) filtered by an overlap-add procedure. An effective implementation of an overlap-add procedure of this type is described in "Digital baseband transmission and recording", Kluwer, 1996 by Bergmans JWM
In one embodiment, the transform function Φ of step 304 is implemented as a phase shift limiter between the current and previous frames. According to this embodiment, the phase change o (k) of each frequency component F (k, n) is calculated compared to the real phase modification y '(k, n-1) applied to the previous sample of the component. corresponding frequency, that is or (k) = y (k, n) - y '(k, n-1).
Subsequently, the desired filter phase component F (k, n) is modified so that the phase shift across the frames is reduced, should the shift result in overlap-add artifacts. According to this embodiment, this is achieved by ensuring that the actual phase difference does not exceed a predetermined threshold c, for example by simply cutting off the phase difference, depending on whether | 5 (fc) | <c (1)
F (k, η -1) ^ *** ««, yes no <sup>v</sup> '
The threshold value c can be a predetermined constant, for example between π / 8 and π / 3 rad. In one embodiment, the threshold c may not be a constant but for example as a function of time, frequency, and / or the like. Furthermore, alternatively to the above strict limit for phase change, other phase change limiting functions can be used.
In general, in the above embodiment, the desired phase shift along subsequent time frames for individual frequency components is transformed by an input-output function P (or (k)) and the response F '(k , n) of real filter is given by
F '(M) = F' (M-1) exp [j P (8 (k))].
(2)
Thus, according to this embodiment, a phase shift transform function P is introduced over subsequent time frames.
In another embodiment of the filter response transformation, the phase limiting procedure is driven by a suitable measure of tonality, for example a prediction procedure as described below. This has the advantage that phase jumps between consecutive frames which occur in noise-like signals can be excluded from the phase shift limiting method according to the invention. This is an advantage, since limiting such phase jumps in noise-like signals would make the noise-like signal sound more tonal which is often perceived as synthetic or metallic.
IS 2 300 567 T3
According to this embodiment, an error 0 (k) = ^ (k, n) - ^ (k, n-1) -m is calculated<sub>k</sub>-h of predicted phase. In this case, m<sub>k</sub> indicates the frequency corresponding to the k-th frequency component and h indicates the jump size in the samples. In this case, the term jump size refers to the difference between two adjacent window centers, that is, half the analysis length for symmetric windows. The above error is assumed below to be included in the interval [-π, + π].
Subsequently, a prediction Pk measure is calculated for the magnitude of phase predictability in the k-th frequency interval according to P<sub>k</sub> = (π - ¡0 (k) |) / ne [0,1], where | · | indicates the absolute value.
Therefore, the above Pk measure provides a value between 0 and 1 corresponding to the magnitude of phase predictability in the k-th frequency range. Yep<sub>k</sub> is close to 1, the underlying signal can be assumed to have a high degree of tonality, that is, it has a substantially sinusoidal waveform. For such a signal, phase jumps are easily perceptible, for example by the listener of an audio signal. Therefore, phase jumps should preferably be eliminated in this case. On the other hand, if the value of Pk is close to 0, it can be assumed that the underlying signal is noisy. For noisy signals, phase jumps are not easily perceived and therefore can be allowed.
Consequently, the phase limiting function is applied if Pk exceeds a predetermined threshold, i.e. Pk> A, resulting in the actual filter response F '(k, n) according to
<img file="ES2300567T3_D0001.tif" />
In this case, A is bounded by the upper and lower limits of P, which are +1 and 0, respectively. The exact value of A depends on the actual implementation. For example, A can be selected between 0.6 and 0.9.
It is understood that, alternatively, any other suitable measure may be used to estimate tonality. In yet another embodiment, the allowed phase jump c described above can be made dependent on a suitable measure of tonality, for example the measure P<sub>k</sub> above, thus allowing larger phase jumps if P<sub>k</sub> it is large and vice versa.
Figure 4 illustrates a de-correlator for use to synthesize the audio signal. The de-correlator comprises an all-pass filter 401 that receives the monaural signal x and a set of spatial parameters P that include the cross-correlation r between channels and a parameter indicative of the channel difference c. It is indicated that the parameter c is related to the difference in level between channels by ILD = k-log (c), where k is a constant, that is, ILD is proportional to the logarithm of c.
Preferably, the all-pass filter comprises a frequency-dependent delay that provides relatively less delay at high frequencies than at low frequencies. This can be achieved by substituting a fixed all-pass filter delay for an all-pass filter comprising a period of a Schroeder phase complex (see for example MR Schroeder, "Synthesis of low-peak-factor signals and binary sequences with low autocorrelation ”, IEEE Transact. Inf. Theor., 16: 85-89, 1970). The decoder further comprises an analysis circuit 402 that receives the spatial parameters from the decoder and extracts the cross-correlation r between channels and the channel difference c. Circuit 402 determines a mixing matrix M (a, 6) as will be described below. The components of the mixing matrix are fed to the transform circuit 403 which also receives the input signal x and the filtered signal H®x. Circuit 403 performs a mixing operation according to
<img file="ES2300567T3_D0002.tif" />
(3) resulting in the output L and R signals.
The correlation between the signals L and R can be expressed as an angle α between vectors representing the signal L and R, respectively, in a space defined by the signals x and H®x, according to r = cos (a). Therefore, any pair of vectors showing the correct angular distance has the specified correlation.
Therefore, a mixing matrix M that transforms the x and H®x signals into L and R signals with a predetermined correlation r can be expressed as follows:
ES 2 300 567 T3 (4) í cos (a / 2) sin (α / 2) Ί ”\ cos (-a / 2) sin (-α / 2) /
Thus, the amount of signal subjected to the all-pass filter depends on the desired correlation. Also, the energy of the all-pass signal component is the same in both output channels (albeit with a 180 ° phase shift).
It is indicated that the case in which the matrix M is given by
<img file="ES2300567T3_D0003.tif" />
that is, the case in which α = 90 ° corresponding to uncorrelated output signals (r = 0) corresponds to a Lauridsen de-correlator.
To illustrate a problem with the matrix of equation (5), a situation with extreme amplitude going towards the left channel is assumed, that is, a case in which a certain signal is present only in the left channel. The desired correlation between the outputs is also assumed to be zero. In this case, the left channel output of the fa ·,. ·. · ,,,.,. ·, Of equation (3) with the mixing margin of equation (5) gives L = V> / 2 (X<sub>+</sub> Η ® X). Therefore, the output consists of the original x signal combined with its all-pass filtered H®x version.
However, this is an undesirable situation, since the all-pass filter usually deteriorates the perception quality of the signal. Furthermore, the sum of the original signal and the filtered signal results in comb filter effects, such as the perceived coloration of the output signal. In this extreme case, the best solution would be for the left output signal to consist of the input signal. In this way the correlation of the two output signals would remain zero.
In situations with more moderate level differences, the preferred situation is that the strongest output channel contains relatively more of the original signal, and that the weakest output channel contains relatively more of the filtered signal. Therefore, in general, it is preferred to maximize the amount of the original signal present at the two outputs together, and to minimize the amount of the filtered signal.
According to this embodiment, this is achieved by introducing a different mixing matrix that includes an additional common rotation:
<img file="ES2300567T3_D0004.tif" />
In this case, β is an additional rotation and C is a scalar matrix that guarantees that the relative level difference between the output signals is equal to c, that is
<img file="ES2300567T3_D0005.tif" />
Inserting the matrix of equation (6) into equation (3) provides the output signals generated by the operation of applying a matrix according to this embodiment:
<img file="ES2300567T3_D0006.tif" />
IS 2 300 567 T3
Therefore, the output L and R signals still have an angular difference α, that is, the correlation between the L and R signals is not affected by scaling the L and R signals according to the desired level difference and the additional rotation by the angle β of both the signal L and the R.
As mentioned above, preferably, the amount of the original signal x should be maximized in the summed output of L and R. This condition can be used to determine the angle β, according to d (L + R) q ax 'which gives the condition :
- c tan (P) = -J ——- tan (a / 2).
+ c
In summary, this application describes a parametric description of the spatial attributes of multichannel audio signals, based on psychoacoustics. This parametric description allows for considerable bit rate reductions in audio encoders, since only a monaural signal has to be transmitted, combined with (quantized) parameters that describe the spatial properties of the signal. The decoder can form the original number of audio channels by applying the spatial parameters. For near CD-quality stereo audio, a bit rate associated with these spatial parameters of 10 kbit / s or less appears sufficient to reproduce the correct spatial impression at the receiving end. Additionally, this bit scale can be scaled down by reducing the spectral and / or temporal resolution of the spatial parameters and / or by processing the spatial parameters using lossless compression algorithms.
It should be noted that the above-mentioned embodiments illustrate rather than limit the invention, and that those skilled in the art will be able to devise many alternative embodiments without departing from the scope of the appended claims.
For example, the invention has been described primarily in connection with one embodiment using the two position indications ILD and ITD / IPD. In alternative embodiments, other position indications may be used. Furthermore, in one embodiment, the ILD, ITD / IPD, and the cross-correlation between channels can be determined as described above, although only the cross-correlation between channels is transmitted along with the monaural signal, thus further reducing the bandwidth / storage capacity required to transmit / store the audio signal. Alternatively, the cross-correlation between channels and one of ILD and ITD / TPD can be transmitted. In these embodiments, the signal is synthesized only from the monaural signal based on the transmitted parameters.
In the claims, any reference symbols in parentheses should not be construed as limiting the claim. The term "comprise" does not exclude the presence of elements or steps other than those listed in a claim. The term "a" or "an" preceding an element does not exclude the presence of a plurality of such elements.
The invention can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the device claim listing various means, several of these means can be realized by one and the same piece of hardware. The mere fact that certain measures are listed in mutually different claims does not indicate that a combination of these measures cannot be used to advantage.
Contents10
8 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8
83 members in 11 offices
Priority claims20
| Document | Office | Kind | Date |
|---|---|---|---|
| 02076588 | European Patent Office (EPO) | A | |
| 02076588 | European Patent Office (EPO) | A | |
| 20020076588 | European Patent Office (EPO) | – | |
| 02077863 | European Patent Office (EPO) | A | |
| 02077863 | European Patent Office (EPO) | A | |
| 20020077863 | European Patent Office (EPO) | – | |
| 02079303 | European Patent Office (EPO) | A | |
| 02079303 | European Patent Office (EPO) | A | |
| 20020079303 | European Patent Office (EPO) | – | |
| 02079817 | European Patent Office (EPO) | A | |
| 02079817 | European Patent Office (EPO) | A | |
| 20020079817 | European Patent Office (EPO) | – | |
| 02077863 | – | – | – |
| 02079303 | – | – | – |
| 02079817 | – | – | – |
| 0371523702076588 | – | – | – |
| EP20020076588 | – | – | – |
| EP20020077863 | – | – | – |
| EP20020079303 | – | – | – |
| EP20020079817 | – | – | – |
Members83
| Document | Office | Kind | |
|---|---|---|---|
| WO03090206A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO03090207A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO03090208A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU2003216682A1 | Australia | A1 | |
| AU2003216686A1 | Australia | A1 | |
| AU2003219426A1 | Australia | A1 | |
| WO2004036549A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU2003219428A1 | Australia | A1 | |
| BR0304540A | Brazil | A | |
| BR0304541A | Brazil | A | |
| BR0304542A | Brazil | A | |
| KR20040101552A | Republic of Korea | A | |
| KR20040102163A | Republic of Korea | A | |
| KR20040102164A | Republic of Korea | A | |
| EP1500082A1 | European Patent Office (EPO) | A1 | |
| EP1500083A1 | European Patent Office (EPO) | A1 | |
| EP1500084A1 | European Patent Office (EPO) | A1 | |
| KR20050049549A | Republic of Korea | A | |
| EP1554716A1 | European Patent Office (EPO) | A1 | |
| CN1647155A | China | A | |
| CN1647156A | China | A | |
| CN1647157A | China | A | |
| JP2005523479A | Japan | A | |
| JP2005523480A | Japan | A | |
| JP2005523624A | Japan | A | |
| US2005226426A1 | United States of America | A1 | |
| CN1689070A | China | A | |
| US2005254446A1 | United States of America | A1 | |
| JP2006503319A | Japan | A | |
| US2006100861A1 | United States of America | A1 | |
| EP1500083B1 | European Patent Office (EPO) | B1 | |
| AT332003T | Austria | T | |
| ATE332003T1 | Austria | T1 | |
| DE60306512D1 | Germany | D1 | |
| EP1500082B1 | European Patent Office (EPO) | B1 | |
| AT354161T | Austria | T | |
| ATE354161T1 | Austria | T1 | |
| ES2268340T3 | Spain | T3 | |
| CN1307612C | China | C | |
| DE60311794D1 | Germany | D1 | |
| CN1312660C | China | C | |
| DE60306512T2 | Germany | T2 | |
| ES2280736T3 | Spain | T3 | |
| DE60311794T2 | Germany | T2 | |
| EP1500084B1 | European Patent Office (EPO) | B1 | |
| EP1881486A1 | European Patent Office (EPO) | A1 | |
| AT385025T | Austria | T | |
| ATE385025T1 | Austria | T1 | |
| DE60318835D1 | Germany | D1 | |
| ES2300567T3This record | Spain | T3 | |
| US2008170711A1 | United States of America | A1 | |
| DE60318835T2 | Germany | T2 | |
| EP1881486B1 | European Patent Office (EPO) | B1 | |
| AT426235T | Austria | T | |
| ATE426235T1 | Austria | T1 | |
| DE60326782D1 | Germany | D1 | |
| ES2323294T3 | Spain | T3 | |
| JP2009271554A | Japan | A | |
| US2009287495A1 | United States of America | A1 | |
| JP4401173B2 | Japan | B2 | |
| KR20100039433A | Republic of Korea | A | |
| CN1647156B | China | B | |
| KR100978018B1 | Republic of Korea | B1 | |
| KR101016982B1 | Republic of Korea | B1 | |
| KR101021076B1 | Republic of Korea | B1 | |
| KR101021079B1 | Republic of Korea | B1 | |
| US7933415B2 | United States of America | B2 | |
| JP4714415B2 | Japan | B2 | |
| JP4714416B2 | Japan | B2 | |
| US2011166866A1 | United States of America | A1 | |
| JP2012161087A | Japan | A | |
| US8331572B2 | United States of America | B2 | |
| JP5101579B2 | Japan | B2 | |
| US8340302B2 | United States of America | B2 | |
| US2013094654A1 | United States of America | A1 | |
| US8498422B2 | United States of America | B2 | |
| JP5498525B2 | Japan | B2 | |
| US8798275B2 | United States of America | B2 | |
| US9137603B2 | United States of America | B2 | |
| BRPI0304541B1 | Brazil | B1 | |
| BRPI0304540B1 | Brazil | B1 | |
| BRPI0304542B1 | Brazil | B1 | |
| DE60311794C5 | Germany | C5 |
Numbers
- Publication
- 2300567
- Publication, DOCDB
- 2300567
- Publication, EPODOC
- ES2300567T
- Application
- 3715237
- Application, DOCDB
- 03715237
- Application, EPODOC
- ES20030715237T
Titles2
- Spanish
- REPRESENTACION PARAMETRICA DE AUDIO ESPACIAL.
- English
- PARAMETRIC REPRESENTATION OF SPACE AUDIO.
Classification
- CPC, 5
- G10L19/008
- H04R5/00
- H04S3/008
- H04S2420/03
- G10L19/02
- IPC, 3
- G10L19 008
- G10L19 02
- H04S3 00