Method for encoding, decoding and compression of audio-type data
Abstract
AN AUDIO SIGNAL IS CODED. THE SIGNAL IS DIVIDED FIRST IN BANDS (600) A STANDARD SIGNAL ELEMENT (608) IS SELECTED FOR EACH BAND, AND ITS QUANTIFIED MAGNITUDE IS USED FOR THE LOCATION LISTS, THE DECODED SIGNAL IS DECODERATED LATER (918, 926 ACCURACY TO QUANTIFY SIGNAL ELEMENTS WITHOUT STANDARD (624).

Term
Term ended
Projected expiry passed 13 January 2013, 13.7 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
56 claims: 2 independent, 54 dependent
- 1ES 2 155 449 T3 REIVINDICACIONES 1. Móetodo para codificar una senñal de tipo audio, caracterizado porque consiste en:someter la senñal de tipo audio a muestreo para obtener muestras de senñal discretas;realizar una primera transformacioón espectral, actuando la primera transformacióon sobre como mónimo algunas de las muestras de senñal para producir un primer conjunto de coeficientes de transformacióon;dividir como mónimo algunos de los coeficientes de transformacióon del primer conjunto de coeficientes de transformacióon en una pluralidad de bandas, teniendo como mónimo una banda una pluralidad de coeficientes de transformacioón adyacentes;seleccionar un coeficiente de transformacióon de cada una de las muóltiples bandas, perteneciendo como mónimo uno de los coeficientes de transformacioón seleccionados a una de las bandas que tiene una pluralidad de coeficientes de transformacióon adyacentes;y realizar una segunda transformacioón espectral, basóandose esta segunda transformacioón en los coeficientes de transformacioón seleccionados y produciendo un segundo conjunto de coeficientes de transformacióon.
- 2Móetodo seguón la reivindicacióon 1, caracterizado porque el primer conjunto de coeficientes de transformacioón se deriva de un bloque obtenido aplicando una ventana a las muestras de senñal.
- 3Móetodo seguón la reivindicacióon 1 oó 2,caracterizado porque adicionalmente comprende la cuantificacióon del segundo conjunto de coeficientes de transformacióon.
- 4Móetodo seguón la reivindicacióon 1 oó 2, caracterizado porque la segunda transformacioón consiste en una transformacióon que actuóa sobre los coeficientes de transformacioón seleccionados.
- 5Móetodo seguón la reivindicacióon 4, caracterizado porque la segunda transformacioón reduce la cantidad media de bits necesarios para codificar los coeficientes de transformacioón seleccionados.
- 6Móetodo seguón la reivindicacióon 4, caracterizado porque la segunda transformacioón consiste en una transformacióon que actuóa sobre las magnitudes de los coeficientes de transformacioón seleccionados.
- 7Móetodo seguón la reivindicacióon 6, caracterizado porque la segunda transformacioón reduce la cantidad media de bits necesarios para codificar las magnitudes de los coeficientes de transformacioón seleccionados.
- 8Móetodo seguón la reivindicacióon 1 oó 2, caracterizado porque adicionalmente comprende la cuantificacioón de los coeficientes de transformacioón seleccionados y porque la segunda transformacioón consiste en una transformacioón basada en los coeficientes de transformacioón seleccionados cuantificados.
- 9Móetodo seguón la reivindicacióon 8, caracterizado porque la segunda transformacioón reduce la cantidad media de bits necesarios para codificar los coeficientes de transformacioón seleccionados cuantificados.
- 10Móetodo seguón la reivindicacioón 8, caracterizado porque la cuantificacioón consiste en cuantificar las magnitudes de los coeficientes de transformacioón seleccionados y porque la segunda transformacioón consiste en una transformacioón basada en las magnitudes cuantificadas de los coeficientes de transformacioón seleccionados.
- 11Móetodo seguón la reivindicacióon 10, caracterizado porque la segunda transformacióon reduce la cantidad media de bits necesarios para codificar las magnitudes cuantificadas de los coeficientes de transformacioón seleccionados.
- 12Móetodo seguón la reivindicacioón 3, 8, 10 u 11, caracterizado porque la cuantificacióon consiste en una transformacióon no lineal.
- 13Móetodo seguón la reivindicacióon 1 oó 2,caracterizado porque adicionalmente comprende un paso en el que se determina si se realiza la segunda transformacióon.
- 14Móetodo seguón la reivindicacióon 1, 2, 8 óo 10, caracterizado porque la primera transformacioón consiste en una de las siguientes:una DCT (Discrete Cosine Transform, Transformacióon de Coseno Discreto) o una TDAC (TimeDomain Aliasing Cancellation, Cancelacióon de Solapamiento de Dominio Temporal).
- 15Móetodo seguón la reivindicacióon 1, 2, 8 óo 10, caracterizado porque la segunda transformacioón consiste en una de las siguientes:una DCT (Discrete Cosine Transform, Transformacióon de Coseno Discreto) o una DFT (Discrete Fourier Transform, Transformacióon Discreta de Fourier).
- 16Móetodo seguón la reivindicacioón 1, 2, 8 oó 10, caracterizado porque adicionalmente comprende la codificacioón de coeficientes de transformacióon como mónimo en algunas de las bandas a base de los coeficientes de transformacioón seleccionados.
- 17Móetodo seguón la reivindicacioón 16, caracterizado porque la codificacioón de coeficientes de transformacioón comprende la asignacioón de bits a los coeficientes de transformacioón.
- 18Móetodo seguón la reivindicacioón 16, caracterizado porque la codificacioón de coeficientes de transformacióon comprende la determinacioón de niveles de reconstruccióon para los coeficientes de transformacioón.
- 19Móetodo seguón la reivindicacióon 1, 2, 8 oó 10, caracterizado porque la seleccióon de un coeficiente de transformacioón de una de las bandas consiste en seleccionar un coeficiente de transformacioón que tenga la mayor amplitud de los coeficientes de transformacióon de la banda.
- 20Móetodo seguón la reivindicacioón 1, 2, 8 oó 10, caracterizado porque la seleccióon de un coeficiente de transformacioón de una de las bandas consiste en seleccionar un coeficiente de transformacióon que tenga un tamanño preseleccionado en relacióon con otros coeficientes de transformacióon de la banda.
- 21Móetodo seguón la reivindicacioón 1, 2, 8, 10, 19 óo 20, caracterizado porque adicionalmente comprende la determinacióon de una divisioón en bandas de como mónimo algunos de los coeficientes de transformacioón del primer conjunto de coeficientes de transformacióon.
- 22Móetodo seguón la reivindicacioón 21, caracterizado porque la determinacióon de una divisioón ES 2 155 449 T3 consiste en determinar una divisióon de tal modo que como mónimo una banda tenga un nuómero de coeficientes de transformacióon que sea una potencia de dos.
- 23Móetodo seguón la reivindicacioón 21, caracterizado porque la determinacioón de una divisióon consiste en determinar una divisióon de tal modo que como mónimo dos bandas incluyan una cantidad diferente de coeficientes de transformacióon.
- 24Móetodo seguón la reivindicacioón 21, caracterizado porque la divisioón es diferente para senales diferentes.
- 25Móetodo seguón la reivindicacióon 21, caracterizado porque la divisióon es diferente para bloques diferentes.
- 26Móetodo seguón la reivindicacioón 21, caracterizado porque adicionalmente comprende la codificacióon de la divisioón determinada.
- 27Móetodo seguón la reivindicacióon 21, caracterizado porque la determinacioón de una divisióon consiste en una determinacioón basada en caracterósticas locales.
- 28Móetodo seguón la reivindicacioón 27, caracterizado porque la determinacióon basada en caracterósticas locales consiste en comenzar una banda nueva cuando la magnitud de coeficientes de transformacioón adyacentes difiere de forma significativa.
- 29Móetodo de decodificacióon caracterizado porque consiste en:recibir una senal de tipo audio codificada, habiendo sido codificada la senal de tipo audio mediante los pasos consistentes en: someter la senal de tipo audio a muestreo para obtener muestras de senal discretas;realizar una primera transformacioón espectral, actuando la primera transformacióon sobre como nóniiiio algunas de las muestras de senal para producir un primer conjunto de coeficientes de transformacioón;dividir como mónimo algunos de los coeficientes de transformacioón del primer conjunto de coeficientes de transformacióon en una pluralidad de bandas, teniendo como mónimo una banda una pluralidad de coeficientes de transformacioón adyacentes;seleccionar un coeficiente de transformacióon de cada una de las muóltiples bandas, perteneciendo como mónimo uno de los coeficientes de transformacioón seleccionados a una de las bandas que tiene una pluralidad de coeficientes de transformacióon adyacentes;y realizar una segunda transformacioón espectral, basaóndose la segunda transformacióon en los coeficientes de transformacioón seleccionados y produciendo un segundo conjunto de coeficientes de transformacióon;y decodificar como mmimo parte de la senal de tipo audio codificada, de un modo apropiadamente inverso al tipo de codificacióon, consistiendo la decodificacióon en la realizacióon de una transformacioón inversa basada en el segundo conjunto de coeficientes de transformacióon.
- 30Móetodo seguón la reivindicacióon 29, caracterizado porque el primer conjunto de coeficientes de transformacióon se deriva de un bloque obtenido aplicando una ventana a las muestras de senal.
- 31Móetodo seguón la reivindicacioón 29 óo 30, caracterizado porque la senal de tipo audio codificada es del tipo codificado mediante la cuantificacióon del segundo conjunto de coeficientes de transformacioón.
- 32Móetodo seguón la reivindicacióon 29 óo 30, caracterizado porque la segunda transformacióon consiste en una transformacioón que actuóa sobre los coeficientes de transformacioón seleccionados.
- 33Móetodo seguón la reivindicacióon 32, caracterizado porque la segunda transformacioón reduce la cantidad media de bits necesarios para codificar los coeficientes de transformacioón seleccionados.
- 34Móetodo seguón la reivindicacioón 32, caracterizado porque la segunda transformacioón consiste en una transformacióon que actuóa sobre las magnitudes de los coeficientes de transformacióon seleccionados.
- 35Móetodo seguón la reivindicacióon 34, caracterizado porque la segunda transformacioón reduce la cantidad media de bits necesarios para codificar las magnitudes de los coeficientes de transformacioón seleccionados.
- 36Móetodo seguón la reivindicacioón 29 oó 30, caracterizado porque la senal de tipo audio codificada es del tipo codificado mediante cuantificacióon de los coeficientes de transformacioón seleccionados, y porque la segunda transformacioón consiste en una transformacióon basada en los coeficientes de transformacioón seleccionados cuantificados.
- 37Móetodo seguón la reivindicacioón 36, caracterizado porque la segunda transformacioón reduce la cantidad media de bits necesarios para codificar los coeficientes de transformacioón seleccionados cuantificados.
- 38Móetodo seguón la reivindicacioón 36, caracterizado porque la cuantificacioón consiste en cuantificar las magnitudes de los coeficientes de transformacioón seleccionados, y porque la segunda transformacioón consiste en una transformacioón basada en las magnitudes cuantificadas de los coeficientes de transformacioón seleccionados.
- 39Móetodo seguón la reivindicacioón 38, caracterizado porque la segunda transformacioón reduce la cantidad media de bits necesarios para codificar las magnitudes cuantificadas de los coeficientes de transformacioón seleccionados.
- 40Móetodo seguón la reivindicacióon 31, 36, 38 óo 39, caracterizado porque la cuantificacióon consiste en una transformacióon no lineal.
- 41Móetodo seguón la reivindicacióon 29 óo 30, caracterizado porque la senal de tipo audio codificada es del tipo codificado mediante un paso que determina si se realiza la segunda transformacióon.
- 42Móetodo seguón la reivindicacioón 29, 30, 36 óo 38, caracterizado porque la primera transformacioón consiste en una de las siguientes:una DCT (Discrete Cosine Transform, Transformacioón de Coseno Discreto) o una TDAC (TimeDomain Aliasing Cancellation, Cancelacióon de Solapamiento de Dominio Temporal).
- 43Móetodo seguón la reivindicacióon 29, 30, 36 óo 38, caracterizado porque la segunda transfor19 ES 2 155 449 T3 macioón consiste en una de las siguientes:una DCT (Discrete Cosine Transform, Transformacióon de Coseno Discreto) o una DFT (Discrete Fourier Transform, Transformacioón Discreta de Fourier).
- 44Móetodo seguón la reivindicacióon 29, 30, 36 oó 38, caracterizado porque la senñal de tipo audio codificada es del tipo codificado mediante codificacioón de coeficientes de transformacioón como mónimo en algunas de las bandas a base de los coeficientes de transformacióon seleccionados.
- 45Móetodo seguón la reivindicacioón 44, caracterizado porque la codificacioón de coeficientes de transformacioón comprende la asignacioón de bits a los coeficientes de transformacióon.
- 46Móetodo seguón la reivindicacióon 44, caracterizado porque la codificacioón de coeficientes de transformacioón comprende la determinacioón de niveles de reconstruccióon para los coeficientes de transformacióon.
- 47Móetodo seguón la reivindicacióon 29, 30, 36 óo 38, caracterizado porque la seleccióon de un coeficiente de transformacióon de una de las bandas consiste en seleccionar un coeficiente de transformacioón que tenga la mayor amplitud de los coeficientes de transformacióon de la banda.
- 48Móetodo seguón la reivindicacioón 29, 30, 36 óo 38, caracterizado porque la seleccióon de un coeficiente de transformacióon de una de las bandas consiste en seleccionar un coeficiente de transformacióon que tenga un tamanño preseleccionado en relacióon con otros coeficientes de transformacióon de la banda.
- 49Móetodo seguón la reivindicacioón 29, 30, 36, 38, 47 oó 48, caracterizado porque la senñal de tipo audio codificada es del tipo codificado mediante determinacioón de una divisioón en bandas de como mónimo algunos de los coeficientes de transformacióon del primer conjunto de coeficientes de transformacioón.
- 50Móetodo seguón la reivindicacioón 49, caracterizado porque la determinacióon de una divisióon consiste en determinar una divisióon de tal modo que como mónimo una banda tenga un nuómero de coeficientes de transformacióon que sea una potencia de dos.
- 51Móetodo seguón la reivindicacioón 49, caracterizado porque la determinacióon de una divisióon consiste en determinar una divisióon de tal modo que como mónimo dos bandas incluyan una cantidad diferente de coeficientes de transformacióon.
- 52Móetodo seguón la reivindicacioón 49, caracterizado porque la divisioón es diferente para senñales diferentes.
- 53Móetodo seguón la reivindicacióon 49, caracterizado porque la divisióon es diferente para bloques diferentes.
- 54Móetodo seguón la reivindicacioón 49, caracterizado porque la senñal de tipo audio codificada es del tipo codificado mediante codificacioón de la divisioón determinada.
- 55Móetodo seguón la reivindicacióon 49, caracterizado porque la determinacióon de una divisioón consiste en una determinacioón basada en caracterósticas locales.
- 56Móetodo seguón la reivindicacioón 55, caracterizado porque la determinacioón basada en caracterósticas locales consiste en comenzar una banda nueva cuando la magnitud de coeficientes de transformacióon adyacentes difiere de forma significativa. NOTA INFORMATIVA:Conforme a la reserva del art. 167.2 del Convenio de Patentes Europeas (CPE) y a la Disposición Transitoria del RD 2424/1986, de 10 de octubre, relativo a la aplicación del Convenio de Patente Europea, las patentes europeas que designen a España y solicitadas antes del 7-10-1992, no producirán ningún efecto en Espana en la medida en que confieran proteccion a productos quámicos y farmaceuticos como tales. Esta informacioán no prejuzga que la patente estáeo no incluáda en la mencionada reserva.
Independent claims56
152 paragraphs in 3 sections, as filed
IS 2 155 449 T3
DESCRIPTION
Method and apparatus for encoding, decoding and compression of audio type data. Field and background of the invention
The present invention relates generally to the field of signal processing, more specifically to data encoding and compression. Specifically, the invention relates to a method and apparatus for encoding and compressing digital data representing audio signals or signals generally having the characteristics of audio signals.
Audio signals are everywhere. They are transmitted as radio signals and as part of television signals. Other signals, such as voice signals, share pertinent characteristics with audio signals, such as the importance of spectral domain representations. For many applications, it is advantageous to store and transmit digitally encoded audio-type data rather than analog. This encoded data is stored on various types of digital media, including compact audio discs, digital audio tape, magnetic discs, computer memory, both random access (RAM) and read-only (ROM), to name a few.
It is advantageous to reduce to a mononym the amount of digital data necessary to adequately characterize an analogue audio-type signal. Reducing the amount of data to the same name leads to a reduction to the same number of the necessary amount of fossil storage media, thus reducing costs and increasing the convenience of any hardware used in conjunction with the data. The reduction to the mononym of the amount of data necessary to characterize a given time portion of an audio signal also allows faster transmission of a digital representation of the audio signal through a given communication channel. This also leads to cost savings, since, compared to uncompressed data, compressed data representing the same time portion of an audio signal can be sent more quickly, or over a communication channel with less width. band. Either of these two consequences is topically less expensive.
The principles of digital audio signal processing are well known and are described in a number of papers, including Watkinson, John, The Art of digital Audio, Focal Press, London (1988). In figure 1 an analog audio signal x (t) is shown schematically. The horizontal axis represents time. The vertical axis shows the amplitude of the signal at time t. The scale of the time axis is in milliseconds, so that in figure 1 approximately two thousandths of a second of an audio signal are represented schematically. A basic first step for memory storage or transmission of the analog audio signal in the form of a digital signal is to sample the signal in discrete signal elements that are subsequently processed.
Figure 2 schematically shows the sampling of the signal x (t). The signal x (t) was evaluated at many discrete moments in time, for example at a frequency of 48 kHz. By the term "sampling" it is meant that the amplitude of the signal x (t) is noted and recorded forty-eight thousand times per second. Therefore, for a period of one msec (1 x 10<sup>-3</sup> sec), the signal x (t) is recorded forty-eight times. The result is a time series x (n) of amplitudes, as shown in Figure 2, with spaces between the amplitudes for the portions of the digital audio signal x (t) that have not been measured. If the sampling frequency is high enough in relation to the variations of the analog signal as a function of time, the magnitudes of the sampling values will generally follow the form of the analog signal. As shown in figure 2, the sampling values follow the signal x (t) quite well.
Figure 4A schematically shows the general lines of a general digital signal processing method. The initial step for obtaining the audio signal is shown at 99 and the sampling step is indicated at 102. Once the signal has been sampled, it is topically transformed from the time domain, the domain of figures 1 and 2, to another domain that facilitates the analysis. Topically, a signal in time can be written as a sum of a series of simple harmonic functions of time, such as cosωt and sinωt, for each of the various harmonic frequencies of ω. The expression of a signal that varies over time in the form of a series of harmonic functions is generally discussed in Feynman, R., Leighton, R., and Sands,
M., The Feynman Lectures on Physics, Addison Wesley Publishing Company, Reading, Massachusetts (1963) Vol. I, §50.
There are several well-known transformation methods (sometimes called "subband" methods). Baylon, David and Lim, Jae, "Transform / Subband Analysis and Synthesis of Signals", pages 540-544, 2ssPA90, Gold Coast, Australia, 27-31 Ag. (1990). One such method is the Time-Domain Aliasing Cancellation (“TDAC”) method. Another of these transformations is known as Discrete Cosine Transform (“DCT”). The transformation is carried out by applying a transformation function to the original signal. An example of a DCT transformation is:
N-1 <sub>π</sub>
X (k) = 7 / 2x (n) cos— k (2n + 1), for 0 <k <N-1 <sub>n = o</sub><sup>2N</sup> = 0 otherwise, where k is the variable frequency and N is topically the number of samples in the window.
The transformation produces a set of amplitude coefficients of a variable different from time, topically the frequency. The coefficients can be real or complex. (When X (k) is a complex value, the present invention can be applied to the real and imaginary parts
ES 2 155 449 T3 of X (k) separately, or to the magnitude and phase parts of X (k) separately, for example. However, in this description it will be assumed that X (k) is a real value). Figure 3 schematically shows a typical plot of a part of the signal x (n) transformed into X (k). If the inverse transform operation is applied to the transformed signal X (k), the original sampling signal x (n) will be obtained.
The transformation is carried out by applying the transformation function to a time interval of the analog sampling signal x (n). The interval (known as a block, "frame") is selected by applying a window in step 104 to ax (n). There are various window application methods. Windows can be applied sequentially or more topically there is an overlap. The window must be consistent with the transformation method, in a casotopic the TDAC method. As shown in Figure 2, a window w1 (n) is applied to x (n) and includes forty-eight samples covering a duration of one msec (1 x 10<sup>-3</sup> sec). (In this case, forty-eight samples have been shown solely for illustrative purposes. In a topical application, a window includes many more than forty-eight samples.) Window w2 (n) applies to the next msec. Windows are typically overlapping, but non-overlapping windows are shown solely for illustrative purposes. The transformation of signals from one domain to another, for example from temporal to frequency, is discussed in many basic texts, including: Oppenheim, AV, and Schafer, RW, Digital Signal Processing, Englewood Cliffs, NJ, Prentice Hall (1975); Rabiner, LR, Gold, B., Theory and Application of Digital Signal Processing, Englewood Cliffs, NJ, Prentice Hall (1975).
The application of the transformation, indicated in step 106 in Figure 4A, to the window of the sampling signal x (n) leads to a set of coefficients for a series of discrete frequencies. Each coefficient of the transformed signal block represents the amplitude of a component of the transformed signal at the indicated frequency. The number of frequency components is topically the same for each block. Obviously, the amplitudes of the components of the corresponding frequencies differed from segment to segment.
As Figure 3 shows, the signal X (k) is a plurality of amplitudes at discrete frequencies. This signal is called here the "spectrum" of the original signal. According to known methods, the next step is to encode the amplitudes for each of the frequencies according to some binary code, and transmit or store the encoded amplitudes.
An important task in signal coding is to assign the fixed number of available bits to the specification of the amplitudes of the coefficients. The number of bits assigned to a coefficient, or to any other signal element, is here called the "number of bits assigned" of said coefficient or signal element. This step is shown in relation to the other steps at 107 of FIG. 4A. In general, a fixed number of bits is available for each block
N. The N number is determined by considerations such as the bandwidth of the communication channel through which the data will be transmitted, or the capacity of storage media, or the amount of error correction required. As mentioned above, each block generates the same number C of coefficients (even though the amplitude of some of the coefficients may be zero).
Therefore, a simple method to allocate the N available bits is to distribute them evenly among the C coefficients, so that each coefficient can be specified by N / C bits. (For the purposes of this description, N / C is assumed to be an integer number). Thus, considering the transformed signal X (k) as shown in figure 3, the coefficient 32, which has an amplitude of approximately one hundred, was represented by a keyword with the same number of bits (N / C) as the coefficient 34, which has a much smaller amplitude, only about ten. According to most coding methods, to specify or encode a number within a larger range requires more bits than to specify a number within a lower range, assuming both are specified with the same precision. For example, encoding integer numbers between zero and one hundred with perfect accuracy using simple binary code requires seven bits, while specifying integer numbers between zero and ten requires four bits. Therefore, if seven bits were assigned to each of the signal coefficients, three bits would be wasted in each coefficient, which could have been specified using only four bits. When there is only a limited number of bits available to allocate among many coefficients, it is better to keep more bits than to waste them. Bit waste can be reduced if the range of values is known exactly.
There are several known methods for assigning the number of bits to each coefficient. One method is described in US-A-4899384. However, all these known methods lead either to a significant wastage of bits, or to a significant sacrifice in the precision of quantization of the coefficient values. One of these methods is described in a document entitled "High-Quality Audio Transform Coding at 128 Kbits / s", Davidson, G., Fielder, L., and Antill, M., from Dolby Laboratories, Inc., ICASSP, pages 1117-1120, April 3-6, Albuquerque, New Mexico (1990), here called the “Dolby document”.
According to this method, the transformation coefficients are grouped to form bands, the widths of the bands being determined by chromic band analysis. The transformation coefficients within a band are converted to a band block floating point representation (exponent and mantissa). The exponents provide an estimate of the log spectral envelope of the examined audio block, and are transmitted as secondary information to the decoder.
The logarotmic spectral envelope is used by a di3 bit assignment routine.
ES 2 155 449 T3 namic, which derives step information for an adaptive coefficient quantizer. Each block is assigned the same number of bits, N. The dynamic bit assignment routine uses only the exponent of the peak spectral amplitude in each band to increase the quantizer resolution for relevant physioaquostic bands. Each band mantissa is quantized to a bit resolution defined by the sum of an approximate fixed bit component and a dynamically assigned fine component. The fixed bit component is typically set regardless of the particular block, but rather taking into account the type of signal and the portion of the frame in question. For example, lower frequency bands can generally receive more bits as a result of the fixed bit component. The dynamically assigned component is based on the peak exponent for the band. The estimated log spectral data is multiplexed with the adaptive mantissa bits for transmission to the decoder.
Therefore, the method makes a general analysis of the maximum amplitude of a coefficient within a band of the signal and uses this general estimation to assign the number of bits to that band. The general estimate only indicates the integral part of the power of two of the coefficient. For example, if the coefficient is seven, the general estimation determines that the maximum coefficient in the band was between 2<sup>2</sup> y2<sup>3</sup> (four and eight), or, if it is twenty-five, that is between 2<sup>4</sup> y2<sup>5 </sup>(sixteen thirty-two). The general estimation (which is inaccurate) causes two problems: the bit allocation is not exact and the allocated bits are not used effectively, since the range of values for a given coefficient is not exactly known. In the procedure described above, each coefficient in a band is specified with the same level of accuracy as other coefficients in the band. Furthermore, the information referring to the maximum amplitude coefficients in the bands is encoded in two stages: first, the exponents are encoded and transmitted as secondary information; second, the mantissa is transmitted along with the mantissas of the other coefficients.
In addition to determining how many bits to assign to each coefficient to encode the amplitude of that coefficient, a coding method must also divide the total amplitude range into a series of amplitude divisions shown in step 108 in Figure 4A, and assign a cocode at each division, in step 109. The number of bits in the code is equal to the number of bits assigned for each coefficient. The divisions are topically referred to as "quantization levels", since the actual amplitudes are quantized at the available levels or "reconstruction levels" after encoding, transmission, or storage and decoding. For example, if there are three bits available for each coefficient, 2 can be identified<sup>3</sup> or eight levels of reconstruction.
Figure 5 shows a simple scheme for assigning a three-bit keyword for each of the eight amplitude regions between 0 and 100. Keyword 000 is assigned to all coefficients whose amplitude transformed, as shown in the figure 3, is between 0 and 12.5. Therefore, all coefficients between 0 and 12.5 are quantized to the same value, topically the mean value of 6.25. Keyword 001 maps to all coefficients between 12.5 and 25.0, all of which are quantized to the value 18.75. Similarly, the keyword 100 is assigned to all coefficients between 50.0 and 62.5, which are all quantized with the value 56.25. Instead of assigning keywords of uniform length to the coefficients, with uniform levels of quantization, the method consisting of assigning keywords of variable length to encode each coefficient, and applying non-uniform levels of quantization to the encoded coefficients, is also known.
It is also useful to determine a masking level. The level of masking is related to the human perception of the aquatic signals. For a given acoustic signal, one can roughly calculate the level of signal distortion (eg quantization noise) that will not be heard or perceived because of the signal. This is useful in various applications. For example, a certain signal distortion can be tolerated without the human listener noticing. Therefore, the masking level can be used to assign the available bits to different coefficients.
The whole basic process of digitizing an audio signal and synthesizing an audio signal from the encoded digital data is shown schematically in Figure 4A, and the basic apparatus is shown schematically in Figure 4B. An audio signal, such as music, voice, traffic noise, etc., is obtained in step 99 by a known device, for example a microphone. The audio signal x (t) is sampled 102, as described above and shown in Figure 2. The sample signal x (n) is squared, 104, and transformed, 106. After the transformation (which can be a subband representation), the bits are assigned, 107, between the coefficients, and the amplitudes of the coefficients are quantized, 108, assigning each of them to a reconstruction level, and these points quantized are encoded, 109, by binary keywords. At this point, the data is transmitted, 112, through a communication channel or to a memory storage device.
The above steps, 102, 104, 106, 107, 108, 109, and 112 take place on hardware that is generally referred to as a "transmitter," as shown at step 150 in FIG. 4B. The transmitter topically includes a signal encoder (also referred to as an encoder) 156 and may include other elements that prepare the encoded signal for transmission over a channel 160. However, all of the above-mentioned steps generally take place in the encoder, which may itself include multiple components.
Finally, the data is received by a receiver 164 at the other end of the data channel 160, or it is retrieved from the me4 device.
ES 2 155 449 T3 memory. As is known, the receiver includes a decoder 166 that can reverse the encoding process of the signal encoder 156 with reasonable precision. The receiver also typically includes other elements, not shown, to reverse the effect of additional elements of the transmitter that prepare the coded signal for transmission on channel 160. Signal decoder 166 is equipped with a keyword table that correlates keywords with reconstruction levels. Data is decoded from binary code to quantized reconstruction amplitude values. An inverse transformation 116 is applied to each set of quantized amplitude values, the result of which is a signal similar to a block of x (n), that is, it is in the time domain and is formed by a discrete number of values, for each inverse transformed result. However, the signal will not be exactly equal to the corresponding block of x (n), due to the quantification in reconstruction levels and the specific representation used. Topically, the difference between the original value and the reconstruction level value cannot be recovered. An inverse transformed block stream 118 is combined and an audio signal 120 is reproduced, using known apparatus, such as a D / A converter and a speaker.
Object of the invention
Therefore, the various objects of the invention include providing a method and apparatus for encoding and decoding digital audio-type signals, which allow efficient bit allocation such that, in general, fewer bits are used to specify coefficients of magnitude smaller than to specify larger coefficients, that provides for a quantification of the amplitude of the coefficients such that the bands that include higher coefficients are divided into levels of reconstruction different from the bands that only include lower coefficients, so that both the lower and higher coefficients can be specified with greater accuracy. that if the same reconstruction levels had been used for all coefficients, that allows an exact estimation of the masking level, that allows an efficient bit allocation based on the masking level; that only locates errors in small portions of the digitized data and, with respect to that data, limits the error to a small and known range, and that reduces to a mononym the need to code coefficients redundantly, all this allowing a very efficient use of the available bits.
Brief description of the invention
According to the invention, a method of encoding and decoding audio-type signals is provided as specified in claims 1 and 29, respectively.
In a first aspect, there is a method to encode a selected aspect of a signal defined by signal elements that are discrete as a monym in one dimension, a method that includes the steps consisting of: dividing the signal into a band as a monym, having as a monym this or these bands a plurality of adjacent signal elements; identifying as a mononym in a band a signal element having a magnitude of a preselected size in relation to other signal elements of said band and designating said signal element as a "pattern" signal element for said band; and encoding the location of as a monym a pattern signal element with respect to its position in said respective band.
In a second aspect, there is a method to decode a code that represents a selected aspect of a signal that is defined by signal elements that are discrete as monym in one dimension, which has been encoded by means of a method that comprises the steps consisting of: divide the signal into a band as a monym, one of this or these bands having as a monym a plurality of adjacent signal elements; identifying as a mononym in a band a signal element having a magnitude with a preselected size in relation to other signal elements of said band and designating said signal element as a "pattern" signal element for said band; coding the location of as a monym a pattern signal element with respect to its position in said respective band; and using a function of said encoded location of said, as mononym one, pattern signal element to encode said selected aspect of said signal; said decoding method comprising the step of translating said encoded aspect of said signal based on a function of the location of said pattern signal element that is appropriately inverse with respect to said function of the location used to encode said selected aspect of said signal.
In a third aspect, there is an apparatus for encoding a selected aspect of a signal that is defined by signal elements that are discrete as a mononym in one dimension, apparatus comprising: means for dividing the signal into a band as a mononym, having as a mononym this or these bands a plurality of adjacent signal elements; means for identifying as a mononym in a band a signal element having a magnitude with a preselected size in relation to other signal elements of said band and means for designating said signal element as a "pattern" signal element for said band; means for encoding the location of a pattern signal element as a monym with respect to its position in said respective band; and means for quantifying the magnitude of said, as a monym one, pattern signal element whose location has been encoded.
In a fourth aspect there is an apparatus for decoding a cocode that represents a selected aspect of a signal that is defined by signal elements that are discrete as a mononym in one dimension, which has been encoded by means of a method that comprises the steps consisting of: dividing the signal in as a monym a band, this or these bands having as a monym a plurality of adjacent signal elements; identify as a monym in a band a signal element that has a magnitude with
ES 2 155 449 T3 a preselected size relative to other signal elements of said band and designating said signal element as a "pattern" signal element for said band; encoding the location of at least one pattern signal element with respect to its position in said respective band; and utilizing a function of said encoded location of said, as mononym one, pattern signal element to encode said selected aspect of said signal; said decoding apparatus comprising means for translating said encoded aspect of said signal based on a function of the location of said pattern signal element that is appropriately inverse with respect to said function of the location used to encode said selected aspect of said signal.
In a fifth aspect, there is a method to encode a signal element selected from a signal that is defined by signal elements that are discrete as mononymous in one dimension, a method that comprises the steps consisting of: dividing the signal into a plurality of bands, having as monym a band a plurality of adjacent signal elements; identifying in each band a signal element having the greatest magnitude of all signal elements in said band, and designating said signal element as a "pattern" signal element for said band; quantifying the magnitude of each pattern signal element to a first degree of accuracy; and assigning said selected signal element a signal element bit assignment that is a function of the quantized magnitudes of said pattern signal elements, said signal element bit assignment being chosen such that the quantization of said signal element Signal selected using said signal element bit allocation takes place up to a second degree of accuracy, which is less than the first degree of accuracy.
In a sixth aspect, there is a method to encode a signal element selected from a signal that is defined by signal elements that are discrete as mononymous in one dimension, a method that comprises the steps consisting of: dividing the signal into a plurality of bands, a band having as a mononym a plurality of adjacent signal elements, one of said bands including said selected signal element; identifying in each band a signal element having the greatest magnitude of all signal elements in said band, and designating said signal element as a "standard" signal element for said band; quantify the magnitude of each pattern signal element only once; and assigning said selected signal element a signal element bit assignment that is a function of the quantized magnitudes of said pattern signal elements.
In a seventh aspect there is a method for decoding a selected signal element that has been encoded by either of the two preferred methods mentioned above, said decoding method comprising the step consisting of translating a keyword generated by the coding method based on a function of the quantized magnitudes of said pattern signal elements that is appropriately inverse with respect to said function of the quantized magnitudes used to assign bits to that selected signal element.
In an eighth aspect there is an apparatus for encoding a signal element selected from a signal that is defined by signal elements that are discrete as mononymous in one dimension, apparatus comprising: means for dividing the signal into a plurality of bands, having as monym a band a plurality of adjacent signal elements, one of said bands including said selected signal element; means for identifying in each band a signal element having the greatest magnitude of all signal elements in said band, and designating said signal element as a "pattern" signal element for said band; means for quantifying the magnitude of each standard signal element to a first degree of accuracy; means for assigning to said selected signal element a signal element bit assignment which is a function of the quantized magnitudes of said standard signal elements, said signal element bit assignment being chosen such that the quantization of said element The signal selected using said signal element bit assignment takes place up to a second degree of accuracy, which is less than the first degree of accuracy.
In a ninth aspect there is an apparatus for decoding a keyword representing a signal element selected from a signal that has been encoded by a method mentioned above, said apparatus comprising means for translating said keyword based on a function of the quantized magnitudes of said pattern signal elements which is appropriately inverse with respect to said function of the quantized magnitudes used to assign bits to said selected signal element.
In a tenth aspect, there is a method to encode a signal element selected from a signal that is defined by signal elements that are discrete as mononymous in one dimension, a method that comprises the steps consisting of: dividing the signal into a plurality of bands a band having as a monym a plurality of adjacent signal elements; identifying in each band a signal element having the greatest magnitude of all signal elements in said band, and designating said signal element as a "pattern" signal element for said band; quantifying the magnitude of each standard signal element to a first degree of accuracy; and assigning said selected signal element a signal element bit assignment that is a function of the quantized magnitudes of said standard signal elements, said signal element bit assignment being chosen such that the quantization of said signal element Signal selected using said signal element bit allocation takes place up to a second degree of accuracy, which is less than the first degree of accuracy.
In an eleventh aspect there is a method for
ES 2 155 449 T3 encoding a signal element selected from a signal defined by signal elements that are discrete in at least one dimension, a method that comprises the steps consisting of: dividing the signal into a plurality of bands, having as a mononym a band a plurality of adjacent signal elements, one of said bands including said selected signal element; identifying in each band a signal element having the greatest magnitude of all signal elements in said band, and designating said signal element as a "pattern" signal element for said band; quantifying the magnitude of each pattern signal element only once; assigning said selected signal element a signal element bit assignment that is a function of the quantized magnitudes of said standard signal elements.
In a twelfth aspect there is a method for decoding a selected signal element that has been encoded by either of the two preferred methods mentioned above, said decoding method comprising the step consisting of translating a keyword generated by the coding method based on a function of the quantized magnitudes of said standard signal elements that is appropriately inverse with respect to said function of the quantized magnitudes used to assign bits to that selected signal element.
In a thirteenth aspect there is an apparatus for encoding a signal element selected from a signal that is defined by signal elements that are discrete in at least one dimension, apparatus comprising: means for dividing the signal into a plurality of bands, having as monym a band a plurality of adjacent signal elements, one of said bands including said selected signal element; means for identifying in each band a signal element having the greatest magnitude of all signal elements in said band, and designating said signal element as a "pattern" signal element for said band; means for quantifying the magnitude of each standard signal element to a first degree of accuracy; means for assigning to said selected signal element a signal element bit assignment which is a function of the quantized magnitudes of said standard signal elements, said signal element bit assignment being chosen such that the quantization of said element Signal selection using said signal element bit allocation takes place up to a second degree of accuracy, which is less than the first degree of accuracy.
In a fourteenth aspect there is an apparatus for decoding a keyword representing a signal element selected from a signal that has been encoded by a method mentioned above, said apparatus comprising means for translating said keyword based on a function of the quantized magnitudes of said pattern signal elements which is appropriately inverse with respect to said function of the quantized magnitudes used to assign bits to said selected signal element.
Brief description of the figures
- Figure 1 schematically shows an audio type signal.
- Figure 2 schematically shows a sampling of an audio type signal.
- Figure 3 schematically shows the spectrum of an audio-type signal transformed from the time domain to the frequency domain.
- Figure 4A schematically shows the digital processing of an audio type signal according to known methods.
- Figure 4B schematically shows the hardware elements of a known digital signal processing system.
- Figure 5 schematically shows the division of the amplitude of coefficients in reconstruction levels, and the assignment of keywords to these, according to known methods of the previous state of the technique.
- Figure 6 schematically shows the division of a spectrum of an audio type signal into frequency bands according to the previous state of the art.
- Figure 7 schematically shows the spectrum of figure 6, after the application of a scaling operation, in addition to standard coefficients designated within bands.
- Figure 7A shows schematically how the pattern coefficients are used to establish an approximate estimate of | X (k) |<sup>to</sup>.
- Figure 8 schematically shows the division of the amplitude of coefficients of different bands at different levels of reconstruction, according to the method of the invention.
- Figure 9A schematically shows an option for assigning reconstruction levels to a coefficient that can only have a positive value.
- Figure 9B schematically shows another option for assigning reconstruction levels to a coefficient that can only have a positive value.
- Figure 10A schematically shows an option for assigning reconstruction levels to a coefficient that can have a positive or a negative value.
- Figure 10B schematically shows another option for assigning reconstruction levels to a coefficient that can have a positive or negative value.
IS 2 155 449 T3
- Figure 11 shows schematically how the magnitudes of the standard coefficients can be used to assign the number of bits for a band.
- Figures 12A, 12B and 12C show schematically the steps of the method of the invention.
- Figures 13A and 13B, schematically show the components of the apparatus of the invention.
Detailed description of preferred embodiments of the invention
A first preferred embodiment of the invention consists of a method for assigning bits to individual coefficients for encoding the magnitude (ie, the absolute value of the amplitude) of these coefficients. According to the method of the invention, an audio signal x (t) is obtained as shown in step 99 in Figure 4A, and sampled with a suitable frequency, such as 48 kHz as shown in step 102, resulting in x (n). The sample signal is framed, as shown in steps 104 and 106, according to a known suitable technique, such as TDAC or DCT, using an appropriate window of a typical size, for example 512 or 1024 samples. It is to be understood that other transformation and framing techniques are also within the scope of the present invention. If no transformation is performed, the invention applies to sample signal elements instead of coefficient signal elements. In fact, the invention is profitably applied to untransformed, sampled audio-type signals. The transformation is not necessary, but only takes advantage of certain structural characteristics of the signal. Therefore, if the transformation step is skipped, it is more difficult to take advantage of the sort. The result is a spectrum of coefficient signal elements in the frequency domain, as shown in Figure 3. As used herein, the term "signal elements" generally means portions of a signal. They can be sample portions of an untransformed signal, or coefficients of a transformed signal, or the entire signal itself. The steps of the method are shown schematically in the flow chart of Figures 12A, 12B and 12C.
An important aspect of the invention method is the method by which the total number of bits N is assigned among the total number of coefficients C. According to the method of the invention, the number of bits assigned is closely correlated with the amplitude of the coefficient a encode.
The first step of the method is to divide the spectrum of transformation coefficients in X (k) into a number B of bands, such as B equal to sixteen or twenty-six. This step is indicated at 600 in FIG. 12A. It is not necessary for each band to include the same number of coefficients. In fact, it may be desirable to include more frequency coefficients in some bands, for example the higher frequency bands, than in others, such as the lower frequency bands. In such a case, it is advantageous to roughly follow the chromatic band result. In figure 6 an example of the X (k) spectrum (having X (k) real values) divided into bands is shown schematically. Other topical spectra may show a more marked difference in the number of coefficients per band, topically with relatively more coefficients in the upper bands than in the lower ones.
If the number of frequency coefficients of each band is not uniform, the configuration of the width of each band must be known or it must be communicated to the decoding elements of the apparatus of the invention. Non-uniform configuration can be set and stored in decoder accessible memory. If, however, the width of the bands varies "on the fly" depending on local characteristics, these variations must be communicated to the decoder, topically by means of an explicit message indicating the configuration.
As shown in Figure 6, the spectrum is divided into many bands b1, b2, ... bB, indicated by a small dark square between the bands. As explained below, it is useful for each band to be composed of a number of coefficients equal to a power of two. At this point it is also possible to ignore frequencies that are not interesting, for example because they are too high to be perceived by the human listener.
Although not necessary for the invention, it may be useful to analyze the spectrum coefficients in a domain in which the spectral magnitudes are compressed by a non-linear transformation, for example by raising each magnitude to a fractional power α, such as 1/2, or a logarotmic transformation. The human auditory system appears to perform some kind of amplitude compression. Furthermore, a non-linear transformation such as amplitude compression tends to lead to a more uniform distribution of the amplitudes, so that a uniform quantizer is more efficient. A non-linear transformation followed by a uniform quantification is an example of the well-known non-uniform quantification.
This nonlinear transformation step is indicated at 602 in FIG. 12A. The transformed spectrum is shown in Figure 7, which differs from Figure 6 on the vertical scale.
In each band of the exponentially scaled spectrum, the coefficient Cb1, Cb2, ... Cbb that has the greatest magnitude (ignoring the sign) is designated "standard coefficient". This step is indicated at 608 in FIG. 12A. The standard coefficients are indicated in Figure 7 by a small rectangle that encloses the upper part of the coefficient marker. (In another preferred embodiment, described below, instead of designating the coefficient with the highest value in the band as the standard coefficient, another coefficient can be designated. This other coefficient can be one that has an intermediate or median amplitude in the band, or a high magnitude, but not the maximum in the band, for example the second or third highest magnitude. The realization in which the coefficient of
ES 2 155 449 T3 maximum magnitude is designated standard coefficient is the predominant example described below and is the one described first).
The method of the invention comprises several embodiments. According to each of them, the magnitude of the standard coefficients is used to efficiently allocate bits between the coefficients, and also to establish the amount and location of the reconstruction levels. These different embodiments are described in detail below and are indicated in Figures 12A and 12B. More specific embodiments include: further dividing the X (k) spectrum into divided bands in step 612; accurately quantify the location and sign of the standard coefficients in step 614; and performing various transformations on these quantized coefficients in steps 616, 618, and 620 before transmitting data to the decoder. However, the basic method of the invention in its broadest execution does not use divided bands, therefore going from the 610 band division decision to the 614 quantification decision step. In the basic method only the magnitude of the bands is used. pattern coefficients, and therefore the method goes from the quantification decision step 614 to the transformation decision of magnitude 622. It is not necessary to transform the magnitudes at this stage and therefore the basic method goes directly to step 624, in which the magnitude of the pattern coefficients is accurately quantified at reconstruction levels.
The magnitude of each standard coefficient is quantified very accurately, in topical cases, more accurately than the magnitude of the non-standard coefficients. In some cases, this exact translation is manifested in the use of more bits to encode a standard (average) coefficient than to encode a non-standard (average) coefficient. However, as explained below with respect to a uniquely pattern transformation step performed in step 622, this may not be the case. In general, the greater accuracy of the patterns (on average) is characterized by less divergence between the original coefficient value and the quantified value, compared to the divergence between the same two values in the case of a non-standard coefficient (on average). ).
After quantization, the pattern coefficients are encoded in keywords in step 626 (Figure 12B) and transmitted to the receiver in step 628. The encoding scheme can be simple, for example applying the digital representation of the level position. of reconstruction in an ordered set of levels of reconstruction, from the lowest to the highest amplitude. Alternatively a more complicated coding scheme can be used, for example using an encrypted code. As in the case of the prior art receiver, the apparatus of the invention includes a receiver having a decoder equipped to reverse the encoding process executed by the encoding apparatus. If a simple coding technique is used, the recipient can simply reverse that technique. Alternatively, a cipher code can be provided that correlates the keywords assigned to the standard coefficients with the reconstruction levels. Since the standard coefficients are quantized very accurately, when the keywords are translated and the coefficients are reconstructed, they are very similar to the original values. (The next step 632 shown in Figure 12B is executed only if one of the transformation steps 616, 618 or 620 of Figure 12A has been carried out. The embodiments in which these steps are carried out are described above. ).
The magnitudes of the precisely quantized pattern coefficients are used to allocate bits among the remaining coefficients in the band. As in this first described embodiment, each standard coefficient is the coefficient with the highest magnitude within the band to which it belongs, it is known that all the other coefficients of the band have a magnitude less than or equal to that of the standard coefficient. Furthermore, the magnitude of the standard coefficient is also known very precisely. Therefore, it is known how many coefficients have to be encoded in the band with the largest amplitude range, the next in range, the smallest, etc. With this knowledge, bits can be efficiently allocated between bands.
Bits can be assigned in many ways. Two important general methods are to assign bits to each band and then to each coefficient within the band, or to assign bits directly to each coefficient without previously assigning bits to each band. Following an embodiment of the first general method, initially the number of bits allocated to each individual band is determined in step 634. In general, the more coefficients there are in a band, the more bits will be necessary to encode all the coefficients in that band. Similarly, with a higher mean magnitude | X (k) |<sup>α</sup> of the coefficients of the band, more bits were required to encode all the coefficients of that band. Therefore, an approximate measure of the "size" of each band is determined, defining "size" in terms of the number of coefficients and the magnitude of the coefficients, and then the available bits are assigned between the bands according to their sizes. Relatively, the larger bands get more bits and the smaller bands get fewer bits.
For example, as shown in Figure 7A, for a very rough estimate it can be assumed that the magnitude of each coefficient is the same as that of the pattern of that band. This is indicated in Figure 7A by a thickly hatched rectangle with a magnitude equal to the absolute value of the amplitude of the standard coefficient. As can be seen by comparing Figure 7 with Figure 7A, to obtain a rough estimate of the size of each band all coefficients are assumed to be positive. Knowing the number of coefficients of each band, an upper limit can be established for the size of the band. In an unofficial sense, this analysis is similar to determining the energy content of the band, compared to the total energy content of the block. Once the relative sizes have been determined, known techniques are applied to allocate the available bits between the bands according to the sizes.
ES 2 155 449 T3 estimates. One of these techniques is described in Lim, JS, Two-Dimensional Signal and Image Processing, Prentice Hall, Englewood Cliffs, New Jersey (1990), p. 598. Experience may also show that it is advantageous to allocate bits between bands assuming that the mean magnitude | X (k) |<sup>α</sup> of each non-pattern coefficient is equal to some other fraction of the magnitude of the pattern, for example half. This is shown in Figure 7A by the thinner hatched rectangles spanning the signal bands. It should be noted that the thickly hatched regions extend down to the frequency axis, although the lower portion was hidden by the less heavily hatched regions.
It is also possible to adjust the band size estimate depending on the number of coefficients (also known as frequency samples) of the band. For example, the more the coefficients, the less likely it is that the mean magnitude is equal to the magnitude of the standard coefficient. In any case, a rough estimate of the size of the band provides an appropriate allocation of bits to that band.
Within each band, the bits are allocated in step 636 between the coefficients. Topically, the bits are allocated uniformly, however any reasonable rule can be applied. It is to be noted that the magnitudes of the standard coefficients have already been quantized, encoded, and transmitted, and do not need to be quantized, encoded, or transmitted again. Following the prior state of the technique described in the Dolby document, aspects of the coefficients used to perform an approximate analysis of the maximum magnitude of a coefficient within a band are encoded in two different stages, in the first with respect to the exponent and in the second with respect to the mantissa.
As mentioned above, instead of first assigning bits between the bands and then assigning bits between the coefficients of each band, the estimation of | X (k) | can also be used.<sup>α </sup>to assign bits to coefficients directly without the intermediate step of assigning bits to bands. Again, the rough estimate | X (k) |<sup>α </sup>is used to provide a rough estimate of the magnitude of each coefficient. As illustrated in Figure 7A, the rough estimate of the magnitude of each coefficient can be the magnitude of the standard coefficient, or half of this magnitude, or any other reasonable method can be applied. (As described below, a more complicated but more useful estimate can be made if information regarding the location of the pattern coefficients is also recorded and encoded). From the estimation of the magnitude of each of the coefficients, an estimate of the magnitude or total size of the signal can be made, as indicated above, and the proportion of the size of the coefficient with respect to the total size is used. as a basis for assigning a number of bits to the coefficient. The general technique is described in Lim, JS, cited above, on p. 598.
Due to the exact quantization of the pattern coefficients, the present invention leads to an appropriate allocation of more bits to the coefficients of each band than in the case of the method described in the Dolby document of the prior state of the art. Consider, for example, the two bands b4 and b5 (figure 8), which have the standard coefficients 742 and 743, respectively, with magnitudes nine and fifteen, respectively. Following the method of the previous state of the technique, each pattern coefficient is approximated quantized by coding the pattern exponent uniquely, and this approximate quantization is used to assign bits to all the coefficients of the band of that pattern. Therefore, the standard coefficient 742, which has a value of nine, was quantified by the exponent "3", since it is between 2<sup>3</sup> y2<sup>4</sup>. Since fifteen is the maximum number that this exponent could have, the pattern coefficient band 742 receives a bit allocation as if the maximum value of any coefficient were fifteen.
Also according to the method of the previous state of the technique, the standard coefficient 743, which has a value of fifteen, was quantified by the exponent "3", since it is also between 2<sup>3</sup> y2<sup>4</sup>. Therefore, the pattern coefficient band 743 also receives a bit allocation as if the maximum value of any coefficient were fifteen. Therefore, although the two bands have considerably different standard coefficients, each coefficient in the band is assigned the same number of bits. By way of illustration, it can be assumed that each coefficient in the two bands is assigned four bits for quantization.
Conversely, according to the method of the invention, as the standard coefficients are quantified very accurately, the standard coefficient 743, which has a value of fifteen, is quantized with fifteen, or with a value very close to fifteen if there are very few bits available. Furthermore, the standard coefficient 742, which has a value of nine, is quantified with nine or with a value very close to nine. Thus, the coefficients of band b4 were assigned a different number of bits than that assigned to the coefficients of band b5. As an illustrative label, it can be assumed that each coefficient of band b5, which has a pattern with a magnitude fifteen, five bits are assigned, while each coefficient in band b4, which has only one pattern with a magnitude nine, is assigned only three bits.
Comparison of the bit allocation of the inventive method with that of the prior state of the art method shows that the allocation according to the inventive method is much more appropriate. For band b5 there are more bits available (five instead of four) so the quantization will be more exact. Fewer bits are used for band b4 (three instead of four), however, as the range is actually less than what the prior state of the art method can determine (nine instead of fifteen), the assignment of bits is more appropriate. Furthermore, since the invention also uses exact pattern quantification to establish reconstruction levels, which the prior state of the art method does not do, the relative accuracy achieved is even higher, such as
ES 2 155 449 T3 is explained below.
Once each coefficient has been assigned its bit quota in step 636, the very exact quantization of the standard coefficients can be used to properly divide the entire range of the band and to assign reconstruction levels in step 638. The figure 8 schematically shows the allocation of reconstruction levels. Patterns 743 and 742 of bands b5 and b4 are shown, along with non-pattern coefficients 748 and 746, the first from band b4 and the last from band b5, both of magnitude five. Continuing with the example considered above, the allocation of reconstruction levels according to the present invention and the prior state of the art method are illustrated. Since, according to the prior state of the art, the coefficients of the two bands have been assigned the same number of bits, four, for the reconstruction levels, each band will have 2<sup>4</sup> or sixteen levels of reconstruction. These reconstruction coefficients are shown schematically by 750 idomanic scales on either side of Figure 8. (Reconstruction levels are illustrated with a short-scale line shown in the center of each reconstruction level).
The levels of reconstruction that were assigned according to the method of the invention are very different from those of the prior state of the art and, in fact, they are different in the two bands. In the example, band b5 was assigned five bits per coefficient, so 2 are available<sup>5</sup> or thirty-two levels of reconstruction to quantify coefficients in this band, which has a pattern of fifteen. These reconstruction levels are shown schematically on the 780 scale. Band b4 was assigned only three bits, so that 2 bits are available.<sup>3</sup> or eight levels of reconstruction for the quantification of the coefficients of this band, which has a pattern of value nine. These levels of reconstruction are shown on the 782 scale.
Comparison of the accuracy of the two methods shows that the method of the invention is more efficient than the method of the prior state of the art. In the case of the b5 band coefficients, the thirty-two levels of reconstruction expected as a result of the five-bit allocation clearly provides more accuracy than the sixteen levels of reconstruction expected as a result of the four-bit allocation of the previous state of the technical. Also, the thirty-two levels of reconstruction are helpful. In the case of the b4 band coefficients, the eight levels of reconstruction predicted as a result of the present invention do not provide as many levels of reconstruction as the sixteen levels predicted by the prior art. However, the eight anticipated reconstruction levels are used, while several of the reconstruction levels from the previous state of the technique (those between nine and fifteen) probably cannot be useful for this band, since neither coefficient is greater than nine. Therefore, although there are technically more reconstruction levels assigned to this band as a result of the prior state of the art method, many of them cannot be used, and the increase in accuracy thereby achieved is small. The bits consumed in assigning the unused reconstruction levels could be better used in the same band by reallocating the reconstruction levels to the exact known range, or in another band (such as band b5, where the maximum range is relatively large).
The arrangement of the boundaries between reconstruction levels and the assignment of reconstruction values to the reconstruction levels within the range may vary to satisfy specific characteristics of the signal. If uniform reconstruction levels are assigned, they can be arranged as shown in Figure 9A, on the 902 scale that covers a range of ten, assigning the standard value to the higher reconstruction level, and assigning a lower value to each lower level. reduced by an even amount, depending on the size of the tier. In such a scheme, no level of reconstruction was set to zero. Alternatively, as shown on the 904 scale, the lower reconstruction level can be set to zero, and each higher level is higher by a uniform amount. In such a case, no reconstruction level was assigned the standard value. Alternatively, and more topically, as shown on the 906 scale, neither the pattern nor the zero are accurately quantified, but each was separated from the reconstruction level at a distance corresponding to a half level of reconstruction.
As in the case of the irregular allocation of bits to coefficients in a band, if the encoder can apply more than a reconstruction scheme, either a signal must be transmitted to the decoder together with the data pertaining to the quantized coefficients indicating which reconstruction scheme To use, either the decoder must be constructed in such a way that in all situations it reproduces the required reconstruction level distribution. This information was transmitted or generated in an analogous way to the way in which the specific information pertaining to the number of coefficients per band was transmitted or generated, as described above.
Rather than evenly dividing the width of the band, it may be advantageous to divide it in step 638, as shown in Figure 9B, specifying the reconstruction levels that include and exactly reconstruct the zero and the pattern coefficient and decompensate the distribution of the other levels of reconstruction are more towards the end of the standard coefficient of the range. Alternatively, the reconstruction levels could be clustered closer to the zero end of the range, if experience shows that this is statistically more likely. Therefore, in general, the levels of quantification can be irregular, adapted to the characteristics of the particular type of signal.
In the previous examples it has been implicitly assumed that the standard coefficient is greater than zero and that all other coefficients are greater than or equal to zero. Although this may
ES 2 155 449 T3 occur, there will be many situations in which one or both of the assumptions are not met. There are several possible methods for specifying the sign of non-standard coefficients. The basic mash consists of expanding the bandwidth range to a range that is twice the magnitude of the standard coefficient, and assigning reconstruction levels in step 638, as shown in Figure 10A. For example, any coefficient in the region between the amplitude values 2.5 and 5.0 was quantized in step 640 as 3.75 and in 642 it was assigned the three-bit keyword "101". Obviously, such an arrangement has only half the precision that would be possible if it were only necessary to quantify positive coefficients. Negative values, such as those between -5.0 and -7.5, were also quantified as -6.25 and assigned the keyword "001".
Instead of evenly distributing the positive and negative values, the positive or negative reconstruction levels can be assigned more finely, as shown in Figure 10B. In this case it will be necessary to give more levels of reconstruction to the positive portion or the negative portion of the range. In Figure 10B, the positive portion has four complete reconstruction levels and part of the reconstruction level centered around zero, while the negative portion has three complete reconstruction levels and part of the reconstruction level centered around zero.
The above examples demonstrate that with very accurate quantification of the patterns, very accurate range information can be established for a particular band. Consequently, the reconstruction levels can be assigned to a particular band more appropriately, so that the reconstructed values are more similar to the original values. The prior art method leads to relatively larger ranges for any given band and therefore less appropriate allocation of reconstruction levels.
With the application of the method of the invention, the estimation of the level of masking also improves with respect to the previous state of the art. The estimation of the masking level is based on an estimate of the magnitude of the coefficients | X (k) |. As already mentioned, in general, for each coefficient, the masking level is a measure of the amount of noise, such as quantization noise, that can be admitted into the signal without being perceived by a human observer. In most applications, higher amplitude signals can tolerate more noise without being perceived. The determination of the masking level is also influenced by other factors in addition to the amplitude, for example the frequency and the amplitudes of surrounding coefficients. Therefore, a better estimate of | X (k) | for any given coefficient naturally leads to a better estimate of an appropriate masking level. Masking level is used to fine-tune the bit allocation to a coefficient. If the coefficient is positioned such that it can tolerate a relatively high amount of quantization noise, the bit allocation takes this into account and can reduce the number of bits that are allocated to a specific coefficient (or band) compared to the amount that would have been applied if the level of masking had not been taken into account.
Once the coefficients are encoded according to the method of the invention, the stream of keywords is transmitted in step 644 to the communication channel, or memory storage device, as in the prior state of the art shown in Figure 4A. in step 112. After transmission, the keywords are transformed back into an audio signal. As shown in FIG. 12C, in step 660 the encoded pattern coefficients are quantized based on the assignment of reconstruction levels to the keywords. The standard coefficients have been quantified very accurately. Therefore, when translating the keywords to the reconstruction levels, the reconstructed pattern coefficients will very accurately reflect the original pattern coefficients.
At step 622 a decision is made as to whether or not to perform an inverse DCT transformation (or other appropriate transformation) to counteract any DCT-type transformations (described below) that may have been applied in steps 616, 618, or 620 in the encoder. If so, the inverse transformation is applied in step 664. If not, the method of the invention continued to step 666, in which the keywords for the non-pattern coefficients of a single block are translated into levels of quantification. There are many possible schemes that are described further below.
The decoder translates the keywords into quantization levels by applying an inversion of the steps followed in the encoder. From the pattern coefficients, the encoder has available the number of bands and the magnitudes of the patterns. The number of non-standard coefficients for each band is also known either from secondary information or from pre-established information. From this, the decoder can set the reconstruction levels (quantity and locations) by applying the same rule applied by the encoder to set the bit assignments and reconstruction levels. If there is only one of these rules, the decoder simply enforces it. If there are more than one, the decoder chooses the appropriate one, based on secondary information or intrinsic characteristics of the standard coefficients. If the keywords have been applied to the reconstruction levels according to a simple ordering scheme, such as the binary representation of the position of the reconstruction level from the lowest arithmetic value to the highest arithmetic value, this scheme is simply reversed to produce the level of reconstruction. If a more complicated scheme is applied, such as cipher code, the decoder must have access to that cipher scheme or code.
The final result is a set of quantized coefficients for each of the frequencies present in the X (k) spectrum. These coefficients will not be exactly the same as the originals,
ES 2 155 449 T3 because part of the information has been lost in the quantification. However, due to the more efficient bit allocation, the best range division, and the best masking estimate, the quantized coefficients are more similar to the originals than the prior-art requantized coefficients would be. (However, the topically reconstituted non-standard coefficients cannot be compared to the original non-standard coefficients as accurately as the reconstituted standard coefficients compared to the original standard coefficients.) After requantification, the effect of raising the block to the fractional power α, such as 1/2, is canceled in step 668 by raising the values to the reciprocal power 1 / α, in this case 2. Next, in step 670, the inverse transformation to the TDAC type transformation applied in step 106 is applied to transform the frequency information back to the time domain. The result is a segment of data, specified at a sample rate of, for example, 48 kHz. In step 672 sequential windows are combined (topically overlapped) and in step 674 an audio synthesis was performed.
In the above description it has been assumed that only the magnitudes of the standard coefficients have been accurately coded in step 614, and that neither the location of the standard coefficient within the band has been coded (i.e., second coefficient from the extreme low frequency of the band, fourth coefficient from the low frequency end of the band, etc.) nor the sign (or phase). By coding the site well, or these two additional data, the coding can be further refined. In reality, the location coding produces considerable savings, since otherwise it would be necessary to code the standard coefficient twice: once to establish the estimate of | X (k) |<sup>α</sup> and another for its contribution to the signal in the form of a coefficient.
If the location of the standard coefficient has not been encoded, it would be necessary to encode its magnitude in the stream of all coefficients, for example in step 624 shown in FIG. 12a. However, if the standard coefficients have been fully encoded with magnitude and location and sign, their code values can simply be transmitted. If the location has not been coded, the apparatus must first transmit the magnitudes of each pattern, for example in step 628 of FIG. 12B. Bits are then assigned to each band and to each coefficient within the band, including the pattern coefficient in step 636. If the pattern location information has not been stored in memory, the system is insensitive to the pattern's special identity and allocates bits to it in step 636, quantizes it to a reconstruction level in step 640, encodes it in step 642 and transmits its amplitude in step 644. Therefore, its amplitude is transmitted twice: first in step 628 and then in step 644.
However, if the location was originally encoded in step 636, when the system prepares to allocate bits to the pattern in step 636, the pattern coefficient would be identified as such, due to its location, and would be ignored, saving roasted the bits necessary to encode their amplitude. Typically, specifying the placement of patterns only improves efficiency if fewer bits are required to specify their placement than to specify their width. In some cases it may be advantageous to encode the locations of certain pattern signal elements, but not all. For example, if a band includes a large number of coefficients, it may not be advantageous to encode the location of the pattern in that band, but it may still be advantageous to encode the location of a standard coefficient in a band that has fewer coefficients. Furthermore, when evaluating the advantage of specifying the location of the standard coefficients, the probable additional computational and perhaps memory loads required in both the encoder and decoder apparatus must also be considered, in view of the bandwidth of the data channel. available. Typically, it is more cost effective to accept higher compute or memory loads than bandwidth loads.
If in step 614 (FIG. 12A) it is decided to accurately quantify the location of the coefficient in the band, some additional bits would be required to specify and encode each standard coefficient. Topically, the number of coefficients that will be in each band is decided before coding the coefficients. This information is topically known by the decoder, although it is also possible to vary this information and include it in the secondary information transmitted by the encoder. Therefore, for each band the location of the standard coefficient can be specified exactly, and it is only necessary to reserve a sufficient number of bits for the location information required by the number of coefficients of the band in question. For this reason, it is advantageous to assign coefficients to each band with a power of two numbering, so as not to waste bits in specifying the location of the standard coefficient.
As mentioned above, a basic method of allocating bits within the band is to allocate an equal number of bits to each non-standard coefficient. However, in some cases this cannot be done, for example when the number of bits available is not an integer multiple multiple of the number of non-standard coefficients. In this case, it is often advantageous to assign more bits to the coefficients that are closest to the standard coefficient (at an in-band location), since experience has shown that in audio-type signals the adjacent coefficients are frequently closer to each other. magnitude than distant coefficients.
There are various other uses for which additional bits can be assigned. For example, preference can be given to coefficients to the left of the standard coefficient, that is, they have a lower frequency than the standard coefficient. This is in consideration of the result13
ES 2 155 449 T3 of the masking. Topically, the impact of a specific frequency component on the masking function occurs with respect to a region of frequency higher than the frequency in question. Therefore, by giving preference to frequency coefficients lower than the pattern (therefore they are to the left of the pattern in a conventional scale such as that shown in Figure 11), the coefficient that influences higher frequency components. In some circumstances, it may even be advantageous to favor these lower frequency coefficients with something other than just the single extra bit available from an odd number of extra bits. For example, additional bits could be assigned to five coefficients on the lower side of the pattern, but only two on the upper side.
Therefore, the exact specification of the location of the standard coefficient within the band also allows a more appropriate allocation of bits between the various non-standard coefficients. With a more appropriate allocation of bits per non-pattern coefficient, the division of the bits into appropriate reconstruction levels is further improved, as described above.
Knowing the location of the pattern coefficients also allows a better approximate estimate of | X (k) |<sup>α</sup>, which in turn allows a better estimation of the masking function. If the locations of the standard coefficients are known, the estimation of | X (k) |<sup>α</sup> it may be as shown in Figure 11, instead of as shown in Figure 7A. Without the location information, all that can be estimated is that the band coefficients are on average less than some fraction of the magnitude of the standard coefficient. However, knowledge of the locations allows the more topically accurate estimation shown in Figure 11, in which each non-pattern coefficient is assigned an estimated value based on the relationship between adjacent patterns. The assumption underlying such an estimate is that the magnitudes of the coefficients do not change much from one coefficient to the next, and therefore non-standard coefficients are generally found along the lines connecting the coefficients. adjacent patterns. Thus, once the most refined estimate for | X (k) |<sup>α</sup>, the estimates of the individual coefficients can be used to apply either of two modes of bit allocation: bit allocation for bands followed by bit allocation for coefficients, or direct bit allocation for coefficients. Furthermore, this more refined estimate can also be used to more conveniently set the masking level. Therefore, the bit allocation, and therefore also the rank allocation, are improved by encoding the location of the patterns.
If the location of each pattern has been specified, you can revert to any patterns that have been encoded without redundancy and increase the accuracy of your encoding if more bits are available than was assumed at the time the patterns were encoded. For example, the band in question may have received a large number of bits because its pattern is very large, but it may not need as great a number of bits to encode the other signal elements if the band has a very small number of signal elements. sign. If the locations are known, more bits can be assigned to specify the amplitude of the pattern coefficient after the first bit-to-pattern assignment process. If the locations are unknown, this cannot be done effectively without redundancy. One method to further specify the magnitude of the pattern was to use the additional bits to encode the difference between the magnitude of the first encoded pattern and the amplitude of the original pattern. Since the decoder apparatus employed the same routines to determine how the bits were allocated as the routines employed by the encoder, the decoder automatically recognized the increased pattern width information appropriately.
More efficient and accurate coding can be achieved by accurately specifying and coding the sign of the standard coefficient (which corresponds to the phase of the signal components at that frequency). If X (k) are real values, only one additional bit per pattern coefficient is needed to encode their sign.
Knowledge of the sign of the pattern coefficient increases the ability of the method to efficiently determine levels of reconstruction within a given band. For example, experience indicates that a band can frequently include more non-standard coefficients that have the same sign as the standard coefficient. For this reason, it may be advantageous to foresee one or two levels of reconstruction more that have this sign.
In general, knowledge of the sign of the pattern does not increase the estimation of the masking effect. The usefulness of the sign information varies depending on what transformation has been used.
Another preferred embodiment of the method of the invention is particularly useful if the number of bands is relatively small. This embodiment involves a further division of each band of the X (k) spectrum into two divided bands in step 612 of FIG. 12A. One divided band includes the standard coefficient and the other does not. Preferably, the divided bands should divide the band approximately in half. The highest magnitude coefficient in the divided band that does not contain the standard coefficient is also selected in step 650 and quantized in step 624. Figure 7 shows the division of two of the bands, bands b2 and b4, in bands divided by a discontinuous vertical line through the centers of these two bands. If this realization is performed, the standard coefficient and the additional coded coefficient are called the higher standard coefficient and the lower standard coefficient, respectively. This step 650 occurs between the selection of the largest standard coefficients in step 608 and the coding of the magnitude of any standard coefficient in step 626.
The magnitudes of the standard coefficients me14
ES 2 155 449 T3 nores are also accurately quantified in step 624. Since they are minor standards, it is known that their magnitude is not greater than that of the larger standard coefficients. This fact can be used to save bits in your encoding.
There are various methods for dividing the entire block into, for example, sixteen bands. One is to divide the segment from the beginning into sixteen bands. Another consists of dividing the entire segment in two, and then dividing each part in two, and so on, the information derived from the first division being more important than the information derived from the second division. Therefore, the use of split bands provides an important information hierarchy. The first division is more important than the second division, which in turn is more important than the next, etc. Thus, it may be advantageous to reserve bits for the most important divisions.
As mentioned above, it may be advantageous to apply a second transformation to the patterns prior to quantization, encoding, and transmission in steps 624, 626, and 628, respectively. This second transformation could be applied to the major and minor patterns, or only to the major or minor patterns. This is because, depending on the nature of the signal, there may be some kind of configuration or organization between the pattern coefficients. As is known, transformations take advantage of a given configuration of the data to reduce the amount of data information needed to accurately define the data. For example, if each standard coefficient were simply twice the magnitude of the previous coefficient, it would not be necessary to quantify, encode, and transmit the magnitudes of all the coefficients. It would only be necessary to encode the magnitude of the first and apply a doubling function to the received coefficient for the number of steps required.
Thus, in step 622, 652 or 654 (depending on which of the magnitude, location and sign aspects are being accurately quantified) it is decided whether or not to apply a second transformation to the standard coefficients according to a known method, such as DCT. If the nature of the data is such that it is likely to provide a more compact encoding mode, another transformation is applied at steps 618, 616, or 620. Figure 12A indicates that the transformation is a DCT transformation, however any transformation that achieves the goal of reducing the amount of data to be transmitted can be used. Other types of suitable transformations include the Discrete Fourier Transformation.
Due to this potential transformation only of the pattern that is not appropriate in all cases, it is concluded that, according to the method of the invention, the greater accuracy with which the pattern coefficients are encoded is the result of allocating more bits to each pattern coefficient (of average (on average) than for each coefficient no standard (on average). This is because applying only the pattern transformation can lead to a significant reduction in the number of bits needed to encode all pattern coefficients and, therefore, any individual (average) pattern coefficients. Obviously, this bit saving is achieved thanks to increased computing requirements, both in encoding and decoding. In some applications, saving bits justified the computational burden, in others it may not. Both cases will be obvious to people with normal technical experience.
If the patterns are transformed twice, they must be inversely transformed back to the frequency domain of X (k) in step 632 to simplify the calculations required for the bit allocation in steps 634, 636 and the designation of levels of reconstruction in step 638, as described above. Alternatively, instead of reverse transformation, the patterns can be stored in decoder memory and recalled before step 634.
During the decoding steps of the method of the invention, the exact mode of translation in step 666 of the non-standard keywords transmitted at the quantization levels will depend on whether split bands have been used, whether or not the location or the location and the sign of the standard coefficients, and how that information has been packaged. If secondary information has been used to transmit control data, the secondary information has to be decoded and applied. If all the necessary information is contained in a memory accessible to the decoder, it is only necessary to translate the keywords according to established algorithms.
For example, an established algorithm may set the number of coefficients per band in the first half of the block to sixteen and the number of coefficients per band in the second half to thirty-two. A rule can also be established to allocate bits uniformly between coefficients within a band, the eventual additional bits being assigned one by one to the first coefficients of the band. If the sign of the standard coefficient is quantified, each coefficient can be divided into levels of reconstruction with an additional level of reconstruction having a sign equal to that of the standard coefficient.
In light of the above detailed description of the method of the invention, the apparatus of the invention was understood from Figure 13A, which shows the transmitting part of the apparatus, and from Figure 13B, which shows the receiving part. The apparatus of the invention can be used in specialized processors or in an appropriately programmed general purpose digital computer.
The TDAC type transformer 802 transforms an audio type signal, such as x (t), into a spectrum such as X (k). (A DCT transformer is also appropriate and is contemplated by the invention). The operator ||<sup>α</sup> scales the spectrum to a domain more relevant to human perception, or when non-uniform quantification is desired. Spectral band divider 806 divides the scaled spectrum into separate bands. The standard coefficient identifier 808 identifies the coefficients of each band that have 15
ES 2 155 449 T3 n in the highest magnitude. The quantizers 801 and 812 quantify the magnitude of the standard coefficients (and perhaps their sign) and, if desired, the location within the band respectively. The DCT 816 transformer applies a DCT or similar transformation to the quantized pattern information if it is determined that there is sufficient structure between the pattern coefficients to justify additional computation. The encoder 818 encodes the information of quantized patterns, regardless of whether the DCT transformer acted on the information, producing a series of keywords that are transmitted through the transmitter 820 on a data channel.
In a preferred embodiment, a band bit mapper 822 collects the information from the pattern magnitude quantizers 810 and uses this information to establish a rough estimate of | X (k) |<sup>α</sup> as shown in Figure 7A, and uses this estimate to allocate the limited number of bits available between the spectrum bands established by the spectral band divider 806. The coefficient bit allocator 824 uses the information from the pattern position and sign quantizers 812 and 814 along with in-band bit allocation to allocate the bits of the band between the coefficients in the band. The non-standard quantifier 826 uses the same information to establish adequate reconstruction levels for each coefficient in the band and to quantify each coefficient. The quantized coefficients are passed to encoder 818, which assigns a keyword to a non-standard coefficient and passes the keywords to transmitter 820 for transmission.
In another preferred embodiment of the apparatus, the band bit allocator may also collect information from the position quantizer pattern 812 to establish the approximate estimate of | X (k) |<sup>α</sup>. The banding bit allocator will establish a rough estimate as shown in Figure 11 if the location information is used and, based on this estimate, allocate bits to the bands.
In another embodiment of the apparatus of the invention, the band bit mapper 822 also collects sign information from the magnitude quantizer 810 and location information from the location quantizer 812 to map bits to the band, as described above with respect to to the method of the invention.
Figure 13B schematically shows the receiver or decoder part of the invention. The receiver 920 receives the keywords of the communication channel. Pattern decoder 918 decodes the pattern data, resulting in quantized data representing the patterns. The reverse DCT transformer 916 counteracts the effect of any DCT type transformations that were applied in step 816, resulting in a set of scaled pattern coefficients whose magnitude is very similar to that of the original scaled pattern coefficients prior to quantization. in the magnitude quantizer 810. The non-pattern decoder 926 receives the keywords representing the non-pattern coefficients and translates these coefficients into reconstructed non-pattern coefficients. As mentioned above in connection with the method, the operation of the decoder 926 depended on the means used to encode the nonpattern information. The operator 904 raises the quantized coefficients in the reconstructed spectrum to the power of 1 / α, to cancel the effect of the operator 804. The inverse transformer 902 applies an inverse transformation to the spectrum to cancel the effect of the TDAC transformer 802 and transform the signal of the frequency domain back to the time domain, resulting in a boxed time domain segment. Combiner 928 combines individual sample windows and synthesizer 930 synthesizes an audio-type signal.
Another preferred embodiment of the encoder omits the band bit allocator and includes only a coefficient bit allocator, which takes the estimate of X (k)<sup>to</sup> and uses it to assign bits directly to coefficients, as described above with respect to the inventive method.
In the previous description of the method and the apparatus it has been assumed that the standard coefficients are the coefficients that have the absolute value of maximum amplitude in the band. It is also advantageous to use a coefficient other than the maximum magnitude as a reference standard with which to measure the other coefficients. For example, although it is believed that using the maximum amplitude coefficient will obtain optimal results, advantageous results could be obtained using a coefficient with an amplitude close to the maximum amplitude, for example the second or third largest. This method is also contemplated by the invention and is considered covered by the appended claims.
The reference standard can also be the coefficient that, among all the magnitudes of the band coefficients, has the magnitude closest to the intermediate or median coefficient of the band. An intermediate value pattern is advantageous in cases where the statistical characteristics of the signal are such that the intermediate or median value contains more information about the total energy of the signal than the maximum value of a band. This would be the case of a topical signal characterized by deviations within a constant range above and below an intermediate value. It would also be necessary to characterize or estimate a range for the magnitude of the deviations. For example, if the intermediate value of a band had a value of more than five, and if from the statistics of that type of signal it is known that the values of this band diverge topically from the intermediate value only by +/- four units, the range is It was set between more than one and more than nine, and the levels of reconstruction were established within this range. As before, the reconstruction levels can be evenly divided, or more concentrated around the intermediate value, or they can be skewed towards either end of the range, depending on the statistical information on that particular class of signal.
Similarly, the pattern coefficient can be the coefficient that has the closest magnitude
ES 2 155 449 T3 to the mean of all the magnitudes of the other coefficients of the band. This average value is useful if it represents an estimate of the energy of the band better than any other value, for example the maximum or intermediate values.
The invention has been described above with respect to a signal divided into a plurality of bands, and it is expected that it will be the application that will benefit the most from the invention. However, the invention is also useful in relation to encoding the amplitudes of a plurality of coefficients in a single individual band. The application of the invention for a signal or signal component in a single individual band follows the same principles as the application for the multiband signals described above. The pattern is selected and accurately quantified, preferably, but not necessarily, by encoding the location and sign of the pattern. The exact quantization of the pattern is used in conjunction with the number of bits available to establish reconstruction levels and allocate bits between non-pattern coefficients. All the considerations discussed above are applicable to the single-band implementation, except that the amount of bits available for the band will be determined and will not depend on the specific data of other bands, if any.
The present invention has many advantages. Bits related to bit allocation, such as the magnitude of the pattern coefficient as well as their locations and signs, would be well protected. Thus, any errors that occur will be located in a particular band and cannot be greater than the magnitude of the pattern coefficient in each band. The standard coefficients will always be represented exactly. The standard width information is not discarded as in some prior state of the art methods, but is used very efficiently for its own direct use and for bit allocation. Relating to the method described in the Dolby document, the invention uses the available bits more efficiently. In the Dolby method, the exponents of the peak spectral values of each band are encoded. In this way, first a rough estimate of the width of a band is made. Subsequently, all coefficients, including the peak coefficient, are encoded and transmitted using a more precise estimate of their magnitude. Therefore, the accuracy of the peak amplitudes is the same as that of other coefficients in the same band. Furthermore, the accuracy of the standard coefficients of the present invention ensures the use of exact ranges for the determination of reconstruction levels, which enables more efficient use of the available bits.
In addition to the above-described specific embodiments of the method and apparatus of the invention, other variations also fall within the intended scope of the claims. Techniques that take into account the perceptual properties of human observers can be incorporated, in addition to estimating the level of masking.
In addition, it can be considered more than one block at a time. For example, in the special case of silences, you can remove the bits from the block where silence occurs and assign them to another block. In less extreme cases it may still be appropriate to dedicate fewer bits to one block than to another. The establishment of the bands can be done "on the fly", including in a band sequential coefficients close to each other, and then starting a new band with a coefficient of significantly different magnitude.
The method and apparatus of the invention can also be used with encoded data of any type, for example with two-dimensional signals. The data does not need to have been transformed. The invention can be applied to samples of time domain x (n), except that in the case of audio the results will not be as good as they would be if the data had been transformed. Transformation is applied topically to data to take advantage of settings within the data. However, the application of transformation is not necessary and, in some cases, when the data tend towards randomness, it is not topically advantageous. In fact, in the case of time domain samples, the coefficients will be elements of the sampling signal with sampling amplitudes of the real sampling signal, rather than some transformation of them in another domain. The method of the invention is applied in the same way, excluding the transformation and inverse transformation steps. Similarly, the apparatus of the invention would not require the direct and inverse transformation operators in that case. (However, it may still be advantageous to perform the pattern transformation alone).
Interaction between blocks can also be performed.
The above description is to be understood as illustrative and not as limiting in any way. Although the invention has been shown and described in particular with reference to preferred embodiments thereof, it is understood by those skilled in the art that various changes in shape and detail can be made without departing from the scope of the invention as defined in the claims that follow.
Contents3
10 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10
24 members in 7 offices
Priority claims10
| Document | Office | Kind | Date |
|---|---|---|---|
| 19920822247 | United States of America | – | |
| 82224792 | United States of America | A | |
| 82224792 | United States of America | A | |
| 19920879635 | United States of America | – | |
| 87963592 | United States of America | A | |
| 87963592 | United States of America | A | |
| 822247 | – | – | – |
| 879635 | – | – | – |
| US19920822247 | – | – | – |
| US19920879635 | – | – | – |
Members24
| Document | Office | Kind | |
|---|---|---|---|
| CA2128216A1 | Canada | A1 | |
| WO9314492A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US5369724A | United States of America | A | |
| EP0628195A1 | European Patent Office (EPO) | A1 | |
| US5394508A | United States of America | A | |
| EP0628195A4 | European Patent Office (EPO) | A4 | |
| US5625746A | United States of America | A | |
| US5640486A | United States of America | A | |
| EP0961414A2 | European Patent Office (EPO) | A2 | |
| EP0961414A3 | European Patent Office (EPO) | A3 | |
| EP0628195B1 | European Patent Office (EPO) | B1 | |
| AT198384T | Austria | T | |
| ATE198384T1 | Austria | T1 | |
| DE69329796D1 | Germany | D1 | |
| ES2155449T3This record | Spain | T3 | |
| DE69329796T2 | Germany | T2 | |
| EP0961414B1 | European Patent Office (EPO) | B1 | |
| AT291771T | Austria | T | |
| ATE291771T1 | Austria | T1 | |
| DE69333786D1 | Germany | D1 | |
| ES2238798T3 | Spain | T3 | |
| DE69333786T2 | Germany | T2 | |
| CA2128216C | Canada | C | |
| USRE40691E | United States of America | E |
1 legal event, as the office reported them to INPADOC
Events
| Event | Code | |
|---|---|---|
| Definitive protectionFG2A | FG2A |
Numbers
- Publication
- 2155449
- Publication, DOCDB
- 2155449
- Publication, EPODOC
- ES2155449T
- Application
- 93903484
- Application, DOCDB
- 93903484
- Application, EPODOC
- ES19930903484T
Titles2
- Spanish
- METODO Y APARATO PARA LA CODIFICACION, DECODIFICACION Y COMPRESION DE DATOS DE TIPO AUDIO.
- English
- METHOD AND APPARATUS FOR THE CODING, DECODING AND COMPRESSION OF AUDIO-TYPE DATA.
Classification
- CPC, 6
- H04B1/667
- G10L19/0204
- G10L19/0208
- G10L19/0212
- G10L25/27
- G10L2021/02163
- IPC, 3
- G10L19 02
- G10L21 02
- H04B1 66