Speech signal quantization using human auditory models in predictive coding systems
Abstract
A SPEAKING COMPRESSION SYSTEM CALLED "TRANSFORM PREDICTIVE CODING", OR TPC SUPPLIES THE CODING OF SPEAKS IN A 7 KHZ WIDE BAND (16 KHZ SAMPLING) IN A SPEED BIT OF BITS BETWEEN 16 AND 32 KB / S KB 1 TO 2 BITS / SAMPLE). THE SYSTEM USES A SHORT AND LONG-TERM PREDICTION TO ELIMINATE REDUNDANCY IN SPEAK. A PREDICTION RESIDUAL IS TRANSFORMED AND CODED ON THE FREQUENCY DOMAIN TO GET PART OF THE KNOWLEDGE OF HUMAN AUDITIVE PERCEPTION. THE TPC ENCODER USES ONLY QUANTIFICATION OF OPEN CIRCUIT AND THEREFORE HAS AN EMINENTLY LOW COMPLEXITY. THE QUALITY OF TPC SPEECH IS ESSENTIALLY TRANSPARENT AT 32 KB / S, VERY GOOD AT 24 KB / S AND ACCEPTABLE AT 16 KB / S.

Term
Term ended
Projected expiry passed 17 September 2016, 10 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
10 claims: 1 independent, 9 dependent
- 1ES 2 174 030 T3 REIVINDICACIONES 1. Procedimiento de codificacion de una senal que representa a una informacióon (s) de voz, comprendiendo el procedimiento las actuaciones siguientes:generacion (10) de una primera senal (a) que representa a una estimacion de la senal que representa a la informacióon de voz;comparación (20) de la senal que representa a la informacion de voz con dicha primera senal, para formar una segunda senal (d) que representa a la diferencia existente entre dichas senales comparadas;determinacióon (50) de una resolucióon de cuantificador;cuantificacioón (60) de una funcioón de la segunda senal, de acuerdo con la resolución determinada del cuantificador;y generación (MUX 70) de una senal codificada que comprende un primer componente (ir) basado en dicha senal cuantificada, caracterizado porque la resolucioón del cuantificador se determina de acuerdo con dicha primera senal y una senal de enmascaramiento de ruido de percepcion, estando dicha senal de enmascaramiento de ruido basada en un modelo de percepcióon humana de audio.
- 2El procedimiento de la reivindicacióon 1, en el que dicha funcion de dicha segunda senal comprende una diversidad de coeficientes de frecuencia, y dicha determinacioón de una resolucióon de cuantificador comprende la determinacióon de la asignacióon de bits para cuantificar los coeficientes de frecuencia respectivos de dicha segunda senal.
- 3El procedimiento de la reivindicacioón 1, en el que dicha determinacioón de una resolucioón de cuantificador comprende la determinacióon de la asignacioón de bits para identificar vectores de cóodigos en un cuantificador de vectores.
- 4El procedimiento de la reivindicacioón 2, en el que dicha asignacioón de bits comprende la asignacióon de cero bits a al menos un coeficiente de frecuencia de dicha segunda senal.
- 5El procedimiento de la reivindicacióon 2, en el que dichos coeficientes de frecuencia son coeficientes de FFT.
- 6El procedimiento de la reivindicacioón 1, en el que dicha generación de una senal codificada tambióen comprende la generacióon de un segundo componente basado en dicha segunda senal, para permitir la determinacióon de dicha resolucioón de cuantificador a partir de dicha senal codificada.
- 7El procedimiento de la reivindicacióon 1, en el que dicha segunda senal comprende informacióon residual de prediccioón lineal.
- 8El procedimiento de la reivindicacióon 1, en el que dicha funcion de dicha segunda senal comprende dicha segunda senal en la que se ha eliminado la informacioón de prediccióon de tono.
- 9El procedimiento de la reivindicacioón 1, en el que dicha funcion de dicha segunda senal comprende una transformacioón a un dominio de frecuencia de dicha segunda senal, despues de que se ha eliminado la informacioón de prediccióon de tono de dicha segunda senal.
- 10El procedimiento de la reivindicacioón 1, en el que dicha senal de enmascaramiento de ruido de percepcióon comprende un umbral de enmascaramiento de ruido, basado en dicha primera senal. NOTA INFORMATIVA:Conforme a la reserva del art. 167.2 del Convenio de Patentes Europeas (CPE) y a la Disposición Transitoria del RD 2424/1986, de 10 de octubre, relativo a la aplicacion del Convenio de Patente Europea, las patentes europeas que designen a España y solicitadas antes del 7-10-1992, no producirán ningún efecto en Espana en la medida en que confieran proteccián a productos quámicos y farmaceuticos como tales. Esta informacioán no prejuzga que la patente estáeo no incluáda en la mencionada reserva.
Independent claims10
119 paragraphs in 5 sections, as filed
ES 2 174 030 T3
DESCRIPTION
Voice signal quantification using human hearing models in predictive coding systems.
Field of the invention
The present invention relates to the compression (encoding) of audio signals, eg speech signals, using a predictive encoding system.
Background of the invention
As taught in the signal compression literature, voice and music waveforms are encoded using very different encoding techniques. Speech coding, such as telephonic bandwidth (3.4 kHz) speech coding at or below 16 kb / s, has been dominated by predictive time domain coders. These encoders use speech production patterns to predict the speech waveforms to be encoded. The predicted waveforms are then subtracted from the real (original) waveforms (to be encoded) to reduce redundancy in the original signal. Redundancy reduction in the signal provides coding gain. Some examples of such predictive speech coders include Adaptive Predictive Coding, Multi-Impulse Linear Predictive Coding, and Cocode Excited Linear Prediction Coding (CELP), all of which are well known in the art of voice signal compression.
On the other hand, wideband music (0-20 kHz) encoding at or above 64 kb / s has been dominated by frequency-domain or sub-band transform encoders. These muosic coders are fundamentally very different from the previously discussed speech coders. This difference is due to the fact that the music sources, unlike the voice sources, are too varied to allow an easy and fast prediction. Consequently, muosic source models are not generally used in muosic encoding. Instead, music encoders use elaborate models of human hearing to uniquely encode the parts of the signal that are perceptually applicable. That is, unlike speech coders that commonly use speech production models, music encoders use sound reception hearing patterns to obtain encoding gain.
In music encoders, hearing models are used to determine the noise masking ability of the music to be encoded. The term "noise masking ability" refers to the amount of quantizing noise that can be introduced into a music signal without a listener noticing the noise. This noise masking capability is then used to set the resolution of the quantizer (eg quantizer step size). Generally, the higher the “pitch” of the music, the poorer the music will be in the masking quantization noise, and therefore the smaller the required quantizer step size, and vice versa. Smaller step sizes correspond to lower encoding gains and vice versa. Some examples of such muosic encoders include AT & T's Audio Perception Encoder (PAC) and the standard ISO MPEG audio encoder.
Between telephonic bandwidth speech coding and wideband muosic coding, there was wideband speech coding, in which the speech signal is sampled at 16 kHz and has a bandwidth of 7 kHz. The advantage of 7 kHz bandwidth voice is that the resulting voice quality is much better than the telephone bandwidth voice quality and still requires a much lower bit rate for encoding than a signal. 20 kHz audio. Among the previously proposed wideband speech coders, some use predictive time domain coding, some use sub-band coding or frequency domain transformation, and some use a mixture of time domain and frequency domain techniques.
The inclusion of perceptual criteria in predictive speech coding, broadband or otherwise, has been limited to the use of a perceptual weighting filter in the context of selecting the best synthesized speech signal, from among a variety of candidate synthesized voice signals. See, for example, US Patent No. Re. 32,580 to Atal et al. Such filters perform a type of noise shaping that is useful for noise reduction in the encoding process. A known coder attempts to improve this technique by employing a perceptual model in the formation of this perceptual weighting filter. See WW Chang et al., "Audio Coding Using a Perceptual Filter Adapted to the Masking Threshold", Proc. IEEE Workshop Voice Coding for Telecommunications, pp. 9-10, October 1993.
Summary of the invention
Despite the efforts described above, none of the known speech or audio encoders use a speech production model for the purpose of signal prediction, in addition to an audition model to set the resolution of the quantizer, according to a signal capacity anaolysis, noise masking.
On the other hand, the present invention combines a predictive coding system with a quantification process that quantifies a signal, based on a noise-masking signal, determined with a human hearing sensitivity model for noise. The output of the predictive coding system is also quantized with a quantizer that has a resolution (for example, step size in a uniform scalar quantizer, or the number of bits used to identify coding vectors in a vector quantizer) that is a they functioned as a noise-masking signal, determined according to an audio perception model.
According to the invention, which is as set forth in claim 1, a signal is generated
ES 2 174 030 T3 representing an estimate (or prediction) of a signal representing speech information. The term "signal representing voice information" is broad enough to refer not only to the voice itself, but also to the derivatives of the voice signal that are normally found in voice coding systems (such as residual signals pitch prediction and linear prediction). The estimated signal is then compared with the original signal, to form a signal that represents the difference between the compared signals. Subsequently, this signal that represents the difference between the compared signals is quantized, according to a masking signal with perception noise, which is generated by an audio human perception model.
An illustrative embodiment of the present invention, referred to as "Transformation Predictive Coding", or TPC, encodes 7 kHz wideband speech at a target bit rate of 16 to 32 kb / s. As its name implies, TPC combines transformation coding and predictive coding techniques into a single encoder. More specifically, the encoder uses linear prediction to eliminate redundancy from the input speech waveform and subsequently uses transform coding techniques to encode the resulting residual prediction. The transformed residual prediction is quantified based on knowledge of human hearing perception, expressed in terms of a hearing perception model, to encode what is audible and discard what is not audible.
An important feature of the illustrative embodiment relates to the way in which the signal's perceptual noise masking ability is determined (for example, the “just perceptible distortion” perception threshold) and the way in which, subsequently , the bit allocation was carried out. Instead of determining a perception threshold using the unquantized input signal, as is done in conventional music encoders, the noise masking threshold and the bit allocation of the realization are determined based on the response of frequency of a quantized prosthesis filter -in the realization, a quantized LPC prosthesis filter. This feature provides the system with the advantage of not having to communicate bit allocation signals, from the encoder to the decoder, in order for the decoder to accurately reproduce the perception threshold and the bit allocation processing necessary to decode the signal. encoded broadband voice information received. Instead, the filter synthesis coefficients, which are being reported for other purposes, are used to conserve the bit rate.
Another important feature of the illustrative embodiment relates to the way the TPC encoder allocates bits between the encoder frequencies and the way the decoder generates a quantized output signal, based on the allocated bits. In certain circumstances, the TPC encoder only allocates bits to a portion of the audio band (for example, bits can be allocated to coefficients between 0 and 4 kHz, only). No bits are assigned to represent coefficients between 4 kHz and 7 kHz, and therefore the decoder does not obtain any coefficients in this frequency range. Such a circumstance occurs when, for example, the TPC encoder has to operate at very low bit rates, for example 16 kb / s. Despite not having any bits representing the signal encoded in the frequency range 4 kHz to 7 kHz, the decoder still has to synthesize a signal within this range if a broadband response is to be provided. According to this characteristic of the realization, the decoder generates, that is, synthesizes coefficient signals in this frequency range, based on other available information, a proportion of an estimate of the signal spectrum (obtained from the LPC parameters) to a noise masking threshold at the frequencies in the range. The phase values for the coefficients are selected randomly. Due to this technique, the decoder can provide a wideband response without the need to transmit speech signal coefficients for the full band.
Potential applications for a wideband voice scrambler include ISDN audio conferencing or video conferencing, multimedia audio, “hi-fi” (“high frequency”) telephony, and simultaneous voice and data (SVD) over link lines. between quadrants using 28.8 kb / s modems or higher speeds.
Brief description of the drawings
Figure 1 represents an illustrative embodiment of the encoder of the present invention.
Figure 2 represents a detailed block diagram of the LPC analysis processor of Figure 1.
Figure 3 represents a detailed block diagram of the pitch prediction processor of Figure 1.
Figure 4 represents a detailed block diagram of the transformation processor of Figure 1.
Figure 5 represents a detailed block diagram of the hearing model and control processor of the quantizer of Figure 1.
Figure 6 represents an attenuation function of a LPC power spectrum, used for the determination of a masking threshold, for allocation of adapter bits.
Figure 7 represents a general bit allocation of the embodiment of the encoder of Figure 1.
Figure 8 represents an illustrative embodiment of the decoder of the present invention.
Figure 9 represents a flow chart illustrating the process performed to determine an estimated masking threshold function.
ES 2 174 030 T3
Figure 10 represents a flow chart illustrating the process carried out to synthesize the magnitude and phase of the fast Fourier transform residual coefficients, for use by the decoder of Figure 8.
Detailed description
A. Introduction to Illustrative Embodiments For clarity of explanation, the illustrative embodiment of the present invention is presented as comprising individual functional blocks (including functional blocks labeled "processors"). The functions that these blocks represent can be provided through the use of shared hardware or specifically dedicated hardware, including, but not limited to, hardware capable of running software. For example, the functions of the processors shown in Figures 1-5 and 8 can be provided by a single shared processor. (The use of the term "processor" should not be construed as referring exclusively to hardware capable of running software).
Illustrative embodiments may comprise digital signal processor (DSP) hardware, such as the AT&T DSP16 or DSP32C, read-only memory (ROM) to store the software that performed the operations discussed later, and random access memory ( RAM) to store the DSP results. Very large scale integration (VLSI) hardware implementations can also be provided, as well as commercial VLSI circuitry, in combination with general purpose DSP circuitry.
Figure 1 depicts an illustrative embodiment of the TPC speech coder of the present invention. The TPC encoder comprises an LPC analysis processor 10, a LPC (or "short-term") prediction error filter 20, a pitch prediction (or "long-term") processor 30, a Transformation 40, an audition model quantizer control processor 50, a residual quantizer 60, and a bitstream multiplexer (MUX) 70.
According to the embodiment, the short term redundancy of an input speech signal, "s", is eliminated by the LPC prediction error filter 20. The resulting LPC prediction residual signal, "d", still has some long-term redundancy, due to pitch periodicity in the voiced voice. The long-term redundancy is then eliminated by the tone prediction processor 30. After pitch prediction, the final prediction residual signal, "e", is transformed into the frequency domain, by the transformation processor 40, which establishes a Fast Fourier Transformation (FFT). The adapter bit allocation is applied by the residual quantizer 60, to allocate bits to the residual FFT prediction coefficients, according to their perception importance, determined by the audition model quantizer control processor 50.
The codebook ondices, which represent (a) the paraometers of the predictor (forecaster) of LPC (ii); (b) the pitch predictor (forecaster) paraometers (ip, ii); (c) the transform gain levels (ig); and (d) the residual quantized prediction (ir), are multiplexed into a stream of bits and transmitted, over one channel, to a decoder, as side information.
The channel may comprise any suitable communication channel, including wireless channels, computer and data networks, telephone networks; and may include memory, such as solid-state memories (eg, semiconductor memories), optical memory systems (such as CD-ROM), magneto memories (eg, disk memories), and the like.
Basically, the TPC decoder reverses the operations performed in the encoder. The decoder decodes the LPC predictor (forecaster) paraometers, pitch predictor paraometers, gain levels, and FFT coefficients of the residual prediction. The decoded FFT coefficients are transformed again in the time domain, by applying an inverse FFT. The resulting decoded residual prediction is then passed through a pitch sonthetic filter and an LPC sonthetic filter to reconstruct the speech signal.
To keep complexity as low as possible, TPC employs open-loop quantization. Open-circuit quantization means that the quantizer tries to minimize the difference between the unquantized parameter and its quantized version, regardless of the effects on the quality of the output speech. This was done in contrast to CELP encoders, for example, where the pitch predictor, gain, and excitation are normally closed-loop quantized. In closed-loop quantization of an encoder parameter, the quantizer's codebook search attempts to minimize distortion in the final reconstructed output voice. Naturally, this generally leads to better quality of the output speech, but at the cost of greater codebook search complexity.
B. Illustrative Encoder Embodiment
1. LPC analysis and prediction
A detailed block diagram of the LPC analysis processor 10 is depicted in Figure 2. Processor 10 comprises an autocorrelation and windowing processor 210; a white noise correction and spectral smoothing processor 215; a Levinson-Durbin recursion processor 220; a bandwidth expansion processor 225; an LPC to LSP conversion processor 230; and an LPC power spectrum processor 235; an LSP quantizer 240; an LSP classification processor 245; an LSP interpolation processor 250; and an LSP to LPC conversion processor 255.
The autocorrelation and windowing processor 210 begins the LPC coefficient generation process. Processor 210 generates autocorrelation coefficients, "r", in a conventional manner, once every 20 ms, of which the LPC coefficients are subsequently calculated, as discussed below. Vow Rabiner,
ES 2 174 030 T3
LR et al., Digital Voice Signal Processing, Prentice-Hall, Inc., Englewood Cliffs, New Jersey, 1978, (Rabiner et al.). The LPC frame size is 20 ms (or 320 speech samples at 16 kHz sampling rate). Each 20 ms frame is further divided into 5 subframes, each 4 ms (or 64 samples) in length. The LPC analytics processor used a 24 ms Hamming window, which was centered on the last 4 ms subframe of the current frame, in a conventional way.
To alleviate poor conditioning, some conventional signal conditioning techniques are employed. Prior to LPC analysis, the spectral smoothing and white noise correction processor 215 applies a spectral smoothing technique (SST) and a white noise correction technique. The SST, well known in the art (Tohkura, Y. et al., "Spectral Smoothing Technique in Parcor Voice AnalysisSynthesis", lEEE Trans. Acust. Voice Signal Processing, ASSP-26: 587-596, December 1978 (Tohkura et al.)) Involves the multiplication of a calculated matrix of autocorrelation coefficients (from processor 210) by a Gaussian window whose Fourier transformation corresponds to a function probability density (pdf) of a Gaussian distribution with a topical deviation of 40 Hz. White noise correction, also conventional (Chen, J.-H., "A Robust 16 kbit / s Low-Delay CELP Voice Encoder, Proc. IEEE Global Communications Conference, p. 1237-1241, Dallas, TX, November 1989) increases the zero lag autocorrelation coefficient (that is, the energy term) by 0.001%.
Next, the coefficients generated by processor 215 are provided to the Levinson-Durbin recursion processor 220, which generates 16 LPC coefficients, "ai" for i = 1, 2, ..., 16 (the order of predictor 20 of LPC is 16) conventionally.
The bandwidth expansion processor 225 multiplies each "ai" by a factor g ', where g' = 0.994, for additional signal conditioning. This corresponds to a 30 Hz bandwidth expansion (Tohkura et al.).
After such bandwidth expansion, the LPC predictor coefficients are converted to Lone Spectral Pair (LSP) coefficients by the LPC to LSP conversion processor 230, in a conventional manner. See Soong, FK et al., "Lonea Spectrum Pair (LSP) and Voice Data Compression", Proc. IEEE Int. Conf. Acust. Voice Signals Processing, póag. 1.10.1-1.10.4, March 1984 (Soong et al.).
Next, vector quantization (VQ) is provided by vector quantizer 240, to quantify the resulting LSP coefficients. The specific VQ technique, employed by the processor 240, is similar to the split VQ proposed in Paliwal, KK et al., "Efficient Quantification of LPC Parameter Vectors at 24 bits / frame", Proc. IEEE Int. Conf. Acust. Voice Signals Processing, póag. 661664, Toronto, Canadaó, May 1991 (Paliwal et al.). The 16-dimensional LSP vector is divided into 7 smaller sub-vectors that have the dimensions of 2, 2, 2, 2, 2, 3, 3, counting from the low-frequency end. Each of the 7 sub-vectors is quantized to 7 bits (ie, using a VQ codebook of 128 code vectors). Therefore, there are seven codebook onxes, ii (1) -ii (7), each ondx being seven bits long, for a total of 49 bits per frame used in the quantization of LPC parameters. These 49 bits are provided to MUX 70 for transmission to the decoder as side information.
Processor 240 searched through the VQ codebook, using a conventional weighted mean square error distortion (WMSE) measure, as described in Paliwal et al. The codebook used is determined with conventional codebook generation techniques, well known in the art. A conventional MSE distortion measure can also be used instead of the WMSE measure, to reduce encoder complexity, without too much degradation in the quality of the output speech.
Typically, the LSP coefficients increase monotonically. However, quantification can result in disorganization of this order. This disorganization results in an unstable LPC synthesis filter in the decoder. To avoid this problem, the LSP sort processor 245 sorts the quantized LSP coefficients, to restore monotonic increasing order and to ensure stability.
The quantized LSP coefficients are used in the last subframe of the current frame. Linear interpolation between these LSP coefficients and those of the last subframe of the above frame, to provide the LSP coefficients for the first four subframes, was performed by the LSP interpolation processor 250, as is conventional. The interpolated and quantized LSP coefficients are then converted back to the LPC predictor coefficients, for use in each subframe, by the LSP to LPC conversion processor 225 in a conventional manner. This was done in both the encoder and the decoder. LSP interpolation is important to maintain uniform reproduction of the output voice. LSP interpolation allows the LPC predictor to update only once in a subframe (4 ms), in a uniform manner. The resulting LPC predictor 20 is used to predict the input signal to the encoder. The difference between the input signal and its predicted version is the residual prediction of LPC, "d".
two. Pitch prediction
The pitch prediction processor 30 comprises a pitch extraction processor 410, a pitch bypass quantizer 415, and a three-tap pitch prediction error filter 420, as depicted in Figure 3. The processor 30 is used to eliminate redundancy in the residual prediction of LPC, "d", due to the pitch periodicity in the voiced voice. The is5
ES 2 174 030 T3 pitch timing used by processor 30 It is only updated once in a frame (once every 20 ms). There are two kinds of paraometers, in pitch prediction, that need to be quantized and transmitted to the decoder: the pitch period corresponding to the period of the quasi-periodic waveform of the voiced voice and the three pitch predictor coefficients (taps) .
The pitch period of the residual LPC prediction is determined by the pitch extraction processor 410, using a modified version of the efficient two-step search technique, discussed in US Patent No. 5,327,520, entitled " Use of Voice Message Encoder / Decoder ”. Processor 410 first passes the LPC residual prediction through a third-order low-pass eloptic filter to limit the bandwidth to approximately 800 Hz, and then performs an 8: 1 decimation of the output of the LPC. low pass filter. The decimated signal correlation coefficients are calculated for time lags ranging from 4 to 35, which corresponds to time lags of 32-280 samples in the domain of the non-decimated signal. Therefore, the allowable range for the tone period is 2 ms to 17.5 ms, or 57 Hz to 500 Hz, in terms of tone frequency. This is sufficient to cover the normal pitch range of essentially all speakers, including low-pitched men and high-pitched boys.
After the processor 410 has calculated the correlation coefficients from the decimated signal, the first major peak of the correlation coefficients that has the least time delay is identified. This is the first step search. Let “t” be the resulting time delay. This value of "t" is multiplied by 8 to obtain the time delay in the domain of the signal without decimating. The resulting time delay, 8t, points to the neighborhoods most likely to be to the true pitch period. To retain the original time resolution in the non-decimating signal domain, a second pitch search step is performed in the range of t-7 to t + 7. The correlation coefficients of the original residual prediction of the untreated LPC, "d", are calculated for the time lags from t-7 to t + 7 (subject to the lower limit of 32 samples and the upper limit of 280 samples). Then, the time delay corresponding to the maximum correlation coefficient in this interval is identified as the final pitch period, "p". This pitch period, "p", is encoded in 8 bits, with a conventional VQ codebook, and the 8-bit codebook index, "ip", is provided to the MUX 70, for transmission to the decoder, as lateral information. To represent the pitch period, eight bits are enough, since there are only 280-32 + 1 = 249 possible integers that can be selected as the pitch period.
The three taps of the pitch predictor or forecaster are jointly determined in quantized form, by the pitch tap quantizer 415. Quantizer 415 comprises a conventional VQ codebook having 64 code vectors, representing 64 possible sets of pitch predictor taps. The energy of the residual pitch prediction, within the current frame, is used as the distortion measure of a search through the codebook. Such a distortion measure provides a higher pitch prediction gain than a simple MSE measure over the same predictor taps. Normally, with this distortion measure, the complexity of the codebook search would be very high if a brute force approach were used. However, quantizer 415 employs an efficient codebook search technique, which is well known in the art (described in US Patent No. 5,327,520) for this distortion measure.
It can be shown that the minimization of the residual energy distortion measure is equivalent to the maximization of an inner product of two 9-dimensional vectors. One of these 9-dimensional vectors only contains correlation coefficients of the residual prediction of LPC. The other 9-dimensional vector only contains the product terms derived from the set of three takes of the pitch predictor that are under evaluation. As such vector is independent of the signal and only depends on the pitch code vector, there are only 64 possible vectors of this type (one for each pitch code vector) and these vectors can be pre-calculated and stored in a table in the VQ codebook. In an actual codebook search, the 9-dimensional vector of the residual correlation of LPC is first computed. Next, the inner product of the resulting vector is calculated with each of the 64 pre-calculated and stored 9-dimensional vectors. The vector in the stored table that provides the maximum inner product is the winner, and the three quantized pitch predictor taps are derived from it. Since there are 64 vectors in the stored table, a 6-bit ondx, "ii", is sufficient to represent the three quantized taps of the pitch predictor. These 6 bits are provided to the MUX 70 for transmission to the decoder, as side information.
The quantized pitch period and pitch predictor taps, determined as discussed above, are used to update pitch prediction error filter 420 once per frame. The quantized pitch period and the pitch predictor taps are used by filter 420 to predict the residual LPC prediction. Then, the predicted residual LPC prediction is subtracted from the actual residual prediction LPC. After the predicted version is subtracted from the unquantized LPC residual, we have the unquantized residual pitch prediction, "e", which was encoded using the transformation coding approach, described later.
3. Transformation coding of residual prediction
The residual pitch prediction signal, "e", is encoded, frame by frame, by the transform processor 40. Figure 4 represents a detailed block diagram of the
ES 2 174 030 T3 processor 40. Processor 40 comprises: an FFT processor 510, a gain processor 520, a gain quantizer 530, a gain interpolation processor 540, and a normalization processor 550.
The FFT processor 510 calculates a conventional 64-point FFT for each subframe of the residual pitch prediction, "e". This size transformation avoids the so-called "pre-echo" distortion, which is well known in the art of audio coding. See Jayant, N. et al., "Signal Compression Based on Models of Human Perception", Proc. IEEE, póag. 1385-1422, October 1993.
to. Gain calculation and quantification
After each 4 ms subframe of the residual prediction is transformed to the frequency domain, by processor 510, gain levels (or RMS) are extracted by gain processor 520, and quantized by the gain quantizer 530, for the different frequency bands. For each of the five subframes of the current frame, two gain values are extracted by processor 520: (1) the RMS value of the first five FFT coefficients of processor 510 as a low frequency gain (0 to 1 kHz) , and (2) the RMS value of the twenty-seventh to twenty-ninth FFT coefficients of the processor 510, as a high frequency gain (4 to 7 kHz). Therefore, 2x5 = 10 gain values per frame are extracted, for use by the gain quantizer 530.
Gain quantizer 530 employs separate quantization schemes for the high-frequency and low-frequency gains, in each frame. For high frequency gains (4-7 kHz), quantizer 530 encodes the high frequency gain of the last subframe of the current frame into 5 bits, using a conventional scalar quantization. This quantized gain is then converted by quantizer 530 into the logarithmic domain in terms of decibels (dB). Since there are only 32 possible levels of quantized gain (with 5 bits), the corresponding 32 logarotic gains are precalculated and stored in a table, and the gain conversion from the linear domain to the log domain was done by searching the table. The quantizer 530 then performed the log-domain linear interpolation between this resulting log gain and the log gain of the last subframe of the last frame. Such interpolation produces an approximation (that is, a prediction) of the log gains for sub-frames 1 through 4. Next, the linear gains of sub-frames 1 through 4, supplied by the gain processor 520, are converted to the logarotmic domain, and the interpolated log gains are subtracted from the results. This produces 4 log-gain interpolation errors, which are grouped into two vectors of 2 dimensions each.
Then, each 2-dimensional log gain interpolation error vector is conventionally quantized to 7 bits, using a simple MSE distortion measure. The two 7-bit codebook onxes, in addition to the 5-bit scalar representing the last subframe of the current frame, are provided to MUX 70, for transmission to the decoder.
Gain quantizer 530 also adds back the 4 quantized log gain interpolation errors to the 4 interpolated log gains, to obtain the quantized log gains. These 4 quantized log gains are then converted back to the linear domain to obtain the 4 quantized high frequency gains for subframes 1 through 4. These high frequency quantized gains, along with the high frequency quantized gain of subframe 5, are provided to the gain interpolation processor 540, for processing as described later.
Gain quantizer 530 quantized the low-frequency gains (0-1 kHz), based on the quantized high-frequency gains and the quantized taps of the pitch predictor. The statistic of the log gain difference, which is obtained by subtracting the high frequency log gain from the low frequency log gain, of the same subframe, was strongly influenced by the pitch predictor. For frames that do not have much pitch periodicity, the log gain difference should have a mean of approximately zero and have a smaller topical deviation. On the other hand, for frames that have strong pitch periodicity, the log gain difference should have a large negative mean and a larger topical deviation. This observation forms the basis of an efficient quantizer for the 5 low-frequency gains of each frame.
For each of the 64 possible quantized sets of pitch predictor takes, the conditional mean and conditional topical deviation of the log gain difference are precalculated using a large speech database. The resulting 64-input tables are then used by the gain quantizer 530 in quantizing the low-frequency gains.
The low frequency gain of the last subframe is quantized as follows. The codebook ondx, obtained while the takes of the pitch predictor are quantized, is used in table lookup operations to extract the conditional mean and conditional topical deviation of the log gain difference for that particular quantized set. of Pitch Predictor Taps. Then the log gain difference of the last subframe is calculated. The conditional mean is subtracted from this unquantified log gain difference and the resulting difference from the log gain, after subtracting the mean, is divided by the conditional topical deviation. This operation basically produces an amount of unit variance, with zero on average, which is quantized to 4 bits by the gain quantizer 530, using scalar quantization.
Then, the quantized value is multiplied by the conditional topical deviation, and the result is added to the conditional mean to obtain a
ES 2 174 030 T3 quantized log gain difference. The quantized high-frequency log gain is then added back in to obtain the quantized low-frequency log gain of the last subframe. The resulting value is then used to perform a linear interpolation of the logarotmic low-frequency gain for subframes 1 to 4. This interpolation occurs between the logarothmic low-frequency quantized gain of the last subframe of the previous frame and the logarotmic low-frequency quantized gain of the last subframe of the current frame.
Then, the 4 low-frequency logarotmic gain interpolation errors are calculated. First, the linear gains provided by the gain processor 520 are converted to the log domain. Next, the interpolated logarotic low-frequency gains are subtracted from the converted gains. Errors resulting from the log gain interpolation are normalized by the conditional topical deviation of the log gain difference. Then, the normalized interpolation errors are grouped into two 2-dimensional vectors. Each of these two vectors is quantized to 7 bits, using a simple MSE distortion measure, similar to the VQ scheme for the high frequency case. The two 7-bit codebook onxes, in addition to the 4-bit scalar representing the last subframe of the current frame, are provided to MUX 70, for transmission to the decoder.
The gain quantizer also multiplies the 4 quantized values by the conditional topical deviation, to restore the original scale, and then adds the interpolated log gain to the result. The resulting values are the quantized logarotic low-frequency gains for subframes 1 through 4. Finally, the 5 quantized low-frequency log gains are converted to the linear domain, for later use by the gain interpolation processor 540.
Gain interpolation processor 540 determines approximate gains for the 1 to 4 kHz frequency band. First, the gain levels for the thirteenth to sixteenth FFT coefficients (3 to 4 kHz) are chosen to be the same as the quantized high-frequency gain. The gain levels for the sixth to twelfth second FFT coefficients (1 to 3 kHz) are then obtained by linear interpolation between the quantized logarothmic low-frequency gain and the quantized high-frequency log gain. The resulting interpolated log gain values are then converted back to the linear domain. Therefore, with the termination of the gain interpolation processor process, each FFT coefficient from 0 to 7 kHz (or FFT coefficients 1st to 29th) has a gain, quantized or interpolated, associated with it. A vector of these gain values is provided to the gain normalization processor 550 for further processing.
Normalization processor 550 normalizes the FFT coefficients generated by FFT processor 510, dividing each coefficient by its corresponding gain. Then, the resulting normalized gain FFT coefficients are arranged to be quantized by the residual quantizer 60.
b. Bit stream
Figure 7 depicts the bit stream of the illustrative embodiment of the present invention. As described above, 49 bits / frame have been allocated to encode LPC parameters, 8 + 6 = 14 bits / frame have been allocated for the 3-socket tone predictor, and 5 + (2x7) +4+ (2x7) = 37 bits / framework for earnings. Therefore, the total number of lateral information bits is 49 + 14 + 37 = 100 bits per 20 ms frame, or 20 bits per 4 ms subframe. Considering that the encoder could be used at one of three different speeds: 16, 24 and 32 kb / s, at a sampling rate of 16 kHz, these three target rates translate to 1, 1.5 and 2 bits / sample, or 64, 96 and 128 bits / subframe, respectively. With 20 bits / frame used for side information, the number of bits that remain for use in encoding the main information (encoding of FFT coefficients) are 44, 76 and 108 bits / frame for the three rates of 16, 24 and 32 kb / s, respectively.
c. Adapter Bit Mapping
In accordance with the principles of the present invention, the adapter bit allocation was carried out to allocate these remaining bits to various parts of the frequency spectrum with different quantization precision, in order to increase the perception quality of the output voice of the device. TPC decoder. This was done using a human sensitivity model for noise in audio signals. Such models are known in the art of perception audio coding. See, for example, Tobias, JV, ed., Foundations of Modern Hearing Theory, Academic Press, New York and London, 1970. See also Schroeder, MR et al., “Optimizing Digital Voice Encoders by Exploiting the Masking Properties of the Human Ear ”, J. Acoust. Soc. Amer., 66: 1647-1652, December 1979 (Schroeder et al.).
The audition model and quantizer control processor 50 comprises an LPC power spectrum processor 510, a masking threshold processor 515, and a bit allocation processor 520. Although the adapter bit allocation could be performed once each subframe, the illustrative embodiment of the present invention performed the bit allocation once per frame, in order to reduce the complexity of its computation.
Instead of using the unquantized input signal, to derive the masking noise threshold and bit allocation, as was done in conventional muosic encoders, the masking noise threshold and bit allocation of the illustrative embodiment are determined based on the frequency response of the quantized LPC synthesis filter (which is often referred to as the “es8
ES 2 174 030 T3 LPC spectrum "). The LPC spectrum can be considered as an approximation of the spectral envelope of the input signal within the 24 ms window of LPC analysis. The LPC spectrum is determined based on the quantized LPC coefficients. The quantized LPC coefficients are provided by the LPC analysis processor 10 to the LPC spectrum processor 510 of the audicon model and quantizer control processor 50. Processor 510 determines the LPC spectrum, as noted below. The quantized LPC filter coefficients (a;) are first transformed by a 64-point FFT. The power of the first 33 FFT coefficients is determined and the reciprocals of these power values are then calculated. The result is the LPC power spectrum that has the frequency resolution of a 64-point FFT.
After the LPC power spectrum is determined, an estimated masking noise threshold is calculated by the masking threshold processor 515. The masking threshold, "TM", is calculated using a modified version of the procedure described in US Patent No. 5,314,457. Processor 515 scales the 33 LPC power spectrum samples from processor 510, using a frequency-dependent attenuation function, empirically determined from subjective listening experiments. As represented in Figure 6, the attenuation function starts at 12 dB for the DC term of the LPC power spectrum, increases to approximately 15 dB between 700 and 800 Hz, then decreases monotone towards high frequencies, and finally decreases to 6 dB at 8000 Hz.
Then, each of the 33 attenuated LPC power spectrum samples is used to grade a "basilar membrane diffusion function" derived for that particular frequency, in order to calculate the masking threshold. A diffusion function for a specific frequency corresponds to the shape of the masking threshold, in response to a single-tone masker signal at that frequency. Equation (5) of Schroeder et al. Describing such diffusion functions in terms of the "loud cough" frequency scale or the critical band frequency scale is incorporated by reference as fully disclosed herein. The grading process begins with the first 33 frequencies of a 64-point FFT, through 0-16 kHz (i.e., 0 Hz, 250 Hz, 500 Hz, ..., 8000 Hz) which are converted to scale. frequencies of "strong cough". Then, for each of the 33 resulting strong cough values, the corresponding diffusion function is sampled in these 33 strong cough values, using equation (5) of Schroeder et al. The resulting 33 broadcast functions are stored in a table that can be run as part of an off-line process. To calculate the estimated masking threshold, each of the 33 diffusion functions is multiplied by the corresponding sample value of the LPC attenuated power spectrum, and the resulting 33 graded diffusion functions are added together. The result is the estimated masking threshold function that is provided to the bit allocation processor 520. Figure 9 depicts the process performed by processor 520 to determine the estimated masking threshold function.
It should be noted that this technique for estimating the masking threshold is not the only technique available.
To keep complexity low, bit allocation processor 520 uses a "greed" technique to allocate bits for residual quantization. The technique is "greedy" in the sense that one bit is assigned at a time to the most "needed" frequency component, regardless of its potential influence on future bit assignments.
At first, when no bits have been assigned yet, the corresponding output speech will be zero, and the encoding error signal is the input speech itself. Therefore, the LPC power spectrum is initially considered to be the coding noise power spectrum. The noise height at each of the 33 64-point FFT frequencies is then estimated, using the masking threshold calculated earlier and a simplified version of the noise height calculation method from Schroeder et al.
The height of the simplified noise at each of the 33 frequencies is calculated by processor 520, as indicated below. First, the critical bandwidth "Bi" is calculated at the frequency "i<sup>esima</sup>”, Using linear interpolation of the critical bandwidth listed in Table 1 of Scharf's book chapter on Tobias. The result is the approximate value of the term df / dx in equation (3) of Schroeder et al. The 33 critical bandwidth values are pre-calculated and stored in a table. Then, for the frequency “i<sup>esima</sup>”, The power Ni of the noise is compared with the threshold M<sub>i</sub> masking. Without<sub>i</sub> <M<sub>i</sub>, the noise height Li is set to zero. If Ni> Mi, then the noise height is calculated as:
Li = Bi ((Ni-Mi) / (1+ (Si / Ni)<sup>2</sup>))<sup>0,25</sup> where Si is the sample value of the LPC power spectrum at frequency i<sup>esima</sup>.
Once the noise height has been calculated, by the processor 520, for the 33 frequencies, the frequency having the maximum noise height is identified and a bit is assigned to this frequency. Then, the noise power at this frequency is reduced by a factor that is determined empirically from the signal-to-noise ratio (SNR) obtained during the design of the VQ codebook to quantify the prediction of residual FFT coefficients. (Illustrative values for the reduction factor are between 4 and 5 dB). The noise height is then updated at this frequency, using the reduced noise power. Next, the maximum in the series of updated noise heights is again identified, and one bit is assigned to the corresponding frequency. This process con9
ES 2 continues until all available bits have been used up.
For the 32 and 24 kb / s TPC encoder each of the 33 frequencies can receive bits during adapter bit allocation. On the other hand, for the 16 kb / s TPC encoder, better voice quality can be achieved if the encoder only allocates bits to the frequency range 0 to 4 kHz (that is, to the first 16 FFT coefficients) and synthesizes the residual FFT coefficients in the higher frequency band from 4 to 8 kHz. The procedure for synthesizing the 4 to 8 kHz FFT residual coefficients will be described later in connection with the illustrative decoder.
Note that since the quantized LPC synthesis coefficients (a;) are also available in the TPC decoder, there is no need to transmit the bit allocation information. This bit allocation information is determined by an exact reproduction of the decoder's audition model quantizer control processor 50. Therefore, the TPC decoder can locally duplicate the encoder's adapter bit allocation operation to obtain such bit allocation information.
d. Quantification of FFT coefficients
Once bit allocation was done, quantizer 60 performed actual quantization of the normalized prediction FFT residual coefficients, "EN". The DC term of the FFT is a real number, and it is scalar quantized if it always receives any bit during bit allocation. The maximum number of bits it can receive is 4. For the second to sixteenth FFT coefficients, a conventional two-dimensional vector quantizer is used to jointly quantize the real and imaginary parts. The maximum number of bits for this 2-dimensional VQ is 6 bits. For the twelfth to thirtieth FFT coefficients, a conventional 4-dimensional vector quantizer is used to quantize the real and imaginary parts of two adjacent FFT coefficients.
C. Illustrative realization of decoder
An illustrative embodiment of the decoder of the present invention is depicted in Figure 8. The illustrative decoder comprises a demultiplexer (DEMUX) 65, an LPC parameter decoder 80, a hearing model dequantizer control processor 90, a dequantizer 70, a reverse transform processor 100, a tone sonhesis filter 110, an LPC synthesis filter 120, connected as shown in Figure 8. As a general proposition, the decoder realization performed the inverse operations of the operations executed by the illustrative main information encoder.
For each frame, the DEMUX 65 separates all the main and side information components from the received bit stream. The main information is provided to the dequantizer 70. The term "dequantize" used herein refers to the generation of a quantized output based on an encoded T3 18 value, such as an ondx. To de-quantify this main information, an adapter bit allocation must be performed to determine the number of main information bits that was associated with each main information quantized transformation coefficient.
The first stage in the adapter bit allocation is the generation of quantized LPC coefficients (on which the allocation depends). As discussed above, seven LSP codebook ondices, ii (1) -ii (7), are communicated over the channel to the decoder to represent the quantized LSP coefficients. The quantized LSP coefficients are synthesized by decoder 80 using a copy of the LSP codebook (previously discussed) in response to the LSP onxes received from DEMUX 65. Finally, the LPC coefficients are derived from the coefficients of LSP in a conventional way.
With the LPC coefficients, "a", synthesized, the hearing model dequantizer control processor 90 determines the bit allocation (based on the quantized LPC paraometers) for each FFT coefficient, in the same way as above. discussed, in relation to the encoder. Once the bit allocation information has been derived, the dequantizer 70 can correctly decode the main FFT coefficient information and obtain the quantized versions of the normalized gain prediction FFT residual coefficients.
For frequencies that do not receive any bits, the decoded FFT coefficients will be zero. The locations of such "spectral gaps" evolve over time, and this can result in a distinct artificial distortion that is completely common to many transform coders. To avoid such artificial distortion, the dequantizer 70 "fills" the spectral gaps with low-level FFT coefficients that have random phases and magnitudes equal to 3 dB below the quantized gain.
For the 32 and 24 kb / s encoders, bit allocation was done for the entire frequency band, as described earlier in the encoder discussion. For the 16 kb / s encoder, the bit allocation was restricted to the 0 to 4 kHz band. The 4 to 8 kHz band is synthesized as follows. First, the ratio between the LPC power spectrum and the masking threshold, or the signal-to-threshold-masking ratio (SMR), is calculated for the frequencies from 4 to 7 kHz. The twenty-seventh to twenty-ninth FFT coefficients (4 to 7 kHZ) are synthesized using phases that are random and magnitude values that are controlled by the SMR. For frequencies that have SMR> 5dB, the magnitude of the residual FFT coefficients is set to 4 dB above the quantized high-frequency gain (RMS value of FFT coefficients in the 4 to 7 kHz band). For frequencies that have SMR <5 dB, the magnitude is 3 dB below the quantized high-frequency gain. From the thirtieth to thirty-third coefficients of FFT, the magnitude falls from
ES 2 174 030 T3 dB at 30 dB below the quantized high-frequency gain, and the phase is again random. Figure 10 illustrates the process that synthesizes the magnitude and phase of the coefficients of
FFT.
Once all the FFT coefficients have been decoded, filled in, or synthesized, they are ready for grading. The grading was performed by the inverse transformation processor 100, which receives (from DEMUX 65) a 5-bit ondx for the high-frequency gain and a 4-bit onyx for the low-frequency gain, each corresponding to the last subframe of the current frame, as well as ondices for the logarotmic gain interpolation errors for the low and high frequency bands of the first four subframes. These gain indices are decoded, and the results are used to obtain the grading factor for each FFT coefficient, as described earlier in the section describing the calculation and quantification of gain. So the FFT coefficients are graded by your individual earnings.
The resulting quantized, gain-graded FFT coefficients are then transformed back to the time domain by inverse transform processor 100 using an inverse FFT. This inverse transformation produces the time domain quantized residual prediction, "d".
Then, the time domain quantized residual prediction, "d," is passed through the pitch synthesis filter 110. Filter 110 adds pitch periodicity to the residue based on a quantized pitch period, "p", to produce "d", the quantized LPC residual prediction. The quantized pitch period is decoded from the 8-bit onyx, "ip", obtained from DEMUX 65. The pitch predictor taps are decoded from the 6-bit onyx, "it," also obtained from DEMUX 65.
Finally, the quantized output speech, "d", is generated by the LPC synthesis filter 120, using the quantized LPC coefficients, "d", obtained from the LPC parameter decoder 80.
D. Discussion
Although various specific embodiments of this invention have been presented and described herein, it is to be understood that these embodiments are merely illustrative of the many possible specific arrangements that can be envisioned applying the principles of the invention. In light of the above description, those skilled in the art can envision other arrangements, numerous and varied, in accordance with these principles, without departing from the scope of the invention, which is defined by the appended claims.
For example, correct voice and music quality can be maintained by uniquely encoding the FFT phase information in the 4 to 7 kHz band for frequencies where SMR> 5 dB. The magnitude is determined in the same way as that of the high-frequency synthesis procedure described near the end of the bit allocation discussion.
Most CELP encoders update the pitch predictor parameters once every 4 to 6 ms, to achieve a more efficient pitch prediction. This is much more frequent than the 20 ms updates of the illustrative embodiment of the TPC encoder. As such, other update rates are also possible, eg every 10 ms.
Other forms of noise height estimation can also be used. Furthermore, instead of minimizing the maximum noise height, the sum of the noise heights for all frequencies can be minimized. The gain quantization scheme, described earlier in the encoder section, has reasonably good coding efficiency and works well for speech signals. An alternative gain quantization scheme is described later. It may not have as good a coding efficiency, but it is considerably simpler and can be more robust for non-voice signals.
The alternative scheme begins with the calculation of a "frame gain", which is the RMS value of the time domain pitch prediction residual signal, calculated over the entire frame. This value is then converted to dB values and quantized to 5 bits with a scalar quantizer. For each subframe, three gain values of the residual FFT coefficients are calculated. The low frequency gain and the high frequency gain are calculated in the same way as above, that is, the RMS value of the first 5 FFT coefficients and the RMS value of the seventh to twenty-ninth FFT coefficients. In addition, the mean frequency gain is calculated as the RMS value of the sixth to sixteenth FFT coefficients. These three gain values are converted to dB values, and the frame gain in dB is subtracted from them. The result is the obtaining of the normalized subframe gains for the three frequency bands.
The normalized gain of the low-frequency subframe is quantized with a 4-bit scalar quantizer. The normalized subframe gains, mid-frequency and high-frequency, are quantized together by a 7-bit vector quantizer. To obtain the quantized gains of the subframe in the linear domain, the frame gain in dB is added back to the quantized version of the normalized subframe gains, and the result is converted back to the linear domain.
Unlike the previous procedure, in which linear interpolation was performed to obtain the gains for the frequency band from 1 to 4 kHz, this alternative procedure does not require such interpolation. Each residual FFT coefficient belongs to one of the three frequency bands from which a dedicated subframe gain has been determined. Each of the three quantized subframe gains in the linear domain is used to normalize or scale all residual FFT coefficients in the frequency band from which the subframe gain has been derived.
ES 2 174 030 T3
Note that this alternative gain quantization scheme takes more bits to specify all gains. Therefore, for a specific bit rate, fewer bits are available to quantize the residual FFT coefficients.
Contents5
4 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4
11 members in 7 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 19950530980 | United States of America | – | |
| 53098095 | United States of America | A | |
| 53098095 | United States of America | A | |
| 96306736 | – | – | – |
| US19950530980 | – | – | – |
Members11
| Document | Office | Kind | |
|---|---|---|---|
| CA2185731A1 | Canada | A1 | |
| EP0764941A2 | European Patent Office (EPO) | A2 | |
| JPH09152900A | Japan | A | |
| MX9604161A | Mexico | A | |
| US5710863A | United States of America | A | |
| EP0764941A3 | European Patent Office (EPO) | A3 | |
| CA2185731C | Canada | C | |
| EP0764941B1 | European Patent Office (EPO) | B1 | |
| DE69621393D1 | Germany | D1 | |
| ES2174030T3This record | Spain | T3 | |
| DE69621393T2 | Germany | T2 |
1 legal event, as the office reported them to INPADOC
Events
| Event | Code | |
|---|---|---|
| Definitive protectionFG2A | FG2A |
Numbers
- Publication
- 2174030
- Publication, DOCDB
- 2174030
- Publication, EPODOC
- ES2174030T
- Application
- 96306736
- Application, DOCDB
- 96306736
- Application, EPODOC
- ES19960306736T
Titles2
- Spanish
- CUANTIFICACION DE SEÑAL DE VOZ UTILIZANDO MODELOS DE AUDICION HUMANA EN SISTEMAS DE CODIFICACION PREDICTIVA.
- English
- QUANTIFICATION OF VOICE SIGNAL USING HUMAN HEARING MODELS IN PREDICTIVE CODING SYSTEMS.
Classification
- CPC, 8
- G10L19/0212
- G10L19/002
- G10L19/06
- G10L25/24
- G10L25/27
- G10L2019/0003
- G10L2019/0011
- G10L2019/0013
- IPC, 5
- G10L19 04
- G10L19 00
- G10L19 02
- G10L19 06
- H03M7 30