Variable rate vocoder
Abstract
Method for compressing voice signal, by means of variable rate coding of frames of digitized voice samples, comprising the following steps: determining a level of voice activity for a frame of digitized voice samples; selecting a coding rate from a set of speeds based on said determined level of voice activity for said frame; encoding said frame according to an encoding format of a set of encoding formats for said selected speed in which each speed has a corresponding different encoding format and in which each encoding format provides a different plurality of parameter signals representing said digitized voice samples (s (n) 9 according to a voice model; and generating for said frame a data packet of said parameter signals, characterized in that: it provides a speed-linked control indicating a preselected encoding rate for said frame; and modifies said selected encoding rate to provide said preselected encoding rate to encode said frame at said preselected encoding rate.

Term
Term ended
Projected expiry passed 3 June 2012, 14.3 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
37 claims: 12 independent, 25 dependent
- 1ES 2 240 252 T3 REIVINDICACIONES 1. Método para la compresión de señal de voz, mediante la codificación de velocidad variable de tramas de muestras de voz digitalizadas, que comprende las etapas siguientes:determinar un nivel de actividad de voz para una trama de muestras de voz digitalizadas;seleccionar una velocidad de codificación a partir de un conjunto de velocidades sobre la base de dicho nivel determinado de actividad de voz para dicha trama;codificar dicha trama según un formato de codificación de un conjunto de formatos de codificación para dicha velocidad seleccionada en el que cada velocidad presenta un formato de codificación diferente correspondiente y en el que cada formato de codificación proporciona una diferente pluralidad de señales de parámetros que representan dichas muestras de voz digitalizadas (s(n)9 según un modelo de voz;y generar para dicha trama un paquete de datos de dichas señales de parámetros, caracterizado porque: proporciona un control ligado a la velocidad que indica una velocidad de codificación preseleccionada para dicha trama;y modifica dicha velocidad de codificación seleccionada para proporcionar dicha velocidad de codificación preseleccionada para codificar dicha trama a dicha velocidad de codificación preseleccionada.
- 2Método según cualquiera de las reivindicaciones anteriores, en el que dicha etapa de determinación de dicho nivel de actividad de voz de trama comprende las etapas siguientes:medir la actividad de voz en dicha trama de muestras de voz digitalizadas;comparar dicha actividad de voz medida con por lo menos un nivel umbral de actividad de voz de un conjunto predeterminado de niveles umbral de actividad;y ajustar de forma adaptable en respuesta a dicha comparación por lo menos uno de dicho por lo menos un nivel umbral de actividad de voz con respecto a un nivel de actividad de una trama anterior de muestras de voz digitalizadas.
- 3Método según la reivindicación 1 ó 2, en el que dicha velocidad preseleccionada es inferior a una velocidad máxima predeterminada, comprendiendo además dicho método las etapas siguientes:proporcionar un paquete de datos adicional;y combinar dicho paquete de datos con dicho paquete de datos adicional dentro de una trama de transmisión para transmisión.
- 4Método según cualquiera de las reivindicaciones anteriores, en el que dicha etapa de provisión de dicho paquete de datos de dichas señales de parámetro comprende:generar un número variable de bits para representar las señales vectoriales de coeficiente predictivo lineal (LPC) de dicha trama de muestras de voz digitalizadas, en el que dicho número variable de bits que representa dichas señales vectoriales LPC está determinado por dicho nivel de actividad de voz apreciado;generar un número variable de bits para representar señales vectoriales tono de dicha trama de muestras de voz digitalizadas, en el que dicho número variable de bits que representa dichas señales vectoriales de tono está determinado por dicho nivel de actividad de voz apreciado;y generar un número variable de bits para representar señales vectoriales de excitación del libro de códigos de dicha trama de muestras de voz digitalizadas, en el que dicho número variable de bits que representa dichas señales vectoriales de excitación del libro de códigos está determinado por dicho nivel de actividad de voz apreciado.
- 5Método según cualquiera de las reivindicaciones anteriores, en el que dicha etapa de codificación de dicha trama comprende:generar para dicha trama un número variable de coeficientes de predicción lineales en el que dicho número variable de dichos coeficientes de predicción lineales está determinado por dicha velocidad de codificación seleccionada;generar para dicha trama un número variable de coeficientes de tono en el que dicho número variable de dichos coeficientes de tono está determinado por dicha velocidad de codificación seleccionada;y ES 2 240 252 T3 generar para dicha trama un número variable de valores de excitación del libro de códigos en el que dicho número variable de dichos valores de excitación de libro de códigos está determinado por dicha velocidad de codificación seleccionada.
- 6Método según cualquiera de las reivindicaciones anteriores, en el que dicha etapa de determinación de un nivel de actividad de voz comprende la suma de los cuadrados de los valores de dichas muestras de voz digitalizadas.
- 7Método según la reivindicación 6, que comprende además la etapa de generación de bits de protección de error para dicho paquete de datos.
- 8Método según la reivindicación 7, en el que dicha etapa de generación de bits de protección de error para dicho paquete de datos en la que el número de dichos bits de protección está determinado por dicha trama de nivel de actividad de voz.
- 9Método según la reivindicación 2, en el que dicha etapa de ajuste adaptable de los niveles umbral de actividad de voz comprende las etapas siguientes:comparar dicha actividad de voz apreciada a dicho por lo menos uno de los umbrales de actividad de voz y aumentar de forma creciente dicho por lo menos uno de los umbrales de actividad de voz hacia el nivel de dicha actividad de voz de trama cuando dicha actividad de voz de trama excede dicho por lo menos uno de dichos umbrales de actividad de voz;y comparar dicha actividad de voz apreciada con dicho por lo menos uno de los umbrales de actividad de voz y disminuir dicho por lo menos uno de los umbrales de actividad de voz al nivel de dicha actividad de voz de trama cuando dicha actividad de voz de trama es inferior a dicho por lo menos uno de los umbrales de actividad de voz.
- 10Método según la reivindicación 9, en el que dicha etapa de selección de una velocidad de codificación está determinada por una señal de velocidad externa.
- 11Método según la reivindicación 7, en el que dicha etapa de generación de protección de error para dicho paquete de datos comprende además determinar los valores de dichos bits de protección de error según un código cíclico de bloques.
- 12Método según cualquiera de las reivindicaciones anteriores, que comprende además la etapa de premultiplicación de dichas muestras de voz digitalizadas (s(n)) mediante una función de división en ventanas predeterminada.
- 13Método según cualquiera de las reivindicaciones anteriores, que comprende además la etapa de conversión de dichos coeficientes LPC a valores de pares de líneas espectrales (LSP).
- 14Método según cualquiera de las reivindicaciones anteriores, en el que dicha trama de entrada de muestras digitalizadas comprende valores digitalizados para aproximadamente veinte segundos de voz.
- 15Método según cualquiera de las reivindicaciones anteriores, en el que dicha trama de muestras digitalizadas comprende aproximadamente 160 muestras digitalizadas.
- 16Método según cualquiera de las reivindicaciones anteriores, en el que dicho paquete de datos de salida comprende:ciento setenta y un bits que comprenden cuarenta bits para los datos LPC, cuarenta bits para los datos de tono, ochenta bits para los datos vectoriales de excitación y once bits para la protección de error cuando dicha velocidad de datos de salida es la velocidad completa;ochenta bits que comprenden veinte bits para la información de LPC, veinte bits para la información de tono y cuarenta bis para los datos vectoriales de excitación cuando la velocidad de los datos de salida es la mitad de la velocidad;cuarenta bits que comprenden diez bits para la información de LPC, diez bits para la información de tono y veinte bits para los datos vectoriales de excitación cuando dicha velocidad de datos salida es un cuarto de la velocidad;y dieciséis bits que comprenden diez bits para la información de LPC y seis bits para la información del vector de excitación cuando la velocidad de datos de salida es un octavo de la velocidad.
- 17Aparato para la compresión de una señal acústica en datos de velocidad variable que comprende:medios (52) para determinar un nivel de actividad de voz para una trama de entrada (10) de muestras digitalizadas de dicha señal acústica;ES 2 240 252 T3 medios (90, 294, 296) para la selección de una velocidad de datos de salida a partir de un conjunto predeterminado de velocidades sobre la base de dicho nivel determinado de actividad de voz en el interior de dicha trama;medios (58, 104, 106, 108) para la codificación de dicha trama según un formato de codificación de un conjunto de formatos de codificación para dicha velocidad seleccionada para proporcionar una pluralidad de señales de parámetros en el que cada velocidad presenta un formato de codificación diferente con cada formato de codificación proporcionando una pluralidad diferente de señales de parámetro que representan dichas muestras de voz digitalizadas (s(n)) según un modelo de voz;y medios (114) para proporcionar a dicha muestra un paquete de datos correspondiente (p(n)) a una velocidad de datos que corresponde a dicha velocidad seleccionada, caracterizado por: medios para proporcionar un control ligado a la velocidad que indica una velocidad de codificación preseleccionada para dicha trama;y medios para modificar dicha velocidad de codificación seleccionada para proporcionar dicha velocidad de codificación preseleccionada para codificar dicha trama y dicha velocidad de codificación preseleccionada.
- 18Aparato según la reivindicación 17, en el que dicho paquete de datos comprende:un número variable de bits para representar las señales vectoriales LPC de dicha trama (10) de muestras de voz digitalizadas (s(n)), en el que dicho número variable de bits para la representación de dichas señales vectoriales LPC está determinado por dicho nivel de actividad de voz;un número variable de bits para representar señales vectoriales de tono de dicha trama (10) de muestras de voz digitalizadas (s(n)), en el que dicho número variable de bits para la representación de dichas señales vectoriales de tono está determinado por dicho nivel de actividad de voz;y un número variable de bits para representar señales vectoriales de excitación del libro de códigos de dicha trama (10) de muestras de voz digitalizadas (s(n)), en el que dicho número variable de bits para la representación de dichas señales vectoriales de excitación del libro de códigos está determinado por dicho nivel de actividad de voz.
- 19Aparato según la reivindicación 17 ó 18, en el que dichos medios para la determinación de dicho nivel de actividad de voz comprenden:medios (202) para la determinación de un valor energético de dicha trama de entrada;medios (204) para la comparación de dicha energía de la trama de entrada con dicho por lo menos un umbral de actividad de voz;y medios (312) para proporcionar una indicación cuando dicha actividad de trama de entrada excede de cada uno correspondiente de dicho por lo menos un umbral de actividad de voz.
- 20Aparato según la reivindicación 20, que comprende además medios para el ajuste de forma adaptable de dicho por lo menos uno de dicho por lo menos un umbral de actividad de voz.
- 21Aparato según la reivindicación 17, en el que dichos medios para la determinación de un nivel de actividad de voz comprenden:medios de elevación al cuadrado para elevar al cuadrado dichas muestras de audio digitalizadas de una trama;y medios de suma para sumar dichos cuadrados de las muestras de audio digitalizadas de una trama.
- 22Aparato según cualquiera de las reivindicaciones 17, 18 ó 19, en el que dichos medios para la determinación de un nivel de actividad de velocidad comprenden:medios (50) para el cálculo de un conjunto de coeficientes predictivos lineales para dicha trama de entrada de las muestras digitalizadas de las señales acústicas;y medios para determinar dicho nivel de actividad de voz según por lo menos uno de dichos coeficientes predictivos lineales.
- 23Aparato según cualquiera de las reivindicaciones 17 a 22, que comprende además medios (236, 238) para proporcionar bits de protección de error para dicho paquete de datos determinado por dicha velocidad de datos de salida seleccionada. ES 2 240 252 T3
- 24Aparato según la reivindicación 24, en el que dichos medios (236, 238) para proporcionar bits de protección de error proporcionan los valores de dichos bits de protección de error según un código cíclico de bloques.
- 25Aparato según cualquiera de las reivindicaciones 17 a 24, que comprende además medios (208) para convertir dichos coeficientes LPC en valores de pares de líneas espectrales (LSP).
- 26Aparato según cualquiera de las reivindicaciones 17 a 25, en el que dicho conjunto de velocidades comprende velocidad completa, la mitad de la velocidad, un cuarto de velocidad y un octavo de velocidad.
- 27Aparato según cualquiera de las reivindicaciones 17 a 26, en el que dicho conjunto de velocidades comprende 8 Kbps, 4 Kbps, 2 Kbps y 1 Kbps.
- 28Aparato según la reivindicación 22, en el que dichos medios para la determinación de un nivel de actividad de voz determinan dicha energía mediante el cálculo de un conjunto de coeficientes predictivos lineales para dicha trama de salida y determinan dicho nivel de actividad de voz según por lo menos uno de dichos coeficientes predictivos lineales.
- 29Aparato según cualquiera de las reivindicaciones 17 a 28, en el que dicha trama de entrada de las muestras de voz digitalizadas comprende la voz digitalizada durante veinte milisegundos aproximadamente.
- 30Aparato según cualquiera de las reivindicaciones 17 a 29, en el que dicha trama de entrada de las muestras digitalizadas comprende 160 muestras digitalizadas.
- 31Aparato según la reivindicación 37, en el que dicho código cíclico de bloques funciona según un generador polinómico de 1 + x 3 + x 5 + x 6 + x 8 + x 9 + x 10 .
- 32Aparato según cualquiera de las reivindicaciones 17 a 31, que comprende además medios (52, 200) para la premultiplicación de dichas muestras digitalizadas mediante una función de división en ventanas predeterminada.
- 33Aparato según la reivindicación 32, en el que dicha función de división en ventanas predeterminada es una ventana de Hamming.
- 34Aparato según cualquiera de las reivindicaciones 17 a 33, en el que dicho paquete de datos de salida (p(n)) comprende:un número variable de bits para representar señales vectoriales LPC de dicha trama de muestras de voz digitalizadas (s(n)), en el que dicho número variable de bits para la representación de dichas señales vectoriales LPC está determinado por dicho nivel de actividad de voz;un número variable de bits para representar señales vectoriales de tono de dicha trama de muestras de voz digitalizadas (s(n)), en el que dicho número variable de bits para la representación de dichas señales vectoriales de tono está determinado por dicho nivel de actividad de voz;y un número variable de bits para representar señales vectoriales de excitación del libro de códigos de dicha trama de muestras de voz digitalizadas (s(n)), en el que dicho número variable de bits para la representación de dichas señales vectoriales de excitación del libro de códigos está determinado por dicho nivel de actividad de voz.
- 35Aparato según la reivindicación 43, en el que dicho paquete de datos de salida comprende además un número variable de bits para la protección de error, en el que dicho número variable de bits para la protección de error está determinado por dicho nivel de actividad de voz.
- 36Aparato según cualquiera de las reivindicaciones 17 a 35, en el que dicho paquete de datos de salida comprende:ciento setenta y un bits que comprenden cuarenta bits para los datos LPC, cuarenta bits para los datos de tono, ochenta bits para los datos vectoriales de excitación y once bits para la protección de error cuando dicha velocidad de datos de salida es la velocidad completa;ochenta bits que comprenden veinte bits para la información de LPC, veinte bits para la información de tono y cuarenta bits para los datos vectoriales de excitación cuando dicha velocidad de datos de salida es la mitad de la velocidad;cuarenta bits que comprenden diez bits para la información de LPC, diez bits para la información de tono y veinte bits para los datos vectoriales de excitación cuando dicha velocidad de datos de salida es un cuarto de la velocidad;y dieciséis bits que comprenden diez bits para la información de LPC y seis bits para la información vectorial de excitación cuando dicha velocidad de datos de salida es un octavo de la velocidad.
- 37Aparato según cualquiera de las reivindicaciones 17 a 36, en el que dichos medios (90, 294, 296) de selección de una velocidad de codificación están determinados por una señal de velocidad externa.
Independent claims37
411 paragraphs in 27 sections, as filed
IS 2 240 252 T3
DESCRIPTION
Variable speed vocoder.
I. Field of the invention
The invention relates to a method and method for voice signal compression, more particularly, the invention relates to a new and improved method and system for voice compression in which the amount of compression varies dynamically while its impact on the quality of the reconstructed voice is minimal.
II. Description of Related Art
Voice transmission by digital techniques has been widely used, particularly in digital radio telephone applications. This, in turn, has raised interest in determining the minimum amount of information that can be sent through the channel, while preserving the perceived quality of the reconstructed voice. If voice is transmitted simply by sampling and digitizing, a data rate of the order of 64 kilobits per second (Kbit / s) is required to obtain the voice quality of conventional analog telephone. However, through the use of speech analysis, followed by correct coding, transmission and resynthesis at the receiver, a significant reduction in data transmission speed can be achieved.
Devices that employ techniques to compress voiced speech by extracting parameters that relate to a human speech generation model are commonly called vocoders. Said devices consist of an encoder that analyzes the incoming voice to extract the pertinent parameters, and a decoder, which resynthesizes the voice using the parameters it receives through the transmission channel. To be accurate, the model must constantly change. Therefore, the speech is divided into time blocks, or analysis frames, during which the parameters are calculated. The parameters of each new frame are then updated.
Among the various types of speech coders, those performing code-excited linear prediction (CELP) coding, stochastic coding, or vector-excited speech coding, constitute a class. An example of a coding algorithm of this particular class can be found in the document "A 4.8 kbps Code Excited Linear Predictive Coder" by Thomas E. Tremain et al., Proceedings of the Mobile Satellite Conference, 1988.
The function of the vocoder is to compress the digitized voice signal into a low bit rate signal, eliminating all inherent natural redundancies. in speech. Usually, the voice presents short-term redundancies, mainly due to the filtering operation of the vocal tract, and long-term redundancies due to the excitation of the vocal tract by the vocal cords. In a CELP encoder, these operations are modeled by two filters, a short-duration formant filter and a long-duration tone filter. Once these redundancies are removed, the resulting residual signal can be modeled as Gaussian white noise, which must also be coded. The basis of this technique is to calculate the parameters of a filter, called the LPC (Linear Production Coding) filter, which performs short-term prediction of the voice waveform using a model of the human vocal tract. In addition, long-term effects related to the tone of the voice are modeled, calculating the parameters of a tone filter that, in essence, models the human vocal cords. Finally, these filters must be excited, and this is done by determining which particular random excitation waveform from a group contained in a codebook results in the closest approximation to the original speech, when the waveform excites the two filters mentioned above. Therefore, the transmitted parameters refer to three items: (1) the LPC filter, (2) the tone filter, and (3) the codebook drive.
Although the use of speech coding techniques favors the objective of trying to reduce the amount of information sent through the channel and at the same time ensure quality reconstructed speech, it is necessary to use other techniques to achieve a greater reduction. A previously used technique to reduce the amount of information sent is the targeting of speech activity. In this technique, no information is transmitted during voice pauses. Although this technique achieves the desired result of data reduction, it suffers from several shortcomings.
In many cases, the quality of the voice is reduced due to the clipping of the initial part of the words. Another problem with channel disconnection during inactivity is that system users perceive the absence of background noise that normally accompanies voice and their assessment of channel quality is as low as that of a normal telephone call. . Another problem with activity selection is that occasional rude noises in the background can activate the transmitter when there is no voice, causing annoying bursts of noise at the receiver.
In an attempt to improve the quality of synthesized speech in speech activity selection systems, comfort noise is added during the decoding process. Although some improvement in quality is achieved by adding comfort noise, the overall quality improvement is not substantial, since the comfort noise does not model the actual background noise from the encoder.
A more preferred technique to perform data compression, and that manages to reduce the information that is
ES 2 240 252 T3 necessary to send, consists of carrying out variable speed speech coding. Since speech inherently contains periods of silence, that is, pauses, the amount of data required to represent those periods can be reduced. Variable rate speech coding exploits this fact most efficiently by reducing the data rate for these periods of silence. Reducing data transmission speed, as opposed to completely interrupting data transmission during periods of silence, overcomes the problems associated with the selection of speech activity while facilitating the reduction of transmitted information. .
The European patent application EP 0 449 043, which is a prior art according to Article 54 (3) and (4) EPC, deserves special attention. European patent application EP 0 449 043 describes a method and apparatus for digitizing speech. The digitization of speech is performed using both signal encoding and source encoding with an encoder for digitizing and a decoder for reconstruction of the speech signal. The speech signal is divided into segments at the encoder and processed into a portion of the segments with as close an approximation to the sample values as possible, with an estimated value being calculated for the pending sample values using known sample values. In the other part of the segments, only parameters for the speech simulation are derived in terms of source encoding. Individual signal elements are processed at varying bit rates, assigned to different modes of operation, and each signal segment is classified as one of the modes of operation. The individual speech segments are thus encoded according to the requirements of higher or lower number of bits, providing a hybrid encoding method that unifies source encoding and signal encoding. This produces, together with the signal quantization of the upstream and downstream phases of signal processing, an average bit rate of 6 kbit / s and a voice quality similar to that of a telephone transmission.
The article entitled “Phonetically-based vector excitation coding of speech at 3.6 kbps” by Shihua Wang et al, speech processing 1, Glasgow, May 23-26, 1989, ICASSP'89, New York, IEEE, US also deserves special attention. , vol. 1 conf. 14, May 23, 1989, pages 49-52, XP000089669. The article describes a phonetic-based segmentation of the voice, which is performed to classify the segments into five classes: beginning, deaf, low-pass voice, steady-state voice, and transient voice. Segment lengths are limited to an integer multiple of a frame unit. For each class of segment, a distinctive coding scheme based on Vector Excitation Coding (VXC) is used.
According to the present invention, there are provided a method for compression of a voice signal as set forth in claim 1 and an apparatus for compressing an acoustic signal as set forth in claim 17. Preferred embodiments of the invention are disclosed in subordinate claims.
Summary of the invention
The invention will become more clearly apparent from the dependent claims.
The object of the present invention is to provide a new and improved method and system for voice compression using a variable rate vocoding technique.
Next, a vocoder will be described which executes a speech coding algorithm of the class of speech coders mentioned above, that is, code-excited linear prediction (CELP) coding, stochastic coding, or excited speech coding vector. The CELP technique alone provides a significant reduction in the amount of data required to represent speech, in a way that, after resynthesis, results in high-quality speech. As mentioned above, the vocoder parameters are updated for each frame. The vocoder of the present invention provides a variable data rate by changing the frequency and accuracy of the model parameters.
The most notable difference of the embodiment from the basic CELP technique is its ability to generate a variable output data rate based on speech activity. The defined structure allows the parameters to be updated less frequently, or with less precision, during speech pauses, and the technique determines an even greater reduction in the amount of information to be transmitted. The phenomenon that is exploited to reduce the data transmission speed is the speech activity factor, which is the average percentage of time during which a given speaker actually speaks in a conversation. For typical two-way telephone conversations, the average data transmission speed is reduced by a factor of 2 or more. During speech pauses, the vocoder only encodes the background noise. At such times, it is not necessary to transmit some of the parameters relative to the human vocal tract model.
The approach mentioned above, called speech activity selection, for limiting the amount of information transmitted during periods of silence, is a technique in which no information is transmitted during moments of silence. As far as reception is concerned, the period can be filled with synthesized “comfort noise”. In contrast, in an embodiment of the present invention that will be described in detail below, a variable rate vocoder transmits data continuously at rates ranging from 8 Kbit / s to approximately 1 Kbit / s. A vocoder that performs continuous data transmission can dispense with "comfort noise" synthesis, and background noise encoding provides a more natural quality to resynthesized speech. Therefore, the present invention represents a significant improvement in quality.
ES 2 240 252 T3 resynthesized speech with respect to the activity selection of speech signals, by facilitating a smooth transition between speech and background.
The present invention further incorporates a new technique for masking the presence of errors. Because the data is intended to be transmitted over a channel that can be noisy, such as a radio link, the data must include errors. Previous techniques using channel coding to reduce the number of errors present may be partially successful in reducing errors. However, channel coding alone does not provide the complete level of error protection necessary to ensure high-quality reconstructed voice. In the variable rate vocoder, which applies voice coding permanently, an error can destroy data related to some interesting voice event, such as the beginning of a word or a syllable. A common problem with vocoders based on Linear Prediction Coding (LPC) is that errors in the parameters relative to the vocal tract model cause sounds that vaguely resemble human sounds, and can change the sound of the original word. enough to confuse the listener. In the following embodiment, errors are masked to reduce their perceptibility by the listener. This error masking provides a drastic reduction in the effect of errors on speech intelligibility.
Because the maximum change that any parameter can undergo is limited to lower values at low speeds, errors in the parameters transmitted at these speeds will affect speech quality less. Since errors at different speeds have different perceived effects on voice quality, the transmission system can be fully exploited to give more protection to higher speed data. Therefore, as an added feature, the present invention provides resistance to channel errors.
By running a variable rate output version of the CELP algorithm, you get speech compression that varies dynamically between 8: 1 and 64: 1, depending on speech activity. The compression factors just mentioned refer to a μ-law input, the compression factors being higher by a factor of 2 for a linear input. Rate determination is made from frame to frame to fully utilize the speech activity factor. Even when less data is generated for speech pauses, the perceived degradation of the resynthesized background noise is minimized. Using the techniques of the present invention, voice of near-trunk type quality can be achieved at a maximum data rate of 8 Kbit / s and an average data rate of the order of 3.5 Kbit / s in normal conversation.
Since the vocoder makes it possible to detect short pauses of speech, a reduction in the speech activity factor is achieved. Rate decisions can be made from frame to frame with no hang time, and consequently the data transmission rate can be reduced for speech pauses that are as short as the frame duration, which is typically 20 ms in the preferred embodiment. Thus, pauses such as those between syllables can be captured. This technique reduces the vocal activity factor to a greater extent than traditionally achieved, since it is possible to encode not only long pauses between phrases, but also shorter pauses at lower rates.
Because rate decisions are made from frame to frame, there is no clipping of the initial part of the speech, as occurs in the speech activity selection system. Clipping of this nature occurs in the speech activity selection system due to the delay between the detection of speech and the restart of data transmission. When speed utilization decisions are based on the results of individual frames, you get voice in which all transitions sound natural.
If the vocoder transmits continuously, background noise from the speaker's environment will be heard permanently at the receiving end, thereby providing a more natural sound during speech pauses. Consequently, the vocoder allows a smooth transition to the background noise. What the listener can hear in the background during conversation will not suddenly transform into synthesized comfort noise during pauses, as in the speech activity selection system.
Since background noise is continuously vocoded for transmission, interesting events in the background can be delivered with complete clarity. In certain cases, the background noise of interest can be encoded even at the highest speed. Full-speed encryption can occur, for example, when someone is talking loudly in the background, or if an ambulance passes near a user on the street. However, slowly or constantly varying background noise will be encoded at low rates.
The use of variable rate speech coding promises to increase the capacity of a code division multiple access (CDMA) based digital cellular telephone system by more than a factor of two. Variable rate and CDMA speech coding are uniquely matched, since with CDMA, inter-channel interference automatically decreases as the data rate through any channel decreases. For comparison, we will consider systems where transmission slots are assigned, such as TDMA or FDMA systems. For one of these systems to take advantage of any drop in data rate, external intervention is required to coordinate the reallocation of unused slots to other users. The delay inherent in such a system determines that the channel can be reassigned only during long pauses of speech. Therefore, the vocal activity factor cannot be fully exploited. Not
However, with external coordination, variable rate speech coding is useful in systems other than CDMA for the other reasons mentioned.
In a CDMA system, the voice quality of the system may degrade slightly at times when additional system capacity is desired. In abstract terms, the vocoder can be thought of as a group of vocoders that operate at different speeds and provide different qualities of speech. Therefore, the speech qualities can be mixed to further reduce the average data transmission speed. Initial experiments show that by mixing full-rate and half-rate speech coded voice, eg by varying the maximum allowed data rate from frame to frame between 8 Kbit / s and 4 Kbit / s, the resulting voice has a quality that is better than the half-speed variable, 4 Kbit / s maximum, but not as good as the full-speed variable, 8 Kbit / s maximum.
It is well known that in most telephone conversations only one person speaks at a time. As an additional feature for full duplex telephone links, a speed interlock can be provided. If one direction of the link transmits at the highest transmission rate, then the other direction is forced to transmit at the slowest rate. An interlock between the two link addresses can guarantee an average utilization of no more than 50% of each link address. However, when the channel is disabled as in the case of speed interlock in speech activity selection, there is no way for a listener to interrupt the speaker to assume the role of speaker in the conversation. The vocoder to be described below easily provides rate interlocking capability by means of control signals that set the speech coding rate.
Finally, it should be noted that by using a variable rate speech coding model, the signaling information can share the channel with speech data with very little effect on speech quality. For example, a high speed frame can be split in two; one half is used to send the lower speed voice data and the other half is used to send the signaling data. In the vocoder of the preferred embodiment, only a slight degradation in speech quality occurs between speech under full-rate speech and that under half-speed speech encoding. Accordingly, vocoding the speech at the lowest rate for shared transmission with other data results in an almost imperceptible difference in speech quality by the user.
Brief description of the drawings
The foregoing and additional features, objects, and advantages of the present invention will become more apparent upon consideration of the following detailed description, illustrated by the accompanying drawings, in which equivalent reference characters are used for equivalent indications and, in which :
Figures 1a-1e graphically illustrate the vocoder analysis frames and subframes for different rates;
Figures 2a-2d are a series of graphs illustrating the vocoder output binary distribution for different rates;
Figure 3 is a generalized block diagram of an example encoder; Figure 4 is a flow chart of an encoder;
Figure 5 is a generalized block diagram of an example decoder;
Figure 6 is a flow chart of a decoder;
Figure 7 is a more detailed functional block diagram of the encoder;
Figure 8 is a block diagram of an exemplary Hamming window and autocorrelation subsystems; Figure 9 is a block diagram of an exemplary rate determination subsystem; Figure 10 is a block diagram of an exemplary LPC analysis subsystem;
Figure 11 is a block diagram of an exemplary LPC-LSP transform subsystem;
Figure 12 is a block diagram of an exemplary LPC quantization subsystem;
Figure 13 is a block diagram of an exemplary LSP interpolation and LSP-LPC transform subsystem;
Figure 14 is a block diagram of the adaptive codebook for pitch search;
ES 2 240 252 T3 Figure 15 is a block diagram of the encoder decoder; Figure 16 is a block diagram of the tone search subsystem; Figure 17 is a block diagram of the codebook search subsystem; Figure 18 is a block diagram of the data packaging subsystem; Figure 19 is a more detailed functional block diagram of the decoder;
Figures 20a-20d are diagrams illustrating subframe decoding parameters and data received by the decoder for different rates;
Figures 21a-21c are diagrams providing a further illustration of the subframe decoding parameters and data received by the decoder for special conditions;
Figure 22 is a block diagram of the LSP inverse quantization subsystem;
Figure 23 is a more detailed block diagram of the decoder with post-filtering and automatic gain control; and Figure 24 is a diagram illustrating the adaptive characteristics of the brightness filter.
Detailed description of the preferred embodiment
A vocoder will now be described in which sounds such as speech and / or background noise are sampled and digitized using well known techniques. For example, the analog signal can be transformed into a digital signal using the standard 8-bit / μ-law format followed by a μ-law / uniform code conversion. Alternatively, the analog signal can be directly converted to a digital signal in a uniform pulse code modulation (PCM) format. Therefore, each sample is represented by a 16-bit word of data. The samples are organized into input data frames, each frame comprising a predetermined number of samples. In the following description, the sample rate considered is 8 kHz. Each frame comprises 160 samples or 20 ms of speech at the 8 kHz sample rate. It should be understood that other sample rates and frame sizes can be used.
The field of speech coding includes many different techniques for speech coding, one of these being the CELP coding technique. A summary of the CELP coding technique is provided in the document "A 4.8 kbps Code Excited Linear Predictive Coder" mentioned above. The present invention implements a form of CELP coding technique to provide a variable rate to encoded voice data, with LPC analysis being performed with a constant number of samples and pitch and codebook searches being performed with variable numbers of samples. depending on the transmission speed. The CELP coding techniques that apply to the present invention are conceptually described with reference to Figures 3 and 5.
In the following description, the speech analysis frames have a duration of 20 ms, which implies that the extracted parameters are transmitted in a burst 50 times per second. Furthermore, the data transmission speed varies approximately between 8 Kbit / s and 4 Kbit / s, 2 Kbit / s and 1 Kbit / s. At full rate (also referred to as rate 1), data transmission takes place at 8.55 Kbit / s and the encoded parameters for each frame use 171 bits including an internal 11-bit CRC (cyclic redundancy check). In the absence of CRC bits, the rate will be 8 Kbit / s. At half speed (also called speed 1/2), the data transmission is carried out at 4 Kbit / s and the encoded parameters for each frame use 80 bits. At quarter speed (also called 1/4 speed), data transmission is carried out at 2 Kbit / s and the encoded parameters for each frame use 40 bits. At eighth speed (also called 1/8 speed), the data transmission is slightly less than 1 Kbit / s and the encoded parameters for each frame use 16 bits.
Figure 1 graphically illustrates an example speech data analysis frame 10 and the relationship of a Hamming window 12 used in LPC analysis. In Figures 2a-2d, the LPC analysis frame and the tone and codebook subframes for the different rates are graphically illustrated. It should be understood that the LPC analysis frame is the same size for all rates.
With reference to the drawings and, in particular, to Figure 1a, the LPC analysis is carried out using the 160 speech data samples of frame 10 that are subjected to window dressing using a Hamming window 12. As illustrated in In Figure 1a, the samples s (n) are numbered from 0 to 159 within each frame. Hamming window 12 is located at a 60 sample offset within frame 10. Therefore, Hamming window 12 starts at 60<sup>to</sup> sample, s (59), from current data frame 10 and ends at sample 59<sup>to</sup>, s (58) of the next data frame 14. Accordingly, the weighted data generated for the current frame, ie frame 10, will also contain data based on data from the next frame, ie frame 14.
IS 2 240 252 T3
Depending on the data rate, searches are performed to calculate the tone filter and codebook drive parameters multiple times with different subframes of data frame 10, as shown in Figures 1b-1e. It should be understood that only one rate is selected for frame 10, so that the tone and codebook searches are performed on subframes of various sizes corresponding to the selected rate, as described below. However, for illustrative purposes, Figures 1b-1e show the structure of the subframes of frame 10 for tone and codebook searches and the various rates allowed for the preferred embodiment.
At all rates, one LPC calculation is performed per frame 10 as illustrated in Figure 1a. As illustrated in Figure 1b, at full rate there are two codebook subframes 18 for each tone 16 subframe. At full rate, four tone updates are performed, one for each of the four tone subframes 16 of 40 duration samples (5 ms). In addition, eight codebook updates are performed at full rate, one for each of the eight codebook subframes 18, of 20 samples duration (2.5 ms).
At half speed, as illustrated in Figure 1c, there are two codebook subframes 22 for each pitch subframe 20. The pitch is updated twice, once for each of the two pitch frames 20, while the book The codebook is updated four times, one for each of the four codebook subframes 22. At quarter speed, as illustrated in Figure 1d, there are two codebook subframes 26 for the single tone subframe 20. The pitch is updated once for the pitch 24 subframe, while the codebook is updated twice, one for each of the two codebook 26 subframes. As illustrated in Figure 1e, at eighth speed , the pitch is not determined and the codebook is updated only once in frame 28 which corresponds to frame 10.
Furthermore, although the LPC coefficients are calculated only once per frame, they are linearly interpolated, in a spectral line pair (LSP) representation, up to four times using the resulting LSP frequencies from the previous frame to roughly calculate the analysis results. LPC with the Hamming window centered on each subframe. The exception is that, at full rate, the LPC coefficients for the codebook subframes are not interpolated. More information about the calculation of LSP frequencies is provided later.
Aside from less frequent pitch and codebook searches being performed at lower rates, fewer bits are allocated for transmission of the LPC coefficients. The number of bits assigned to the different rates is shown in Figures 2a-2d. Each of Figures 2a-2d represents the number of vocoder-encoded data bits allocated to each of the 160 speech sample frames. In Figures 2a-2d, the number of the respective LPC block 30a-30d is the number of bits used at the corresponding rate to encode the short-term LPC coefficients. The number of bits used to encode the LPC coefficients at the rates, full, half, fourth and eighth are respectively 40, 20, 10 and 10.
To perform variable rate coding, the LPC coefficients are first transformed into spectral line pairs (LSP) and the resulting LSP frequencies are individually encoded using DPCM encoders. The order of LPC is 10, that is, there are 10 LSP frequencies and 10 independent DPCM encoders. The bit allocation for the DPCM encoders is made according to Table I.
TABLE I
<td rowspan="2"></td><td colspan="10">DPCM ENCODER NUMBER</td>
<td> 1</td><td> 2</td><td> 3</td><td> 4</td><td> 5</td><td> 6</td><td> 7</td><td> 8</td><td> 9</td><td> 10</td>
<td>SPEED. 1</td><td> 4</td><td> 4</td><td> 4</td><td> 4</td><td> 4</td><td> 4</td><td> 4</td><td> 4</td><td> 4</td><td> 4</td>
<td>SPEED. 1/2</td><td> 2</td><td> 2</td><td> 2</td><td> 2</td><td> 2</td><td> 2</td><td> 2</td><td> 2</td><td> 2</td><td> 2</td>
<td>SPEED. 1/4</td><td> 1</td><td> 1</td><td> 1</td><td> 1</td><td> 1</td><td> 1</td><td> 1</td><td> 1</td><td> 1</td><td> 1</td>
<td>SPEED. 1/8</td><td> 1</td><td> 1</td><td> 1</td><td> 1</td><td> 1</td><td> 1</td><td> 1</td><td> 1</td><td> 1</td><td> 1</td>
In both the encoder and decoder, the LSP frequencies are converted back to LPC filter coefficients prior to use in tone and codebook searches.
With respect to the tone search, the tone update is calculated four times at full rate, one for every quarter of the speech frame, as illustrated in Figure 2a. For each full-rate pitch update, 10 bits are used to encode the new pitch parameters. Pitch updates are performed a variable number of times for the other rates shown in Figures 2b-2d. As the speed decreases, the number of pitch updates also decreases. Figure 2b illustrates the half rate pitch updates that are calculated twice, once for each half speech frame. Similarly, Figure 2c illustrates quarter rate tone updates that are calculated once for each complete speech frame. As for full rate, 10 bits are used to encode the new tone parameters for each half and quarter speed tone update. However, as illustrated in Figure 2d, for eighth speed no pitch update is calculated, since this rate is used to encode frames when the voice present is null or almost null and there are no pitch redundancies.
IS 2 240 252 T3
In each 10-bit pitch update, 7 bits represent pitch delay and 3 bits represent pitch gain. The pitch delay is limited to values between 17 and 143. The pitch gain is linearly quantized between 0 and 2 for representation by the 3-bit value.
In connection with the codebook search, as illustrated in Figure 2a, at full rate the codebook update is calculated eight times, once for every eighth of the speech frame. For each full rate codebook update, 10 bits are used to encode the new codebook parameters. Codebook updates are performed a variable number of times at the rates shown in Figures 2b-2d. However, as the speed decreases, so does the number of codebook updates. Figure 2b illustrates half rate codebook updates that are calculated four times, once for every quarter of the speech frame. Figure 2c illustrates the quarter rate codebook updates that are calculated twice, once for each half of the speech frame. As for full speed, 10 bits are used to encode the new codebook parameters for each half speed and quarter speed tone update. Finally, Figure 2d illustrates eighth rate codebook updates that are only calculated once for each complete speech frame. It should be noted that at eighth speed 6 bits are transmitted; 2 of which are representative of codebook gain and the remaining 4 are random bits. More information about bit assignments for codebook updates is provided later.
The bits allocated for the codebook updates represent the data bits necessary to vectorially quantize the pitch prediction residue. For full, half, and quarter rates, each codebook update consists of 7 codebook index bits plus 3 codebook gain bits for a total of 10 bits. The codebook gain is encoded using a differential pulse code modulation (DPCM) encoder operating in the logarithmic domain. Although a similar bit arrangement can be used for eighth speed, it is preferable to use an alternative pattern. At eighth speed, the codebook gain is represented by 2 bits, while 4 randomly generated bits are used with the received data as the seed for the pseudo-random number generator that replaces the codebook.
With respect to the block diagram of the encoder illustrated in Figure 3, the LPC analysis is carried out in an open loop mode. For each frame of input speech samples s (n), the coefficients LPC0 (α<sub>1</sub>α<sub>10</sub>) as will be described later, by LPC 50 analysis / quantification, for use in formant synthesis filter 60.
However, the tone search calculation is performed in a closed-loop mode which is often referred to as the analysis-by-synthesis procedure. However, in the running example, a new hybrid closed-loop / open-loop technique is used to drive the search for the pitch. In pitch search, encoding is carried out by selecting parameters that minimize the root mean square error between the input speech and the synthesized speech. For simplicity, this part of the description will not discuss speed. However, more detailed information about the effect of the selected speed on pitch and codebook searches is provided below.
In the conceptual embodiment illustrated in Figure 3, the perceptual weighting filter 52 is characterized by the following equations:
W (z) =
A (z)
A (z | u) (1) where
A (z) = 1 - Σ a¡z<sup>-i </sup>1 = 1 (2) the formant prediction filter and μ, a perceptual weighting parameter, which in the following description equals 0.8. The tone synthesis filter 58 is characterized by the following equation:
_ 1
P (z) 1 - bz<sup>-L</sup> (3)
The formant synthesis filter 60, a weighted filter described below, is characterized by the following equation:
H (z) =
A (z)
W (z) =
A (z | u) (4)
IS 2 240 252 T3
The input speech samples s (n) are weighted by the perceptual weighting filter 52, and the weighted speech samples x (n) are provided to a sum input of the adder 62. The perceptual weighting is used to weight the error. at frequencies where there is less signal strength. It is at these low signal power frequencies that noise is most perceptually noticeable. The synthesized speech samples x '(n) are passed from the formant synthesis filter 60 to a subtraction input of the adder 62 where they are subtracted from the x (n) samples. The sample difference obtained from adder 62 is input to mean square error element (MSE) 64 where it is squared and added. Results from MSE element 64 are provided to minimization element 66 which generates values for pitch delay L, pitch gain b, codebook index I, and codebook gain.
In the minimization element 66, all possible values of L, the pitch delay parameter of P (z), are input into the pitch synthesis filter 58 along with the value c (n) of the multiplier 56. During the pitch pitch search there is no contribution from the codebook, that is, c (n) = 0. The values of L and b that minimize the weighted error between the input voice and the synthesized voice are chosen by the minimization element 66. The tone synthesis filter 58 generates and supplies the p (n) value to the formant synthesis filter 60. Once the tone delay L and the tone gain b for the tone filter have been found, the search is performed. codebook in a similar way.
It should be understood that Figure 3 is a conceptual representation of the analysis-by-synthesis approach. In the exemplary embodiment, the filters are not used in the typical closed-loop feedback configuration. The feedback connection is broken during the search and is replaced by an open-loop formant residue described later.
The minimization element 66 then generates values for the codebook index I and the codebook gain G. The values obtained from the codebook 54, selected from a plurality of Gaussian random vector values according to the codebook index of code I, are multiplied at multiplier 55 by the codebook gain G to generate the sequence of values c (n) used in tone synthesis filter 58. The codebook index I and the codebook gain G that are chosen for transmission are those that minimize the root mean square error.
It should be noted that both the input speech and the synthesized speech are perceptually weighted W (z) by the perceptual weighting filter 52 and the weighting function included in the formant synthesis filter 60, respectively. The formant synthesis filter 60 therefore is actually a weighted formant synthesis filter, combining the weighting function of equation 1 with the typical characteristic of the formant prediction filter to provide the weighted formant synthesis function. from equation 3.
It should be understood that alternatively, the perceptual weighting filter 52 may be positioned between the adder 62 and the MSE element 64. In this case, the formant synthesis filter 60 will have the normal filter characteristic of
TO(-).
Figure 4 illustrates a flow chart of the steps relating to speech coding with the encoder of Figure 3. For descriptive purposes, the steps relating to rate decision are included in the flow chart of Figure 4. Digitized speech samples are obtained (block 80) from the sampling circuits from which the LPC coefficients are then calculated (block 82). In the calculation of LPC coefficients, the Hamming window and autocorrelation techniques are used. For the frame of interest, an initial rate decision is made (block 84) based on frame energy.
To efficiently encode the LPC coefficients into a small number of bits, the LPC coefficients are transformed into spectral line pair (LSP) frequencies (block 86) and then quantized (block 88) for transmission. Optionally, a further speed determination may be made (block 90), the speed being increased if the quantization of the LSP coefficients for the initial speed is deemed insufficient (block 92).
For the first pitch subframe of the speech frame being analyzed, the LSP frequencies are interpolated and transformed into LPC coefficients (block 94) for use in the direction of the pitch search. When searching for the tone, the codebook excitation is set to zero. In the pitch search (blocks 96 and 98) which is a synthesis analysis procedure as described above, for each possible pitch delay L, the synthesized voice is compared with the original voice. For each value of L, an integer value is determined, the optimal pitch gain b. Of the groups of L and b values, the group of optimal L and b values provides the least perceptually weighted root mean square error between the synthesized voice and the original voice. For the optimal values of L and b determined for that tone subframe, the value of b is quantized (block 100) for transmission along with the corresponding value of L. In an alternative execution of the tone search, the value of b and L they can be quantized values that participate in the pitch search, these quantized values being used to drive the pitch search. Therefore, in this run, it will no longer be necessary to quantize the selected b value after the pitch search (block 100).
IS 2 240 252 T3
For the first codebook subframe of the speech frame being analyzed, the LSP frequencies are interpolated and transformed into LPC coefficients (block 102), for use in the codebook search direction. In the exemplary description, however, at full rate, the LSP frequencies are only interpolated down to the pitch subframe level. This interpolation and transformation stage is carried out for both the codebook search and the pitch search, due to the difference in size of the pitch and codebook subframes for each speed, except for speed 1 / 8, as this is irrelevant, since no pitch data is calculated. In codebook search (blocks 104 and 106), the optimal pitch delay L and gain b values in the pitch synthesis filter are used to compare, for each possible codebook index I, the voice synthesized with the original voice. For each value of I, an integer value is determined; the optimal G-codebook gain. Of the groups of I and G values, the group of optimal I and G values provides the smallest error between the synthesized voice and the original voice. For the determined optimal values of I and G for said book code subframe, the G value is quantized (block 108), for transmission together with the corresponding I value. On the other hand, in an alternative execution of the book search of code, the quantization of the G-values may be carried out as part of the codebook search, these quantized values being used in the direction of the code search. In this alternate execution, quantization of the selected G-value after the codebook search is no longer necessary (block 108).
After the codebook search, the encoder decoder runs with the optimal I, G, L, and b values. Encoder decoder execution reconstructs the encoder filter memories for use in future subframes.
Next, a check is performed (block 110) to determine whether the codebook subframe whose analysis has just completed is the last codebook subframe in the group of codebook subframes corresponding to the tone subframe for which the search for tone is destined. In other words, it is determined whether there are more codebook subframes left that correspond to the tone subframe. In the exemplary embodiment, there are only two codebook subframes per tone subframe. If it is determined that there is another codebook subframe remaining that corresponds to the tone frame, steps 102-108 are repeated for that codebook subframe.
In the event that there are no more codebook subframes remaining corresponding to the tone frame, a check will be performed (block 112) to determine if any tone subframes remain within the speech frame being analyzed. If another pitch subframe remains in the current speech frame being analyzed, steps 94110 are repeated for each pitch subframe and corresponding codebook subframes. When all calculations have been completed for the current speech frame being analyzed, the representative values of the LPC coefficients for the speech frame, the delay L and the pitch gain b for each pitch subframe, and the I index and the codebook gain G for each codebook subframe is packed for transmission (block 114).
Referring to Figure 5, a decoder block diagram is illustrated in which the received values for the LPC coefficients (a,), the pitch delays and gains (L and b), and the codebook gains and indices ( I and G) are used to synthesize the voice. Again, in Figure 5, as in Figure 3, speed information is not taken into account to simplify the description. The data rate information can be sent as supplementary information and, in certain cases, can be obtained at the channel demodulation stage.
The decoder consists of a codebook 130 which is provided with the received codebook indices or, for eighth speed, the random seed. The output of codebook 130 is provided to one input of multiplier 132, while the other input of multiplier 132 receives the codebook gain G. The output of multiplier 132 is provided along with delay L and gain b of tone to the tone synthesis filter 134. The output of the pitch synthesis filter 134 is provided along with the LPC coefficients a, to the formant synthesis filter 136. The output of the formant synthesis filter 136 is provided to the adaptive post-filter 138 where it is filtered and provided as reconstructed speech. . As described below, a version of the decoder runs in the encoder. The encoder decoder does not include an adaptive post filter 138, but instead includes a perceptual weighting filter.
Figure 6 is a flow chart corresponding to the operation of the decoder of Figure 5. In the decoder, the speech is reconstructed from the received parameters (block 150). In particular, the received value of the codebook index is input to the codebook that generates a vector of the code or output value of the codebook (block 152). The multiplier receives the code vector along with the received G codebook gain and multiplies these values (block 154), the resulting signal being provided to the pitch synthesis filter. It should be noted that the codebook gain G is reconstructed by decoding and inverse quantizing the received DPCM parameters. The tone synthesis filter is provided with the received tone delay L and gain b values, along with the multiplier output signal, to allow filtering of the multiplier output (block 156).
The values obtained after filtering the codebook vector by the pitch synthesis filter are entered into the formant synthesis filter. Also, the formant synthesis filter is provided with the LPC coefficients a, for use in filtering the output signal of the tone synthesis filter (block 158). The LPC coefficients are reconstructed in the decoder for interpolation by decoding the received DPCM parameters at quantized LSP frequencies, inverse quantization of the LSP frequencies, and transformation of the frequencies.
IS 2 240 252 T3
LSP in LPC coefficients a¡. The output of the formant synthesis filter is provided to the adaptive postfilter in which quantization noise is masked and in which the reconstructed speech is subjected to gain control (block 160). Reconstructed speech is obtained (block 162) for analog conversion.
With reference to the block diagram illustration of Figures 7a and 7b, more information is provided about the speech coding techniques of the present invention. In Figure 7a, each of the frames of the digitized speech samples is provided to a Hamming window subsystem 200, in which the input speech is windowed prior to calculation of the autocorrelation coefficients in the autocorrelation 202.
The Hamming window subsystem 200 and the autocorrelation subsystem 202 are illustrated in an exemplary embodiment in Figure 8. The Hamming window subsystem 200 consists of a look-up table 250, which is typically read-only memory (ROM ) of 80x16 bits, and a multiplier 252. For each speed, the speech window is centered between the 139 samples<sup>to</sup> and 140<sup>to</sup> of each analysis frame having a length of 160 samples. The window for calculating the autocorrelation coefficients is thus offset by 60 frames with respect to the analysis frame.
Windowing is done using a ROM table containing 80 of the 160 W values<sub>H</sub>(n), since the Hamming window is symmetric about the center. The Hamming window shifting is accomplished by offsetting the ROM address pointer 60 positions relative to the first sample of an analysis frame. These values are multiplied with simple precision with the corresponding input speech samples by multiplier 252. Suppose that s (n) is the input speech signal in the analysis window. The voice signal subjected to poisoned s<sub>w</sub>(n) is defined by:
s<sub>W</sub>(n) = s (n + 60) W<sub>H</sub>(n) for 0 <= n <= 79 (5) <sup>Y</sup> s<sub>W</sub>(n) = s (n + 60) W<sub>h</sub>(159 - n) for 80 <= n <= 159 (6)
Examples of hexadecimal values of the contents of lookup table 250 are provided in Table II. These values are interpreted as two's complement numbers having 14 fractional bits, the table being read from left to right and top to bottom.
TABLE II
<td>0x051f</td><td>0x0525</td><td>0x0536</td><td>0x0554</td><td>0x057d</td><td>0x05b1</td><td>0x0512</td><td>0x063d</td>
<td>0x0694</td><td>0x0616</td><td>0x0764</td><td>0x07dc</td><td>0x085e</td><td>0x08ec</td><td>0x0983</td><td>0x0a24</td>
<td>OxOadO</td><td>0x0b84</td><td>0x0c42</td><td>0x0d09</td><td>0x0dd9</td><td>OxOebO</td><td>0x0190</td><td>0x1077</td>
<td>0x1166</td><td>0x125b</td><td>0x1357</td><td>0x1459</td><td>0x1560</td><td>0x166d</td><td>0x1771</td><td>0x1895</td>
<td>0x19af</td><td>0x1acd</td><td>0x1bee</td><td>0x1d11</td><td>0x1e37</td><td>0x115e</td><td>0x2087</td><td>0x21bO</td>
<td>0x22da</td><td>0x2403</td><td>0x252d</td><td>0x2655</td><td>0x277b</td><td>0x28a0</td><td>0x29c2</td><td>0x2ae1</td>
<td>0x2bfd</td><td>0x2d15</td><td>0x2e29</td><td>Ox2f39</td><td>0x3043</td><td>0x3148</td><td>0x3247</td><td>0x3331</td>
<td>0x3431</td><td>0x351c</td><td>0x3600</td><td>0x36db</td><td>0x37al</td><td>0x387a</td><td>0x393d</td><td>0x3916</td>
<td>0x3aa6</td><td>0x3b4c</td><td>0x3be9</td><td>0x3c7b</td><td>0x3d03</td><td>0x3d80</td><td>0x3df3</td><td>0x3e5b</td>
<td>0x3eb7</td><td>0x3109</td><td>0x3141</td><td>0x3í89</td><td>0x3fb8</td><td>0x3idb</td><td>0x3113</td><td>0x3fff</td>
The autocorrelation subsystem 202 consists of a register 254, a multiplexer 256, a shift register 258, a multiplier 260, an adder 262, a circular shift register 264, and a buffer 266. Every 20 ms, voice samples are obtained windowed s<sub>W</sub>(n) and are blocked in register 254. In the sample s<sub>W</sub>(0), the first sample of an LPC analysis frame, shift registers 258 and 264 are set to 0. For each new sample s<sub>W</sub>(n), the multiplexer 256 receives a new sample selection signal that allows the input of the sample from register 254. The new sample s<sub>W</sub>(n) is also passed to multiplier 260 where it is multiplied by the sample s<sub>W</sub>(n-10), which is the last position SR10 of the shift register 258. The resulting value is added in the adder 262 with the value of the last position CSR11 of the circular shift register 264.
Shift registers 258 and 260 are iteratively shifted once, substituting s<sub>W</sub>(n-1) for sW (n) in the first position SR1 of shift register 258 and substituting the value that was previously present in position CSR10. Upon iterative shift of shift register 258, the new sample select signal is removed from the input of multiplexer 256, thereby allowing the sample to s<sub>W</sub>(n-9) currently in position SR10 of shift register 260 enters multiplexer 256.
IS 2 240 252 T3
In circular shift register 264, the value previously in position CSR11 is shifted to the first position CSR1. Once the new sample selection signal is withdrawn from the multiplexer, shift register 258 prepares to provide a circular shift of shift register data like that of circular shift register 264.
Shift registers 258 and 264 are iteratively shifted 11 times in total for each sample, thereby performing 11 multiplication / accumulation operations. Once 160 samples have been iteratively input, the autocorrelation results, which are contained in circular shift register 264, are iteratively transmitted to buffer 266 as R (0) -R (10) values. All shift registers are cleared, and the procedure is repeated for the next frame of windowed speech samples.
Referring again to Figure 7a, when the autocorrelation coefficients for the speech frame have already been calculated, the rate determination subsystem 204 and the LPC analysis subsystem 206 use these data to calculate, respectively, a transmission rate. raster data and LPC coefficients. Since these operations are independent of each other, they can be calculated in any order or even simultaneously. For explanatory purposes, the speed determination will be described first.
The speed determination subsystem 204 has two functions:
(1) determine the current frame rate, and (2) calculate a new approximate value for the background noise level. The speed of the current analysis frame is initially determined based on the energy of the current frame, the previous estimate of the background noise level, the previous speed, and the speed command of the control microprocessor. The new background noise level is calculated using the previous background noise level calculation and the current frame energy.
The vocoder uses an adaptive threshold setting technique for speed determination. Along with the change in background noise occurs the change of thresholds that are used for speed selection. In the exemplary embodiment, three thresholds are calculated to determine a preliminary speed selection RT<sub>p</sub>. The thresholds are the quadratic functions of the background noise calculation shown below:
T1 (B) = -5.544613 (10<sup>-6</sup>) B<sup>2</sup> + 4.047152 B + 363.1293; (7)
T2 (B) = -1.529733 (10<sup>-5</sup>) B<sup>2</sup> + 8.750045 B + 1136.214; (8)<sup>Y</sup>
T3 (B) = -3.957050 (10<sup>-5</sup>) B<sup>2</sup> + 18.89962 B + 3346.789 (9) where B is the previous background noise calculation.
The energy of the frame is compared to the three thresholds T1 (B), T2 (B) and T3 (B). If the frame energy is below the three thresholds, the lowest transmission rate (1 Kbit / s) is selected, that is, the rate 1/8, in which RT<sub>p</sub> = 4. If the frame energy is below two thresholds, the second transmission rate (2 Kbit / s) is selected, that is, the 1/4 rate, at which RT<sub>p</sub> = 3. If the frame energy is below a threshold only, the third transmission rate (4 Kbit / s) is selected, that is, rate 1/2, where RTp = 2. If the frame energy is above the three thresholds, the highest transmission rate (8 Kbit / s) is selected, that is, rate 1, where RTp = 1.
Preliminary speed RT<sub>p</sub> can then be modified based on the final speed of the previous frame RT<sub>r</sub>. If the preliminary velocity RT<sub>p</sub> is less than the final speed of the previous frame minus one (RT<sub>r</sub>-1), an intermediate speed RT is established<sub>m</sub>, where RT<sub>m</sub>= (RT<sub>r</sub>-1). This modification procedure determines that the speed drops slowly when a high energy signal transitions to a low energy signal. However, if the initial speed selection is greater than or equal to the previous speed minus one (RT<sub>r</sub>-1), the intermediate speed RTm is set to the same value as the preliminary speed RTp, that is, RTm = RTp. In this situation, the speed increases immediately, therefore, when there is a transition from a low energy signal to a high energy signal.
Finally, the intermediate speed RTm is further modified by speed limit commands from a microprocessor. If the speed RTm is higher than the highest speed allowed by the microprocessor, the initial speed RT, is set to the highest possible value. Similarly, if the intermediate velocity RT<sub>m</sub> is less than the lowest speed allowed by the microprocessor, the initial speed RT, is set to the lowest allowed value.
In certain cases, it may be desirable to encode the entire speech at a rate determined by the microprocessor. Rate limit commands can be used to set the desired frame rate by setting
ES 2 240 252 T3 the maximum and minimum speeds allowed at the desired speed. Speed limit commands can be used in special speed control situations such as speed interlock, and “fade-burst” transmission, both described below.
Figure 9 provides an execution example of the rate decision algorithm. To start the calculation, register 270 is preloaded with the value 1 that is provided to adder 272. Circular shift registers 274, 276, and 278 are loaded respectively with the first, second, and third coefficients of the quadratic equations of threshold (7) - (9). For example, the last, middle, and first positions of circular shift register 274 are respectively loaded with the first coefficient of the equations from which T1, T2, and T3 are calculated. Likewise, the last, middle, and first positions of the circular shift register 276 are loaded respectively with the second coefficient of the equations with which T1, T2, and T3 are calculated. Finally, the last, middle and first positions of the circular shift register 278 are respectively loaded with the constant term of the equations with which T1, T2 and T3 are calculated. In each of the circular shift registers 274, 276, and 278, the value is obtained from the last position.
When calculating the first threshold T1, the background noise calculation of the previous frame B is squared by multiplying the value by itself at multiplier 280. The B value<sup>2</sup> resulting is multiplied by the first coefficient, -5.544613 (10 <sup>6</sup>), which is obtained from the last position of the circular shift register 274. This resulting value is added in the adder 286 with the product of the background noise B and the second coefficient, 4.047152, obtained from the last position of the shift register. circular shift 276, of multiplier 284. The value obtained from adder 286 is then added at adder 288 with the constant term, 363.1293, obtained from the last position of circular shift register 278. The output of adder 288 is the calculated value of T1.
The calculated value of T1 obtained from adder 290 is subtracted in adder 288 from the energy value of frame E<sub>F </sub>which, in the following description, is the R (0) value of the linear domain, provided by the autocorrelation subsystem.
In an alternative implementation, the frame energy Ef can also be represented in the logarithmic domain in dB, where it is roughly calculated by the logarithm of the first autocorrelation coefficient R (0) normalized by the effective length of the window:
Ef = lOlogjo
R (0)
La / 2 (10) where L<sub>TO</sub> the length of the autocorrelation window. It should also be understood that speech activity can also be measured from various other parameters, including pitch prediction gain or G-formant prediction gain.<sub>to</sub>:
G<sub>to</sub> = praise
AND<sup>(10)</sup>
AND*<sup>0</sup>) (11) where E<sup>(10)</sup> the energy of the prediction residual after 10<sup>to</sup> iteration and E<sup>(0)</sup>, the energy of the initial LPC prediction residual, described below with respect to LPC analysis, which is equal to R (0).
From the output of adder 290, the complement of the sign bit of the resulting two's complement difference is extracted by comparator or limiter 292 and provided to adder 272, where it is added with the output of register 270. Thus Therefore, if the difference between R (0) and T1 is positive, register 270 is increased by one. If the difference is negative, register 270 remains the same.
The circular registers 274, 276 and 278 move iteratively, obtaining at their output the coefficients of the equation for T2, that is, equation (8). The procedure of calculating the threshold value T2 and comparing it with the energy of the frame is repeated as described in connection with the procedure for the threshold value T1. The circular registers 274, 276 and 278 are iteratively shifted again, obtaining the coefficients of the equation for T3, that is, equation (9), at their output. The calculation of the threshold value T3 and the comparison with the frame energy have already been described previously. After the three threshold calculations and comparisons have been performed, register 270 will contain the initial velocity calculation RT ,. Preliminary velocity calculation RT<sub>p</sub> is provided to the speed down logic 294. The logic 294 is also given the final speed of the previous frame RT<sub>r</sub> from the LSP frequency quantization subsystem that is stored in register 298. Logic 296 calculates the value (RT<sub>r</sub>-1) and, at the output, gives the highest value among the preliminary speed calculation RT<sub>p</sub> and the value (RT<sub>r</sub>-1). The RT value<sub>m</sub> is provided to the logic of the speed limiter 296.
As mentioned above, the microprocessor provides speed limit commands to the vocoder, in particular to logic 296. In a digital signal processor execution, this command is received in logic 296 before the LPC analysis part the encoding procedure has finished. Logic 296 ensures that the speed does not exceed the speed limits and modifies the RT value<sub>m</sub> if it exceeds the limits. If the RTm value is within the permitted speed range, logic 296 provides it as the initial speed value.
IS 2 240 252 T3
RT¡. The initial velocity value RT, is passed from logic 296 to LSP quantization subsystem 210 of Figure 7a.
The background noise calculation mentioned above is used in the calculation of the adaptive speed thresholds. For the current frame, the background noise calculation from the previous frame B is used to set the rate thresholds for the current frame. However, for each frame, the background noise calculation is updated for use in determining the rate thresholds for the next frame. The new background noise calculation B 'is determined in the current frame based on the background noise calculation of the previous frame B and the energy of the current frame E<sub>F</sub>.
When determining the new background noise calculation B 'for use during the next frame (such as the background noise calculation of the previous frame B) two values are calculated. The first value V1 is simply the energy of the current frame E<sub>F</sub>. The second V value<sub>2</sub> is the largest of B + 1 and KB, where K = 1.00547. To prevent the second value from getting too high, it is forced to stay below a constant high M = 160,000. The smaller of the two V values is chosen<sub>1</sub> and V<sub>2</sub> as a new calculation of background noise B '.
Mathematically,
Vi = R (0) (12)
V<sub>2</sub> = min (160000, max (KB, B + 1)) (13) and the new calculation of background noise B 'is:
B '= min (V1, V2) (14) where min. (x, y) the minimum of x and y, and max. (x, y), the maximum of x and y.
Figure 9 also shows an exemplary execution of the background noise calculation algorithm. The first value V1 is simply the energy of the current frame Ef supplied directly to an input of the multiplexer 300.
The second V2 value is calculated from the KB and B + 1 values, which are calculated first. When the KB and B + 1 values are calculated, the background noise calculation from the previous frame B stored in register 302 is passed to the adder 304 and multiplier 306. It should be noted that the background noise calculation from the previous frame B stored in register 302 for use in the current frame is equal to the recalculation of background noise B 'performed in the previous frame. Adder 304 is also provided with an input value of 1 to add to the value B to generate the term B + 1. The multiplier 304 is also provided with the input value K to multiply with the value B to generate the term KB. Terms B + 1 and KB are passed respectively from adder 304 and multiplier 306 to independent inputs of multiplexer 308 and adder 310.
Adder 310 and comparator or limiter 312 are used to select the larger of the B + 1 and KB terms. Adder 310 subtracts the B + 1 term from KB and provides the resulting value to comparator or limiter 312. Limiter 312 provides a control signal to multiplexer 308 to select the larger of the B + 1 and KB terms as output. The selected term B + 1 or KB passes from multiplexer 308 to limiter 314, which is a saturation type limiter, which provides the selected value if it is less than the constant value M, or the value M if it is greater than the M value. The output of limiter 314 is provided as a second input to multiplexer 300 and as an input to adder 316.
Likewise, the adder 316 receives in another input the frame energy value E<sub>F</sub>. Adder 316 and comparator or limiter 318 are used to select the smallest value between the E value<sub>F</sub> and the term provided by limiter 314. Adder 316 subtracts the frame energy value from the value provided by limiter 314 and passes the resulting value to comparator or limiter 318. Limiter 318 provides a control signal to multiplexer 300 to select the lesser of the E value<sub>F</sub> and the output of limiter 314. The selected value provided by multiplexer 300 is passed as a new B 'background noise calculation to register 302 where it is stored for use during the next frame as a previous frame B background noise calculation.
Referring again to Figure 7, each of the autocorrelation coefficients R (0) -R (10) passes from the autocorrelation subsystem 202 to the LPC 206 analysis subsystem. The LPC coefficients are calculated in the LPC 206 analysis subsystem, in the perceptual weighting filter 52 and in the formant synthesis filter 60.
The LPC coefficients can be obtained by the autocorrelation procedure using Durbin recursion as indicated in the document Digital Processing of Speech Signals, by Rabiner and Schafer, Prentice-Hall, Inc., 1978. This technique it is an efficient calculation procedure to obtain the LPC coefficients. The algorithm can be expressed by the following equations:
IS 2 240 252 T3
AND<sup>(0)</sup> = R (0), i = 1; ki = | r (í) -Σ a (<sup>i</sup>’<sup>1)</sup>R (i - J)} / E<sup>(i-1)</sup>; α. ' = ki;
j = aJ<sup>i-1)</sup> K<sub>i</sub>to*-<sup>-1)</sup> for 1 <= J <= i - 1;
AND<sup>(i)</sup> = (1 - k<sub>i</sub><sup>2</sup>)AND<sup>(i-1)</sup>;
If i <10, then return to equation (16) with i = i + 1 (15) (16) (17) (18) (19) (20)
The ten LPC coefficients are designated by the labels α /<sup>10)</sup>, for 1 <j <10.
Before encoding the LPC coefficients, the stability of the filter must be ensured. Filter stability is achieved by radially scaling the filter poles inward by a small amount, thereby reducing the magnitude of the peak frequency responses while widening the bandwidth of the peaks. This technique is commonly called bandwidth extension. For a more detailed description of this technique, see the article "Spectral Smoothing in PARCOR Speech Analysis-Synthesis" by Tohkura et al., ASSP Transactions, December 1978 In the present case, the bandwidth expansion can be carried out efficiently by scaling each LPC coefficient. Therefore, as set forth in Table III, each of the resulting LPC coefficients is multiplied by a corresponding hexadecimal value to give the final output LPC coefficients α<sub>1</sub> - α<sub>10</sub> of the LPC 206 analysis subsystem. It should be noted that the values presented in Table III are hexadecimal and that the 15 fractional bits are provided in two's complement notation. Thus, the value 0x8000 represents -1.0 and the value 0x7333 (or 29491) represents 0.899994 = 29491/32768.
TABLE III α1 = α1<sup>(10)</sup>0x7333 α2 = α2<sup>(10)</sup>0x67ae α3 = α<sub>3</sub><sup>(10)</sup>0x5d4f α4 = α4<sup>(10)</sup>0x53fb α5 = α<sub>5</sub><sup>(10)</sup>0x4b95 α6 = α6<sup>(10)</sup>0x4406 α7 = α<sub>7</sub><sup>(10)</sup>0x3d38 α8 = α8<sup>(10)</sup>0x3719 α9 = α<sub>9</sub><sup>(10)</sup>0x3196 α<sub>10</sub> = α<sub>10</sub><sup>(10)</sup>0x2ca1
The operations are preferably carried out with double precision, that is, with 32-bit divisions, multiplications and additions. Double precision accuracy is preferred to maintain the dynamic range of autocorrelation functions and filter coefficients.
In Figure 10, a block diagram of an exemplary embodiment of the LPC 206 subsystem is shown, executing equations (15) - (20) above. The LPC subsystem 206 consists of three circuit parts, a main calculation circuit 330 and two buffer update circuits 332 and 334 that are used to update the registers of the main calculation circuit 330. The calculation begins by loading first. the R (1) -R (10) values in buffer 340. To begin the calculation, register 348 is preloaded with the value R (1) by means of multiplexer 344. The register is initialized with R (0) by means of multiplexer 350, buffer 352 (containing 10 a¡<sup>(i-1)</sup> values) is initialized with only zeros via multiplexer 354, the buffer (containing 10 aj (i) values) is initialized with all zeros via multiplexer 358, and i is set to 1 for the computation cycle. For clarity, the counters for i and j and other computational controls are not shown, as those skilled in the art of digital logic design are highly trained to carry out the design and integration of these types of logic circuits.
The value aj<sup>(i-1)</sup> is obtained from buffer 356 to calculate the term k<sub>i</sub>AND<sup>(i-1)</sup>indicated in equation (14).
IS 2 240 252 T3
Each R (ij) value is obtained from buffer 340 for multiplication with the value σ /<sup>1-1)</sup> at multiplier 360. Each resulting value is subtracted in adder 362 from the value in register 346. The result of each subtraction is stored in register 346 from where the next term is subtracted. There are i-1 multiplications and accumulations in the i-th cycle, as indicated by the summation term in equation (14). At the end of this cycle, the value in register 346 is divided by the divisor 264 by the value E<sup>(i-1)</sup>of register 348 to provide the ki value.
The value k, is then used in the buffer update circuit 332 to calculate the value E<sup>(i)</sup> as in equation (19) above, which is used as E value<sup>(i-1)</sup> during the next cycle of calculating k,. The current cycle value k¡ is multiplied by itself at multiplier 366 to obtain the value k¡<sup>2</sup>. The value k<sup>2</sup> is then subtracted from the value of 1 in adder 368. The result of this sum is multiplied in multiplier 370 with the value E<sup>(¡)</sup> from register 348. The resulting value E<sup>(¡)</sup> entered into register 348 via multiplexer 350 for storage as value E<sup>(</sup>¡<sup>-1)</sup> for the next cycle.
The value k, is then used to calculate the value α, '<sup>1</sup>'as in equation (15). In this case, the value k¡ is entered into the buffer 356 by means of the multiplexer 358. Also, the value k¡ is used in the buffer update circuit 334 to calculate the values ja from the values aj<sup>(</sup>¡<sup>-1)</sup> as in equation (18). The values currently stored in buffer 352 are used to calculate the values "¡!<sup>α)</sup>. As indicated in equation (18), there are i-1 calculations in the i-th cycle. In iteration i = 1, such calculations are not required. For each value of j of the i-th cycle, a value of α} is calculated<sup>ω</sup>. When calculating each value of Oj<sup>(</sup>¡<sup>)</sup>, each value of a, -j<sup>(</sup>¡<sup>-1)</sup> it is multiplied in multiplier 372 with the value k¡ to pass it to adder 374. In adder 374, the value k¡a¡-j<sup>(</sup>¡<sup>-1)</sup> is subtracted from the value Oj<sup>(</sup>¡<sup>-1)</sup> which is also input to adder 374. The result of each multiplication and addition is provided as a value of j to buffer 356 via multiplexer 358.
Once the α values have been calculated, '<sup>1</sup>'and j for the current cycle, the values just calculated and stored in buffer 356 are passed to buffer 352 via multiplexer 354. Values stored in buffer 356 are stored in corresponding locations in buffer 352 In this way, the buffer 352 is updated for the calculation of the value k, of cycle i + 1.
It is important to note that the data Oj<sup>(</sup>¡<sup>-1)</sup> generated at the end of a previous cycle are used during the current cycle to generate updates j for the next cycle. Data from the previous cycle must be preserved to fully generate updated data for the next cycle. In this manner, the two buffers 356 and 352 are used to retain this previous cycle data until the updated data has been fully generated.
The above description refers to a parallel transfer of data from buffer 356 to buffer 352 until the calculation of the updated values is completed. This execution ensures that old data is preserved throughout the new data calculation process, with no loss of old data until it has been fully used, as in a single buffer arrangement. The run described is one of several runs available that achieve the same result. For example, buffers 352 and 356 can be multiplexed such that, after calculating the value k, for a current cycle from the values stored in a first buffer, updates are stored in the second buffer for use. during the next calculation cycle. In this next cycle, the value k, is calculated from the values stored in the second buffer. The values of the second buffer and the value k, are used to generate updates for the next cycle, these updates being stored in the first buffer. This alternation of buffers allows the retention of the values of the current calculation cycle, from which the updates are generated, and at the same time, the storage of the update values without overwriting the current values that are necessary for generate the updates. Using this technique, the delay associated with calculating the k, value for the next cycle can be minimized. Therefore, the updates for the multiplications / accumulations of the calculation of k, can be carried out at the same time as the next value of Oj is calculated<sup>(</sup>¡<sup>-1)</sup>.
The ten LPC Oj coefficients<sup>(10)</sup>, stored in buffer 356 after the end of the last calculation cycle (i = 10), are scaled to arrive at the corresponding final LPC coefficients. Scaling is accomplished by providing a scale select signal to multiplexers 344, 376, and 378 so that scale values stored in look-up table 342, hexadecimal values from Table III, are selected to be provided. through multiplexer 344. The values stored in the look-up table 342 are extracted iteratively in sequence and introduced into the multiplier 360. Likewise, the multiplier 360 receives through the multiplexer 376 the «j<sup>(10)</sup> Sequentially obtained values from register 356. The scaled values are provided from multiplier 360 via multiplexer 378 as output to the LPC-LSP transformation subsystem 208 (Figure 7).
To efficiently encode each of the ten scaled LPC coefficients in a reduced number of bits, the coefficients are transformed into spectral line pair frequencies as described in the article "Line Spectrum Pair (LSP) and Speech Data Compression" (" Spectral Line Pair (LSP) and Voice Data Compression ”), from Soong and Juang, ICASSP '84. Next, the calculation of the LSP parameters is shown in equations (21) and (22) together with Table IV.
IS 2 240 252 T3
The LSP frequencies are the ten roots between 0 and π of the following equations:
Ρ (ω) = cos 5ω + p<sub>1</sub> cos 4ω + ... + p<sub>4</sub> cos ω + p<sub>5</sub>/2; (21)
Q / ω) = cos 5ω + qi cos 4ω + ... + q<sub>4</sub> cos ω + q<sub>5</sub>/two; (22) in which the p-values<sub>n</sub> and what<sub>n</sub> for n = 1,2, 3,4 and 5 they are defined repetitively in Table IV.
TABLE IV
<td>p1 = - (α + α1ο) - 1</td><td>q1 = - (α - α) + 1</td>
<td><sup>p</sup>2 = <sup>- (α</sup>2 + <sup>α</sup>9) <sup>- p</sup>1</td><td><sup>what</sup>2 = <sup>- (α</sup>2<sup>-α</sup>9) + <sup>what</sup>1</td>
<td><sup>p</sup>3 = <sup>- (α</sup>3 + <sup>α</sup>8) <sup>- p</sup>2</td><td><sup>what</sup>3 = <sup>- (α</sup>3<sup>-α</sup>8) + <sup>what</sup>2</td>
<td><sup>p</sup>4 = <sup>- (α</sup>4 + <sup>α</sup>7 ) <sup>- p</sup>3</td><td><sup>what</sup>4 = <sup>- (α</sup>4<sup>-α</sup>7) + <sup>what</sup>3</td>
<td><sup>p</sup>5 = <sup>- (α</sup>5 + <sup>α</sup>6) <sup>- p</sup>4</td><td><sup>what</sup>5 = <sup>- (α</sup>5<sup>-α</sup>6) + <sup>what</sup>4</td>
In Table IV, the α values<sub>1</sub>, ..., α<sub>10</sub> are the scaled coefficients resulting from the LPC analysis. For simplicity, the ten roots of equations (21) and (22) are scaled by a value between 0 and 0.5. A property of LSP frequencies is that, if the LPC filter is stable, the roots of the two functions alternate; that is, the lowest root, ω<sub>1</sub>, is the lowest root of P (p), the next lowest root, ω<sub>2</sub>, is the lowest root of Q / ω), and so on. Of the ten frequencies, the odd frequencies are the roots of P (p), and the even frequencies are the roots of Q / ω).
The root search is carried out as described below. First, the coefficients p and q are calculated with double precision by summing the LPC coefficients as shown above. Then, every π / 256 radians, the evaluation of P (p) is carried out and these values are then evaluated to check sign changes, which indicate a root in that subzone. If a root is found, then a linear interpolation is made between the two boundaries of this zone to roughly calculate the location of the root. The existence of a root Q is guaranteed between each pair of roots P (the fifth root Q is between the fifth root P and π), due to the ordering property of the frequencies. A binary search is performed between each pair of P roots to determine the location of the Q roots. For ease of execution, each root P is roughly computed using the closest π / 256 value and the binary search is performed between these approximate computations. If no roots are found, the previous unquantized values of the LSP frequencies from the last frame in which the roots were found are used.
In Figure 11, an example of the execution of the circuits used to generate the LSP frequencies is illustrated. The operation described above requires a total of 257 possible cosine values between 0 and π, which are stored with double precision in a look-up table, the cosine look-up table 400, which is accessed by counter 402 of module 256. For each value of j entered in lookup table 400, an output of cos ω, cos 2 ω, cos 3 ω, cos 4 ω, and cos 5 ω is provided, where:
ω = jn / 256 (23) where j is a counter value.
The values of cos ω, cos 2 ω, cos 3 ω and cos 4 ω obtained from the look-up table 400 are entered in a respective multiplier 404,406,408 and 410, while the value of cos 5 ω is entered directly in the adder 412. These values are multiplied in a respective multiplier 404,406,408 and 410 with a respective value of the p-values<sub>4</sub>, p<sub>3</sub>, p<sub>2</sub> And p<sub>1</sub> entered therein by means of multiplexers 414, 416, 418 and 420. The values resulting from this multiplication are also entered into adder 412. Furthermore, the value p<sub>5</sub> is provided through multiplexer 422 to multiplexer 424, the constant value being 0.5, ie 1/2, also provided to multiplier 424. The resulting value obtained from multiplier 424 is provided as another input to adder 412. The multiplexers 414-422 select between p-values<sub>1</sub>-p<sub>5</sub> oq<sub>1</sub>-q<sub>5</sub>, in response to a p / q coefficient selection signal, to use the same circuits to calculate either the P / ω) or the Q / ω) values. The circuits to generate the p-values<sub>1</sub> -p<sub>5</sub> oq<sub>1</sub> -q<sub>5</sub> not shown, but easily executed using a series of adders to add and subtract the LPC coefficients and p-values<sub>1</sub>-p<sub>5</sub> oq<sub>1</sub> -q<sub>5</sub>, along with registers to store the p-values<sub>1</sub>-p<sub>5</sub> or what<sub>1</sub> -q<sub>5</sub>.
The adder 412 sums the input values to provide the output value P / ω) or Q / ω) as appropriate. To facilitate the description, the case of the P / ω) values will be considered, the Q / ω) values being calculated in a similar way using the q values<sub>1</sub>-q<sub>5</sub>. The current value of P / ω) is obtained from adder 412 and stored in register 426. The preceding value of P / ω), previously stored in register 426, is shifted to register 428. The sign bits of the current and previous values of P / ω) are exclusive-ORed at exclusive-OR gate 430 to give a zero crossing or sign change indication, in the form of an enable signal that is sent to linear interpolator 434. The current and previous value of P / ω) are also passed from registers 426 and 428 to linear interpolator 434,
ES 2 240 252 T3 which is sensitive to the enable signal, to interpolate the point between the two values of P (m) where the zero crossing occurs. This linear interpolation fractional value result, that is, the distance from the j-1 value, is provided to the buffer 436 along with the j value of the counter 256. Gate 430 also provides the enable signal to buffer 436 which allows storage of the value j and the corresponding fractional value FVj.
The fractional value is subtracted from the j value when it is entered into the adder 438 from the buffer 436 or, alternatively, it can be subtracted from it when it is entered into the buffer 436. On the other hand, a register from line j can be used entered into buffer 436 so that the value j-1 is entered into buffer 436, the fractional value being entered into it as well. The fractional value can be added to the j-1 value either before storage in register 436 or after it is output. In either case, the combined value of j + FVj or (j-1) + FVj is passed to the divisor 440 where it is divided by the constant input value of 512. The division operation can be performed simply by changing the binary location of the point at the representative binary word. This division operation provides the scaling necessary to arrive at an LSP frequency between 0 and 0.5.
Each evaluation function of P (m) or Q (m) requires 5 cosine queries, 4 double precision multiplications, and 4 additions. Typically calculated roots only have a precision of about 13 bits, and are stored with single precision. The LSP frequencies are provided to the LSP 210 quantization subsystem (Figure 7) for quantization.
Once the LSP frequencies have been calculated, they must be quantized for transmission. Each of the ten LSP frequencies is roughly centered around an offset value. It should be noted that the LSP frequencies approach the offset values when the input speech has uniform spectral characteristics and short-term prediction cannot be performed. Offsets are subtracted at the encoder, and a simple DPCM quantizer is used. At the decoder, the runout is added back. Table V shows the negative hexadecimal values of the offset value, for each LSP frequency, ω<sub>1</sub> -ω<sub>10</sub>, provided by the LPC-LSP transformation subsystem. Again, the values given in Table V are in two's complement notation with 15 fractional bits. The hexadecimal value 0x8000 (or -32768) represents -1.0. Therefore, the first value in Table V, the value 0xfa2f (or -1489) represents -0.045441 = -1489/32768.
TABLE V
<td>LSP frequency</td><td>Negative runout value</td>
<td>ω1</td><td>0xfa2f</td>
<td>ω2</td><td>0xf45e</td>
<td>ω3</td><td>0xee8c</td>
<td>ω4</td><td>0xe8bb</td>
<td>ω5</td><td>0xe2e9</td>
<td>ω6</td><td>0xdd18</td>
<td>ω7</td><td>0xd746</td>
<td>ω8</td><td>0xd175</td>
<td>ω9</td><td>0xcba3</td>
<td>ω10</td><td>0xc5d2</td>
The predictor used in the subsystem is 0.9 times the quantized LSP frequency of the previous frame stored in a subsystem buffer. This decay constant of 0.9 is inserted so that the channel errors eventually disappear.
The quantizers used are linear, but vary in dynamic range and step size with speed. Also, in high-speed frames, more bits are transmitted for each LSP frequency, and therefore the number of quantization levels is speed-dependent. Table VI shows the bit allocation and dynamic range of quantization for each frequency at each of the rates. For example, at speed 1, ω<sub>1</sub> it is quantized uniformly using 4 bits (ie, in 16 levels) with the highest quantization level being 0.025 and the lowest being -0.025.
IS 2 240 252 T3
TABLE VI
<td>VELOCITY</td><td>Complete</td><td>Half</td><td>Bedroom</td><td>Eighth</td>
<td>ω1</td><td> 4:+0,025</td><td> 2:+0,015</td><td> 1:+0,01</td><td> 1:+0,01</td>
<td>ω2</td><td> 4:+04</td><td> 2:+0,015</td><td> 1:+0,01</td><td> 1:+0,015</td>
<td>ω3</td><td> 4:+07</td><td> 2:+0,03</td><td> 1:+0,01</td><td> 1:+0,015</td>
<td>ω4</td><td> 4:+07</td><td> 2:+0,03</td><td> 1:+0,01</td><td> 1:+0,015</td>
<td>ω5</td><td> 4:+06</td><td> 2:+0,03</td><td> 1:+0,01</td><td> 1:+0,015</td>
<td>ω6</td><td> 4:+06</td><td> 2:+0,02</td><td> 1:+0,01</td><td> 1:+0,015</td>
<td>ω7</td><td> 4:+05</td><td> 2:+0,02</td><td> 1:+0,01</td><td> 1:+0,01</td>
<td>ω8</td><td> 4:+05</td><td> 2:+0,02</td><td> 1:+0,01</td><td> 1:+0,01</td>
<td>ω9</td><td> 4:+04</td><td> 2:+0,02</td><td> 1:+0,01</td><td> 1:+0,01</td>
<td>ω10</td><td> 4:+04</td><td> 2:+0,02</td><td> 1:+0,01</td><td> 1:+0,01</td>
<td>Total</td><td>40 bit</td><td>20 bits</td><td>10 bits</td><td>10 bits</td>
If the quantization ranges for the speed chosen by the speed decision algorithm are not wide enough or a slope overflow occurs, the speed is raised to the next highest speed. The speed continues to increase until the dynamic range is accommodated or full speed is reached. An example block diagram illustration of one execution of the optional speed climb technique is provided in Figure 12.
Figure 12 is a block diagram illustrating an exemplary implementation of the LSP 210 quantization subsystem that includes the up-rate circuits. In Figure 12, the LSP frequencies of the current frame are passed from divider 440 (Figure 11) to register 442, where they are stored to be provided during a rate-up determination in the next frame. The LSP frequencies of the previous frame and the LSP frequencies of the current frame are respectively passed from register 440 and divider 440 to a rate-up logic 442 for a rate-up determination of the current frame. The speed climb logic 442 also receives the initial speed decision, along with the speed limit commands from the speed determination subsystem 204. To determine if a speed increase is necessary, logic 442 compares the LSP frequencies of the previous frame with the LSP frequencies of the current frame, based on the sum of the square of the difference between the LSP frequencies of the current frame. and the previous plot. The resulting value is then compared to a threshold value which, if exceeded, indicates that a speed increase is necessary to ensure high-quality encoding of speech. When the threshold value is exceeded, logic 442 increases the initial speed by one speed level to provide an output of the final speed to always be used by the encoder.
In Figure 12, the LSP frequency values ωι -ω<sub>ί0</sub> they are entered one at a time into adder 450 along with the corresponding offset value. The offset value is subtracted from the entered LSP value and the result is passed to adder 452. Adder 452 also receives as input a predictor value, an LSP value corresponding to the previous frame multiplied by a decay constant. The predictor value is subtracted from the output of adder 450 by adder 452. The output of adder 452 is provided as input to quantizer 454.
Quantizer 454 consists of limiter 456, minimum dynamic range lookup table 458, reverse step size lookup table 460, adder 462, multiplier 464, and bit mask 466. Quantization is performed at quantizer 454, first determining whether the input value is within the dynamic range of quantizer 454. The input value is provided to limiter 456 which limits the input value to the upper and lower limits of the dynamic range if the input exceeds the limits provided by the look-up table 458. The look-up table 458 provides the stored limits, according to Table VI, to limiter 456 in response to input of speed and frequency index LSP i. The value obtained from the limiter 456 is input into the adder 462 where it is subtracted from the minimum of the dynamic range, provided by the look-up table 458. The value obtained from the look-up table 458 is again determined by the speed and the frequency index LSP i, according to the minimum dynamic range values (regardless of their sign) set forth in Table VI. For example, the value in lookup table 458 for (full speed, ωι) is 0.025.
Next, the output of adder 462 is multiplied at multiplier 464 by a value selected from look-up table 460. Look-up table 460 contains values corresponding to the inverse of the step size for each LSP value of each speed, depending on the values set forth in Table VI. The value obtained from the look-up table 460 is selected by the speed and the frequency index LSP i. For each speed and frequency index LSP i, the value stored in lookup table 460 is the amount ((2<sup>n</sup>-1) / dynamic range), where n is the number of
ES 2 240 252 T3 bits representing the quantized value. Also, for example, the value in lookup table 460 for (speed 1, ωι) is (15 / 0.05) or 300.
The output of multiplier 464 is a value between 0 and 2<sup>n</sup>-1 which is provided to bit mask 466. Bit mask 466, in response to speed and LSP frequency index, extracts the appropriate number of bits from the input value according to Table VI. The bits extracted are the n integer bits of the input value to provide a limited bit output Δω ,. The Δω values are the quantized differential coding centered quantized LSP frequencies that are transmitted through the representative channel of the LPC coefficients.
The Δω value is also applied as feedback through a predictor consisting of the inverse quantizer 468, the adder 470, the buffer 472, and the multiplier 474. The inverse quantizer 468 consists of the step size lookup table 476 , the minimum dynamic range lookup table 478, the multiplier 480, and the adder 482.
The value Δω is entered in the multiplier 480 together with a value selected in the look-up table 476. The look-up table 476 contains the values corresponding to the step size of each LSP value for each of the speeds, according to the values shown. in Table VI. The value obtained from the look-up table 476 is selected by the speed and the frequency index LSP i. For each speed and frequency index LSP i, the value stored in lookup table 460 is the quantity (dynamic range / 2<sup>n</sup>-1), where n is the number of bits that represent the quantized value. The multiplier 480 multiplies the input values and provides an output to the adder 482.
The adder 482 receives as another input a value from the look-up table 478. The value obtained from the look-up table 478 is determined by the speed and the frequency index LSP, according to the values of the minimum dynamic range (regardless of the sign of the themselves) set forth in Table VI. Adder 482 adds the minimum dynamic range value provided by look-up table 478 with the value obtained from multiplier 480, the resulting value obtained being passed to adder 470.
Adder 470 receives as another input the predictor value obtained from multiplier 474. These values are added in adder 470 and stored in ten-word storage buffer 472. Each previous frame value obtained from buffer 472 during the current frame is multiplied, at multiplier 474, by the constant 0.9. Predictor values obtained from multiplier 474 are provided to adders 452 and 470 as described above.
In the current frame, the value stored in buffer 472 is the reconstructed LSP value from the previous frame minus the offset value. Similarly, in the current frame, the value obtained from adder 470 is the reconstructed LSP value of the current frame from which the offset has also been subtracted. In the current frame, the outputs of buffer 472 and adder 470 are provided, respectively, to adders 484 and 486, where the offset is added to the values. The values obtained from the adders 484 and 486 are, respectively, the reconstructed LSP frequency values of the previous frame and the reconstructed LSP frequency values of the current frame. LSP smoothing is carried out at the lowest speeds according to the following equation:
Smoothed LSP = a (current LSP) + (1-a) (previous LSP) (24) where a = 0 for full speed; a = 0.1 for half speed; a = 0.5 for quarter speed; ya = 0.85 for eighth speed.
The values ω \<sub>Μ</sub> reconstructed LSP frequencies of the previous frame (f-1) and the values ω '<sub>σ</sub> of reconstructed LSP frequencies of the current frame (f) are obtained from quantization subsystem 210 and passed to tone subframe LSP interpolation subsystem 216 and codebook subframe LSP interpolation subsystem 226. The quantized values of frequencies LSP Δω, are passed from LSP quantization subsystem 210 to data assembler subsystem 236 for transmission.
The LPC coefficients used in the weighting filter and the formant synthesis filter described below are suitable for the tone subframe being encoded. For the tone subframes, the interpolation of the LPC coefficients is carried out once for each tone subframe as indicated in Table VII:
IS 2 240 252 T3
TABLE VII
Speed 1:
<td></td><td>ωί = 0.75ω i, f<sub>-</sub>1 + 0.25ω i, f</td><td>for tone 1 subframe</td>
<td></td><td>ω<sub>i</sub> = 0.5ω \<sub>Μ</sub> + 0.5ω \<sub>ί</sub></td><td>for tone 2 subframe</td>
<td></td><td>ωi = 0.25ω í, f<sub>-</sub>1 + 0.75ω i, f</td><td>for tone 3 subframe</td>
<td></td><td>ωi = ω '<sub>σ</sub></td><td>for tone 4 subframe</td>
<td>Speed 1/2:</td><td>ω<sub>i</sub> = 0.625ω \<sub>Μ</sub> + 0.375ω \<sub>ί</sub></td><td>for tone 1 subframe</td>
<td></td><td>ωi = 0.125ω \<sub>Μ</sub> + 0.875ω \<sub>4</sub></td><td>for tone 2 subframe</td>
<td>Speed 1/4:</td><td>ω<sub>i</sub> = 0.625ω \<sub>Μ</sub> + 0.375ω \<sub>ί</sub></td><td>for tone 1 subframe</td>
<td>1/8 speed:</td><td></td><td></td>
Tone search is not performed.
The tone subframe counter 224 is used to keep track of the tone subframes for which the tone parameters are calculated, the counter output being provided to the tone subframe LSP interpolation subsystem 216 for use in interpolation. Tone Subframe LSP. The tone subframe counter 224 also provides an output indicating completion of the tone subframe for the selected rate to the data packing subsystem 236.
Figure 13 illustrates an exemplary implementation of the tone subframe LSP interpolation subsystem 216 to interpolate the lSp frequencies for the relevant tone subframe. In Figure 13, the previous and current LSP frequencies ω \<sub>Μ</sub> and ω '<sub>σ</sub> they are passed, respectively, from the LSP quantization subsystem to multipliers 500 and 502 where they are multiplied, respectively, by a constant provided by memory 504. The memory 504 stores a group of constant values and, in accordance with an input of the number of tone subframes of a tone subframe counter, which will be described later, provides an output of constants such as those set out in Table VII for its multiplication with the previous and current frame LSP values. The outputs of multipliers 500 and 502 are added, in adder 506, to provide the LSP frequency values for the tone subframe according to the equations in Table VII. For each tone subframe, once the interpolation of the LSP frequencies has been carried out, an inverse LSP-LPC transformation is performed to obtain the current coefficients of A (z) and the perceptual weighting filter. The interpolated LSP frequency values are therefore provided to the LSP-LPC transformation subsystem 218 of Figure 7.
The LSP-LPC transformation subsystem 218 converts the interpolated LSP frequencies back into LPC coefficients for use in speech resynthesis. Again, the above-mentioned article "Line Spectrum Pair (LSP) and Speech Data Compression" by Soong and Juang describes in detail the algorithm implemented in the present invention in the transformation process and indicates how it can be deduced. The calculation aspects allow to express P (z) and Q (z) in terms of the LSP frequencies by means of the equations:
P (-) = (1 + - <sup>1</sup>HI (1 - 2cos (w2i-1) - <sup>1</sup> + - <sup>2</sup>) i = 1 where ω<sub>i</sub> the roots of the polynomial P '(odd frequencies), and
Q (-) = (1 - - <sup>1</sup>HI (1 - 2cos ^ 2i) - <sup>1</sup> + - <sup>2</sup>) i = 1 where ω<sub>i</sub> the roots of the polynomial Q '(even frequencies), and
P (-) + Q (-) (25) (26) (27)
A (z)
The calculation is carried out, first calculating the values 2cos ^) for all the odd frequencies i. This calculation is performed using a fifth-order Taylor series expansion of the cosine around zero (0) with simple precision. A Taylor expansion around the nearest point on the table of cosines might be more power-accurate, but the precision provided by expansion around 0 is sufficient and does not involve an excessive amount of computation.
Next, the coefficients of the polynomial P are calculated. The coefficients of a product of polynomials is the
ES 2 240 252 T3 convolution of the sequences of coefficients of the individual polynomials. Next, we calculate the convolution of the 6 sequences of z polynomial coefficients from equation (25) above, {1, - 2cos (w<sub>1</sub>), 1}, {1, -2cos (w<sub>3</sub>), 1} ... {1, -2cos (w<sub>9</sub>), 1} and {1, 1}.
Once the polynomial P has been calculated, the same procedure is repeated for the polynomial Q, in which the 6 sequences of z coefficients of the polynomial of equation (26) above, {1, -2cos (w<sub>2</sub>), 1}, {1, -2cos (w<sub>4</sub>), 1} ... {1, -2cos (w<sub>10</sub>), 1}, and {1,1} and the appropriate coefficients are added and divided by 2, that is, shifted 1 bit, to generate the LPC coefficients.
Figure 13 also shows in detail an execution example of the LSP-LPC transformation subsystem. Circuit part 508 calculates the value of -2cos (w¡) from the input value of ω ,. Circuit portion 508 consists of buffer 509, adders 510 and 515, multipliers 511, 512, 514, 516, and 518, and registers 513 and 515. When the values of -2cos © ¡) are calculated, the registers 513 and 515 are reset. Since this circuit calculates sin (pi), we first subtract ω ,, at adder 510, from the constant input value π / 2. This value is squared by the multiplier 511, and then the values (π / 2-ω,) are calculated in sequence<sup>2</sup>, (π / 2-ω,)<sup>4</sup>, (π / 2-ω,)<sup>6</sup> and (π / 2-ω,)<sup>8</sup> using multiplier 512 and register 513.
The coefficients of the Taylor series expansion c [1] -c [4] are entered in sequence in multiplier 514 together with the values obtained from multiplier 512. The values obtained from multiplier 514 are entered in adder 515 where they are added with the output of register 516 to provide the output c [1] (π / 2-ω,)<sup>2</sup> + c [2] (π / 2-ω,)<sup>4</sup> + c [3] (π / 2-ω,)<sup>6</sup> + c [4] (π / 2-ω,)<sup>8</sup> to multiplier 517. The input to multiplier 517 of register 516 is multiplied in multiplier 517 by the output (π / 2-ω,) of adder 510. The output of multiplier 517, that is, the value cos © ¡), is multiply at the multiplier 518 by the constant -2 to provide the output -2cos © ¡). The value -2cos © ¡) is provided to circuit part 520.
Circuit portion 520 is used in computing the coefficients of the polynomial P. Circuit portion 520 consists of memory 521, multiplier 522, and adder 523. The set of memory locations P (1) ... P (11) are set to 0 except for P (1) which sets to 1. The old indexed values -2cos © ¡) are introduced in the multiplier 524 to carry out the convolution of (1, -2cos © ¡), 1) where 1 <i <5, 1 <j <2i + 1, P (j) = 0 for j <1. The part of the 520 circuit is doubled (not shown) to calculate the coefficients of the polynomial Q. The resulting new final values of P (1) -P (11) and Q (1) -Q (11) are given to the part of circuit 524.
The circuit portion 524 is provided with ten LPC coefficients a,, with i being a value between 1 and 10, to finish the tone subframe calculation. The circuit portion 524 consists of the buffers 525 and 526, the adders 527, 528, and 529, and the bit divider or shifter 530. The final P (i) and Q (i) values are stored in the buffers 525 and 526. The values P (i) and P (i + 1) are added in the adder 527, while the corresponding values Q (i) and Q (i + 1) are subtracted in the adder 528, for 1 <i <10. The output from adders 527 and 528, respectively P (z) and Q (z), is input to adder 529 where it is added and provided as the value (P (z) + Q (z)). The output of the adder is divided by two by shifting the bits one position. Each bit-shifted value of (P (z) + Q (z)) / 2 is an output LPC coefficient. The tone subframe LPC coefficients are provided to the tone search subsystem 220 of Figure 7.
Also, the LSP frequencies are interpolated for each codebook subframe determined by the selected rate, except for the full rate. Interpolation is calculated in the same way as tone subframe LSP interpolations. The codebook subframe LSP interpolations are computed in the codebook subframe LSP interpolation subsystem 226 and provided to the LSP-LPC transform subsystem 228 where the transform is computed similarly to the LSP-LPC transform subsystem 218.
As described in relation to Figure 3, pitch search is an analysis-by-synthesis technique, in which encoding is performed by selecting parameters that minimize the error between the input speech and the synthesized speech using those parameters. . In the search for pitch, the voice is synthesized using the pitch synthesis filter whose response is expressed in equation (2). Every 20 ms, the speech frame is subdivided into a number of tone subframes which, as described above, depends on the data rate chosen for the frame. Once for each pitch subframe, the parameters b and L are calculated, that is, the pitch gain and delay, respectively. In the present exemplary embodiment, the pitch delay L ranges from 17 to 143 and, for transmission reasons, L = 16 is reserved for the case where b = 0.
The speech coder uses a perceptual noise weighting filter in the manner stated in equation (1). As mentioned above, the purpose of the perceptual weighting filter is to weight the error at lower power frequencies to reduce the effect of error related noise. The perceptual weighting filter is derived from the short-term prediction filter found earlier. The LPC coefficients used in the weighting filter, and the formant synthesis filter described below, are the appropriate interpolated values for the subframe being encoded.
When performing analysis-by-synthesis operations, a copy of the speech decoder / synthesizer is used in the encoder. The shape of the synthesis filter used in the speech coder is obtained by means of the
ES 2 240 252 T3 equations (3) and (4). Equations (3) and (4) correspond to a decoder's speech synthesis filter followed by the perceptual weighting filter, thus called a weighted synthesis filter.
The pitch search is carried out under the assumption of a zero contribution from the codebook in the current frame, that is, G = 0. For each possible pitch delay, L, the voice is synthesized and compared with the original voice . The error between the input speech and the synthesized speech is weighted by the perceptual weighting filter before its root mean square error (MSE) is calculated. The goal of this is to choose values of L and b, out of all the possible values of L and b, that minimize the error between the perceptually weighted voice and the perceptually weighted synthesized voice. The minimization of the error can be expressed by the following equation:
Lp-1
MSE, Σ (x (n) - x '(n))<sup>2</sup>
Lp n = 0 (28) where LP is the number of samples of the pitch subframe which, in the exemplary embodiment, is 40 for a full rate pitch subframe. The tone gain, b, is calculated which minimizes the MSE. These calculations are repeated for all allowed values of L, and the values of L and b that generate the minimum MSE for the tone filter are chosen.
The optimal pitch delay calculation includes the formant residue (p (n) in Figure 3) for the time between n = -L<sub>max</sub> yn = (L<sub>P</sub>-L<sub>min</sub>) -1, where L<sub>max</sub> the maximum pitch delay value, L<sub>min</sub> the minimum pitch delay value and LP the length of the pitch subframe for the selected speed, and where n = 0 is the start of the pitch subframe. In the example embodiment L<sub>max</sub> = 143 and 1 ,,,,,,, = 17. Using the numbering model provided in Figure 14, for speed 1/4, n = -143 an = 142, for speed 1/2, n = -143 an = 62 and for speed 1, n = -143 to n = 22. For n <0, the formant residue is simply the output of the tone filter from the previous tone subframes, which is kept in the memory of the tone filter and it is called a residue of closed-loop formants. For n> 0, the formant residue is the output of a formant analysis filter having a filter characteristic of A (z) in which the input is the speech samples of the current analysis frame. For n> 0, the formant residue is called the open-loop formant residue and will be exactly p (n) if the tone filter and codebook make a perfect prediction on this subframe. Referring to Figures 14-17, more information will be provided about calculating the optimal pitch delay, from the associated formant residue values.
The tone search is performed with respect to 143 reconstructed samples of closed-loop formant residues, p (n) for n <0, plus LP-Lmin unquantized samples of open-loop formant residues, po (n) for n > 0. Efficiently and gradually, the search that is fundamentally an open-loop search in which L is small, and therefore most of the residue samples used are n> 0, becomes a search that is mainly a search closed-loop where L is large, and therefore all residual samples used are n <0. For example, using the numbering model provided in Figure 14 at full speed, in which the pitch subframe consists of 40 speech samples, the pitch search begins using the set of formant residue samples numbered n = - 17 to n = 22. In this model, from n = -17 to n = -1, the samples are residue samples of closed-loop formants, while from n = 0 to n = 22, the samples are residue samples of open-loop formants. The next group of formant residue samples used in determining the optimal pitch delay are the numbered samples from n = -18 to n = 21. Again, from n = -18 to n = -1, the samples are residue samples of closed-loop formants, while from n = 0 to n = 21, the samples are residue samples of open-loop formants. This process continues with the groups of samples until the pitch delay is obtained for the last group of formant residue samples, n = -143 to n = -104.
As described above in relation to equation (28), the goal is to minimize the error between x (n), the perceptually weighted voice minus the zero input response (ZIR) of the weighted formant filter, and x ' (n), the perceptually weighted synthesized voice that has no memory allocated in the filters, with respect to all possible values of L and b, given a zero contribution from the stochastic codebook (G = 0). Equation (28) can be rewritten in relation to b as follows:
Lp-1
MSE = - 2 (x (n) - by (n))<sup>2</sup>
LP n = 0 (29) in which, y (n) = h (n) * p (n - L) for 0 <n <L<sub>P</sub> - 1 (30) where y (n) is the synthesized voice weighted with the pitch delay L when b = 1, and h (n) is the impulsive response of the weighted formant synthesis filter that has the filter characteristic according to equation (3 ).
IS 2 240 252 T3
The minimization procedure is equivalent to maximizing the E value<sub>L</sub>, (Exy)<sup>2</sup><sup>AND</sup>yy (31) where,
Lp-1
Exy = Σ x (n) y (n) n = 0 (32) y,
Lp-1
Eyy = y (n) y (n) n = 0 (33)
The optimal b for the given L turns out to be:
bL
Former <sup>AND</sup>and (34)
This search is repeated for all allowed values of L. The optimal b is limited to positive values, and therefore an L that results in a negative Exy value is ignored in the search. Finally, the delay, L, and the gain, b, of the tone are chosen for the transmission, increasing the maximum EL.
As mentioned above, x (n) is actually the perceptually weighted difference between the input voice and the ZIR of the weighted formant filter, because for the recursive convolution, discussed later in equations (35) - ( 38), the assumption is that filter A (z) always starts with 0 in filter memory. However, the real case is not that of the filter that starts with a 0 in the filter memory. In short, the filter will have a state that persists from the previous subframe. In performance, the effects of the initial state are subtracted from the perceptually weighted voice at the start. In this way, it is only necessary to calculate for each L the response ap (n) of the steady-state filter A (z), with all the memories initially set to 0, being able to use recursive convolution. It is only necessary to calculate this value of x (n) once, but it is necessary to calculate y (n), the response to zero state of the formant filter at the output of the tone filter, for each delay L. The calculation of each y (n) includes many redundant multiplications, which do not need to be calculated for each delay. The recursive convolution procedure described below is used to minimize the necessary calculations.
In relation to recursive convolution, the value yL (n) is defined by the value y (n), where:
yL (n) = h (n) * p (n - L) 17 <L <143 (35) o, yL (n) = Σ h (i) p (n - L - i) 17 <L <143 ( 36)
From equations (32) and (33) it can be seen that:
yL (0) = p (-L) h (0) (37) y<sub>L</sub>(n) = y<sub>L-1</sub>(n - 1) + p (-L) h (n) 1 <n <L<sub>P</sub>, 17 <L <143 (38)
In this way, once the initial convolution for y17 (n) has been performed, the rest of the convolutions can be performed recursively, greatly reducing the number of calculations required. In the example provided above for speed 1, the value y<sub>17</sub>(n) is calculated by equation (36) using the group of formant residue samples numbered from n = -17 to n = 22.
Referring to Figure 15, the encoder includes a duplicate of the decoder of Figure 5, the decoder subsystem 235 of Figure 7, in the absence of the adaptive postfilter. In Figure 15, the input to pitch synthesis filter 550 is the product of the codebook value cI (n) and the codebook gain G. Samples of
ES 2 240 252 T3 residues of formants provided p (n) are introduced into the synthesis filter of formants 552 where they are filtered and provided as reconstructed speech samples s' (n). The reconstructed speech samples s (n) are subtracted from the corresponding input speech samples s (n) in the adder 554. The difference between the samples s (n) 'and s (n) is input into the perceptual weighting filter 556. Regarding the tone synthesis filter 550, the formant synthesis filter 552 and the perceptual weighting filter 556, each of these filters contains a memory of the filter state, where M<sub>P</sub> 550, M tone synthesis filter memory<sub>to</sub> the memory of the formant synthesis filter 552 and MW the memory of the perceptual weighting filter 556.
Filter status M<sub>to</sub> of the formant synthesis filter 552 of the decoder subsystem is provided to the tone search subsystem 220 of Figure 7. In Figure 16, the status of the M filter is provided<sub>to</sub> to calculate the zero input response (ZIR) of filter 560 which calculates the ZIR of formant synthesis filter 552. The calculated ZIR value is subtracted from the input speech samples s (n) in adder 562, where the result weighted by the perceptual weighting filter 564. The output of the perceptual weighting filter 564, x<sub>p</sub>(n), is used as weighted input voice in equations (28) - (34), where x (n) = x<sub>p</sub>(n).
Again, referring to Figures 14 and 15, the pitch synthesis filter 550 illustrated in Figure 14 provides the closed-loop and open-loop formant residue samples, calculated in the manner described above, to the adaptive codebook. 568 which, in essence, is a memory to store them. The closed-loop formant residue is stored in memory portion 570, while the open-loop formant residue is stored in memory portion 572. The samples are stored according to the example numbering model described above. The closed-loop formant residue is organized as described above in relation to usage for each pitch delay L search. The open loop formant residue is calculated from the input speech samples s (n) of each pitch subframe using the formant analysis filter 574 using the memory Ma of the formant synthesis filter 552 of the subsystem decoder to calculate p-values<sub>or</sub>(n). The p values<sub>or</sub>(n) for the current pitch subframe are shifted through a series of delay elements 576 before being provided to memory portion 572 of adaptive codebook 568. Open-loop formant residues are stored with the first sample waste generated numbered 0 and the last numbered 142.
Referring to Figure 16, the impulsive response h (n) of the formant filter is calculated in filter 566 and passed to shift register 580. As noted above in relation to the impulsive response of the formant filter h ( n), equations (29) - (30) and (35) - (38), these values are calculated for all tone subframes in the filter. To further reduce the computational requirements of the tone filter subsystem, the impulsive response of the formant filter h (n) is truncated to 20 samples.
Shift register 580 along with multiplier 582, adder 584, and shift register 586 are configured to perform recursive convolution between the h (n) values of shift register 580 and the c (m) values of the book. adaptive code 568, as described. This convolution operation is performed to find the zero-state response (zSr) of the formant filter to the input from the tone filter memory, assuming that the tone gain is set to 1. With the operation of the tone circuits convolution, n is iteratively shifted from Lp to 1 for each m, while m is iteratively shifted from (L<sub>p</sub>17) -1 to -143. In register 586, data is not transmitted when n = 1, and data is not blocked when n = L<sub>p</sub>. Data is output from convolution circuits when m <-17.
After the convolution circuits are the correlation and comparison circuits that perform the search to find the optimal pitch delay L and pitch gain b. Correlation circuits, also called mean square error (MSE) circuits, calculate the autocorrelation and cross-correlation of the ZSR with the perceptually weighted difference between the ZIR of the formant filter and the input speech, that is, x (n ). Using these values, the correlation circuitry calculates the optimal pitch gain b value for each pitch delay value. Correlation circuits consist of shift register 588, multipliers 590 and 592, adders 594 and 596, registers 598 and 600, and divisor 602. In correlation circuits, calculations determine that n is iteratively shifted from L<sub>p</sub> a 1, while m iteratively shifts from (L<sub>p</sub>-17) -1 to-143.
The correlation circuits are followed by the comparison circuits that carry out the comparisons and store the data to determine the optimal value of the delay L and the gain b of tone. The comparison circuits consist of multiplier 604, comparator 606, registers 608, 610, and 612, and quantizer 614. The comparison circuits provide for each pitch subframe the values of L and b that minimize the error between the synthesized speech and the input speech. The value of b is quantized in eight levels by quantizer 614 and is represented by a 3-bit value, an additional level being inferred, level b = 0, when L = 16. These L and b values are provided to codebook search subsystem 230 and data buffer 222. These values are provided via data packing subsystem 238 or data buffer 222 to decoder 234 for use in searching for the tone.
Like pitch search, codebook search is an analysis-by-synthesis coding system, in which coding is performed by selecting parameters that minimize the error between the input speech and the synthesized speech using the parameters. For 1/8 speed, the pitch gain is set to zero.
IS 2 240 252 T3
As described above, every 20 ms, the speech frame is subdivided into a number of codebook subframes which, as indicated, depends on the data rate chosen for the frame. The G and I parameters, the gain and the codebook index, respectively, are calculated once per codebook subframe. In calculating these parameters, the LSP frequencies for the subframe, except for full rate, are interpolated into the codebook 226 subframe LSP interpolation subsystem in a manner similar to that described for the subframe LSP interpolation subsystem. pitch 216. Interpolated LSP frequencies from codebook subframes are also converted to LPC coefficients by LSPLPC transform subsystem 228 for each codebook subframe. The codebook subframe counter 232 is used to keep track of the codebook subframes for which the codebook parameters are calculated, with the counter output being provided to the codebook subframe LSP interpolation subsystem. code 226 for use in codebook subframe LSP interpolation. Also, codebook subframe counter 232 provides an output, indicating completion of a codebook subframe for the selected rate, to tone subframe counter 224.
The excitation codebook consists of 2<sup>M</sup> code vectors that are constructed from a Gaussian white random sequence of unit variance. There are 128 entries in the codebook for M = 7. The codebook is recursively arranged so that each code vector differs from the adjacent code vector in a sample; that is, the samples in a code vector are shifted one position so that a new sample enters one end and another sample falls through the other. Consequently, a recursive codebook can be stored as a linear sort that has length 2<sup>M</sup> + (L<sub>C</sub>-1), where L<sub>C</sub> the length of the codebook subframe. However, to simplify execution and conserve memory space, a circular 2-inch codebook is used.<sup>M </sup>length samples (128 samples).
To reduce calculations, the Gaussian values in the codebook are clipped down the center. Initially, the values are chosen using a variance 1 Gaussian white procedure. Then, any value with a magnitude less than 1.2 is set to zero. And in this way, about 75% of the values are effectively set to zero, generating a pulse codebook. This central codebook clipping reduces the number of multiplications required to perform the recursive convolution of the codebook search by a factor of 4, since multiplication by zero is not necessary. The codebook used in the current run is provided below in Table VIII.
TABLE VIII
<td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x2afe</td><td>0x0000</td><td>0x0000</td><td>0x0000</td>
<td>0x41 gives</td><td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x0000</td>
<td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x3bb3</td><td>0x0000</td><td>0x363e</td>
<td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x417d</td><td>0x0000</td>
<td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x0000</td>
<td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x9dfe</td><td>0x0000</td><td>0x0000</td>
<td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x0000</td>
<td>0x0000</td><td>0xc58a</td><td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x0000</td>
<td>0x0000</td><td>0xc8db</td><td>0xd365</td><td>0x0000</td><td>0x0000</td><td>0xd6a8</td><td>0x0000</td><td>0x0000</td>
<td>0x0000</td><td>0x3e53</td><td>0x0000</td><td>0x0000</td><td>0xd5ed</td><td>0x0000</td><td>0x0000</td><td>0x0000</td>
<td>0xd08b</td><td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x3d14</td><td>0x396a</td><td>0x0000</td>
<td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x4ee7</td><td>0xd7ca</td><td>0x0000</td>
<td>0x0000</td><td>0x438c</td><td>0x0000</td><td>0x0000</td><td>0xad49</td><td>0x30b1</td><td>0x0000</td><td>0x0000</td>
<td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x0000</td>
<td>0x0000</td><td>0x0000</td><td>0x3fcd</td><td>0x0000</td><td>0x0000</td><td>0xd187</td><td>0x2e16</td><td>0xd09b</td>
<td>0xcb8d</td><td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x0000</td><td>0x32ff</td>
Again, the speech coder uses a perceptual noise weighting filter as indicated in equation (1) including a weighted synthesis filter as indicated in equation (3). For each codebook index, I, the voice is synthesized and compared to the original voice. The error is weighted by the perceptual weighting filter before the MSE is calculated.
As stated above, the objective is to minimize the error between x (n) and x '(n) with respect to all possible values of I and G. The minimization of the error can be expressed by the following equation:
Lc-1
MSE = - Σ (x (n) - x '(n))<sup>1 2</sup>
Le n = 0 (39)
ES 2 240 252 T3 where L<sub>C</sub> the number of samples in the codebook subframe. Equation (38) can be rewritten in relation to G as:
Lc-1
MSE = - Σ (x (n) - Gy (n))<sup>:</sup>
Le n = 0 (40) being y deduced by convolving the impulsive response of the formant filter with the I-th code vector, assuming that G = 1. Minimizing the MSE is, in turn, equivalent to maximizing:
Ei = (Exy)<sup>:</sup> , 2 yy (41) where 'xy yy
Le-1 = Σ x (n) y (n) n = 0
Le-1 = y (n) y (n) n = 0 (42) (43)
The optimal G for the given I is found by the following equation
<img file="ES2240252T3_D0001.tif" />
yy (44)
This search is repeated for all allowed values of I. Unlike the search for tone, the optimal gain, G, can be positive or negative. Finally, the codebook index I and the codebook gain G are chosen for transmission, maximizing E<sub>I</sub>.
Again, it should be noted that it is only necessary to calculate x (n) once, that is, the perceptually weighted difference between the input speech and the ZIR of the weighted pitch and formant filters. However, for each index I, it is necessary to calculate y (n), that is, the zero-state response of the pitch and formant filters of each code vector. Because a circular codebook is used, the recursive convolution procedure described for pitch search can be used to minimize the calculations required.
Referring again to Figure 15, the encoder includes a duplicate of the decoder of Figure 5, the decoder subsystem 235 of Figure 7, in which the filter states are calculated, where M<sub>p</sub> 550, M tone synthesis filter memory<sub>to</sub> Formant synthesis filter memory 552 and M<sub>w</sub> the memory of the perceptual weighting filter 556.
The filter states M<sub>p</sub> and M<sub>to</sub> of the tone and formant synthesis filters 550 and 552 (Figure 15) of the decoder subsystem, are provided to the codebook search subsystem 230 of Figure 7. In Figure 17, the filter states M<sub>p</sub> and M<sub>to</sub> are provided to Zero Impulse Response (ZIR) filter 620 which calculates the ZIR of the formant synthesis and tone filters 550 and 552. The calculated ZIR of the formant synthesis and tone filters is subtracted from the samples of input voice s (n) in adder 622, the result being weighted by perceptual weighting filter 624. The output of perceptual weighting filter 564, x<sub>c</sub>(n), is used as the weighted input voice in the above MSE equations (39) - (44), where x (n) = x<sub>c</sub>(n).
In Figure 17, the impulsive response h (n) of the formant filter is computed in filter 626 and provided to shift register 628. The impulsive response of the formant filter h (n) is computed for each workbook subframe. code. To further reduce the computational requirements, the impulsive response h (n) of the formant filter is truncated at 20 samples.
Shift register 628 along with multiplier 630, adder 632, and shift register 634 are configured to perform recursive convolution between the h (n) values of shift register 628 and the c (m) values of the book. code 636 containing the codebook vectors described above. This convolution operation is carried out to find the zero-state response (ZSR) of the formant filter to each code vector, assuming that the codebook gain is set to 1. With the operation of the convolution circuits, n iteratively shifts from L<sub>C</sub> a 1 for each m, while m is iteratively shifted
ES 2 240 252 T3 from 1 to 256. In register 586, data is not transmitted when n = 1 and data is not blocked when n = L<sub>C</sub>. The data is output from the convolution circuits when m <1. It should be noted that the convolution circuits must be initialized to drive the recursive convolution operation by iteratively shifting m times the size of the subframe before starting the correlation and comparison circuits. that follow the convolution circuits.
The correlation and comparison circuits direct the present codebook search to provide the values of the codebook index I and the gain of codebook G. The correlation circuits, also called mean square error (MSE) circuits, calculate autocorrelation and cross correlation ada of the ZSR with the perceptually weighted difference between the ZIR of the pitch and formant filters, and the input voice x '(n). That is, the correlation circuits calculate the value of the codebook gain G for each codebook index value I. The correlation circuits consist of shift register 638, multipliers 640 and 642, adders 644 and 646, registers 648 and 650, and divisor 652. In correlation circuits, calculations determine that n is iteratively shifted from L<sub>C</sub> a 1, while m iteratively shifts from 1 to 256.
The correlation circuits are followed by the comparison circuits that perform the comparisons and data storage to determine the optimal value of the index I and the codebook gain G. The comparison circuits consist of multiplier 654, comparator 656, registers 658, 660, and 662, and quantizer 664. The comparison circuitry provides for each codebook subframe the I and G values that minimize the error between the synthesized speech and the input speech. The G codebook gain is quantized in quantizer 614 which DPCM encodes the values, during quantization, in a manner similar to the quantization and encoding of LSP frequencies with offset subtraction described in relation to Figure 12. These I and G values are then provided to data buffer 222.
The DPCM quantization and encoding of the G codebook gain is calculated according to the following equation:
Quantified Gi = 20 log G¡ - 0.45 (20 log G¡-1 + 20 log G) (45) where 20 log G¡-1 and 20 log G¡-2 are the respective values calculated for the immediately preceding frame ( i-1) and the frame preceding the immediately preceding frame (i-2).
The LSP, I, G, L, and b values along with the rate are provided to the data packing subsystem 236, where the data is arranged for transmission. In one embodiment, the LSP, I, G, L, and b values along with the rate can be provided to the decoder 234 via the data packing subsystem 236. In another embodiment, these values may be provided via data buffer 222 to decoder 234 for use in finding the tone. However, in the following description, a codebook sign bit protection is employed in the data packaging subsystem 236 that may affect the codebook index. Therefore, this protection must be taken into account if the I and G data are provided directly from the data buffer 222.
In the data packaging subsystem 236, the data can be packed in various formats for transmission. Figure 18 illustrates the functional elements of the data packing subsystem 236. The data packing subsystem 236 consists of the pseudo-random generator (PN) 670, the cyclic redundancy check (CRC) calculation element 672, the protection logic data 674 and data combiner 676. The PN 670 generator receives the speed and, for eighth speed, generates a 4-bit random number that is provided to the data combiner 676. The CRC element 672 receives the codebook gain and LSP values along with the speed and , for full rate, generates an internal 11-bit CRC code that is provided to data combiner 676.
The data combiner 674 receives the random number, the CRC code, and along with the speed and LSP, I, G, L, and b values from the data buffer 222 (Figure 7b) provides an output to the data processor subsystem. transmission channel 234. In the execution where the data is provided directly from the data buffer 222 to the decoder 234 at a minimum rate, the 4-bit number of the PN generator is passed from the PN 670 generator, via the data combiner 676, to the decoder. 2. 3. 4. At full rate, the CRC bits are included along with the frame data obtained from data combiner 674, while at eighth rate, the codebook index value is excluded and replaced by the 4-bit random number.
It is preferable to provide protection for the codebook gain sign bit. The purpose of protecting this bit is to make the vocoder decoder less sensitive to one-bit errors in this bit. If the sign bit changes due to an undetected error, the book code index will point to a vector unrelated to the optimum. In the unprotected error situation, the negative of the optimal vector will be selected, a vector that is essentially the worst possible vector to use. The protection model used here ensures that a one-bit error in the gain sign bit does not cause the selection of the negative of the optimal vector in the error situation. The data protection logic 674 receives the index and the codebook gain and examines the sign bit of the gain value. If the sign bit of the gain value is found to be negative, the value 89 (modulo 128) is added to the associated codebook index. The codebook index, whether or not it is modified, is provided by data protection logic 674 to data combiner 676.
IS 2 240 252 T3
It is preferable that at full rate the most perceptually sensitive bits of the compressed voice packet data are protected, for example, by an internal CRC (cyclic redundancy check). Eleven additional bits are used to carry out this error detection and correction function which is capable of correcting any errors in the protected block. The protected block consists of the most significant bit of the 10 LSP frequencies and the most significant bit of the 8 codebook gain values. If an uncorrectable error occurs in this block, the packet is rejected and a clear operation is declared, described later. In the other cases, the tone gain is set to zero, but the rest of the parameters are used as they are received. In the exemplary embodiment, a cyclic code is chosen that has a generator polynomial of:
g (x) = 1 + x<sup>3</sup> + x<sup>5</sup> + x<sup>6</sup> + x<sup>8</sup> + x<sup>9</sup> + x<sup>10</sup> (46) which provides a cyclic code (31,31). However, it should be understood that other generator polynomials can be used. To make this code a (32, 31) code, a global parity bit is added at the end. Since there are only 18 bits of information, the first 3 digits of the code word are set to zero and are not transmitted. This technique provides additional protection and; thus, if the syndrome indicates an error in these positions, it means that it is an incorrigible error. The codification of a cyclic code in a systematic way involves the calculation of parity bits according to: x10 u (x) modulo g (x), where u (x) is the polynomial of the message.
At the decoding end, the syndrome is computed as the remainder by dividing the received vector by g (x). If the syndrome does not indicate an error, the packet is accepted regardless of the state of the global parity bit. If the syndrome indicates an error, the error is corrected if the state of the global parity bit is not check. If the syndrome indicates more than one error, the packet is rejected. More information about such an error protection model and the syndrome calculation can be found in section 4.5 of Lin and Costello's “Error Control coding: Fundamentals and Applications” document.
In a CDMA cellular phone system implementation, data is provided by data combiner 674 to transmission channel data processor subsystem 238 for data packaging for transmission in 20 ms data transmission frames. In a transmission frame in which the vocoder is ready for full rate, 192 bits are transmitted for an effective bit rate of 9.6 Kbit / s. The transmission frame in this case consists of a mixed mode bit used to indicate the type of mixed frame (0 = voice only, 1 = voice and data / signaling), 160 bits of vocoder data along with 11 bits of internal CRC ; 12 external or frame CRC bits and 8 tail or level bits. At half speed, 80 bits of vocoder data are transmitted along with 8 frame CRC bits and 8 tail bits for an effective bit rate of 4.8 Kbit / s. At quarter speed, 40 bits of vocoder data are transmitted along with 8 tail bits for an effective bit rate of 2.4 Kbit / s. Finally, at eighth speed, 16 bits of vocoder data are transmitted along with 8 tail bits for an effective bit rate of 1.2 Kbit / s.
The pending US patent application serial number 07 / 543,496, filed on June 25, 1990 and entitled "SYSTEM AND METHOD FOR GENERATING SIGNAL WAVEFORMS IN A CDMA CELLULAR TELEPHONE SYSTEM" in a CDMA cellular telephone system), granted to the assignee of the present invention, provides further information about the modulation employed in a CDMA system in which the vocoder of the present invention will be employed. In this system, at rates other than full, a model is used in which the data bits are arranged in groups, the groups of bits being pseudo-randomly located within the 20 ms data transmission frame. It should be understood that it is possible to easily employ other frame rates and bit representations than those presented for illustrative purposes here, in connection with the performance of the vocoder and the CDMA system, thus making other implementations available for the vocoder and other system applications.
In the CDMA system, and also applicable to other systems, the processor subsystem 238 can interrupt the transmission of vocoder data from frame to frame to transmit other data such as signaling data or other non-voice information data. This particular type of transmission situation is called a "space-burst." Processor subsystem 238 essentially replaces the vocoder data with the desired transmission data for the frame.
Another situation may arise where it is desired to transmit both vocoder data and other data during the same data transmission frame. This particular type of transmission situation is called "burst attenuation." In a "fade-burst" transmission, the vocoder receives rate limit commands that set the final rate of the vocoder at the desired rate, eg, half rate. The half-rate encoded vocoder data is provided to processor subsystem 238, which inserts the additional data along with the vocoder data for the data transmission frame.
An additional feature provided for full duplex telephone links is rate interlock. If one direction of the link is transmitting at the highest transmission rate, then the other direction of the link is forced to transmit at the lowest rate. Even at the slowest speed, enough intelligibility is available for the active talker to realize that they have been interrupted and stop speaking, thereby allowing the other direction of the link to assume the role of active talker. Also, if the active speaker continues to speak during an interruption attempt, they probably will not perceive a quality degradation because their own
ES 2 240 252 T3 voice "interferes" with the ability to perceive quality. Again, using the rate limit commands, the vocoder can adapt to the speech coding of the speech at a lower than normal rate.
It should be understood that rate limit commands can be used to set the maximum vocoder rate to a rate lower than full rate when additional capacity is needed in the CDMA system. In a CDMA system where a common frequency spectrum is used for transmission, the signal from one user shows up as interference to the other users of the system. The user capacity of the system is thus limited by the total interference caused by the users of the system. As the level of interference increases, typically due to an increase in users on the system, users experience a degradation in quality due to the increase in interference.
Each user's contribution to CDMA interference is a function of the users' data transmission rate. By adapting the vocoder for speech coding at a lower than normal rate, the encoded data is transmitted at the corresponding reduced data rate, thereby decreasing the level of interference caused by the user. Therefore, the capacity of the system can be greatly increased by voice coding at a lower rate. When system demand increases, user vocoders can be controlled by the system controller or cell base station to reduce the encoding speed. The quality of the vocoder determines that there is little, if any, perceptible difference between full-speed and half-speed coded speech. Consequently, the effect on the quality of communications between system users when voice is voice coded at low speed, for example at medium speed, is less significant than that caused by an increasing level of interference resulting from a largest number of users in the system.
Accordingly, various models can be used to set individual vocoder rate limits for lower than normal speech coding rates. For example, all users in a cell can be controlled to encode speech at half speed. This action significantly reduces system interference, with a negligible effect on the quality of communications between users, while providing a considerable increase in capacity for additional users. Until the total system interference has increased to the level of degradation due to additional users, it will not affect the quality of communications between users.
As indicated above, the encoder includes a copy of the decoder to apply the analysis-by-synthesis technique to the encoding of the frames of the speech samples. As illustrated in Figure 7, decoder 234 receives the L, b, I, and G values either through data packaging subsystem 238, or through data buffer 222 to reconstruct the synthesized speech and compare it with the input voice. The outputs of the decoder are the values Mp, Ma and Mw described above. The use of decoder 234 in the encoder and in reconstructing the synthesized speech at the other end of the transmission channel will be jointly described with reference to Figures 19-24.
Figure 19 is a flow chart for a decoder execution example. Due to the common structure of the decoder executed in the encoder and the one executed in the receiver, these executions are described together. The description relative to Figure 19 refers mainly to the decoder at the end of the transmission channel, since the data received there must be previously processed in the decoder, while the appropriate data (speed, I , G, L, and b) directly from data packaging subsystem 238 or data buffer 222. However, the basic function of the decoder is the same for both the encoder and the decoder execution.
As indicated in connection with Figure 5, for each codebook subframe, the codebook vector indicated by the codebook index I is extracted from the stored codebook. The vector is multiplied by the codebook gain G and then filtered by the tone filter of each tone subframe to obtain the formant residue. This formant residue is filtered by the formant filter and then passed through an adaptive formant postfilter and brightness postfilter, and automatic gain control (AGC), to generate the output speech signal.
Although the length of the pitch and codebook subframe varies, decoding is carried out in blocks of 40 samples for ease of execution. First, the received compressed data is unpacked into codebook gains, codebook indices, pitch gains, pitch delays, and LSP frequencies. LSP frequencies should be processed through their respective inverse quantizers and DPCM decoders as described in relation to Figure 22. Similarly, codebook gain values should be processed similarly to LSP frequencies, except for as far as runout is concerned. Also, the pitch gain values are inverse quantized. The parameters of each decoding subframe are provided below. In each decoding subframe, 2 groups of codebook parameters (G and I), 1 group of pitch parameters (b and L), and 1 group of LPC coefficients are needed to generate 40 output samples. Figures 20 and 21 illustrate examples of subframe decoding parameters for the various rates and other frame conditions.
For full rate frames, there are 8 groups of received codebook parameters and 4 groups of received tone parameters. LSP frequencies are interpolated four times to provide 4 groups of frequencies
IS 2 240 252 T3
LSP. The received parameters and the corresponding subframe information are listed in Figure 20a.
For half-rate frames, each group of the four received codebook parameters is repeated once, each group of the two received tone parameters is repeated once. The LSP frequencies are interpolated three times to provide 4 groups of LSP frequencies. The received parameters and the corresponding subframe information are listed in Figure 20b.
For quarter-rate frames, each group of the two received codebook parameters is repeated four times, and the group of tone parameters is also repeated four times. The LSP frequencies are interpolated once to provide 2 groups of LSP frequencies. The received parameters and corresponding subframe information are listed in Figure 20c.
For eighth rate frames, the received codebook parameter set is used for the entire frame. There is no pitch parameter present for eighth speed frames and the pitch gain is simply set to zero. The LSP frequencies are interpolated once to provide 1 group of LSP frequencies. The received parameters and corresponding subframe information are listed in Figure 20d.
Sometimes voice packets can be left blank for the CDMA cell or mobile station to transmit signaling information. When the vocoder receives a blank frame, it continues with a slight modification in the parameters of the previous frame. The codebook gain is set to zero. The pitch delay and gain of the previous frame are used as the pitch delay and gain of the current frame, but the gain is limited to one or less. The LSP frequencies from the previous frame are used as is, without interpolation. It should be noted that the encoding end and the decoding end are still in sync and that the vocoder can recover from a blank frame very quickly. The received parameters and the corresponding subframe information are listed in Figure 21a.
In the event that a frame is lost due to a channel error, the vocoder attempts to mask the error by keeping a fraction of the energy of the previous frame and making a smooth transition to the background noise. In this case, the pitch gain is set to zero, a random codebook is selected using the codebook index from the previous frame plus 89, and the codebook gain is 0.7 times the codebook gain. code from the previous subframe. It should be noted that the number 89 is not used for any particular reason, but is just a convenient way to select a pseudo-random codebook vector. The LSP frequencies of the previous frame are forced to decrease towards their off-center values according to:
ω<sub>i</sub> = 0.9 (ω above - offset value of ω ^ + offset value of ω<sub>i</sub> (47)
The offset values of the LSP frequencies are shown in Table 5. The received parameters and corresponding subframe information are listed in Figure 21b.
If the rate cannot be determined at the receiver, the packet is rejected and an erase operation is declared. However, if the receiver determines that the frame is most likely transmitted at full rate, albeit with errors, the action described below is taken. As described above for full rate, the most perceptually sensitive bits of the compressed voice packet data are protected by an internal CRC. In the decoding area, the syndrome is calculated as the remainder by dividing the received vector by g (x), from equation (46). If the syndrome does not indicate an error, the packet is accepted regardless of the state of the global parity bit. If the syndrome indicates an error, the error is corrected if the state of the global parity bit is not check. If the syndrome indicates more than one error, the packet is rejected. If an uncorrectable error occurs in this block, the packet is rejected and a delete operation is declared. In other cases, the tone gain is set to zero, but the rest of the parameters are used as they are received with corrections, as illustrated in Figure 21c.
The postfilters used in this run were first described in the document “Real-Time Vector APC Speech Coding At 4800 BPS with Adaptive postfiltering” by JH Chen et al., Proc. ICASSP, 1987. Since the speech formants are perceptually more important than the spectral valleys, the postfilter slightly boosts the formants to improve the perceptual quality of the coded speech. This is done by scaling the poles of the formant synthesis filter radially toward the origin. However, an all-pole postfilter generally introduces a spectral tilt that results in damping of the filtered speech. The spectral tilt of this all-pole postfilter is reduced by adding zeros that have the same phase angles as the poles, but smaller radii, resulting in a postfilter as follows:
H (-) =
Α (- / ρ)
Α (- / σ) (48) where A (z) is the formant prediction filter and the ρ and σ values are the scale factors set at 0.5 and 0.8, respectively.
IS 2 240 252 T3
An adaptive brightness filter is added to further compensate for the spectral tilt introduced by the formant postfilter. The brightness filter is as follows:
B (z) =
- κζ <sup>1 </sup>1 + κζ<sup>-1</sup> (49) the value of κ (the coefficient of this one-shot filter) being determined by means of the mean value of the LSP frequencies that provides an approximate value of the change in the spectral inclination of A (z).
To avoid large gain shifts as a result of post-filtering, an AGC loop is executed to scale the speech output so that it has approximately the same energy as the speech that has not been post-filtered. Gain control is performed by dividing the sum of the squares of the 40 samples fed into the filter by the sum of the squares of the 40 samples drawn from the filter to obtain the inverse filter gain. The square root of this gain factor is then smoothed:
Smoothed β = 0.2 current β + 0.98 previous β (50) and then the filter output is multiplied by this smoothed inverse gain to generate the output voice.
In Figure 19, the channel data along with the rate, whether transmitted with the data or obtained by other means, is provided to the data packaging subsystem 700. In an exemplary implementation for a CDMA system, a decision of The speed that can be obtained from the error rate is the data received when it is decoded at each of the different speeds. In data unpacking subsystem 700, at full rate, an error CRC is performed, the result of this verification being provided to subframe data unpacking subsystem 702. Subsystem 700 provides an indication of frame conditions anomalous frames such as blank frames, frame erasure, or erroneous frames with data usable to subsystem 702. Subsystem 700 provides the rate along with frame parameters I, G, L, and b to subsystem 702. When codebook index I and G gain values are provided, the sign bit of the gain value is checked at subsystem 702. If the sign bit is negative, the value 89 (modulo 128) is subtracted from the associated codebook index. Furthermore, in the subsystem, the codebook gain is inverse quantized and DPCM encoded, while the pitch gain is inverse quantized.
In addition, subsystem 700 provides the speed and LSP frequencies to LSP interpolation / inverse quantization subsystem 704. Subsystem 700 further provides frame blank, frame erasure, or bad frame indication with usable data to subsystem 704. Decoding subframe counter 706 provides an indication of the value of subframe counter i and j to subsystems 702 and 704.
In subsystem 704, the LSP frequencies are inversely quantized and interpolated. Figure 22 illustrates an execution of the inverse quantization part of subsystem 704, the interpolation part being practically identical to that described in relation to Figure 12. In Figure 22, the inverse quantization portion of subsystem 704 is shown consisting of inverse quantizer 750, identical in construction to inverse quantizer 468 of Figure 12 and similar in operation. The output of inverse quantizer 750 is provided as an input to adder 752. The other input of adder 752 is provided as an output of multiplier 754. The output of adder 752 is provided to register 756, where it is stored and provided for multiplication with the constant 0.9 in multiplier 754. The output of adder 752 is also provided to adder 758, where the value of the runout is added back to the LSP frequency. The ordering of the LSP frequencies is ensured by logic 760 which forces the LSP frequencies to have a minimum separation. Generally, the need to force separation does not arise unless there is a transmission error. The LSP frequencies are then interpolated as described in relation to Figure 13 and in relation to Figures 20a-20d and 21a-21c.
Referring again to Figure 19, memory 708 is coupled to subsystem 704 to store the previous frame LSP frequencies, ω<sub>σ-1</sub>, and can also be used to store the offset values Ηω ,. These previous frame values are used in interpolation for all rates. Under blank frame conditions, frame erasure or erroneous frames with usable data, the above LSP frequencies are used ω<sub>σ-1 </sub>according to the graph of Figures 21a-21c. In response to a blank frame indication from subsystem 700, subsystem 704 retrieves the previous frame LSP frequencies stored in memory 708 for use in the current frame. In response to a frame erasure indication, subsystem 704 again retrieves the previous frame LSP frequencies from memory 708 along with the offset values to calculate the current frame LSP frequencies as described. When this calculation is carried out, the stored offset value is subtracted from the LSP frequency of the previous frame in an adder, the result being multiplied in the multiplier by a constant value of 0.9 and this result being added in the adder to the value of stored runout. In response to a bad frame indication with usable data, the LSP frequencies are interpolated in the same way as for full rate if the CRC is successful.
The LSP frequencies are provided to the LSP-LPC transformation subsystem 710, where the LSP frequencies
ES 2 240 252 T3 are converted back to LPC values. Subsystem 710 is substantially identical to the LSP-LPC transformation subsystems 218 and 228 of Figure 7 described in relation to Figure 13. The LPC coefficients a, are then provided to the formant filter 714 and to the formant post-filter 716. Also, the mean value of the LSP frequencies across the subframe in the LSP averaging subsystem 712 is calculated and provided to the adaptive brightness filter 718 as the κ value.
Subsystem 702 receives the I, G, L, and b parameters for the frame from subsystem 700, along with the rate or abnormal frame condition indication. Also, subsystem 702 receives from subframe counter 706 the counts j for each count i of each decoding subframe 1-4. Subsystem 702 is also coupled to memory 720 which stores previous frame values of G, I, L, and b, for use under abnormal frame conditions. Subsystem 702, under normal frame conditions, except in eighth speed, provides the codebook index value Ij to codebook 722, the codebook gain value Gj to multiplier 724, and the delay values L and gain b of tone to tone filter 726, according to Figure 20a-20d. For eighth speed, since no value is sent for the codebook index, a packet seed, which is the 16-bit parameter value (Figure 2d) for eighth speed, is provided to codebook 722 along with a speed indication. For abnormal frame conditions, the values are provided from subsystem 702 according to Figures 21a-21c. In addition, for eighth speed, an indication is provided to codebook 722 as described in connection with Figure 23.
In response to a blank frame indication from subsystem 700, subsystem 702 retrieves pitch delay L and gain b values from the previous frame, although here the gain is limited to one or less, stored in memory 708 , to be used in the decoding subframes of the current frame. Also, no codebook index I is provided and the codebook gain G is set to zero. In response to a frame erase indication, subsystem 702 also retrieves the subframe codebook index of the previous frame from memory 720 and adds, in the adder, the value 89. The subframe codebook gain The previous frame is multiplied by the multiplier by the constant 0.7, to generate the respective G-values of the subframes. No pitch delay value is provided and the pitch gain is set to zero. In response to a bad frame indication with usable data, the index and codebook gain are used as in a full rate frame, provided the CRC is satisfactory, and no pitch delay value is provided and the Tone gain is set to zero.
As described in connection with the analysis-by-synthesis technique encoder decoder, the codebook index I is used as the starting address for the codebook value to be provided to the multiplier 724. The book gain value Code number is multiplied at multiplier 724 by the codebook output value 722, the result being provided to tone filter 726. The tone filter 726 uses the input tone gain b and delay L values to generate the residual of formants that is provided to the formant filter 714. In the formant filter 714, the LPC coefficients are used to filter the residual of formants and rebuild the voice. At the receiver decoder, the reconstructed speech is again filtered by the formant postfilter 716 and the adaptive brightness filter 718. The AGC loop 728 is used at the output of the formant filter 714 and the formant post-filter 716, the output of which is multiplied by the multiplier 730 by the output of the adaptive brightness filter 718. The output of the multiplier 730 is voice reconstructed which is then converted to analog speech using known techniques and presented to the listener. In the encoder's decoder, the perceptual weighting filter is placed at the encoder's output to update its memories.
In Figure 22, more details of the execution of the decoder itself are illustrated. Encoder 722 of Figure 22 consists of a memory 750 similar to that described with reference to Figure 17. However, for purposes of explanation, Figure 22 illustrates a slightly different approach to memory 750 and address addressing. herself. Codebook 722 further consists of a switch 752, a multiplexer 753, and a pseudo-random number (PN) generator 754. The switch 752 is codebook index responsive to signal the location of the index address of the memory 750, as indicated with reference to Figure 17. The memory 750 is a circular memory, where the switch 752 points to the initial memory location, the values being shifted through memory for output. The codebook values are obtained from memory 750 through switch 752 as input to multiplexer 753. Multiplexer 753 is full speed, half speed and quarter speed sensitive to provide an output of the values provided, through switch 752, to the codebook gain amplifier, multiplier 724. Multiplexer 753 is also sensitive to the eighth speed indication to select the PN 754 generator output as the 722 codebook output for the 724 multiplier.
To maintain high quality speech in CELP encoding, the encoder and decoder must have the same values stored in their internal filter memories. This is done by transmitting the codebook index, so that the decoder and encoder filters are driven by the same sequence of values. However, for the highest quality speech, these sequences consist mostly of zeros with some peaks distributed between them. This type of excitation is not optimal for background noise coding.
When background noise is encoded, at the lowest data rate, a pseudo-random sequence can be executed to drive the filters. To ensure that the filter memories are the same in both the encoder and the decoder, the two pseudo-random sequences must be the same. It is necessary to transmit a seed to the receiver's decoder anyway. Since there are no additional bits that can be used to
ES 2 240 252 T3 send the seed, the bits of the transmitted packet can be used as seed, as if they constituted a number. It is possible to carry out this technique, because, at low speed, the exact same CELP analysis by synthesis structure is used to determine the gain and codebook index. The only difference is that the codebook index is discarded and instead the encoder filter memories are updated using a pseudo-random sequence. Therefore, the seed for arousal can be determined after the analysis is done. To ensure that the packets themselves are not iteratively and periodically shifted between a group of binary patterns, four random bits are inserted into the eighth-speed packet in place of the codebook index values. Therefore, the seed of the packet is the 16-bit value indicated in Figure 2d.
The PN 754 generator is built using well known techniques and can be executed by various algorithms. The algorithm used is of the type described in the article “DSP chips can produce random numbers using proven algorithm” by Paul Mennen, EDN, January 21, 1991. The package of transmitted bits is used as seed (from subsystem 700 of Figure 18) to generate the sequence. In one run, the seed is multiplied by the value 521, adding the value 259 to the result. From the resulting value, the least significant bits are used as a signed 16-bit number. This value is then used as a seed to generate the next codebook value. The sequence generated by the PN generator is normalized to have a variance of 1.
Each value obtained from codebook 722 is multiplied at multiplier 724 by the codebook gain G provided during the decoding subframe. This value is provided as input to the adder 756 of the tone filter 726. The tone filter 726 further consists of the multiplier 758 and memory 760. The tone delay L determines the position of a tap of memory 760 that is passed to the multiplier 758. The output of memory 760 is multiplied at multiplier 758 by the pitch gain value b, with the result being passed to adder 756. The output of adder 756 is provided to a memory input 760 which is a series of elements of delay, such as a shift register. The values are scrolled through memory 760 (in the direction indicated by the arrow) and fed to the selected tap output determined by the value of L. Since the values are shifted through memory 760, values older than 143 shifts are rejected. The output of adder 756 is also provided as input to formant filter 714.
The output of adder 756 is provided to one input of adder 762 of formant filter 714. Formant filter 714 further consists of multiplier group 764a-764j and memory 766. The output of adder 762 is provided as input to the memory 766 which is also constructed as a series of tapped delay elements such as a shift register. The values are scrolled through memory 766 (in the direction indicated by the arrow) and discarded at the end. Each element has a socket that provides the stored value as an output to the corresponding multiplier of multipliers 764a-764j. Each of the multipliers 764a-764j also receives the corresponding LPC coefficient of the LPC coefficients α<sub>1</sub> - α<sub>10</sub> to multiply it by the output of memory 766. The output of adder 762 is provided as the output of formant filter 714.
The output of formant filter 714 is provided as input to formant postfilter 716 and AGC subsystem 728. Formant postfilter 716 consists of adders 768 and 770, along with memory 772 and multipliers 774a-774j, 776a-776j , 780a-780j and 782a-782j. As the values scroll through memory 772, they are provided by corresponding taps for multiplication by the scaled LPC coefficient values and their sum in adders 768 and 770. The output of the formant postfilter 716 is provided as an input to the adaptive brightness filter 718.
Adaptive brightness filter 718 consists of adders 784 and 786, registers 788 and 790, and multipliers 792 and 794. Figure 24 is a graph illustrating the characteristics of the adaptive brightness filter. The output of the formant postfilter 716 is provided to the adder 784 as one of its inputs, while the other input comes from the output of the multiplier 792. The output of adder 784 is provided to register 788, stored for one cycle, and provided for the next cycle to multipliers 792 and 794, along with the -κ value provided by LSP averager 712 of Figure 19. The output of the multipliers 792 and 794 is provided to adders 784 and 786. The output of adder 786 is provided to AGC subsystem 728 and shift register 790. Register 790 is used as a delay line to ensure coordination in the data provided by formant filter 714 to AGC subsystem 728, and provided to adaptive brightness filter 718 via formant postfilter 716.
The AGC subsystem 728 receives data from the formant postfilter 716 and the adaptive brightness filter 718 to scale the output speech energy to approximately the input speech energy into the formant postfilter 716 and the adaptive brightness filter 718. The AGC subsystem 728 consists of multipliers 798, 800, 802, and 804, adders 806, 808, and 810, registers 812, 814, and 816, divisor 818, and square root element 820. The output of 40 samples from formant postfilter 716 is squared at multiplier 798 and added in an accumulator, consisting of adder 806 and register 812, to generate the value "x". Similarly, the 40-sample output of adaptive brightness filter 718, taken before register 790, is squared at multiplier 800 and added in an accumulator, consisting of adder 808 and register 814, to generate the "y" value. The value "y" is divided by the value "x" in divisor 816 to give the inverse gain of the filters. The square root of the inverse gain factor is obtained at item 818, the result being smoothed. The smoothing operation is carried out by multiplying the current value of gain G by the constant value 0.02 in multiplier 802, where
ES 2 240 252 T3 this result added in adder 810 to the result of multiplying by 0.98 the previous gain calculated using register 820 and multiplier 804. The output of filter 718 is then multiplied by the inverse gain smoothed in the multiplier 730 to provide the reconstructed output voice. The output speech is then converted to analog speech using the various well known conversion techniques to provide it to the user.
It should be understood that the embodiment of the present invention disclosed herein is only an exemplary embodiment, and that variants of the embodiment with equivalent functionality can be made. The present invention can be implemented in a digital signal processor under the control of a suitable program that provides the functional operation disclosed herein for encoding the speech samples and decoding the encoded speech. In other embodiments, the present invention may take the embodiment of an application specific integrated circuit (ASIC) using well known very large scale integration (VLSI) techniques.
The preceding description of the preferred embodiments is provided to enable those skilled in the art to use the present invention. The various modifications to these embodiments will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other embodiments without using inventiveness. Therefore, the present invention is not intended to be limited to the embodiments shown herein, but various modifications and changes should be made to the present invention without departing from the scope of the present invention as defined in the appended claims.
Contents27
23 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23
119 members in 21 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 19910713661 | United States of America | – | |
| 71366191 | United States of America | A | |
| 71366191 | United States of America | A | |
| 71366101103640 | – | – | – |
| US19910713661 | – | – | – |
Members119
| Document | Office | Kind | |
|---|---|---|---|
| MX9202808A | Mexico | A | |
| CA2102099A1 | Canada | A1 | |
| CA2483296A1 | Canada | A1 | |
| CA2483322A1 | Canada | A1 | |
| CA2483324A1 | Canada | A1 | |
| CA2568984A1 | Canada | A1 | |
| CA2635914A1 | Canada | A1 | |
| WO9222891A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU2186592A | Australia | A | |
| ZA924082B | South Africa | B | |
| CN1071036A | China | A | |
| NO934544D0 | Norway | D0 | |
| NO934544L | Norway | L | |
| FI935597A | Finland | A | |
| FI935597A7 | Finland | A7 | |
| EP0588932A1 | European Patent Office (EPO) | A1 | |
| JPH06511320A | Japan | A | |
| BR9206143A | Brazil | A | |
| US5414796A | United States of America | A | |
| HUT70719A | Hungary | A | |
| IL113986D0 | Israel | D0 | |
| IL113987D0 | Israel | D0 | |
| IL113988D0 | Israel | D0 | |
| IL102146A | Israel | A | |
| AU671952B2 | Australia | B2 | |
| AU6089396A | Australia | A | |
| IL113986A | Israel | A | |
| IL113987A | Israel | A | |
| IL113988A | Israel | A | |
| AU1482597A | Australia | A | |
| US5657420A | United States of America | A | |
| CN1159639A | China | A | |
| CN1167309A | China | A | |
| RU2107951C1 | Russian Federation | C1 | |
| AU693374B2 | Australia | B2 | |
| US5778338A | United States of America | A | |
| HU215861B | Hungary | B | |
| HK1014796A1 | Hong Kong, China | A1 | |
| AU711484B2 | Australia | B2 | |
| SG70558A1 | Singapore | A1 | |
| EP1107231A2 | European Patent Office (EPO) | A2 | |
| FI20011508A | Finland | A | |
| FI20011508A7 | Finland | A7 | |
| FI20011508L | Finland | L | |
| FI20011509A | Finland | A | |
| FI20011509A7 | Finland | A7 | |
| FI20011509L | Finland | L | |
| EP1126437A2 | European Patent Office (EPO) | A2 | |
| EP0588932B1 | European Patent Office (EPO) | B1 | |
| AT208945T | Austria | T | |
| ATE208945T1 | Austria | T1 | |
| EP1107231A3 | European Patent Office (EPO) | A3 | |
| EP1126437A3 | European Patent Office (EPO) | A3 | |
| EP1162601A2 | European Patent Office (EPO) | A2 | |
| DE69232202D1 | Germany | D1 | |
| JP2002023796A | Japan | A | |
| DK0588932T3 | Denmark | T3 | |
| ES2166355T3 | Spain | T3 | |
| EP1162601A3 | European Patent Office (EPO) | A3 | |
| JP2002202800A | Japan | A | |
| DE69232202T2 | Germany | T2 | |
| EP1239456A1 | European Patent Office (EPO) | A1 | |
| CN1091535C | China | C | |
| CN1381956A | China | A | |
| CN1398052A | China | A | |
| CN1112673C | China | C | |
| JP3432822B2 | Japan | B2 | |
| CN1119796C | China | C | |
| JP2004004897A | Japan | A | |
| CN1492395A | China | A | |
| EP1126437B1 | European Patent Office (EPO) | B1 | |
| AT272883T | Austria | T | |
| ATE272883T1 | Austria | T1 | |
| DE69233397D1 | Germany | D1 | |
| JP3566669B2 | Japan | B2 | |
| DK1126437T3 | Denmark | T3 | |
| HK1064785A1 | Hong Kong, China | A1 | |
| ES2225321T3 | Spain | T3 | |
| CN1196271C | China | C | |
| EP1107231B1 | European Patent Office (EPO) | B1 | |
| AT294441T | Austria | T | |
| ATE294441T1 | Austria | T1 | |
| DE69233502D1 | Germany | D1 | |
| JP2005182075A | Japan | A | |
| DE69233397T2 | Germany | T2 | |
| NO319559B1 | Norway | B1 | |
| CN1220334C | China | C | |
| ES2240252T3This record | Spain | T3 | |
| DE69233502T2 | Germany | T2 | |
| JP3751957B2 | Japan | B2 | |
| JP2006079107A | Japan | A | |
| CA2102099C | Canada | C | |
| EP1675100A2 | European Patent Office (EPO) | A2 | |
| JP2006221186A | Japan | A | |
| CN1286086C | China | C | |
| FI20061121A | Finland | A | |
| FI20061121A7 | Finland | A7 | |
| FI20061122A7 | Finland | A7 | |
| FI20061122L | Finland | L | |
| CN1909059A | China | A |
Numbers
- Publication
- 2240252
- Publication, DOCDB
- 2240252
- Publication, EPODOC
- ES2240252T
- Application
- 1103640
- Application, DOCDB
- 01103640
- Application, EPODOC
- ES20010103640T
Titles2
- Spanish
- VOCODIFICADOR DE VELOCIDAD VARIABLE.
- English
- VARIABLE SPEED VOCODIFIER.
Classification
- CPC, 14
- H04L1/0057
- G10L19/005
- G10L19/012
- G10L19/12
- G10L19/22
- G10L19/24
- G10L25/78
- G10L2025/786
- H04B1/66
- H04J3/1688
- H04L1/0014
- H04L1/0017
- H04L1/0041
- H04L1/0046
- IPC, 11
- G10L19 00
- G10L19 04
- G10L
- G10L19 038
- G10L19 24
- G10L25 93
- H03M7 30
- H03M7 36
- H04B1 66
- H04J3 16
- H04L1 00