A speech communication system and method for handling lost frames
Abstract
A voice communication system (151) comprising: a decoder and an encoder (159) that processes speech frames and determines a delay parameter of tonal degree for each speech frame; a transmitter coupled to the encoder (159) that transmits the delay parameter of the tonal degree for each speech frame; the decoder comprising: a receiver that receives the delay parameters of the tonal degree from the transmitter, on a frame by frame basis; a control logic coupled to the receiver to resynthesize the speech signal based in part on the tonal degree delay parameters; a lost frame detector that detects if a frame has not been received by the receiver; characterized in that the decoder additionally comprises: a frame recovery logic that, when the lost frame detector detects a lost frame, uses the tonal degree delay parameters of a plurality of previously received frames, in order to extrapolate a tonal degree delay parameter for the lost frame a temporary adaptive codebook store that contains a total excitation for the first frame that follows the lost frame, total excitation including a quantized excitation component of adaptive codebook; wherein the frame recovery logic uses the tonal degree delay parameter of the first frame following the lost frame to adjust the previously set tonal degree delay parameter for the lost frame; and wherein the total excitation stored as an adaptive codebook excitation is extracted for the first frame that follows the lost frame, and where the frame recovery logic uses the tonal degree delay parameter of the first frame that Follow the lost plot to adjust the quantized excitation component of the adaptive codebook.

Term
Term ended
Projected expiry passed 9 July 2021, 5.2 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
10 claims: 2 independent, 8 dependent
- 1ES 2 325 151 T3 IS 2 325 151 T3 CLAIMS REIVINDICACIONES 1. A speech communication system (151) comprising:a decoder and an encoder (159) that processes speech frames and that determines a pitch delay parameter for each speech frame;1. Un sistema (151) de comunicación vocalque comprende: un descodificador y un codificador (159) que procesa tramas de habla y que determina un parámetro de retardo de grado tonal para cada trama de habla;a transmitter coupled to the encoder (159) that transmits the pitch delay parameter for each speech frame;un transmisor acoplado al codificador (159) que transmite el parámetro de retardo del grado tonal para cada trama de habla;comprendiendo el descodificador: comprising the decoder: a receiver that receives the pitch delay parameters from the transmitter, on a frame-by-frame basis;un receptor que recibe los parámetros de retardo del grado tonal desde el transmisor, en base trama a trama;control logic coupled to the receiver to resynthesize the speech signal based in part on the pitch delay parameters;una lógica de control acoplada al receptor para resintetizar la señal de habla basándose en parte en los parámetros de retardo de grado tonal;a lost frame detector that detects if a frame has not been received by the receiver;un detector de tramas perdidas que detecta si una trama no ha sido recibida por el receptor;caracterizado porque el descodificador comprende adicionalmente: characterized in that the decoder additionally comprises: a frame recovery logic that, when the lost frame detector detects a lost frame, uses the pitch delay parameters of a plurality of previously received frames to extrapolate a pitch delay parameter for the lost frame an adaptive codebook buffer containing total excitation for the first frame that follows the lost frame, total drive including an adaptive codebook quantized drive component;una lógica de recuperación de tramas que, cuando el detector de tramas perdidas detecta una trama perdida, utiliza los parámetros de retardo de grado tonal de una pluralidad de tramas recibidas anteriormente, a fin de extrapolar un parámetro de retardo de grado tonal para la trama perdida un almacén temporal de libro de códigos adaptable que contiene una excitación total para la primera trama que sigue a la trama perdida, incluyendo la excitación total un componente de excitación cuantizada de libro de códigos adaptable;en el que la lógica de recuperación de tramas utiliza el parámetro de retardo de grado tonal de la primera trama que sigue a la trama perdida para ajustar el parámetro de retardo del grado tonal anteriormente fijado para la trama perdida;y en el que la excitación total almacenada como excitación de libro de códigos adaptable se extrae para la primera trama que sigue a la trama perdida, y en donde la lógica de recuperación de tramas utiliza el parámetro de retardo de grado tonal de la primera trama que sigue a la trama perdida para ajustar el componente de excitación cuantizada de libro de códigos adaptable. wherein the frame recovery logic uses the pitch delay parameter of the first frame following the lost frame to adjust the previously set pitch delay parameter for the lost frame;and wherein the total excitation stored as adaptive codebook excitation is extracted for the first frame that follows the lost frame, and wherein the frame recovery logic uses the pitch delay parameter of the first frame that follows the lost frame to adjust the adaptive codebook quantized drive component.
- 8A method for encoding or decoding speech in a communication system (151) comprising the encoding steps of:providing a speech signal on a frame-by-frame basis, where each frame includes a plurality of subframes;determining a parameter for each frame, based on the speech signal;transmit parameters on a frame-by-frame basis;8. Un procedimiento para codificar o descodificar el habla en un sistema (151) de comunicación que comprende las etapas de codificación de: proporcionar una señal de habla en base trama a trama, donde cada trama incluye una pluralidad de subtramas;determinar un parámetro para cada trama, basándose en la señal de habla;transmitir parámetros en base trama a trama;ES 2 325 151 T3 and the decoding steps of: receiving parameters on a frame by frame basis;detect if a frame containing the parameter has been lost;ES 2 325 151 T3 y las etapas de descodificación de: recibir parámetros en base trama a trama;detectar si una trama que contiene el parámetro se ha perdido;caracterizado porque las etapas de descodificación comprenden adicionalmente: characterized in that the decoding steps further comprise: gestionar el parámetro perdido para la trama perdida, si se detecta que se ha perdido una trama, utilizando parámetros de retardo de grado tonal de una pluralidad de tramas anteriormente recibidas para extrapolar un parámetro de retardo de grado tonal para la trama perdida;proporcionar un almacén temporal de libro de códigos adaptable que contiene una excitación total para la primera trama que sigue a la trama perdida, incluyendo la excitación total un componente de excitación cuantizada de libro de códigos adaptable;managing the lost parameter for the lost frame, if it is detected that a frame has been lost, using pitch delay parameters from a plurality of previously received frames to extrapolate a pitch delay parameter for the lost frame;providing an adaptive codebook buffer containing a total drive for the first frame that follows the lost frame, the total drive including an adaptive codebook quantized drive component;using the pitch delay parameter of the first frame following the lost frame to adjust the previously set pitch delay parameter for the lost frame;utilizar el parámetro de retardo de grado tonal de la primera trama que sigue a la trama perdida para ajustar el parámetro de retardo de grado tonal anteriormente fijado para la trama perdida;extracting the total drive stored as an adaptive codebook drive for the first frame that follows the lost frame;extraer la excitación total almacenada como una excitación de libro de códigos adaptable para la primera trama que sigue a la trama perdida;adjusting the adaptive codebook quantized drive component using the pitch delay parameter of the first frame following the lost frame;and use the parameters to reproduce the speech signal. ajustar el componente de excitación cuantizada de libro de códigos adaptable utilizando el parámetro de retardo de grado tonal de la primera trama que sigue a la trama perdida;y utilizar los parámetros para reproducir la señal de habla.
Independent claims2
171 paragraphs in 7 sections, as filed
IS 2 325 151 T3
DESCRIPTION
Voice communication system and procedure to manage lost frames.
The field of the present invention relates generally to speech encoding and decoding in speech communication systems and, more specifically, to a method and apparatus for handling erroneous or lost frames.
To model basic speech sounds, speech signals are sampled over time and stored in frames as a discrete wave to be processed digitally. However, in order to increase the efficient use of speech communication bandwidth, speech is encoded before being transmitted, especially when speech is intended to be transmitted under limited bandwidth constraints. Numerous algorithms have been proposed for the various aspects of speech coding. For example, an analysis-by-synthesis coding approach can be carried out on a speech signal. When encoding speech, the speech encoding algorithm attempts to represent the characteristics of the speech signal in a way that requires less bandwidth. For example, the speech coding algorithm seeks to eliminate redundancies in the speech signal. A first stage is to eliminate short-term correlations. One type of signal coding technique is linear predictive coding (LPC). Using an LPC approach, the value of the speech signal at any specific time is modeled as a linear function of the previous values. Using an LPC approach, short-term correlations can be reduced and efficient representations of the speech signal can be determined by estimating and applying certain prediction parameters to represent the signal. The LPC spectrum, which is an envelope of short-term correlations in the speech signal, can be represented, for example, by LSFs (line spectral frequencies). After removal of short-term correlations in a speech signal, a residual LPC signal remains. This residual signal contains periodicity information that needs to be modeled. The second stage in eliminating redundancies in speech is modeling the periodicity information. Periodicity information can be modeled using pitch grade prediction. Certain portions of speech have periodicity, while other portions do not. For example, the sound "aah" has periodicity information, while the sound "shhh" does not have any periodicity information.
By applying the LPC technique, a conventional source encoder operates on the speech signals to extract modeling and parameter information, to be encoded for communication to a conventional source decoder via a communication channel. One way to encode parameter and modeling information into a smaller volume of information is to use quantization. Quantizing a parameter involves selecting the closest entry, in a table or codebook, to represent the parameter. So, for example, a parameter of 0.125 can be represented by 0.1 if the codebook contains 0,0,0,1,0,2,0.3, and so on. Quantization includes scalar quantization and vector quantization. In scalar quantization, the entry in the table or codebook that is the closest approximation to the parameter is selected, as described above. In contrast, vector quantization combines two or more parameters and selects the entry in the table or codebook that is closest to the combined parameters. For example, vector quantization can select the codebook entry that is closest to the difference between the parameters. A codebook used to vectorly quantize two parameters at once is often referred to as a two-dimensional codebook. An n-dimensional codebook quantizes n parameters at a time.
The quantized parameters can be packed into data packets that are transmitted from the encoder to the decoder. In other words, once encoded, the parameters representing the input speech signal are transmitted to a transceiver. Thus, for example, LSFs can be quantized and the index for a codebook can be converted into bits and transmitted from the encoder to the decoder. According to the embodiment, each packet may represent a portion of a speech signal frame, a speech frame, or more than one speech frame. At the transceiver, a decoder receives the encoded information. Because the decoder is configured to understand how speech signals are encoded, the decoder decodes the encoded information to reconstruct a signal for playback, which sounds to the human ear like original speech. However, it may be unavoidable that at least one data packet is lost during transmission and that the decoder does not receive all the information sent by the encoder. For example, when speech is being transmitted from one cell phone to another cell phone, data can be lost when reception is weak or noisy. Therefore, the transmission of the encoded modeling and parameter information to the decoder requires a way of correcting or adjusting, by the decoder, the lost data packets. While the prior art describes certain ways of adjusting for data packet loss, such as extrapolation to try and guess what the information was in the lost packet, these procedures are limited, so improved procedures are needed.
In addition to the information from the LSFs, other parameters transmitted to the decoder may be lost. In CELP (Code Excited Linear Prediction) speech coding, for example, there are two types of gain that are also quantized and transmitted to the decoder. The first type of gain is the G gain<sub>p</sub> tonal grade, also known as the adaptive codebook gain. Adaptive codebook gain is also referred to, even herein, by the subscript "a" instead of the subscript "p". The second type of gain is the G gain<sub>c</sub> fixed codebook. Speech coding algorithms have quantized parameters including adaptive codebook gain and fixed codebook gain. Other parameters, for example, may include pitch delays, representing the periodicity of speech with speech. If the speech coder classifies the speech signals, the classification information about the speech signal can also be transmitted to the decoder. For an Enhanced Speech Encoder / Decoder that classifies speech and works
ES 2 325 151 T3 in different modalities, see US Patent Application Serial No. 09 / 574,396, entitled "A New Speech Gain Quantization Strategy", Conexant File No. 99RSS313, filed May 19, 2000.
Because these, and other, parameter information is sent by imperfect transmission means to the decoder, some of these parameters are lost, or are never received by the decoder. For speech communication systems that transmit an information packet per speech frame, a lost packet results in a lost information frame. In order to reconstruct or estimate lost information, prior art systems have tried different approaches, depending on the missing parameter. Some approaches simply use the parameter from the previous frame that has actually been received by the decoder. These prior art approaches have their drawbacks, inaccuracies, and problems. Thus, there is a need for an improved way to correct or adjust for loss of information in order to recreate a speech signal as close as possible to the original speech signal.
Certain prior art speech communication systems do not transmit a fixed codebook drive from the encoder to the decoder in order to save bandwidth. Instead, these systems have a local Gaussian time-series generator that uses an initial fixed seed to generate a random excitation value and then updates that seed each time the system encounters a frame that contains silence or background noise. In this way, the seed changes for each noise frame. Because the encoder and decoder have the same Gaussian time series generator, which uses the same seeds in the same sequence, they generate the same random drive value for the noise frames. However, if a noise frame is lost and is not received by the decoder, the encoder and decoder use different seeds for the same noise frame, thus losing their synchronicity.
Therefore, there is a need for a speech communication system that does not transmit fixed codebook drive values to the decoder, but maintains synchronicity between the encoder and decoder when a frame is lost during transmission.
An apparatus and method for carrying out the compression of speech signals is known from the International Patent Application published under publication number WO 92/22891. In the event that a frame is lost due to a channel error, the speech coder attempts to mask this error by keeping a fraction of the energy of the previous frame, and making a smooth transition to background noise.
It is known from the International Patent application published under publication number WO 99/66494, a lost frame recovery technique for LPC-based parametric coding systems, which employs interpolation of parameters of good, previous and subsequent frames.
It is an object of the invention to create an improved speech communication system that is capable of generating more accurate estimates for lost information in a lost data packet, for example, to more precisely manipulate lost information, such as pitch delay. .
This is achieved by the apparatus according to claim 1 and the method according to claim 8. Further advantageous embodiments are the claims in dependent claims 2a7 and 9a10, respectively.
Other aspects, advantages, and novel features of the present invention will become apparent from the following Detailed Description of a Preferred Embodiment, when considered in conjunction with the accompanying figures.
Fig. 1 is a functional block diagram of a speech communication system with a source encoder and a source decoder.
Fig. 2 is a more detailed functional block diagram of the speech communication system of Fig. 1.
FIG. 3 is a functional block diagram of an exemplary first stage, a speech preprocessor, of the source encoder used by one embodiment of the speech communication system of FIG. 1.
FIG. 4 is a functional block diagram illustrating an exemplary second stage of the source encoder used by an embodiment of the speech communication system of FIG. 1.
FIG. 5 is a functional block diagram illustrating an exemplary third stage of the source encoder used by an embodiment of the word communication system of FIG. 1.
FIG. 6 is a functional block diagram illustrating an exemplary fourth stage of the source encoder used by one embodiment of the speech communication system of FIG. 1, to process non-periodic speech (mode 0).
IS 2 325 151 T3
FIG. 7 is a functional block diagram illustrating an exemplary fourth stage of the source encoder used by an embodiment of the speech communication system of FIG. 1, to process periodic speech (mode 1).
FIG. 8 is a block diagram of an embodiment of a speech decoder for processing encoded information from a speech encoder constructed in accordance with the present invention.
Fig. 9 illustrates a hypothetical example of received frames and a lost frame.
Fig. 10 illustrates a hypothetical example of received frames and a lost frame, as well as the minimum gaps between the LSFs assigned to each frame in a prior art system, and of a speech communication system constructed according to present invention.
Fig. 11 illustrates a hypothetical example showing how a prior art speech communication system allocates and uses pitch delay and pitch delay delta information for each frame.
Fig. 12 illustrates a hypothetical example showing how a speech communication system, constructed in accordance with the present invention, allocates and uses pitch lag and pitch lag delta information for each frame.
Fig. 13 illustrates a hypothetical example showing how a speech decoder constructed in accordance with the present invention assigns adaptive gain parameter information for each frame when there is a lost frame.
Fig. 14 illustrates a hypothetical example showing how a prior art encoder uses seeds to generate a random drive value for each frame containing silence or background noise.
Fig. 15 illustrates a hypothetical example showing how a prior art decoder uses seeds to generate a random drive value for each frame containing silence or background noise, and loses synchronicity with the encoder if there is a lost frame.
FIG. 16 is a flow chart showing exemplary non-periodic style speech processing in accordance with the present invention.
FIG. 17 is a flow chart showing exemplary periodic style speech processing in accordance with the present invention.
Detailed description of a preferred embodiment
A general description of the total speech communication system is set forth first, and then a detailed description of an embodiment of the present invention is provided.
Fig. 1 is a schematic block diagram of a speech communication system illustrating the general use of a speech encoder and decoder in a communication system. A speech communication system 100 transmits and reproduces speech over a communication channel 103. Although it may comprise, for example, a cable, fiber, or optical medium link, the communication channel 103 typically comprises, at least in part, a radio frequency link that often must support multiple simultaneous switchboards that they require shared bandwidth resources such as those found in cell phones.
A storage device may be coupled with communication channel 103 to temporarily store speech information for delayed playback, eg, to perform answering machine functions, electronic voice mail, etc. Also, the communication channel 103 could be replaced by such a storage device in a single device embodiment of the communication system 100 that, for example, only records and stores speech for subsequent playback.
In particular, a microphone 111 produces a speech signal in real time. The microphone 111 supplies the speech signal to an A / D converter 115 (analog to digital). The analog-to-digital converter 115 converts the analog speech signal to digital form and then delivers the digitized speech signal to a speech encoder 117.
Speech encoder 117 encodes digitized speech using a mode selected from a plurality of encoding modes. Each of the plurality of encoding modalities employs specific techniques that attempt to optimize the quality of the resulting reproduced speech. By operating in any of the plurality of modes, the speech encoder 117 produces a series of modeling and parameter information (eg. eg, "speech parameters") and supplies the speech parameters to an optional channel encoder 119.
Optional channel encoder 119 coordinates with channel decoder 131 to supply speech parameters on communication channel 103. Channel decoder 131 forwards speech parameters to speech decoder 133. By operating in a mode corresponding to that of speech encoder 117, the
ES 2 325 151 T3 speech decoder 133 attempts to recreate the original speech from the speech parameters, as precisely as possible. The speech decoder 133 supplies the reproduced speech to a d / a (digital to analog) converter 135, so that the reproduced speech can be heard through a loudspeaker 137.
FIG. 2 is a functional block diagram illustrating an exemplary communication device of FIG. 1. A communication device 151 comprises both a speech encoder and decoder for simultaneous speech capture and playback. Typically, within a single enclosure, the communication device 151, for example, could comprise a cellular telephone, a portable telephone, a computer system, or some other communication device. Alternatively, if some memory element is provided to store encoded speech information, the communication device 151 could comprise an answering machine, a recorder, a voice mail system, or other communication memory device.
A microphone 155 and an A / D converter 157 supply a digital speech signal to a coding system 159. The encoding system 159 encodes the speech and supplies the resulting speech parameter information to the communication channel. The supplied speech parameter information may be directed to another communication device (not shown) at a remote location.
As the speech parameter information is received, a decoding system 165 performs speech decoding. The decoding system supplies the speech parameter information to a D / A converter 167, where the analog speech output can be reproduced on a speaker 169. The end result is the reproduction of sounds, as similar as possible to speech. originally captured.
The coding system 159 comprises both a speech processing circuit 185, which carries out the speech coding, and an optional channel processing circuit 187 which carries out the optional channel coding. Similarly, decoding system 165 comprises a speech processing circuit 189 that performs speech decoding and an optional channel processing circuit 191 that performs channel decoding.
Although the speech processing circuit 185 and the optional channel processing circuit 187 are illustrated separately, they may be combined, in part, or in whole, into a single unit. For example, speech processing circuit 185 and channel processing circuitry 187 may share a single DSP (digital signal processor) and / or other processing circuitry. Similarly, speech processing circuit 189 and optional channel processing circuit 191 may be entirely independent, or may be partially or fully combined. In addition, the combinations, total or partial, can be applied to speech processing circuits 185 and 189, to channel processing circuits 187 and 191, to processing circuits 185, 187, 189 and 191, or others, depending on convenient. Furthermore, each or all of the circuits that control aspects of decoder and / or encoder operation may be referred to as control logic and implemented, for example, by a microprocessor, a microcontroller, a CPU (central processing unit ), an ALU (arithmetic / logic unit), a coprocessor, an ASIC (application-specific electronic circuit), or any other kind of circuit and / or software.
Both the encoding system 159 and the decoding system 165 use a memory 161. The speech processing circuit 185 uses a fixed codebook 181 and an adaptive codebook 183 from a word memory 177 during the speech encoding process. the fountain. Similarly, speech processing circuit 189 uses fixed codebook 181 and adaptive codebook 183 during the source decoding process.
Although speech memory 177, as illustrated, is shared by speech processing circuits 185 and 189, one or more speech memories may be separately assigned to each of processing circuits 185 and 189. Memory 161 also contains software used by processing circuits 185, 187, 189, and 191 to perform various functions required in source encoding and decoding processes.
Before discussing the details of an embodiment of the speech coding improvement, an overview of the general speech coding algorithm is provided at this point. The improved word encoding algorithm mentioned in this specification may, for example, be the eX-CELP (Extended CELP) algorithm which is based on the CELP model. Details of the eX-CELP algorithm are set forth in a US patent application issued by the same applicant, Conexant Systems, Inc., and previously incorporated herein by reference: Provisional US Patent Application Serial No. 60 / 155,321, entitled “4 bits / s Speech Coding”, Conexant File No. 99RSS485, registered on September 22, 1999.
In order to achieve toll quality at a low bit rate (such as 4 kilobits per second), the improved speech encoding algorithm departs somewhat from the strict wave-matching criteria of traditional CELP algorithms, and strives to capture the perceptually important characteristics of the input signal. To do this, the improved speech encoding algorithm analyzes the input signal according to certain characteristics such as the degree of noise-like content, the degree of spiky content, the degree of speech content, the degree of non-speech content, the evolution of the magnitude spectrum, the evolution of the energy contour, the evolution of the periodicity, etc., and uses this information to control the weighting during the
ES 2 325 151 T3 encoding and quantization. The philosophy is to accurately represent perceptually important features and allow relatively larger errors on less important features. As a result, the improved word coding algorithm focuses on perceptual matching rather than wave matching. The focus on perceptual matching results in satisfactory speech reproduction, due to the assumption that, at 4 kbits per second, the waveform matching is not accurate enough to accurately capture all the information in the input signal. Consequently, the Enhanced Speech Encoder performs some prioritization to achieve improved results.
In a particular embodiment, the improved speech encoder uses a frame size of 20 milliseconds, that is, 160 samples per second, each frame being divided into two or three subframes. The number of subframes depends on the mode of the subframe processing. In this specific embodiment, one of two modes can be selected for each speech frame: Mode 0 and Mode 1. It is important that the way the subframes are processed depends on the mode. In this specific embodiment, Mode 0 uses two subframes per frame, where the size of each subframe is 10 milliseconds long, or contains 80 samples. Also, in this exemplary embodiment, Mode 1 employs three subframes per frame, where the first and second subframes are 6.625 milliseconds long, or contain 53 samples, and the third subframe is 6.75 milliseconds long, or contains 54 samples. In both Modes, a lead margin of 15 milliseconds can be used. For both Modes 0 and 1, a 10th order Linear Prediction (LP) model can be used to represent the spectral envelope of the signal. The LP model can be encoded in the Line Spectral Frequency (LSF) domain using, for example, a delayed decision multi-stage switched predictive vector quantization scheme.
Mode 0 employs a traditional word encoding algorithm, such as a CELP algorithm. However, Mode 0 is not used for all speech frames. Instead, Mode 0 is selected to manipulate the frames of all speech other than "newspaper style" speech, as discussed in more detail below. For convenience, "periodic-style" speech is referred to here as periodic speech, and all other speech is "non-periodic" speech. Such "non-periodic" speech includes transition frames where typical parameters, such as pitch correlation and pitch pitch lag, change rapidly, and frames whose signal is predominantly noise-like. Mode 0 splits each frame into two subframes. Mode 0 encodes tonal grade delay once per subframe and has a two-dimensional vector quantizer to co-encode tonal grade gain (i.e., adaptive codebook gain) and fixed codebook gain once per subplot. In this exemplary embodiment, the fixed codebook contains two pulse codebooks and one Gaussian codebook; the two pulse code subbooks have, respectively, two and three pulses.
Mode 1 departs from the traditional CELP algorithm. Mode 1 handles frames that contain periodic speech, which typically have high periodicity and are often well represented with a flat pitch span. In this specific embodiment, Mode 1 uses three subframes per frame. The pitch lag is encoded once per frame prior to subframe processing, as part of pitch pre-processing, and the interpolated pitch span is derived from this delay. The three tonal grade gains of the subframes exhibit very stable behavior and are jointly quantized using prevector quantization, based on a mean squares error criterion prior to closed-loop subframe processing. The three reference pitch gains, which are not quantized, are derived from weighted speech and are a by-product of frame-based pitch pre-processing. Using the pre-quantized pitch gains, traditional CELP subframe processing is performed, except that the three fixed codebook gains are left unquantized. The three fixed codebook gains are jointly quantized after subframe processing, which is based on a delayed decision approach, using a moving average energy prediction. The three subframes are subsequently synthesized with fully quantized parameters.
The way in which the processing mode is selected for each speech frame, based on the classification of the speech contained in the frame, and the innovative way in which periodic speech is processed, enables quantization of the gain with significantly less amount bits, without any significant sacrifice in the perceptual quality of speech. Details of this way of processing speech are provided below.
Figs. 3-7 are functional block diagrams illustrating a multi-stage coding approach employed by one embodiment of the speech coder illustrated in Figs. 1 and 2. In particular, Fig. 3 is a functional block diagram illustrating a speech preprocessor 193 comprising the first stage of the multi-stage encoding approach: Fig. 4 is a functional block diagram illustrating the second stage; Figs. 5 and 6 are functional block diagrams describing Mode 0 of the third stage; and Fig. 7 is a functional block diagram describing Mode 1 of the third stage. The speech coder, comprising coder processing circuitry, typically operates under software instructions to carry out the following functions.
The input speech is read and temporarily stored in frames. Returning to the speech preprocessor 193 of FIG. 3, an input speech frame 192 is provided to a silence enhancer 195 which determines whether the speech frame is pure silence, that is, only the "noise of silence" is present. . Speech enhancer 195 adaptively detects, on a frame basis, whether the current frame is simply "silence noise." If signal 192 is "silence noise", speech enhancer 195 skews the signal to the zero level of signal 192. Otherwise, if signal 192 is not "silence noise", speech enhancer 195 does not modify signal 192. Speech enhancer 195 clears the silence portions of clean speech for very low-level noise and thus improves the perceptual quality of speech
ES 2 325 151 T3 clean. The effect of the speech enhancement function becomes especially remarkable when the input speech originates from an A-law source; that is, the input has gone through A-law encoding and decoding immediately prior to processing by the present word encoding algorithm. Because the A-law amplifies the sample values around 0 (p. eg, -1, 0, +1) to -8, or +8, A-law amplification could transform inaudible silence noise into clearly audible noise. After processing by speech enhancer 195, the speech signal is provided to a high pass filter 197.
The high pass filter 197 removes frequencies below a certain cutoff frequency and allows frequencies above the cutoff frequency to pass to a noise attenuator 199. In this specific embodiment, the high pass filter 197 is identical to the high pass input filter of the ITU-T (International Telecommunication Union Telecommunication Standard) word coding standard G.729. That is, it is a second order zero pole filter with a cutoff frequency of 140 Hertz (Hz). Of course, the high pass filter 197 need not be one such filter, and can be constructed to be any kind of suitable filter known to those of ordinary skill in the art.
The noise attenuator 199 performs a noise suppression algorithm. In this specific embodiment, the noise attenuator 199 performs a weak attenuation of the noise, of a maximum of 5 decibels (dB) of the ambient noise, in order to improve the estimation of the parameters by the speech coding algorithm. Specific procedures for improving silence, constructing a high pass filter 197, and attenuating noise may employ any of a number of techniques known to those of ordinary skill in the art. The output of speech preprocessor 193 is preprocessed speech 200.
Of course, the silence enhancer 195, high pass filter 197, and noise attenuator 199 may be replaced by any other device, or modified in a manner known to those of ordinary skill in the art, and suitable for the specific application.
Returning to Fig. 4, a functional block diagram of common frame-based processing of a speech signal is provided. In other words, Fig. 4 illustrates the processing of a speech signal frame by frame. This frame processing occurs regardless of mode (eg, Modes 0 or 1) before the mode-dependent processing 250 is performed. The pre-processed speech 200 is received by a perceptual weighting filter 252 which is used to add emphasis to the valley areas and de-emphasize the peak areas of the pre-processed speech signal 200. The perceptual weighting filter 252 may be replaced by any other device, or modified in a manner known to those of ordinary skill in the art, and suitable for the specific application.
An LPC analyzer 260 receives the pre-processed speech signal 200 and estimates the short-term spectral envelope of the speech signal 200. The LPC analyzer 260 extracts the LPC coefficients from the characteristics that define the speech signal 200. In one embodiment, three 10th order LPC analyzes are performed for each frame. They focus on the middle third, the final third, and the lead margin of the plot. The LPC analysis for the leading margin is recycled to the next frame as the LPC analysis focused on the first third of the frame. Thus, for each frame, four sets of LPC parameters are generated. The LPC analyzer 260 can also quantize the LPC coefficients, for example, in a line spectral frequency domain (LSF). The quantization of the LPC coefficients can be scalar or vector quantization, and can be performed in any suitable domain, in any manner known in the art.
A classifier 270 obtains information about the characteristics of preprocessed speech 200 by looking at, for example, the absolute maximum of the frame, the reflection coefficients, the prediction error, the LSF vector from the LPC analyzer 260, the tenth order autocorrelation, the recent pitch lag and recent pitch gains. These parameters are known to those of ordinary skill in the art and, for that reason, are not explained further here. Classifier 270 uses estimation to control other aspects of the encoder, such as signal-to-noise ratio estimation, tonal grade estimation, classification, spectral smoothing, energetic smoothing, and gain normalization. Again, these aspects are known to those of ordinary skill in the art and, for that reason, are not explained further here. A brief summary of the classification algorithm is provided below.
The classifier 270, with the aid of the pitch grade preprocessor 254, classifies each frame into one of six classes, based on the dominant characteristic of the frame. The classes are (1) Silence / Background noise; (2) Speech without a Noise-Like Voice; (3) No voice; (4) Transition (includes rush); (5) Not Static with Voice; and (6) Static with Voice. The classifier 270 can employ any approach to classify the input signal into periodic signals and non-periodic signals. For example, the classifier 270 may take the pre-processed speech signal, the pitch lag, and the correlation of the second half of the frame, and other information, as input parameters.
Various criteria can be used to determine whether speech is considered to be periodic. For example, speech can be considered periodic if speech is a static signal with voice. Periodic speech may be considered by some people to include static speech with voice and non-static speech with voice, but for the purposes of this specification, periodic speech includes static speech with voice. Also, periodic speech can be flat and static speech. A speech with voice is considered "static" when the speech signal does not change more than a certain amount within a frame. Such a speech signal is more likely to have a well-defined energetic contour. A speech signal is
ES 2 325 151 T3 “flat” if the gain G<sub>p</sub> that speech's adaptive codebook is greater than a threshold value. For example, if the threshold value is 0.7, a speech signal in a subframe is considered to be flat if its gain G<sub>p</sub> adaptive codebook is greater than 0.7. Nonperiodic speech, or speechless speech, includes nonvocal speech (eg, fricatives such as the "shhh" sound), transitions (eg, rush, drift), background noise, and silence.
More specifically, in the exemplary embodiment, the speech coder initially derives the following parameters:
Spectral Slope (estimate of the first reflection coefficient 4 times per frame):
Z<sub>Sk</sub>(n) -<sub>Sk</sub>(nl)
K (k) = ^ -iTi ---------- k = 0,1, ... 3. (1)
EskW n = 0 in which L = 80 is the window over which the reflection coefficient is calculated and s<sub>k</sub>(n) is the k-th segment, given by
5, (/?) = Ί (λ-40 - 20 + / i) ^ (n), λ = 0,1, ... 79, (2) where w<sub>h</sub>(n) is an 80 sample Hamming window and s (0), s (1), ..., s (159) is the current frame of the pre-processed speech signal.
Absolute Maximum (trace the absolute maximum of the signal, 8 estimates per frame):
/ (k) = max {s (n) '. zi = / ?, (fc), / ?, (£) + l, ..., / 7<sub>r</sub>(Á) -1}, k = 0,1, ..., 7 (3) where n<sub>s</sub>(k) yn<sub>c</sub>(k) is the starting point and the ending point, respectively, for the search for the k-th maximum at time k. 160/8 raster samples In general, the segment length is 1.5 times the pitch period and the segments overlap. Thus, a flat contour of the amplitude envelope can be obtained.
The Spectral Slope, Absolute Maximum, and Tonal Degree Correlation parameters form the basis for the classification. However, further processing and analysis of the parameters is carried out prior to the classification decision. Parameter processing initially applies the weighting to all three parameters. Weighting, to some extent, removes the background noise component in the parameters by subtracting the contribution from the background noise. This provides a parameter space that is "independent" of any background noise and thus is more uniform and improves the robustness of the classification in background noise.
The moving averages of the noise tonal grade period energy, noise spectral decay, noise absolute maximum, and noise tonal grade correlation are updated eight times per frame, based on the following equations, equations 4 to 7 The following parameters defined by Equations 4 to 7 are estimated / sampled eight times per frame, providing a fine temporal resolution of the parameter space:
Moving average of the energy of the tonal degree period of the noise:
<E<sub>N</sub> p (k)> = cq <E<sub>N</sub> p (kl)> + (the<sub>l</sub>) -Ep (k), (4) where E<sub>N</sub>,<sub>p</sub>(k) is the normalized energy of the tonal degree period at time k.160 / 8 samples of the raster. The segments on which the energy is calculated may overlap, as the tonal degree period typically exceeds 20 samples (160 samples / 8). Moving average noise spectral decline:
<K<sub>H</sub>(k)> = a ,. <jr<sub>w</sub>(Jt-l)> + (1 -a<sub>1</sub>) - «- (imod2). (3)
Moving average of absolute maximum noise:
<X<sub>N</sub> (*)> = a, · <(¿-1)> + (1 - a,) ·% (k). (6)
IS 2 325 151 T3
Moving average of the noise tonal degree correlation:
<img file="ES2325151T3_D0001.tif" />
in which R<sub>p</sub> is the input pitch correlation for the second half of the frame. The adaptation constant is adaptive, although the typical value is «= 0.99.
«I
The ratio between the background noise and the signal is calculated according to:
<img file="ES2325151T3_D0002.tif" />
The parametric noise attenuation is limited to 30 dB, i.e.
<img file="ES2325151T3_D0003.tif" />
The noise-free set of parameters (weighted parameters) is obtained by removing the noise component according to the following Equations 10 to 12:
Weighted spectral decline estimate:
(k) = <(k mod 2) - / (A) <k <sub>v</sub>(á)>.
Weighted absolute maximum estimate:
Xw (k) = Z (k) - y (k) - <χ<sub>Ν</sub> (k)>.
(10) (11)
Weighted tonal grade correlation estimate:
<img file="ES2325151T3_D0004.tif" />
The evolution of the weighted decline and the weighted maximum is calculated according to the following Equations 13 and 14, respectively, as the slope of the first-order approximation:
<img file="ES2325151T3_D0005.tif" />
IS 2 325 151 T3
Once the parameters in Equations 4 through 14 are updated for the eight sample points in the frame, the following frame-based parameters are calculated from the parameters in Equations 4 through 14:
Weighted Maximum Tonal Degree Correlation:
<img file="ES2325151T3_D0006.tif" />
Averaged maximum tonal grade correlation:
Λΐ> 'Σ<sup>Λ</sup>-./*~<sup>7</sup>+0. of)<sup>0</sup> ("OR
Moving average of the weighted average tonal degree correlation:
<RZ M> = <sup>to</sup>2 <RZ -1)> + (1 - * 2) RZ '(17) where m is the frame number already<sub>2</sub> = 0.75 is the adaptation constant.
Normalized standard deviation of tonal degree delay:
<img file="ES2325151T3_D0007.tif" />
in which L<sub>p</sub>(m) is the input pitch lag and p<sub>Lp</sub>(m) is the mean of the pitch lag over the three previous frames, given by <sup>2</sup>
μ. (m) = -E (L<sub>p</sub>(m-2 + I). (19)
- * 3 LO '
Minimum weighted spectral decline:
<img file="ES2325151T3_D0008.tif" />
Minimum weighted spectral decline moving average:
<img file="ES2325151T3_D0009.tif" />
Average weighted spectral decline:
<img file="ES2325151T3_D0010.tif" />
Minimum slope of weighted decline:
<img file="ES2325151T3_D0011.tif" />
IS 2 325 151 T3
Cumulative slope of the weighted spectral decline:
3κ Γ = Σ<sup>0ί</sup>-Ο-<sup>7 + Ζ</sup>>·(24) /-0
Maximum slope of the weighted maximum:
dz7 <sup>= max</sup>bz. (* ~<sup>7</sup> + Φ = 0,1, ..., 7 (25)
Cumulative slope of the weighted maximum:
Z = XdZw (* ~<sup>7 + /</sup>) (26) * /. Or
The parameters given by Equations 23, 25, and 26 are used to mark where a frame is likely to contain a commit, and the parameters given by Equations 16 to 18, and 20 to 22, are used to mark where a plot is dominated by vocal speech. Based on initial marks, earlier marks, and other information, the frame is classified into one of six classes.
A more detailed description of how the classifier 270 classifies preprocessed speech 200 is described in a US patent application assigned to the present assignee, Conexant Systems, Inc., and previously incorporated herein by reference: US Provisional Patent Application Serial No. 60 / 155,321, entitled “4 kbits / s Speech Coding”, Conexant File No. 99RSS485, filed on September 22, 1999.
The LSF quantizer 267 receives the LPC coefficients from the LPC analyzer 260 and quantizes the LPC coefficients. The purpose of LSF quantization, which can be any known quantization procedure, including scalar or vector quantization, is to represent the coefficients with fewer bits. In this specific embodiment, the LSF quantizer 267 quantizes the 10th order LPC pattern. The LSF 267 quantizer can also smooth out LSFs to reduce undesirable fluctuations in the spectral envelope of the LPC synthesis filter. The LSF 267 quantizer sends the quantized coefficients A<sub>what</sub>(z) 268 to the word encoder subframe processing portion 250. The subframe processing portion of the speech encoder is mode dependent. Although the LSF domain is preferred, quantizer 267 can quantize LPC coefficients in a domain other than LSF.
If tonal grade preprocessing is selected, the weighted speech signal 256 is sent to tonal grade preprocessor 254. The pitch preprocessor 254 cooperates with the open loop pitch estimator 272 to modify weighted speech 256 so that its pitch information can be more accurately quantized. The pitch preprocessor 254 may, for example, use known compression or dilation techniques on pitch cycles to improve the ability of the word encoder to quantize pitch gains. In other words, the tonal grade preprocessor 254 modifies the weighted speech signal 256 in order to better match the estimated tonal grade tracking and thus more precisely tune the coding model, while producing perceptually indistinguishable reproduced speech. . If the encoder processing circuitry selects a preprocessing mode, the pitch preprocessor 254 performs the pitch pitch preprocessing of the weighted speech signal 256. The pitch preprocessor 254 deforms the weighted speech signal 256 to match the interpolated pitch values that will be generated by the decoder processing circuitry. When tonal grade pre-processing is applied, the warped speech signal is called a modified weighted speech signal 258. If the pitch preprocessing mode is not selected, the weighted speech signal 256 passes through the pitch preprocessor 254 without pitch preprocessing (and, for convenience, is still referred to as the "modified weighted speech signal 258". Tonal grade preprocessor 254 may include a wave interpolator whose function and implementation are known to those of ordinary skill in the art. The wave interpolator can modify certain irregular transition segments using known wave lead-lag interpolation techniques, in order to improve the regularities and suppress the irregularities of the speech signal. The pitch gain and pitch correlation for the weighted signal 256 are estimated by the pitch preprocessor 254. The open-loop pitch estimator 272 extracts information about pitch characteristics from weighted speech 256. The pitch information includes pitch lag and pitch gain information.
Tonal grade preprocessor 254 also interacts with classifier 270 via open loop tonal grade estimator 272 to refine classifier 270's classification of the speech signal. Because the tonal grade preprocessor 254 obtains additional information about the speech signal, the additional information
ES 2 325 151 T3 can be used by classifier 270 in order to fine-tune its classification of the speech signal. After performing pitch preprocessing, pitch preprocessor 254 outputs pitch tracking information 284 and unquantized pitch gains 286 to the mode-dependent subframe processing portion 250 of the speech encoder. .
Once the classifier 270 classifies the pre-processed speech 200 into one of a plurality of possible classes, the classification number of the pre-processed speech signal 200 is sent to the mode selector 274 and the mode-dependent subframe processor 250 as information. 280 control. The mode selector 274 uses the classification number to select the mode of operation. In this specific embodiment, the classifier 270 classifies the pre-processed speech signal 200 into one of six possible classes. If the pre-processed speech signal 200 is static speech (eg, called "periodic" speech), the mode selector 274 sets the mode 282 to Mode 1. Otherwise, the mode selector 274 sets the mode. Mode 282 in Mode 0. The mode signal 282 is sent to the mode dependent subframe processing portion 250 of the speech encoder. The mode information 282 is added to the bit stream that is transmitted to the decoder.
The labeling of speech as "periodic" and "non-periodic" should be interpreted with some caution in this specific embodiment. For example, frames encoded using Mode 1 are those that maintain high pitch correlation and high pitch gain throughout the frame, based on pitch tracking 284 derived from only seven bits per frame. Consequently, the selection of Mode 0 instead of Mode 1 could be due to an inaccurate representation of tonal degree tracking 284, with only seven bits, and not necessarily due to the absence of periodicity. Thus, signals encoded using Mode 0 may very well contain periodicity, although not well represented by only seven bits per frame for tonal grade tracking. Therefore, Mode 0 encodes the pitch trace with seven bits twice per frame, for a total of fourteen bits per frame, in order to more adequately represent the pitch trace.
Each of the functional blocks in Figs 3 and 4, and the other Figs in this specification, are not necessarily discrete structures, and can be combined with another functional block, or more, as desired.
The speech encoder mode dependent subframe processing portion 250 operates in two modes, Mode 0 and Mode 1. Figs. 5 and 6 provide functional block diagrams of Mode 0 subframe processing, while Fig. 7 illustrates the functional block diagram of Mode 1 subframe processing of the third stage of the speech encoder. Fig. 8 illustrates a block diagram of a word decoder corresponding to the enhanced speech encoder. The word decoder performs a reverse mapping of the bit stream to the algorithmic parameters, followed by a mode dependent synthesis. A more detailed description of these figures and modalities is provided in a US patent application assigned to the same awardee, Conexant Systems, Inc .; the entire application was previously incorporated herein by reference, United States Patent Application Serial No. 09 / 574,396, entitled "A NEW SPEECH GAIN QUANTIZATION STRATEGY", File Conexant No. 99RSS312, filed May 19, 2000.
The quantized parameters representing the speech signal can be packetized and then transmitted in data packets from the encoder to the decoder. In the exemplary embodiment described below, the speech signal is analyzed frame by frame, where each frame may have at least one subframe, and each data packet contains information for one frame. Thus, in this example, the parametric information for each frame is transmitted in an information packet. In other words, there is a packet for each frame. Of course, other variations are possible and, depending on the embodiment, each packet could represent a portion of a frame, more than one speech frame, or a plurality of frames.
LSF
An LSF (line spectral frequency) is a representation of the LPC spectrum (that is, the short-term envelope of the speech spectrum). LSFs can be considered as particular frequencies, in which the spectrum of the word is sampled. If, for example, the system uses a 10th order LPC, there would be 10 LSFs per frame. There must be a minimum spacing between consecutive LSFs, so that they do not create quasi-unstable filters. For example, if f<sub>i</sub> is the i-th LSF, and is equal to 100 Hz, the (i + 1) -th LSF, f<sub>i + 1</sub>, must be at least f<sub>i</sub> + the minimum spacing. For example, if f, = 100 Hz and the minimum spacing is 60 Hz, f<sub>i + 1</sub> must be at least 160 Hz and can be any frequency greater than 160 Hz .. The minimum spacing is a fixed number that does not vary between frame and frame, and is known to both the encoder and the decoder, so that they can cooperate. .
Suppose the encoder uses predictive encoding to encode LSFs (as opposed to non-predictive encoding), which is necessary to achieve speech communication at low bit rates. In other words, the encoder uses the quantized LSF of a previous frame (s) to predict the LSF of the current frame. The error between the predicted LSF and the true LSF of the current frame, which the encoder derives from the LPC spectrum, is quantized and transmitted to the decoder. The decoder determines the predicted LSF of the current frame in the same way that the encoder did. Then, knowing the error that was transmitted by the encoder, the decoder can calculate the true LSF of the current frame. However, what happens if a frame containing LSF information is lost? Returning to Fig. 9, suppose that the encoder transmits frames 0 to 3, but that the decoder only receives frames 0, 2 and 3. Frame 1 is the lost or "erased" frame. If the current plot
ES 2 325 151 T3 is lost frame 1, the decoder does not have the error information that is necessary to calculate the true LSF. As a result, prior art systems did not calculate the true LSF and instead set the LSF as the LSF of the previous frame, or the average LSF of a certain number of previous frames. The problems with this approach are that the LSF of the current frame may be too inaccurate (compared to the true LSF) and that subsequent frames (that is, frames 2 and 3 in the example in Fig. 9) use an LSF inaccurate of frame 1 to determine its own LSFs. Consequently, extrapolation of the LSF error introduced by a lost frame contaminates the accuracy of the LSFs of subsequent frames.
In an exemplary embodiment of the present invention, an improved speech decoder includes a counter that counts the number of good frames that follow the lost frame. Fig. 10 illustrates an example of the minimum spacings of the LSFs associated with each frame. Suppose a good 0 frame is received by the decoder, but frame 1 is lost. According to the prior art approach, the minimum spacing between the LSFs was a fixed number (60 Hz in Fig. 10), which does not change. In contrast, when the enhanced speech decoder notices a dropped frame, it increases the minimum spacing of that frame, in order to avoid creating a quasi-unstable filter. The magnitude of the increase in this "adaptive and controlled LSF spacing" depends on which increase in spacing would be best for that specific case. For example, the enhanced speech decoder can consider how the signal energy (or signal power) has evolved over time, how the signal's frequency content (spectrum) has evolved over time , and the counter, to determine what value the minimum missing frame spacing should be set to. A person of average skill in the art could run simple experiments to determine what minimum spacing value would be satisfactory to use. An advantage of analyzing the speech signal and / or its parameters to derive a suitable LSF is that the resulting LSF can be closer to the true (but lost) LSF of that frame.
Adaptive Codebook Excitation (Tonal Degree Delay)
Total arousal e<sub>T</sub>, composed of adaptive codebook drive and fixed codebook drive, is described by the following equation:
ey = gp * Cxp + ge * <sup>and</sup>xc (<sup>27</sup>) in which g<sub>p</sub> and g<sub>c</sub> are, respectively, the quantized adaptive codebook gain and the fixed codebook gain, and e<sub>xp</sub> ye<sub>xc</sub> they are adaptive codebook drive and fixed codebook drive. A temporary store (also called the adaptive codebook temporary store) preserves eT and its components from the previous frame. Based on the delay parameter of the pitch in the current frame, the speech communication system selects an e<sub>T</sub> temporary warehouse and uses it as e<sub>xp</sub> for the current plot. Values for g<sub>p</sub>, g<sub>c </sub>ye<sub>xc</sub> are obtained from the current frame. The values of e<sub>xp</sub>, g<sub>p</sub>, g<sub>c</sub> ye<sub>xc</sub> they are then entered into the formula to calculate an eT for the current frame. The calculated eT and its components are stored for the current frame in the temporary store. The process is repeated, so the e<sub>T</sub> stored is then used as e<sub>xp</sub> for the next plot. Thus, the feedback nature of this encoding approach (which is replicated by the decoder) is evident. Because the information in the equation is quantized, the encoder and decoder are in sync. Note that the buffer is a type of adaptive codebook (but is different from the adaptive codebook used for gain drives).
Fig. 11 illustrates an example of pitch delay information transmitted by a prior art speech system for four frames 1 to 4. The prior art encoder would transmit pitch delay for the current frame and a delta value, where the delta value is the difference between the pitch lag of the current frame and the pitch lag of the previous frame. The EVRC (Enhanced Variable Velocity Encoder) standard specifies the use of the delta of the tonal degree delay. Thus, for example, the information packet concerning frame 1 would include the pitch delay L1 and the delta (L1 - L0), where L0 is the pitch delay of the preceding frame 0; the information packet concerning frame 2 would include the pitch-grade delay L2 and the delta (L2-L1); the information packet concerning frame 3 would include the pitch-grade delay L3 and the delta (L3-L2); and so on. Note that the pitch lags of the adjacent frames could be equal, so the delta values could be zero. If frame 2 has been lost and has never been received by the decoder, the only information available about the pitch delay at the time of frame 2 is the pitch delay L1, because the previous frame 1 has not been lost . The loss of information from the tonal grade L2 delay and the delta (L2 - L1) has created two problems. The first problem is how to estimate an exact tonal degree delay L2 for the lost frame 2. The second problem is how to prevent the error in estimating the tonal degree delay L2 from creating errors in subsequent frames. Some prior art systems do not attempt to solve any of the problems.
In attempting to solve the first problem, some prior art systems use the tonal degree delay L1 of the good frame 1 above as an estimated tonal degree delay L2 'for the lost frame 2, even though any difference between the estimated delay L2 'of tonal degree and the true delay L2 of tonal degree would be wrong.
The second problem is how to prevent the error in estimating the estimated pitch delay L2 'from creating errors in subsequent frames. Recall that, as discussed earlier, the pitch lag of the screen
ES 2 325 151 T3 n is used to update the adaptive codebook buffer, which, in turn, is used by subsequent frames. The error between the estimated pitch lag L2 'and the true pitch lag L2 would create an error in the adaptive codebook buffer, which would then create an error in subsequent received frames. In other words, the error in the estimated pitch lag L2 'may result in loss of synchronicity between the adaptive codebook buffer, from the encoder's point of view, and the adaptive codebook buffer. from the decoder's point of view. As a further example, during the processing of the lost current frame 2, the prior art decoder would use the estimated tonal degree delay L2 'as the tonal degree delay L1 (which probably differs from the true tonal degree delay L2) to get e<sub>xp</sub> for frame 2. Using the wrong pitch delay, therefore, selects the e<sub>xp</sub> failed for frame 2, and this error propagates through all subsequent frames. To solve this problem in the prior art, when frame 3 is received by the decoder, the decoder then has the tonal degree delay L3 and the delta (L3 - L2) and can thus calculate inversely what the true delay should have been. Tonal grade L2. The true pitch L2 delay is simply the pitch L3 delay minus the delta (L3 - L2). Thus, the prior art decoder would correct the adaptive codebook buffer that is used by frame 3. Because the lost frame 2 has already been processed with the estimated pitch delay L2 ', it is too late to fix lost plot 2.
Fig. 12 illustrates a hypothetical case of frames to demonstrate the operation of an exemplary embodiment of an improved speech communication system with both problems, due to lost pitch delay information. Suppose frame 2 is lost and frames 0, 1, 3, and 4 are received. During the time the encoder is processing lost frame 2, the enhanced decoder can use the frame's pitch-grade L1 delay 1 above. Alternatively and preferably, the improved decoder may perform an extrapolation based on the tonal degree delay, or delays, of the preceding frame (s), to determine an estimated tonal degree delay L2 ', which may result in a more accurate estimate than the tonal degree delay L1. Thus, for example, the decoder may use the pitch delays L0 and L1 to extrapolate the estimated pitch delay L2 '. The extrapolation procedure can be any extrapolation procedure, such as a curve fitting procedure that assumes a smooth contour of the past tonal degree to estimate the lost pitch delay L2, one that uses an average of the previous degree lags. tonal, or any other extrapolation procedure. This approach reduces the number of bits that are transmitted from the encoder to the decoder, because it is not necessary to transmit the delta value.
To solve the second problem, when the enhanced decoder receives frame 3, the decoder has the correct pitch-grade delay L3. However, as explained above, the adaptive codebook buffer used by frame 3 may be wrong, due to any extrapolation errors in estimating pitch lag L2 '. The improved decoder seeks to correct that errors in estimating pitch lag L2 'in frame 2 affect frames after frame 2, but without having to transmit pitch lag delta information. Once the improved decoder obtains the pitch-grade delay L3, it employs an interpolation procedure, such as a curve fitting procedure, to adjust or fine-tune its previous estimate of the pitch-grade delay L2 '. Knowing the pitch lags L1 and L3, the curve fitting procedure can estimate L2 'more accurately than when the pitch lag L3 was unknown. The result is a tonal grade tuned delay L2 ”, which is used to adjust or correct the adaptive codebook buffer for use by frame 3. More specifically, the tonal grade tuned delay L2 "is used to adjust or correct the adaptive codebook quantized excitation in the adaptive codebook store. Consequently, the improved decoder reduces the number of bits that must be transmitted by fine-tuning the pitch-grade delay L2 'in a way that is satisfactory for most cases. Thus, in order to reduce the effect of any error in the estimation of the pitch delay L2 on the frames received subsequently, the improved decoder may use the pitch delay L3 of the next frame 3 and the pitch delay L1 of the next frame. the previously received frame 1, to fine-tune the previous estimate of the pitch delay L2, assuming a flat contour of the pitch. The accuracy of this estimation approach, based on the pitch delays of the preceding and successive received frames of the lost frame, can be very good, because pitch contours are generally flat for spoken speech.
Profits
During transmission of frames from the encoder to the decoder, a lost frame also results in lost gain parameters, such as the gain g<sub>p</sub> adaptive codebook and gain g<sub>c </sub>fixed codebook. Each frame contains a plurality of subframes, where each subframe has gain information. Thus, the loss of a frame results in lost gain information for each subframe of the frame. Speech communication systems have to estimate the gain information for each subframe of the lost frame. The gain information for one subframe may differ from that for another subframe.
Prior art systems have taken various approaches to estimating the gains for subframes of the lost frame, such as using the gain of the last subframe of the previous good frame as the gains of each subframe of the lost frame. Another variation was to use the gain of the last subframe of the previous good frame as the gain of the first subframe of the lost frame and gradually attenuate this gain before it is used as the gains of the next subframes of the lost frame. In other words, for example, if each frame has four subframes and frame 1 is received, but frame 2 is lost, the gain parameters in the last subframe of received frame 1 are used as the gain parameters of the first plot subplot
ES 2 325 151 T3 loss 2, the gain parameters are then decremented by a certain amount and are used as the gain parameters of the second subframe of the lost frame 2, the gain parameters are then decremented by a certain amount and are used as the gain parameters of the third subframe of the lost frame 2, and the gain parameters are further decremented and used as the gain parameters of the last subframe of the lost frame 2. Yet another approach was to examine the gain parameters of the subframes of a fixed number of previously received frames to calculate the average gain parameters, which are then used as the gain parameters of the first subframe of the lost frame 2, where the parameters The gain parameters could be gradually decreased and used as the gain parameters of the remaining subframes of the lost frame. And yet another approach was to derive median gain parameters by examining the subframes of a fixed number of previously received frames and use the median values as the gain parameters of the first subframe of the lost frame 2, where the gain parameters could gradually decrease and used as the gain parameters of the remaining subframes of the lost frame. Notably, prior art approaches did not perform separate retrieval procedures for adaptive codebook gains and fixed codebook gains; they used the same recovery procedure on both types of profit.
The improved speech communication system can also address lost gain parameters due to a lost frame. If the speech communication system distinguishes between periodic-style speech and non-periodic-style speech, the system can handle the lost gain parameters for each type of speech differently. Additionally, the improved system manages adaptive codebook lost earnings differently than it handles fixed codebook lost earnings. Let us first examine the case of non-periodic style speech. To determine an estimated profit g<sub>p</sub> adaptive codebook, the enhanced decoder calculates an average g<sub>p</sub> of the subframes of an adaptive number of previously received frames. The pitch delay of the current frame (ie, the lost frame), which has been estimated by the decoder, is used to determine the number of previously received frames to examine. In general, the greater the pitch delay, the greater the number of previously received frames to use to calculate an average gp. Therefore, the improved decoder uses a tonal grade sync averaging approach to estimate the adaptive codebook gain gp for non-periodic style speech. The improved decoder then computes a beta β indicating how good the prediction of g has been<sub>p</sub>, based on the following formula:
β = adaptive codebook excitation energy / total energy e<sub>T</sub> excitation = II g<sub>P</sub> * exp ||<sup>2</sup> / (|| g<sub>p</sub> E<sub>xp</sub>||<sup>2</sup> + U & ♦ ejj<sup>2</sup>) (28) β ranges from 0 to 1, and represents the percentage effect of adaptive codebook excitation energy on total excitation energy. The larger β, the greater the effect of the adaptive codebook drive energy. Although unnecessary, the improved decoder preferably treats non-periodic style speech and periodic style speech differently.
Fig. 16 illustrates an example decoder processing flow chart for non-periodic style speech. Step 1000 determines if the current frame is the first lost frame after receiving a frame (ie, a "good" frame). If the current frame is the first lost frame after a good frame, step 1002 determines whether the current subframe processed by the decoder is the first subframe of a frame. If the current subframe is the first subframe, step 1004 calculates an average gp for a number of previous subframes, where the number of subframes depends on the pitch lag of the current subframe. In an exemplary embodiment, if the pitch lag is less than or equal to 40, the average gp is based on two previous subframes; if the pitch delay is greater than 40, but less than or equal to 80, the g<sub>p</sub> average is based on four previous subframes; if the pitch delay is greater than 80, but less than or equal to 120, the g<sub>p</sub> average is based on previous six subframes; and if the pitch delay is greater than 120, the g<sub>p</sub> Average is based on eight previous subframes. Of course, these values are arbitrary and can be set to any other values, depending on the length of the subframe. Step 1006 determines if the maximum value of β exceeds a certain threshold. If the maximum value of β exceeds a certain threshold, step 1008 sets the fixed codebook gain gc, for all subframes of the lost frame, to zero and sets g<sub>p</sub>, for all subframes of the lost frame, at an arbitrarily high number, such as 0.95, instead of the previously determined average gp. The arbitrarily high number indicates a good voice signal. The arbitrarily high number at which the value of gp of the current subframe of the lost frame is set can be based on a number of factors including, but not limited to, the maximum value of β of a certain number of frames. above, the spectral decay of the previously received frame and the energy of the previously received frame.
Otherwise, if the maximum value of β does not exceed a certain threshold (that is, a previously received frame contains the speech rush), step 1010 sets the g<sub>p</sub> of the current subframe of the lost frame at the minimum between (i) the previously determined average gp and (ii) the arbitrarily high number selected (eg, 0.95). Another alternative is to set the gp of the current subframe of the lost frame based on the spectral decay of the frame.
ES 2 325 151 T3 previously received, the energy of the previously received frame, and the minimum between the g<sub>p</sub> previously determined average and the arbitrarily high number selected (eg, 0.95). In the case where the maximum value of β does not exceed a certain threshold, the gain g<sub>c</sub> is based on the gain-adjusted fixed codebook drive energy in the previous subframe and the fixed codebook drive energy in the current subframe. Specifically, the gain-adjusted fixed codebook drive energy in the previous subframe is divided by the fixed codebook drive energy in the current subframe, take the square root of the result, and multiply by an attenuation fraction, and is set as the value of gc, as shown in the following formula:
g<sub>c</sub> = attenuation factor * square root (|| g<sub>c</sub> * e<sub>xc</sub> || ¡_i <sup>2</sup> / || e<sub>xc</sub>||¡<sup>2</sup>) (29)
Alternatively, the decoder may derive the gc for the current subframe from the lost frame based on the ratio of the energy of the previously received frame to the energy of the current lost frame.
Returning to step 1002, if the current subframe is not 1<sup>to</sup> subframe, step 1020 sets the g<sub>p</sub> of the current subframe of the lost frame by a value that is attenuated or reduced from the g<sub>p</sub> of the previous subplot. Each g<sub>p</sub> of the remaining subframes is set to an additionally attenuated value from the g<sub>p</sub> of the previous subplot. The g<sub>c</sub> of the current subframe is calculated in the same way as in step 1010 and formula 29.
Returning to step 1000, if this is not the first lost frame after a good frame, step 1022 calculates the gc of the current subframe in the same way as in step 1010 and the formula 29. Step 1022 also sets the gp of the current subframe of the lost frame by a value that is attenuated or reduced from the gp of the previous subframe. Because the decoder estimates the g<sub>p</sub> and the g<sub>c</sub> otherwise, the decoder can estimate them more accurately than prior art systems.
Let us now examine the case of periodic-style speech, based on the example flowchart illustrated in Fig. 17. Because the decoder can apply different approaches to estimate g<sub>p</sub> and g<sub>c</sub> for periodic-style speech and non-periodic-style speech, estimation of gain parameters may be more accurate than prior art approaches. Step 1030 determines if the current frame is the first lost frame after receiving a frame (ie, a "good" frame). If the current frame is the first lost frame after a good frame, step 1032 sets g<sub>c</sub> to zero for all subframes of the current frame and sets g<sub>p</sub> to an arbitrarily high number, such as 0.95, for all subframes of the current frame. If the current frame is not the first lost frame after a good frame (eg, it is the 2nd frame lost, the 3rd frame lost, etc.), step 1034 sets g<sub>c</sub> to zero for all subframes of the current frame and sets gp to a value that is attenuated from the gp of the previous subframe.
Fig. 13 illustrates a case of frames to demonstrate the operation of the improved speech decoder. Suppose that frames 1, 3, and 4 are good (that is, received) frames, while frames 2, and 5 through 8 are lost frames. If the current lost frame is the first lost frame after a good one, the decoder sets gp to an arbitrarily high number (such as 0.95) for all subframes of the lost frame. Returning to Fig. 13, this would apply to lost frames 2 and 5. The g<sub>p</sub> of the first lost frame 5 is gradually dimmed to set the gp of the other lost frames 6 to 8. Then, for example, if gp is set to 0.95 for the lost frame 5, the gp could be set to 0.9 for lost frame 6 and 0.85 for lost frame 7, and 0.8 for lost frame 8. For gcs, the decoder calculates the average gp from the previously received frames and, if this average gp exceeds a certain threshold, gc is set to zero for all sub-frames of the lost frame. If the average gp does not exceed a certain threshold, the decoder uses the same approach of setting gc for non-periodic style signals described above to set gc here.
After the decoder estimates the missing parameters (e.g., LSF, pitch delays, gains, classification, etc.) in a lost frame and synthesizes the resulting speech, the decoder can adjust the energy of the synthesized speech from the frame lost to frame energy previously received by extrapolation techniques. This can further improve the reproduction accuracy of the original speech, despite the lost frames.
Seed to Generate Fixed Codebook Excitations
In order to save bandwidth, it is not necessary for a speech encoder to transmit a fixed codebook drive to the decoder during periods of background noise or silence. Instead, both the encoder and decoder can randomly generate an excitation value locally, using a Gaussian time series generator. Both the encoder and decoder are configured to generate the same random drive value in the same order. As a result, because the decoder can locally generate the same random excitation value that the encoder generated for a given noise frame, it is not necessary for the excitation value to be transmitted from the encoder to the decoder. To generate a random excitation value, the Gaussian time series generator uses an initial seed to generate the first random excitation value, and then the generator updates the seed with a new value. The generator then uses the updated seed to generate the next random drive value and updates the seed with yet another value. Fig. 14 illustrates a hypothetical frame case to illustrate how a Gaussian time series generator in
ES 2 325 151 T3 a word encoder uses a seed to generate a random drive value and then updates that seed to generate the next random drive value. Suppose frames 0 and 4 contain a speech signal, while frames 2, 3, and 5 contain silence or background noise. In finding the first noise frame (ie frame 2), the encoder uses the initial seed (called "seed 1") to generate a random drive value, to be used as the fixed codebook drive for that frame. . For each sample in that frame, the seed is changed to generate a new fixed codebook drive. Thus, if a frame were sampled 160 times, the seed would change 160 times. In this way, at the moment the next noise frame (noise frame 3) is found, the encoder uses a second, and different, seed (i.e., seed 2) to generate the random drive value for that plot. Although technically the seed for the first sample of the second frame is not the "second" seed, because the seed has changed for each sample of the first frame, the seed for the first sample of the second frame is referred to here as seed 2 for the sake of of comfort. For noise frame 4, the encoder uses a third seed (different from the first and second seeds). To generate the random excitation value for the noise frame 6, the Gaussian time series generator could either restart with the seed or continue with the seed 4, depending on the implementation of the speech communication system. By configuring the encoder and decoder to update the seed in the same way, the encoder and decoder can generate the same seed and thus the same random drive values in the same order. However, a lost frame destroys this synchronicity between the encoder and the decoder in prior art speech communication systems.
Fig. 15 illustrates the hypothetical case presented in Fig. 14, but from the decoder point of view. Suppose that frame 2 of noise has been lost and that frames 1 and 3 are received by the decoder. Because the noise frame 2 has been lost, the decoder assumes that it was of the same type as the previous frame (ie, a speech frame). Having assumed the wrong hypothesis about the lost noise frame 2, the decoder assumes that the noise frame 3 is the first noise frame, when in fact it is the second noise frame found. Because the seeds are updated for every sample of every noise frame found, the decoder would erroneously use seed 1 to generate the random drive value for noise frame 3, when seed 2 should have been used. The lost frame, therefore, has resulted in the lost synchronicity between the encoder and the decoder. Since frame 2 is a noise frame, it is not significant that the decoder uses seed 1, while the encoder has used seed 2, as the result is noise other than the original noise. The same is true for frame 3. However, the error in seed values is significant in its impact on subsequently received frames containing speech. For example, let's focus on frame 4 of speech. Seed-based locally generated Gaussian excitation is used to continuously update the adaptive codebook buffer of frame 3. When frame 4 is processed, adaptive codebook drive is extracted from the adaptive codebook buffer of frame 3, based on information such as the pitch delay in frame 4. Because the encoder has used seed 3 to update the adaptive codebook buffer of frame 3, and the decoder is using seed 2 (the wrong seed!) To update the codebook buffer Frame 3 adaptive, the difference in the refresh of the frame 3 adaptive codebook buffer could create a quality problem on frame 4 in some cases.
The improved speech communication system constructed in accordance with the present invention does not use a fixed initial seed or then update that seed each time the system encounters a noise frame. Instead, the enhanced encoder and decoder derive the seed, for a given frame, from the parameters in that frame. For example, spectrum information, energy and / or gain information in the current frame could be used to generate the seed for that frame. For example, the bits representing the spectrum (say 5 bits b1, b2, b3, b4, b5) and the bits representing energy (say 3 bits c1, c2, c3) could be used to form a string b1, b2 , b3, b4, b5, c1, c2, c3, whose value is the seed. As a numerical example, suppose the spectrum is represented by 01101 and the energy is represented by 011; then the seed is 01101011. Certainly other alternative methods of deriving a seed of information in the frame are possible and are within the scope of the invention. Consequently, in the example of Fig. 15, where the noise frame 2 has been lost, the decoder will be able to derive a seed for the noise frame 3 which is the same seed derived by the encoder. Thus, a lost frame does not destroy the synchronicity between the encoder and the decoder.
While embodiments and implementations of the subject invention have been shown and described, it should be apparent that many more embodiments and implementations are within the scope of the subject invention. Accordingly, the invention is not to be restricted, except in light of the claims and their equivalents.
Contents7
20 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20
93 members in 13 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 20000617191 | United States of America | – | |
| 61719100 | United States of America | A | |
| 61719100 | United States of America | A | |
| 03018041617191 | – | – | – |
| US20000617191 | – | – | – |
Members93
| Document | Office | Kind | |
|---|---|---|---|
| WO0122402A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU7486200A | Australia | A | |
| WO0191112A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU5542201A | Australia | A | |
| WO0207061A2 | World Intellectual Property Organization (WIPO) | A2 | |
| AU6627801A | Australia | A | |
| KR20020033819A | Republic of Korea | A | |
| EP1214706A1 | European Patent Office (EPO) | A1 | |
| TW493161B | Taiwan Province of China | B | |
| WO0207061A3 | World Intellectual Property Organization (WIPO) | A3 | |
| KR20030001523A | Republic of Korea | A | |
| JP2003513296A | Japan | A | |
| EP1301891A2 | European Patent Office (EPO) | A2 | |
| KR20030040358A | Republic of Korea | A | |
| US6574593B1 | United States of America | B1 | |
| BR0014212A | Brazil | A | |
| US6581032B1 | United States of America | B1 | |
| US6604070B1 | United States of America | B1 | |
| EP1338003A1 | European Patent Office (EPO) | A1 | |
| CN1441950A | China | A | |
| US6636829B1 | United States of America | B1 | |
| CN1451155A | China | A | |
| US2003200092A1 | United States of America | A1 | |
| EP1363273A1 | European Patent Office (EPO) | A1 | |
| CN1468427A | China | A | |
| KR20040005970A | Republic of Korea | A | |
| JP2004504637A | Japan | A | |
| JP2004510174A | Japan | A | |
| US6735567B2 | United States of America | B2 | |
| US6757649B1 | United States of America | B1 | |
| JP2004206132A | Japan | A | |
| CN1516113A | China | A | |
| EP1214706B1 | European Patent Office (EPO) | B1 | |
| AT272885T | Austria | T | |
| ATE272885T1 | Austria | T1 | |
| US6782360B1 | United States of America | B1 | |
| DE60012760D1 | Germany | D1 | |
| AU2001255422B2 | Australia | B2 | |
| BR0110831A | Brazil | A | |
| US2004260545A1 | United States of America | A1 | |
| EP1214706B9 | European Patent Office (EPO) | B9 | |
| KR100488080B1 | Republic of Korea | B1 | |
| KR20050061615A | Republic of Korea | A | |
| CN1212606C | China | C | |
| RU2257556C2 | Russian Federation | C2 | |
| DE60012760T2 | Germany | T2 | |
| EP1577881A2 | European Patent Office (EPO) | A2 | |
| EP1577881A3 | European Patent Office (EPO) | A3 | |
| RU2262748C2 | Russian Federation | C2 | |
| US6959274B1 | United States of America | B1 | |
| US6961698B1 | United States of America | B1 | |
| JP2005338872A | Japan | A | |
| JP2006011464A | Japan | A | |
| CN1722231A | China | A | |
| KR100546444B1 | Republic of Korea | B1 | |
| EP1301891B1 | European Patent Office (EPO) | B1 | |
| AT317571T | Austria | T | |
| ATE317571T1 | Austria | T1 | |
| CN1245706C | China | C | |
| CN1252681C | China | C | |
| DE60117144D1 | Germany | D1 | |
| US7054809B1 | United States of America | B1 | |
| CN1267891C | China | C | |
| EP1338003B1 | European Patent Office (EPO) | B1 | |
| DE60117144T2 | Germany | T2 | |
| AT343199T | Austria | T | |
| ATE343199T1 | Austria | T1 | |
| DE60123999D1 | Germany | D1 | |
| US7191122B1 | United States of America | B1 | |
| US2007136052A1 | United States of America | A1 | |
| KR100742443B1 | Republic of Korea | B1 | |
| US7260522B2 | United States of America | B2 | |
| KR100754085B1 | Republic of Korea | B1 | |
| US2007255559A1 | United States of America | A1 | |
| JP4137634B2 | Japan | B2 | |
| JP4176349B2 | Japan | B2 | |
| JP4222951B2 | Japan | B2 | |
| US2009043574A1 | United States of America | A1 | |
| EP1363273B1 | European Patent Office (EPO) | B1 | |
| AT427546T | Austria | T | |
| ATE427546T1 | Austria | T1 | |
| DE60138226D1 | Germany | D1 | |
| US2009177464A1 | United States of America | A1 | |
| EP2093756A1 | European Patent Office (EPO) | A1 | |
| ES2325151T3This record | Spain | T3 | |
| US7593852B2 | United States of America | B2 | |
| US7660712B2 | United States of America | B2 | |
| EP2093756B1 | European Patent Office (EPO) | B1 | |
| US8620649B2 | United States of America | B2 | |
| US2014119572A1 | United States of America | A1 | |
| BRPI0014212B1 | Brazil | B1 | |
| US10181327B2 | United States of America | B2 | |
| US10204628B2 | United States of America | B2 |
Numbers
- Publication
- 2325151
- Publication, DOCDB
- 2325151
- Publication, EPODOC
- ES2325151T
- Application
- 3018041
- Application, DOCDB
- 03018041
- Application, EPODOC
- ES20030018041T
Titles2
- Spanish
- SISTEMA DE COMUNICACION VOCAL Y PROCEDIMIENTO PARA GESTIONAR TRAMAS PERDIDAS.
- English
- VOCAL COMMUNICATION SYSTEM AND PROCEDURE FOR MANAGING LOST SECTIONS.
Classification
- CPC, 7
- G10L19/005
- G10L19/08
- G10L19/07
- G10L19/083
- G10L25/90
- G10L2019/0012
- G10L19/04
- IPC, 8
- G10L13 00
- G10L19 005
- G10L19 04
- H03M7 30
- H03M7 36
- H04B14 04
- H04L1 00
- H04M1 00