A speech communication system and method for handling lost frames
Abstract
This record has no abstract on file.
Term
Term ended
Expired 9 July 2021, 5.2 years ago.
- Priority
- Filed
- Granted
- Expired
- Today
23 claims: 3 independent, 20 dependent
- 1音声通信システムのためのデコーダであって、デコードされるべき音声信号のパラメータを受信する受信機を含み、パラメータはフレームごとに受信され、パラメータは各フレームに対する線スペクトル周波数(LSF)を含み、前記デコーダはさらに、受信機に結合され、パラメータをデコードするための、および音声信号を再合成するための制御ロジックを含み、制御ロジックは連続するフレームのLSF間に必要な最小差を示す最小間隔を含み、前記デコーダはさらに、紛失フレームを検出する紛失フレーム検出器と、紛失フレーム検出器が紛失フレームを検出すると、紛失フレームの最小間隔を、前に受信されたフレームの最小間隔よりも大きい第1の値に設定するフレーム回復ロジックとを含む、デコーダ。
- 2紛失フレーム検出器は制御ロジックの一部である、請求項1に記載のデコーダ。
- 3フレーム回復ロジックは、紛失フレームの後で受信されたフレームの最小間隔を第2の値に設定し、第2の値は、紛失フレームの直前に受信されたフレームの最小間隔よりも大きく、紛失フレームの最小間隔よりも小さい、請求項1に記載のデコーダ。
- 4フレーム回復ロジックは、紛失フレームの後で受信された第2のフレームの最小間隔を第3の値に設定し、第3の値は、紛失フレームの最小間隔よりも小さい、またはそれに等しい、請求項3に記載のデコーダ。
- 5フレーム回復ロジックは、紛失フレームの後で受信された第2のフレームの最小間隔を第3の値に設定し、第3の値も、紛失フレームの後で受信された第1のフレームの最小間隔よりも小さい、またはそれに等しい、請求項4に記載のデコーダ。
- 6紛失フレームに続いて受信されたフレームの数をカウントするカウンタをさらに含み、カウントは受信フレームの最小間隔の値を求める、請求項1に記載のデコーダ。
- 7紛失フレームに続いて受信されたフレームの数をカウントするカウンタをさらに含み、カウントは受信フレームの最小間隔の値を求める、請求項3に記載のデコーダ。
- 8フレーム回復ロジックは、音声信号のエネルギに少なくとも部分的に基づいて、紛失フレームの最小間隔を設定する、請求項1に記載のデコーダ。
- 9フレーム回復ロジックは、音声信号の周波数スペクトルに少なくとも部分的に基づいて、紛失フレームの最小間隔を設定する、請求項1に記載のデコーダ。
- 10フレーム回復ロジックは、音声信号のエネルギに少なくとも部分的に基づいて、紛失フレームの最小間隔を設定する、請求項3に記載のデコーダ。
- 11フレーム回復ロジックは、音声信号の周波数スペクトルに少なくとも部分的に基づいて、紛失フレームの最小間隔を設定する、請求項3に記載のデコーダ。
- 12フレーム回復ロジックは、音声信号の周波数スペクトルにも少なくとも部分的に基づいて、紛失フレームの最小間隔パラメータを設定する、請求項10に記載のデコーダ。
- 13フレーム回復ロジックは、音声信号のエネルギにも少なくとも部分的に基づいて、紛失フレームの最小間隔パラメータを設定する、請求項11に記載のデコーダ。
- 14フレーム回復ロジックが紛失フレームの紛失パラメータを設定した後で、デコーダは紛失フレームからの音声を再合成し、合成された音声のエネルギを調整して、前に受信されたフレームから合成された音声のエネルギをマッチングさせる、請求項1に記載のデコーダ。
- 15フレーム回復ロジックが紛失フレームの紛失パラメータを設定した後で、デコーダは紛失フレームからの音声を再合成し、合成された音声のエネルギを調整して、前に受信されたフレームから合成された音声のエネルギをマッチングさせる、請求項3に記載のデコーダ。
- 16フレーム回復ロジックが紛失フレームの紛失パラメータを設定した後で、デコーダは紛失フレームからの音声を再合成し、合成された音声のエネルギを調整して、前に受信されたフレームから合成された音声のエネルギをマッチングさせる、請求項9に記載のデコーダ。
- 17デコードする方法であって、音声信号のパラメータをフレームごとに受信するステップを含み、パラメータは各フレームに対する線スペクトル周波数(LSF)を含み、前記方法はさらに、パラメータをフレームごとにデコードして音声信号を再現するステップを含み、デコードするステップは、連続するフレームのLSF間に必要な最小差を示す最小間隔を用いており、前記方法はさらに、紛失フレームを検出するステップと、紛失フレームの最小間隔を、前に受信されたフレームの最小間隔よりも大きい第1の値に設定するステップとを含む、方法。
- 18紛失パラメータは、紛失フレームに対する線スペクトル周波数の最小間隔を表わす、請求項17に記載の方法。
- 19取扱うステップは、紛失フレームの最小間隔パラメータを、前に受信されたフレームの最小間隔パラメータよりも大きい、またはそれに等しい第1の値に設定する、請求項17に記載の方法。
- 20設定するステップは、紛失フレームの後で受信されたフレームの最小間隔を第2の値に設定し、第2の値は、紛失フレームの直前に受信されたフレームの最小間隔よりも大きいかまたはそれに等しく、紛失フレームの最小間隔よりも小さいかまたはそれに等しい、請求項17に記載の方法。
- 21第1の値は音声信号の周波数スペクトルに少なくとも部分的に基づいている、請求項17に記載の方法。
- 22第1の値は音声信号のエネルギに少なくとも部分的に基づいている、請求項17に記載の方法。
- 23設定するステップが紛失フレームの紛失パラメータを設定した後で紛失フレームからの音声を再合成するステップと、合成された音声のエネルギを調整して、前に受信されたフレームから合成された音声のエネルギをマッチングさせるステップとをさらに含む、請求項17に記載の方法。
Independent claims23
122 paragraphs, as filed
[0001] [Incorporation by Citation] The following U.S. patent applications are incorporated herein by reference in their entirety and become part of this application.
[0002] US Patent Application No. 09 / 156,650, "Speech Encoder Using Gain Normalization That Combines Open And Closed Loop Gains", Conexant ( Conexant) Case No. 98RSS399, filed September 18, 1998.
[0003] US Provisional Patent Application No. 60 / 155,321, "4 kbits / s Speech Coding", Connectant Case No. 99RSS485, filed September 22, 1999.
[0004] US Patent Application No. 09 / 574,396, "A New Speech Gain Quantization Strategy," Connexant Case No. 99RSS312, filed May 19, 2000.
Background of the Invention The field of the invention generally relates to voice encoding and decoding in voice communication systems, and more specifically to methods and devices for handling incorrect or lost frames.
[0006] To model a basic audio sound, the audio signal is sampled over time and stored in a frame as a discrete waveform to be digitally processed. However, in order to increase the efficient use of voice communication bandwidth, voice is encoded before it is transmitted, especially if the voice is to be transmitted under limited bandwidth constraints. Numerous algorithms have been proposed for various aspects of speech coding. For example, a synthetic analysis coding technique may be applied to a voice signal. When encoding speech, speech coding algorithms attempt to characterize the speech signal in a way that requires less bandwidth. For example, speech coding algorithms seek to eliminate redundancy in speech signals. The first step is to remove the short-term correlation. One type of speech coding technique is Linear Predictive Co-coding (LPC). When using the LPC technique, the audio signal value at any particular time is modeled as a linear function of the previous value. Short-term correlations can be reduced by using LPC techniques, and efficient audio signal display can be determined by estimating and applying certain predictive parameters to represent the signal. The LPC spectrum, which is the envelope of short-term correlation in an audio signal, may be represented, for example, by LSF (Line Spectral Frequency). After removing the short-term correlation in the audio signal, the LPC residual signal remains. This residual signal contains periodic information that needs to be modeled. The second step in removing redundancy in speech is to model periodic information. Periodic information may be modeled by using pitch prediction. Some parts of the voice have periodicity, while others do not. For example, the sound "aah" has periodic information, but the sound "shhh" does not have periodic information.
[0007] When applying the LPC technique, a conventional source encoder operates on an audio signal to extract modeling and parameter information to be encoded in order to communicate with the conventional source decoder via a communication channel. One way to encode modeling and parameter information into less information is to use quantization. Parameter quantization involves selecting the closest entry in a table or codebook to represent the parameter. Thus, for example, a parameter of 0.125 may be represented by 0.1 if the codebook contains 0, 0.1, 0.2, 0.3, and so on. Quantization includes scalar quantization and vector quantization. Scalar quantization selects an entry in a table or codebook that is the closest approximation to a parameter, as described above. Vector quantization, on the other hand, combines two or more parameters and selects the entry in the table or codebook that is closest to the combined parameter. For example, vector quantization may select the entry in the codebook that is closest to the difference between the parameters. The codebook used to vector-quantize two parameters at once is often referred to as a two-dimensional codebook. The n-dimensional codebook quantizes n parameters at once.
[0008] The quantized parameters may be packaged in packets of data transmitted from the encoder to the decoder. In other words, once encoded, the parameters representing the input audio signal are transmitted to the transceiver. So, for example, the LSF may be quantized and the index to the codebook may be converted to bits and sent from the encoder to the decoder. Depending on the embodiment, each packet may represent a portion of a frame of a voice signal, a voice frame, or more than a voice frame. In the transceiver, the decoder receives the encoded information. Since the decoder is configured to know how to encode the audio signal, the decoder decodes the encoded information and restores the signal for playback that sounds like the original audio to the human ear. .. However, it may be unavoidable that at least one packet of data is lost in transit and the decoder may not receive all of the information sent by the encoder. For example, when audio is being sent from one mobile phone to another, data may be lost if the reception is poor or noisy. Therefore, sending encoded modeling and parameter information to the decoder requires a way for the decoder to correct or adjust for lost data packets. Prior art describes some methods of adjusting for lost packets of data, such as by extrapolation trying to guess what the information in the lost packets was, but these methods are limited and improved. A method is needed.
[0009] In addition to the LSF information, other parameters sent to the decoder may be lost. For example, in CELP (Code Excited Linear Prediction) speech coding, there are two types of gain that are also quantized and sent to the decoder. The first type of gain is pitch gain G<sub>P</sub>It is also known as adaptive codebook gain. Adaptive codebook gains, including here, are sometimes referred to with the subscript a instead of the subscript p. The second type of gain is fixed codebook gain G<sub>C</sub>Is. The speech coding algorithm has quantized parameters including adaptive codebook gain and fixed codebook gain. Other parameters may include, for example, a pitch lag that represents the periodicity of the generated speech. When the voice encoder classifies the voice signal, the classification information regarding the voice signal may also be transmitted to the decoder. For improved voice encoders / decoders that classify voice and operate in different modes, US Patent Application No. 09 / 574,396, "New Voice Gain Quantization Strategies," Conexant Case No. 99RSS312, incorporated earlier. See the May 19, 2000 application.
[0010] Since these and other parameter information is sent to the decoder through an incomplete transmission medium, some of these parameters are lost or never received by the decoder. For voice communication systems that transmit one packet of information per frame of voice, lost packets result in lost frames of information. Prior art systems have tried different techniques to recover or estimate lost information, depending on the lost parameters. Some techniques simply use parameters from the previous frame actually received by the decoder. These prior art techniques have drawbacks, errors, and problems. For this reason, there is a demand for improved methods of correcting or adjusting lost information so as to reproduce the audio signal as close as possible to the original audio signal.
[0011] Some prior art voice communication systems do not transmit fixed codebook excitation from the encoder to the decoder in order to save bandwidth. Instead, these systems use an initial fixed seed to generate a random excitation value, and then update that seed each time the system encounters a frame with silence or background noise, a local Gaussian time series. Has a generator. Therefore, the seed changes for each noise frame. Since the encoder and decoder have the same Gauss time series generator with the same sequence and the same seed, they generate the same random excitation value for the noise frame. However, if the noise frame is lost and not received by the decoder, the encoder and decoder use different seeds for the same noise frame, thereby losing their simultaneity. Therefore, although the fixed codebook excitation value is not transmitted to the decoder, there is a demand for a voice communication system that maintains simultaneity between the encoder and the decoder when a frame is lost during transmission.
[0012] Various distinct aspects of the invention can be found in voice communication systems and methods that have an improved way of dealing with information lost during transmission from an encoder to a decoder. In particular, improved voice communication systems can generate more accurate estimates of lost information in data lost packets. For example, an improved voice communication system can more accurately handle lost information such as LSF, pitch lag (or adaptive codebook excitation), fixed codebook excitation, and / or gain information. In one embodiment of a voice communication system that does not send a fixed codebook excitation value to the decoder, the improved encoder / decoder is the same for a given noise frame, even if the previous noise frame is lost during transmission. Random excitation values can be generated.
A first distinct aspect of the invention is to set the minimum spacing between LSFs to an increased value and then reduce the value for subsequent frames in a controlled and adaptive manner. It is a voice communication system that handles lost LSF information.
A second distinct aspect of the present invention is a voice communication system that estimates lost pitch lag by extrapolating from the pitch lag of a plurality of previous receive frames.
A third distinct aspect of the invention is to receive the pitch lag of the next receive frame and use a curve that fits between the pitch lag of the previous receive frame and the pitch lag of the next receive frame to make the lost frame. A voice communication system that fine-tunes the pitch lag estimation for and adjusts or corrects the adaptive codebook buffer before use by subsequent frames.
A fourth distinct aspect of the present invention is in a voice communication system that estimates the loss gain parameter of periodic-like speech, as opposed to estimating the loss gain parameter of periodic-like speech. is there.
A fifth distinct aspect of the present invention is a voice communication system that estimates lost adaptive codebook gain parameters, as opposed to estimating lost fixed codebook gain parameters.
A sixth distinct aspect of the present invention is the loss of aperiodic-like voice loss frames based on the average adaptation codebook gain parameter of the subframes of the frames received prior to the adaptation number. Adaptive Codebook A voice communication system that determines gain parameters.
A seventh distinct aspect of the invention is based on the average adaptive codebook gain parameter of the subframes of the frames received before the adaptive number and the ratio of the adaptive codebook excitation energy to the total excitation energy. , A voice communication system that determines the lost adaptive codebook gain parameters of aperiodic-like lost frames.
The eighth distinct aspect of the present invention is the average adaptive codebook gain parameter of the subframe of the frame received before the adaptive number, the ratio of the adaptive codebook excitation energy to the total excitation energy, received before. A voice communication system that determines the lost adaptive codebook gain parameter of an aperiodic-like lost voice frame based on the spectral tilt of the frame and / or the energy of the previously received frame.
A ninth distinct aspect of the present invention is a voice communication system that sets the lost adaptive codebook gain parameter of the aperiodic-like lost voice frame to an arbitrarily large number.
A tenth distinct aspect of the present invention is a voice communication system that sets the lost fixed codebook gain parameter to zero for all subframes of aperiodic-like lost voice frames.
The eleventh distinct aspect of the present invention is the loss of the current subframe of the aperiodic-like voice loss frame, based on the ratio of the energy of the previously received frame to the energy of the lost frame. It is a voice communication system that determines fixed codebook gain parameters.
A twelfth distinct aspect of the present invention is the lost fixed codebook gain parameter of the current subframe of the lost frame, based on the ratio of the energy of the previously received frame to the energy of the lost frame. A voice communication system that determines and then attenuates that parameter to set the lost fixed codebook gain parameter for the remaining subframes of the lost frame.
A thirteenth distinct aspect of the present invention is to arbitrarily increase the lost adaptive codebook gain parameter of the first frame of periodic-like speech that will be lost after the received frame. It is a voice communication system to be set.
A fourteenth distinct aspect of the present invention is to arbitrarily increase the lost adaptive codebook gain parameter of the first frame of periodic-like audio that will be lost after the received frame. A voice communication system that sets and then attenuates its parameters to set the lost adaptive codebook gain parameters for the remaining subframes of the lost frame.
A fifteenth distinct aspect of the present invention is the lost fixation of periodic-like audio loss frames when the average adaptive codebook gain parameter of multiple previously received frames exceeds a threshold. A voice communication system that sets the codebook gain parameter to zero.
A sixteenth distinct aspect of the present invention is that if the average adaptive codebook gain parameter of a plurality of previously received frames does not exceed the threshold value, then the energy of the previously received frame for the lost frame. A voice communication system that determines the lost fixed codebook gain parameter of the current subframe of a periodic-like lost voice frame based on the energy ratio.
A seventeenth distinct aspect of the present invention is the energy of the previously received frame relative to the energy of the lost frame if the mean adaptive codebook gain parameter of the previously received frames exceeds the threshold. Determine the lost fixed codebook gain parameter of the current subframe of the lost frame based on the ratio of, then attenuate that parameter to the lost fixed codebook gain of the remaining subframes of the lost frame. It is a voice communication system that sets parameters.
An eighteenth distinct aspect of the present invention is a voice communication system that randomly generates a fixed codebook excitation for a given frame by using a seed whose value is determined by the information in that frame. ..
A nineteenth distinct aspect of the present invention is voice communication in which the lost parameters in the lost frame are estimated, the voice is synthesized, and then the energy of the synthesized voice is matched to the energy of the previously received frame. It is a decoder.
[0032] A twentieth distinct aspect of the invention is any of the separate aspects described above, either individually or in some combination.
Further distinct aspects of the invention can also be found in methods of encoding and / or decoding audio signals that practice either of the above separate aspects individually or in certain combinations.
Other aspects, advantages, and novel features of the invention will be apparent from the detailed description of the following preferred embodiments, along with the accompanying drawings.
[Detailed Description of Preferred Examples] First, a general description of the entire voice communication system will be described, and then examples of the present invention will be described in detail.
[0036] FIG. 1 is a schematic block diagram of a voice communication system showing a general use example of a voice encoder and decoder in a communication system. The voice communication system 100 transmits and reproduces voice over the communication channel 103. Communication channel 103 may include, for example, wire, fiber, or optical links, but typically includes at least partially radio frequency links, which require shared bandwidth resources that can be seen on mobile phones. It often has to support a large number of simultaneous voice exchanges.
A storage device is coupled to the communication channel 103 to temporarily store voice information for later reproduction or reproduction, such as performing an answering machine function or voice mail. Similarly, the communication channel 103 can be replaced with a storage device in a single device embodiment of communication system 100 that only records and stores voice for later playback, for example.
[0038] Specifically, the microphone 111 generates an audio signal in real time. The microphone 111 passes the audio signal to the A / D (analog to digital) converter 115. The A / D converter 115 converts the analog audio signal into a digital format, and then passes the digitized audio signal to the audio encoder 117.
[0039] The voice encoder 117 encodes digitized voice using one of a plurality of encoding modes selected. Each of the multiple encoding modes uses a particular technique that attempts to optimize the quality of the resulting reproduced speech. While operating in any of a plurality of modes, the audio encoder 117 generates a set of modeling and parameter information (eg, "audio parameters") and passes the audio parameters to any channel encoder 119.
[0040] The arbitrary channel encoder 119 cooperates with the channel decoder 131 to send voice parameters via the communication channel 130. The channel decoder 131 sends audio parameters to the audio decoder 133. While operating in a mode corresponding to the mode of the audio encoder 117, the audio decoder 133 attempts to reproduce the original audio as accurately as possible from the audio parameters. The audio decoder 133 passes the reproduced audio to the D / A (digital to analog) converter 135, and the reproduced audio can be heard from the speaker 137.
FIG. 2 is a functional block diagram showing an example of the communication device of FIG. The communication device 151 includes both an audio encoder and a decoder for simultaneously capturing and reproducing audio. The communication device 151, typically in a single housing, may include, for example, a cell phone, mobile phone, computing system, or other communication device. Alternatively, if a memory element for storing encoded voice information is provided, the communication device 151 may include an answering machine, a recording device, a voice mail system, or other communication memory device.
The microphone 155 and the A / D converter 157 pass the digital audio signal to the encoding system 159. The encoding system 159 performs voice encoding and passes the resulting voice parameter information to the communication channel. The passed voice parameter information can be directed to another communication device (not shown) at a remote location.
[0043] When the voice parameter information is received, the decoding system 165 performs voice decoding. The decoding system can pass the audio parameter information to the D / A converter 167 and send an analog audio output from the speaker 169. The end result is a sound that is as similar to the original captured voice as possible.
[0044] The encoding system 159 includes both an audio processing circuit 185 that performs audio encoding and an arbitrary channel processing circuit 187 that performs arbitrary channel encoding. Similarly, the decoding system 165 includes an audio processing circuit 189 that performs audio decoding and an arbitrary channel processing circuit 191 that performs channel decoding.
Although the audio processing circuit 185 and the arbitrary channel processing circuit 187 are illustrated separately, they can be partially or wholly combined into a single unit. For example, the audio processing circuit 185 and the channel processing circuit 187 may share a single DSP (digital signal processor) and / or other processing circuits. Similarly, the audio processing circuit 189 and any channel processing circuit 191 may be completely separate or may be partially or wholly combined. In addition, whole or partial combinations can be applied to speech processing circuits 185 and 189, channel processing circuits 187 and 191 and processing circuits 185, 187, 189 and 191 or otherwise as appropriate. Is. In addition, each or all of the circuits that control the behavior of the decoder and / or encoder are sometimes referred to as control logic, such as microprocessors, microcontrollers, CPUs (Central Processing Units), ALUs (Arithmetic Logic Units). , Coprocessors, ASICs (Dedicated Integrated Circuits), or any other type of circuit and / or software.
[0046] Both the encoding system 159 and the decoding system 165 use the memory 161. The voice processing circuit 185 uses the fixed codebook 181 and the adaptive codebook 183 of the voice memory 177 during the source encoding process. Similarly, the speech processing circuit 189 uses the fixed codebook 181 and the adaptive codebook 183 during the source decoding process.
[0047] The illustrated voice memory 177 is shared by voice processing circuits 185 and 189, but one or more separate voice memories may be assigned to each of the processing circuits 185 and 189. Memory 161 further includes software used by processing circuits 185, 187, 189 and 191 to perform various functions required for source encoding and decoding.
[0048] Before discussing examples of speech coding improvements in detail, here is an overview of the voice encoding algorithm as a whole. The improved voice encoding algorithm referred to herein can be, for example, an eX-CELP (extended CELP) algorithm based on the CELP model. Details of the eX-CELP algorithm have been assigned to the same assignee, Conexant Systems Incorporated, and are incorporated herein by reference in a U.S. patent application, namely Conexant Case No. 99RSS485, filed on September 22, 1999, " It is discussed in US Provisional Patent Application No. 60 / 155,321 entitled "4 kilobits / second speech coding".
To achieve call quality at low bit rates (eg 4 kilobits per second), improved voice encoding algorithms depart from the rigorous waveform matching criteria of traditional CELP algorithms and depart from the input signal. Attempts to acquire perceptually important features. To do this, the improved voice encoding algorithm has a noise-like content degree, a spike-like content degree, a voiced content degree, an unvoiced content degree, an amplitude spectrum expansion, an energy contour expansion, etc. The input signal is analyzed according to several features, such as periodic evolution, and this information is used to control weighting during encoding and quantization. The principle here is to accurately represent perceptually important features and tolerate relatively large errors for less important features. As a result, the improved speech encoding algorithm focuses on perceptual matching instead of waveform matching. The focus on perceptual matching results in satisfactory speech reproduction, which is based on the premise that waveform matching is not accurate enough to faithfully capture all the information in the input signal at 4 kilobits per second. Accordingly, the improved audio encoder performs some prioritization to achieve the improved result.
[0050] In one particular embodiment, the improved audio encoder uses a frame size of 20 ms, or 160 samples per second, and each frame is divided into two or three subframes. The number of subframes depends on the mode of subframe processing. In this particular embodiment, one of two modes, mode 0 and mode 1, can be selected for each audio frame. It is important that the way subframes are processed depends on the mode. In this particular embodiment, mode 0 uses two subframes per frame, where the size of each subframe is in the duration of 10 ms, or contains 80 samples. Similarly, in this example, mode 1 uses three subframes per frame, where the first and second subframes have a duration of 6.625 ms or contain 53 samples and a third. The subframe has a duration of 6.75 ms or contains 54 samples. You can use 15ms preemption in both modes. For both mode 0 and mode 1, a linear prediction (LP) model on the 10th order can be used to represent the spectral envelope of the signal. The LP model can be encoded in the line spectral frequency (LSF) region, for example, by using a delayed multi-stage predictive vector quantization scheme.
Mode 0 operates traditional voice encoding algorithms such as the CELP algorithm. However, mode 0 is not used for all audio frames. Mode 0 is selected to handle all frames of audio except "periodic" audio, as discussed in more detail later. For convenience, a "periodic-like" voice is referred to as a periodic voice, and all other voices are "aperiodic" voices. Such "aperiodic" audio includes transition frames in which typical parameters such as pitch correlation and pitch lag change rapidly, and frames in which the signal is mostly noise-like. Mode 0 divides each frame into two subframes. Mode 0 encodes the pitch lag once per subframe and also has a two-dimensional vector quantizer, which produces pitch gain (ie adaptive codebook gain) and fixed codebook gain once per subframe. Encode together. In this embodiment the fixed codebook includes two pulse subcodebooks and one Gauss subcodebook. These two pulse subcodebooks have two and three pulses, respectively.
Mode 1 is different from the traditional CELP algorithm. Mode 1 deals with frames containing periodic audio, which typically has high periodicity and is often represented by smoothed pitch areas. In this particular embodiment, mode 1 uses three subframes per frame. The pitch lag is encoded once per frame prior to the subframe processing as part of the pitch preprocessing, from which the interpolated pitch area is derived. The three pitch gains of the subframe behave very stably and are quantized together using pre-vector quantization based on the mean square error criterion prior to the closed loop subframe processing. The three quantized reference pitch gains are derived from the weighted speech and are a by-product of frame-based pitch preprocessing. Traditional CELP subframe processing is performed using pre-quantized pitch gain, but the three fixed code book gains remain unquantized. These three fixed codebook gains are quantized together after subframe processing, which is based on a delayed determination method that uses moving average prediction of energy. The three subframes are then combined with the fully quantized parameters.
[0053] The perceptual quality of speech is significantly sacrificed by the mode of selecting a processing mode for each speech frame based on the classification of the speech contained within the frame and the innovative method of processing periodic speech. Gain quantization is possible with significantly fewer bits. Details of this aspect of processing speech are described below.
[0054] FIGS. 3 to 7 are functional block diagrams illustrating a multi-stage encoding method used by an embodiment of the voice encoder illustrated in FIGS. 1 and 2. Specifically, FIG. 3 is a functional block diagram illustrating an audio preprocessor 193 that includes a first stage of a multi-stage encoding technique. FIG. 4 is a functional block diagram illustrating the second stage. 5 and 6 are functional block diagrams showing mode 0 of the third stage. FIG. 7 is a functional block diagram showing mode 1 of the third stage. A voice encoder includes an encoder processing circuit and typically operates under software instructions to perform the following functions.
[0055] The input voice is read and buffered into a frame. The input audio frame 192 goes to the audio preprocessor 193 of FIG. 3 and is given to the silence enhancer 195, which determines if the audio frame is purely silent, that is, if there is only "silence noise". To do. The voice enhancer 195 adaptively detects whether the current frame is pure "silence" on a frame basis. If the signal 192 is "silent noise", the voice enhancer 195 brings the signal to level 0 of the signal 192. Conversely, if the signal 192 is not "silent noise", the voice enhancer 195 does not make any changes to the signal 192. Speech Enhancer 195 cleans the silence of clean speech due to extremely low levels of noise, thus improving the perceptual quality of clean speech. The effect of the speech enhancement function is particularly noticeable when the input speech is derived from the A-law source, that is, when the input passes through A-law encoding and decoding immediately before processing by this speech coding algorithm. Since the A law amplifies sample values near 0 (for example, -1, 0, +1) to -8 or +8, the A law amplifies inaudible silence noise to be clearly audible. Can be changed to. After processing by the audio enhancer 195, the audio signal is applied to the high frequency pass filter 197.
[0056] The high frequency pass filter 197 removes frequencies below a certain cutoff frequency and allows frequencies above the cutoff frequency to pass through the noise attenuator 199. In this particular embodiment, the high pass filter 197 is identical to the input high pass filter of the ITU-T G.729 voice coding standard. That is, it is a second order pole 0 filter with a cutoff frequency of 140 hertz (Hz). As a matter of course, the high frequency pass filter 197 does not have to be such a filter, and may be configured by any kind of filter known to those skilled in the art as long as it is suitable.
The noise attenuator 199 executes a noise suppression algorithm. In this particular embodiment, the noise attenuator 199 performs weak noise attenuation of environmental noise up to 5 dB (dB) to improve parameter estimation by the voice encoding algorithm. Any of a number of techniques known to those of skill in the art may be used as a particular method for improving silence, constructing a high pass filter 197, and attenuating noise. As the output of the voice preprocessor 193, the preprocessed voice 200 is obtained.
Of course, the silence enhancer 195, the high frequency pass filter 197 and the noise attenuator 199 may be replaced with any other device known to those of skill in the art and suitable for a particular application, or in such an embodiment. It can be transformed with.
[0059] With reference to FIG. 4, a functional block diagram of general frame-based processing of audio signals is shown. In other words, FIG. 4 illustrates the processing of audio signals on a frame-by-frame basis. This frame processing is performed before the mode-dependent processing 250 is performed regardless of the mode (for example, mode 0 or 1). The preprocessed speech 200 is received by the perceptual weighting filter 252, which acts to emphasize the valley area and leave the peak area of the preprocessed speech signal 200 unemphasized. The perceptual weighting filter 252 may be replaced with any other device known to those of skill in the art and suitable for a particular application, or may be modified in such an manner.
[0060] The LPC analyzer 260 receives the preprocessed audio signal 200 and estimates the short-term spectral envelope of the audio signal 200. The LPC analyzer 260 extracts the LPC coefficient from the characteristics that define the voice signal 200. In one embodiment, three 10th order LPC analyzes are performed for each frame. These analyzes are centered on the middle third, last third, and preemption of the frame. The LPC analysis for preemption is reused in the next frame as an LPC analysis centered on the first third of the frame. In this way, four sets of LPC parameters are generated for each frame. The LPC analyzer 260 can further quantize the LPC coefficients into, for example, the linear spectral frequency (LSF) region. The quantization of the LPC coefficient is scalar or vector quantization, and may be performed in any suitable region by any method known in the art.
[0061] The classifier 270 is to examine, for example, the absolute maximum value of the frame, the reflectance coefficient, the prediction error, the LSF vector from the LPC analyzer 260, the autocorrelation of the tenth order, the recent pitch lag, and the recent pitch gain. So get the information about the characteristics of the preprocessed audio 200. These parameters are known to those of skill in the art and will not be described further herein. The classifier 270 uses this information to control other elements of the encoder, such as signal-to-noise ratio, pitch estimation, classification, spectrum smoothing, energy smoothing, and gain normalization. These aspects are also known to those of skill in the art and will not be described further here. A brief overview of the classification algorithm is given below.
[0062] With the help of the pitch preprocessor 254, classifier 270 classifies each frame into one of six classes according to the dominant characteristics of the frame. These classes are (1) silent / background noise, (2) noise / silent-like voice, (3) silent, (4) transition (including start), (5) non-stationary voiced, and (6) stationary. Voiced. The classifier 270 may use any method to classify the input signal into a periodic signal and a non-periodic signal. For example, classifier 270 can use preprocessed audio signals, late frame correlation and pitch lag, and other information as input parameters.
[0063] Various criteria can be used to determine if speech is considered periodic. For example, if the voice is a steady and voiced signal, the voice can be considered periodic. Some people may think that stationary voiced speech and non-stationary voiced speech are included in periodic voice, but here the periodic voice includes stationary voiced speech. Further, the periodic voice can be a smoothed and steady voice. A voiced voice is considered "steady" if the voice signal does not change more than a certain amount within the frame. Such audio signals are more likely to have well-defined energy contours. Audio Adaptive Codebook Gain G<sub>P</sub>If is above the threshold, the audio signal is "smooth". For example, if the threshold is 0.7, the audio signal in the subframe will have its adaptive codebook gain G.<sub>P</sub>If is above 0.7, it is considered smooth. Aperiodic or unvoiced speech includes unvoiced speech (eg, fricatives such as shhh sounds), transitions (eg, start, end), background noise, and silence.
[0064] More specifically, in an exemplary embodiment, the audio encoder first derives the following parameters. Spectral tilt (estimation of the first reflectance coefficient 4 times per frame) [0065] [Equation 1]<img file="JP4137634B2_D0001.tif" />[0066] Here, L = 80 is the window from which the reflectance coefficient is calculated, and s.<sub>k</sub>(n) is [0067] [number 2]<img file="JP4137634B2_D0002.tif" />The kth segment given by [0068], where w<sub>h</sub>(n) is a humming window of 80 samples, and s (0), s (1), ... s (159) are the current frames of the preprocessed audio signal. Absolute maximum value (tracking of maximum absolute signal value, estimation 8 times per frame) [0069] [Equation 3]<img file="JP4137634B2_D0003.tif" />[0070] Here n<sub>s</sub>(k) and n<sub>e</sub>(k), respectively, the time k · 160/8 samples of the frame is the point of beginning and end for finding the maximum value of the k-th in the Le. Generally, the segment length is 1.5 times the pitch period and segment overlap. In this way, the smoothed contour of the amplitude envelope can be obtained.
Spectral gradients, absolute maximums and pitch correlation parameters form the basis for classification. However, additional parameter processing and analysis is performed prior to the classification decision. First, parameter processing applies weighting to three parameters. Weighting removes the background noise component within the parameter in a sense by reducing the contribution from the background noise. This provides a more uniform parameter space that is "independent" of any background noise, thus improving the strength of the classification against background noise.
[0072] The noise pitch period energy run intermediate, the noise spectral gradient, the absolute maximum noise value, and the noise pitch correlation are updated eight times per frame according to equations 4-7 below. The following parameters specified by equations 4 to 7 are estimated / sampled 8 times per frame, which gives a fine time decomposition of the parameter space. Noise pitch Period Energy run Intermediate [0073] [Equation 4]<img file="JP4137634B2_D0004.tif" />[0074] Here E<sub>N, p</sub>(k) is the normalized energy of the pitch period in the frame time k · 160/8 sample. Since the pitch period typically exceeds 20 samples (160 samples / 8), the energy-calculated segments can overlap. Run middle of noise spectrum gradient [0075] [Equation 5]<img file="JP4137634B2_D0005.tif" />[0076] The run intermediate of the absolute maximum value of noise [0077] [Number 6]<img file="JP4137634B2_D0006.tif" />[0078] Run middle of noise pitch correlation [0079] [Equation 7]<img file="JP4137634B2_D0007.tif" />[0080] Here R<sub>P</sub>Is the input pitch correlation for the second half of the frame. Adaptation constant α<sub>1</sub>Is adaptive, but typical values are α<sub>1</sub>= 0.99. The background noise to signal ratio is calculated by the following formula.
[0081] [Number 8]<img file="JP4137634B2_D0008.tif" />The noise attenuation of the parameter is limited to 30 dB, i.e. as follows.
[0083] [Number 9]<img file="JP4137634B2_D0009.tif" />[0084] A noise-free parameter set (weighted parameter) is obtained by removing the noise component according to the following equations 10 to 12. Estimating the weighted spectral gradient [0085] [Equation 10]<img file="JP4137634B2_D0010.tif" />[0086] Estimating the weighted absolute maximum [0087] [Equation 11]<img file="JP4137634B2_D0011.tif" />Estimating Weighted Pitch Correlation [0089] [Equation 12]<img file="JP4137634B2_D0012.tif" />The weighted slope and the weighted maximum expansion are calculated according to Equations 13 and 14, respectively, as the approximate gradients of the first order.
[0091] [Number 13]<img file="JP4137634B2_D0013.tif" />[0092] [Number 14]<img file="JP4137634B2_D0014.tif" />[0093] Once the parameters of equations 4 to 14 are updated for the eight sample points of the frame, the following parameters based on the frame are calculated from the parameters of equations 4 to 14. Weighted maximum pitch correlation [0094] [Equation 15]<img file="JP4137634B2_D0015.tif" />Weighted Average Pitch Correlation [0096] [Equation 16]<img file="JP4137634B2_D0016.tif" />[0097] Run intermediate of weighted average pitch correlation [0098] [Equation 17]<img file="JP4137634B2_D0017.tif" />[0099] Here, m is a frame number, and α<sub>2</sub>= 0.75 is an adaptive constant. Normalized standard deviation of pitch lag [0100] [Equation 18]<img file="JP4137634B2_D0018.tif" />[0101] Here L<sub>p</sub>(m) is the input pitch lag, μ<sub>Lp</sub>(m) is the middle of the pitch lag over the past three frames given by the following equation.
[0102] [Number 19]<img file="JP4137634B2_D0019.tif" />[0103] Weighted minimum spectral gradient [0104] [Equation 20]<img file="JP4137634B2_D0020.tif" />Run middle of weighted minimum spectral gradient [0106] [Equation 21]<img file="JP4137634B2_D0021.tif" />[0107] Weighted mean spectral gradients [0108] [Equation 22]<img file="JP4137634B2_D0022.tif" />Minimum slope of weighted slope [0110] [Equation 23]<img file="JP4137634B2_D0023.tif" />Cumulative Gradient of Weighted Spectral Gradients [0112] [Equation 24]<img file="JP4137634B2_D0024.tif" />[0113] Maximum gradient of weighted maximum value [0114] [Equation 25]<img file="JP4137634B2_D0025.tif" />Cumulative Gradient of Weighted Maximum Values [0116] [Equation 26]<img file="JP4137634B2_D0026.tif" />The parameters given in Equations 23, 25 and 26 were used to mark whether the frame could contain a start and were given in Equations 16-18, 20-22. The parameters are used to mark whether voiced speech may be dominant in the frame. Based on initial marks, past marks and other information, frames fall into one of six classes.
A more detailed description of how the classifier 270 classifies preprocessed audio 200 is transferred to the same assignee, Conexant Systems Incorporated, a US patent application incorporated herein by reference. That is, it is described in Conexant Case No. 99RSS485, filed on September 22, 1999, US Provisional Patent Application No. 60 / 155,321 entitled "4 Kilobit / s Speech Coding".
[0119] The LSF quantizer 267 receives the LPC coefficient from the LPC analyzer 260 and quantizes the LPC coefficient. LSF quantization can be any known quantization method, including scalar or vector quantization, and the purpose of this quantization is to represent the coefficients in fewer bits. In this particular embodiment, the LSF quantizer 267 quantizes the 10th order LPC model. In addition, the LSF quantizer 267 can smooth out the LSF to reduce unwanted variations in the spectral envelope of the LPC synthesis filter. LSF Quantizer 267 has a quantized coefficient A<sub>q</sub>(z) 268 is sent to the subframe processing part 250 of the audio encoder. The subframe processing part of the audio encoder depends on the mode. Although LSF is preferred, the quantizer 267 can also quantize the LPC coefficient into regions other than the LSF region.
If pitch preprocessing is selected, the weighted audio signal 256 is sent to the pitch preprocessor 254. The pitch preprocessor 254 can work with the open loop pitch estimator 272 to make changes to the weighted speech 256, thus more accurately quantizing its pitch information. For example, the pitch preprocessor 254 can use known compression or decompression techniques for pitch cycles to improve the ability of voice encoders to quantize pitch gain. In other words, the pitch preprocessor 254 modifies the weighted speech signal 256 to better match the estimated pitch tracks, thus fitting the coded model more accurately, while perceptually indistinguishable reproduced speech. Bring. When the encoder processing circuit selects the pitch preprocessing mode, the pitch preprocessor 254 performs pitch preprocessing of the weighted audio signal 256. The pitch preprocessor 254 distorts the weighted audio signal 256 to match the interpolated pitch values that would be generated by the decoder processing circuitry. When pitch preprocessing is applied, the distorted audio signal is referred to as the modified and weighted audio signal 258. If no pitch preprocessing mode is selected, the weighted audio signal 256 passes through the pitch preprocessor 254 without pitch preprocessing (also referred to as "modified and weighted audio signal" 258 for convenience). The pitch preprocessor 254 may include a waveform interpolator, the functionality and implementation of which is known to those of skill in the art. The waveform interpolator can modify a certain irregular transition segment by using a known forward / reverse waveform interpolation technique, thereby increasing the regularity of the audio signal and suppressing the irregularity. The pitch gain and pitch correlation for the weighted signal 256 are estimated by the pitch preprocessor 254. The open loop pitch estimator 272 extracts information about the pitch characteristics from the weighted speech 256. Pitch information
[0121] The pitch preprocessor 254 further interacts with the classifier 270 through the open loop pitch estimator 272 to further refine the classification of the audio signal by the classifier 270. Since the pitch preprocessor 254 obtains additional information about the voice information, the classifier 270 can use this additional information to fine-tune the classification of the voice signal. After performing the pitch preprocessing, the pitch preprocessor 254 outputs the pitch track information 284 and the unquantized pitch gain 286 to the mode-dependent subframe processing portion 254 of the audio encoder.
Once the classifier 270 classifies the preprocessed voice 200 into one of a plurality of possible classes, the classification number of the preprocessed voice signal 200 is the mode selector 274 and the mode dependent subframe. It is sent to the processor 250 as control information 280. The mode selector 274 selects the operation mode using the classification number. In this particular embodiment, the classifier 270 classifies the preprocessed audio signal 200 into one of six possible classes. If the preprocessed speech signal 200 is a steady, voiced speech (eg, called "periodic" speech), the mode selector 274 sets mode 282 to mode 1. Otherwise, mode selector 274 sets mode 282 to mode 0. The mode signal 282 is sent to the mode-dependent subframe processing portion 250 of the audio encoder. Mode information 282 is added to the bitstream sent to the decoder.
The naming of audio as "periodic" and "aperiodic" should be interpreted with some caution in this particular embodiment. For example, a frame encoded using mode 1 is a frame that maintains high pitch correlation and high pitch gain throughout the frame, based on pitch track 284 derived from only 7 bits per frame. Therefore, the choice of mode 0 instead of mode 1 may be due to an inaccurate representation of pitch track 284 with only 7 bits, not necessarily due to lack of periodicity. .. Therefore, a signal encoded using mode 0 may contain periodicity, if not well represented by only 7 bits per frame for the pitch track. Therefore, mode 0 encodes the pitch track at twice the 7 bits per frame, that is, a total of 14 bits per frame, in order to better represent the pitch track.
Each of the functional blocks of FIGS. 3-4, and the other figures herein, need not have a separate structure and may be combined with one or more additional functional blocks if desired. ..
The mode-dependent subframe processing portion 250 of the audio encoder operates in two modes, mode 0 and mode 1. FIGS. 5 to 6 show a functional block diagram of mode 0 subframe processing, and FIG. 7 shows a functional block diagram of mode 1 subframe processing of the third stage of the voice encoder. FIG. 8 shows a block diagram of a voice decoder corresponding to an improved voice encoder. The audio decoder performs an inverse mapping of the bitstream to the algorithm parameters, followed by mode-dependent synthesis. A more detailed description of these numbers and modes is a U.S. patent application transferred to the same assignee, Conexant Systems, Inc., ie, Conexant Case No. 99RSS312, filed May 19, 2000, "New Voice Gain. It is described in US Patent Application No. 09 / 574,396 entitled "Quantization Policy", the entire application of which is incorporated herein by reference.
[0126] The quantized parameter representing the audio signal is packetized and transmitted as a data packet from the encoder to the decoder. In the following embodiments, the audio signal is analyzed on a frame-by-frame basis, each frame having at least one subframe, and each data packet may contain information about one frame. Therefore, in this example, the parameter information for each frame is transmitted as an information packet. In other words, there is one packet for each frame. Of course, other variants are possible, and depending on the embodiment, each packet may represent part of a frame, more than an audio frame, or multiple frames.
【0127】<u style="single">LSF</u>LSF (Line Spectral Frequency) is a representation of the LPC spectrum (ie, the short-term envelope of the audio spectrum). LSF can be thought of as a particular frequency at which the audio spectrum is sampled. For example, if the system uses a 10th order LPC, there will be 10 LSFs per frame. Minimal spacing should be provided between successive LSFs to prevent them from resulting in semi-unstable filters. For example, f<sub>i</sub>If is the i-th LSF and is equal to 100 Hz, then the (i + 1) th LSF or f<sub>I + 1</sub>Is at least f<sub>i</sub>+ Must be the minimum interval. For example, f<sub>i</sub>If = 100Hz and the minimum interval is 60Hz, f<sub>I + 1</sub>Must be at least 160Hz and can be any frequency above 160Hz. The minimum spacing is a fixed number that does not change from frame to frame, and is also known to both encoders and decoders, which allows both to work together.
It is assumed that the encoder uses the predictive coding required to achieve voice communication at a low bit rate (rather than the non-predictive coding) to code the LSF. In other words, the encoder uses the quantized LSF of the previous frame to predict the LSF of the current frame. The error between the true LSF of the current frame that the encoder derives from the LPC spectrum and the predicted LSF is quantized and sent to the decoder. The decoder finds the predicted LSF of the current frame in the same way as the encoder. The decoder can then calculate the true LSF of the current frame by knowing the error transmitted by the encoder. But what if the frame containing the LSF information is lost? Suppose that the encoder sends frames 0-3 and the decoder receives only frames 0, 2 and 3 with reference to Figure 9. Frame 1 is a lost or "erased" frame. If the current frame is lost frame 1, the decoder does not have the error information needed to calculate the true LSF. As a result, prior art systems do not calculate the true LSF, but instead set the LSF to the LSF of the previous frame, or the average LSF of a number of previous frames. The problem with this technique is that the LSF of the current frame is too inaccurate (compared to the true LSF), and subsequent frames (ie frames 2 and 3 in the example in Figure 9) frame to find their own LSF. There is a risk of using an inaccurate LSF of 1. Therefore, the LSF extrapolation error caused by the loss of the frame impairs the accuracy of the LSF of the subsequent frame.
[0129] In an embodiment of the invention, the improved audio decoder includes a counter that counts the number of good frames following the lost frame. Figure 10 illustrates the minimum LSF spacing associated with each frame. Suppose good frame 0 is received by the decoder and frame 1 is lost. In the prior art method, the minimum interval between LSFs was a fixed number (60 Hz in Figure 10) that did not change. In contrast, when an improved audio decoder notices a lost frame, the decoder avoids introducing a semi-unstable filter by increasing the minimum spacing of this frame. The amount of increase in this "controlled adaptive LSF interval" depends on which interval increase is best for that particular case. For example, an improved audio decoder takes into account how the energy of a signal (or the power of a signal) evolves over time, and how the frequency content (spectrum) of the signal evolves over time. However, by further considering the counter, it is possible to determine the value at which the minimum interval of lost frames should be set. One of ordinary skill in the art will be able to perform simple experiments to determine which minimum interval value is sufficient for use. One advantage of analyzing the audio signal and / or its parameters to derive a suitable LSF is that the resulting LSF will be closer to the true (but lost) LSF of this frame. ..
【0130】<u style="single">Adaptive codebook excitation (pitch lag)</u>Total excitation consisting of adaptive codebook excitation and fixed codebook excitation e<sub>T</sub>Is described by the following formula.
[0131] [Number 27]<img file="JP4137634B2_D0027.tif" />[0132] Here g<sub>p</sub>And g<sub>c</sub>Are the quantized adaptive and fixed codebook gains, respectively, and e<sub>xp</sub>And e<sub>xc</sub>Is an adaptive codebook excitation and a fixed codebook excitation. The buffer (also known as the adaptive codebook buffer) is an e from the preceding frame.<sub>T</sub>And its components are retained. Based on the pitch lag parameters of the current frame, the voice communication system e from the buffer<sub>T</sub>Select and make this e about the current frame<sub>xp</sub>Used as. g<sub>p</sub>, G<sub>c</sub>And e<sub>xc</sub>The value for is obtained from the current frame. Then e<sub>xp</sub>, G<sub>p</sub>, G<sub>c</sub>And e<sub>xc</sub>In the formula about the current frame e<sub>T</sub>Is calculated. E calculated for the current frame<sub>T</sub>And its components are stored in the buffer. Repeat this process, then buffered e<sub>T</sub>The next frame about e<sub>xp</sub>Used as. Thus, the feedback nature of this encoding technique (which is repeated by the decoder) is clear. The information in the equation is quantized so that the encoder and decoder are synchronized. Note that the buffer is a type of adaptive codebook (although it is different from the adaptive codebook used for gain excitation).
FIG. 11 illustrates pitch lag information for four frames 1-4 transmitted by a prior art voice system. Prior art encoders transmit the pitch lag and delta value for the current frame, where the delta value is the difference between the pitch lag of the current frame and the pitch lag of the previous frame. The EVRC (Enhanced Variable Rate Codec) standard specifies the use of data pitch lag. So, for example, the information packet for frame 1 would include pitch lag L1 and delta (L1-L0), where L0 is the pitch lag of the preceding frame 0, and the information packet for frame 2 would include pitch lag L2 and delta (L1-L0). L2-L1) will be included, and the information packet for frame 3 will include pitch lag L3 and delta (L3-L2), and so on. Note that the pitch lags of adjacent frames are equal, so the delta value may be 0. If frame 2 is lost and not received by the decoder, the only information available about pitch lag at frame 2 is pitch lag L1 because the previous frame 1 is not lost. Loss of pitch lag L2 and delta (L2-L1) information caused two problems. The first question is how to estimate the exact pitch lag L2 for the lost frame 2. The second issue is how to prevent an error in estimating the pitch lag L2 from causing an error in subsequent frames. Some prior art systems do not address either problem.
[0134] In an attempt to solve the first problem, some prior art systems use the good pitch lag L1 from the front frame 1 as the estimated pitch lag L2'for the lost frame 2, but with the estimated pitch lag L2'and true. Any difference from the pitch lag L2 will result in an error.
The second question is how to prevent an error at the estimated pitch lag L2'causing an error in subsequent frames. Recall that, as already discussed, the pitch lag of frame n is used to update the adaptive codebook buffer, which in turn is used by subsequent frames. An error between the estimated pitch lag L2'and the true pitch lag L2 will cause an error in the adaptive codebook buffer, which in turn will cause an error in later received frames. In other words, the error at the estimated pitch lag L2'may result in loss of simultaneity between the adaptive codebook buffer from the encoder's point of view and the adaptive codebook buffer from the decoder's point of view. As a further example, the prior art decoder uses pitch lag L1 (which is probably different from true pitch lag L2) as the estimated pitch lag L2'during processing of the current lost frame 2 to e for frame 2.<sub>xp</sub>Will be regained. Therefore, the wrong e in frame 2 due to the use of the wrong pitch lag<sub>xp</sub>Is selected and this error propagates throughout subsequent frames. To solve this prior art problem, when frame 3 was received by the decoder, the decoder now has pitch lag L3 and delta (L3-L2), thus deciding what the true pitch lag L2 should have been. Can be calculated backwards. The true pitch lag L2 is simply the pitch lag L3 minus the delta (L3-L2). Prior art decoders may thus be able to correct the adaptive codebook buffer used by frame 3. It is no longer too late to correct the lost frame 2 because the lost frame 2 has already been processed with the estimated pitch lag L2'.
FIG. 12 shows a hypothetical example of a frame to show the behavior of an embodiment of an improved voice communication system that addresses both problems due to loss of pitch lag information. Suppose frame 2 is lost and frames 0, 1, 3 and 4 are received. The improved decoder can use the pitch lag L1 from the previous frame 1 while the decoder handles the lost frame 2. Alternatively or preferably, the improved decoder can extrapolate based on the pitch lag of the previous frame to obtain the estimated pitch lag L2', which results in a more accurate estimate than the pitch lag L1. Thus, for example, the decoder can use pitch lags L0 and L1 to extrapolate the estimated pitch lag L2'. The extrapolation method may be any extrapolation method. For example, in order to estimate the lost pitch lag L2, a method of fitting a curve assuming a pitch contour smoothed from the past, or a method of using the average of the past pitch lags. , Or any other extrapolation method. This technique reduces the number of bits transmitted from the encoder to the decoder because it is not necessary to transmit the delta value.
To solve the second problem, when the improved decoder receives frame 3, the decoder has the correct pitch lag L3. However, as mentioned above, the adaptive codebook buffer used by frame 3 may be incorrect due to extrapolation errors in estimating the pitch lag L2'. The improved decoder attempts to prevent errors in estimating the pitch lag L2'of frame 2 from affecting frames after frame 2 without transmitting delta pitch lag information. Once the pitch lag L3 is obtained, the improved decoder adjusts or fine-tunes the previous estimation of pitch lag L2'using interpolation methods such as curve fitting. Since the pitch lags L1 and L3 are known, the curve fitting method can estimate L2'more accurately than if the pitch lag L3 was not known. The result is a fine-tuned pitch lag L2 , which is used to adjust or correct the adaptive codebook buffer for use by frame 3. More specifically, the fine-tuned pitch lag L2 is used to adjust or correct the quantized adaptive codebook excitation in the adaptive codebook buffer. Thus, the improved decoder reduces the number of bits to be transmitted and, in most cases, fine-tunes the pitch lag L2'in a satisfactory manner. Thus, in order to reduce the effect of any error in estimating pitch lag L2 on later received frames, the improved decoder combines pitch lag L3 in next frame 3 with pitch lag L1 in previously received frame 1. Use to fine-tune the previous estimation for pitch lag L2, assuming smoothed pitch contours. The accuracy of this estimation method, which is based on the pitch lag of the received frames preceding and following the lost frame, can be quite good because the pitch contours are generally smooth for voiced speech.
【0138】<u style="single">gain</u>Adaptive codebook gain g as a result of frame loss while transmitting frames from the encoder to the decoder<sub>p</sub>And fixed codebook gain g<sub>c</sub>Gain parameters such as are also lost. Each frame contains a plurality of subframes, and each subframe has gain information. Therefore, as a result of the loss of the frame, the gain information in each subframe of the frame is also lost. The voice communication system needs to estimate the gain information for each subframe of the lost frame. The gain information in one subframe may differ from the gain information in another subframe.
Prior art systems use a variety of techniques to estimate the gain for a lost frame subframe, such as using the gain from the last subframe of a good previous frame as the gain for each subframe of the lost frame. Was taken. In another variant, the gain from the last subframe of the good previous frame is used as the gain of the first subframe of the lost frame, this gain is gradually attenuated, and then this is used as the gain of the next subframe of the lost frame. Used as. In other words, for example, if each frame has four subframes, frame 1 is received and frame 2 is lost, the gain parameter in the last subframe of received frame 1 is the first of the lost frames 2. Used as the gain parameter of the subframe, then decremented by a certain amount and used as the gain parameter of the second subframe of the lost frame 2, and then decremented again to the third sub of the lost frame 2. It is used as the gain parameter of the frame, and the gain parameter is further reduced and used as the gain parameter of the last subframe of the lost frame 2. Yet another approach is to look at the gain parameters of the previously received fixed number of frames subframes to calculate the average gain parameter, which is then used as the gain parameter of the first subframe of the lost frame 2. Here, the gain parameter can be gradually reduced and used as the gain parameter of the remaining subframes of the lost frame. Yet another approach derives the intermediate gain parameter by examining the previously received fixed number of frame subframes and using the intermediate value as the gain parameter of the first subframe of the lost frame 2, where The gain parameter can be gradually reduced and used as the gain parameter for the remaining subframes of the lost frame. Notably, the prior art method did not use different recovery methods for adaptive and fixed codebook gains, but used the same recovery method for both types of gain.
The improved voice communication system can also handle gain parameters that are lost due to the loss of the frame. If the voice communication system differentiates periodic-like voice from non-periodic-like voice, the system can handle lost gain parameters differently for each type of voice. In addition, the improved system treats lost adaptive codebook gains differently from lost fixed codebook gains. First, consider the case of aperiodic-like voice. Estimated adaptive codebook gain g<sub>p</sub>To find the average g of subframes of the adaptive number of frames previously received, the improved decoder<sub>p</sub>To calculate. The pitch lag of the current frame (ie the lost frame) estimated by the decoder is used to determine the number of previously received frames to examine. Generally, the larger the pitch lag, the more the average g<sub>p</sub>The number of frames received before it should be used to calculate is large. Thus, the improved decoder uses the pitch-synchronized averaging technique to adapt the chordbook gain g for aperiodic-like speech.<sub>p</sub>To estimate. The improved decoder then calculates beta β based on the following equation, which is g<sub>p</sub>Shows how good the prediction is.
[0141] [Number 28]<img file="JP4137634B2_D0028.tif" />[0142] β changes from 0 to 1, and the effect of adaptive codebook excitation energy on total excitation energy is expressed as a percentage. The larger β, the greater the effect of adaptive codebook excitation energy. It is preferable that the improved decoder treats aperiodic-like speech and periodic-like speech differently, but this is not essential.
FIG. 16 illustrates a flow chart of decoder processing for aperiodic-like speech. Step 1000 determines if the current frame is the first frame lost after receiving a frame (ie, a "good" frame). If the current frame is the first lost frame after a good frame, step 1002 determines if the current subframe being processed by the decoder is the first subframe of the frame. If the current subframe is the first subframe, step 1004 is the average g for a number of previous subframes.<sub>p</sub>Where the number of subframes depends on the pitch lag of the current subframe. In an exemplary example, if the pitch lag is 40 or less, the average g<sub>p</sub>Is based on two previous subframes. If the pitch lag is greater than 40 and less than 80, then the average g<sub>p</sub>Is based on four pre-subframes. If the pitch lag is greater than 80 and less than 120, then the average g<sub>p</sub>Is based on 6 pre-subframes. If the pitch lag is greater than 120, then the average g<sub>p</sub>Is based on eight previous subframes. Of course, these values are arbitrary and may be set to any other value depending on the length of the subframe. Step 1006 determines if the maximum value β exceeds a certain threshold. If the maximum value β exceeds a certain threshold, step 1008 is a fixed codebook gain g for all subframes of the lost frame.<sub>c</sub>Set to zero and g for all lost frame subframes<sub>p</sub>, The average g obtained above<sub>p</sub>Instead of, set it to an arbitrarily large number, such as 0.95. This arbitrarily large number indicates a good vocal signal. G of the current subframe of the lost frame<sub>p</sub>Any large number for which is set can be based on several factors, including the maximum value β of a certain number of previous frames, the spectral gradient of the previously received frame, and the previously received frame. It includes, but is not limited to, energy.
Conversely, if the maximum value β does not exceed a certain threshold (ie, the previously received frame contains the start of audio), step 1010 is g of the current subframe of the lost frame.<sub>p</sub>(I) The average g obtained on<sub>p</sub>, And (ii) a number of arbitrarily selected sizes (eg 0.95), set to a minimum. Instead, g of the current subframe of the lost frame<sub>p</sub>The spectral gradient of the previously received frame, the energy of the previously received frame, and the mean g obtained above.<sub>p</sub>It can also be set based on a minimum of and a number of arbitrarily selected sizes (eg 0.95). Fixed codebook gain g if maximum β does not exceed a certain threshold<sub>c</sub>Is based on the gain scaled fixed codebook excitation energy in the previous subframe and the fixed codebook excitation energy in the current subframe. Specifically, the energy of the gain scaling fixed codebook excitation in the previous subframe is divided by the energy of the fixed codebook excitation in the current subframe, and the result is multiplied by the decay fraction to find its square root. G shown in the following formula<sub>c</sub>Set to.
[0145] [Number 29]<img file="JP4137634B2_D0029.tif" />[0146] Instead, the decoder g for the current subframe of the lost frame based on the ratio of the energy of the previously received frame to the energy of the current lost frame.<sub>c</sub>Can be derived.
Returning to step 1002, if the current subframe is not the first subframe, step 1020 is g of the current subframe of the lost frame.<sub>p</sub>, The g of the previous subframe<sub>p</sub>Set to a value that is attenuated or reduced from. Each g of the remaining subframes<sub>p</sub>Is the g of the previous subframe<sub>p</sub>Is set to a further attenuated value from. Current subframe g<sub>c</sub>Is calculated in the same way as in step 1010 and equation 29.
Returning to step 1000, if the current frame is not the first lost frame after a good frame, step 1022 g in the current subframe in the same way as in steps 1010 and equation 29.<sub>c</sub>Is calculated. Step 1022 also g of the current subframe of the lost frame<sub>p</sub>, The g of the previous subframe<sub>p</sub>Set to attenuated and reduced values from. Decoder is g<sub>p</sub>And g<sub>c</sub>Because they estimate differently, the decoder can estimate these more accurately than the prior art system.
[0149] Next, the case of periodicity-like voice will be examined according to the flowchart illustrated in FIG. The decoder g for periodic-like and aperiodic-like speech<sub>p</sub>And g<sub>c</sub>Gain parameter estimates will be more accurate than prior art methods, as different methods can be applied to estimate. Step 1030 determines if the current frame is the first frame lost after receiving a frame (ie, a "good" frame). If the current frame is the first lost frame after a good frame, step 1032 is g<sub>c</sub>Set to zero for all subframes of the current frame, g<sub>p</sub>Is set to an arbitrarily large number, such as 0.95, for all subframes of the current frame. If the current frame is not the first lost frame after a good frame (eg second lost frame, third lost frame, etc.), step 1034 is g<sub>c</sub>Set to zero for all subframes of the current frame, g<sub>p</sub>, The g of the previous subframe<sub>p</sub>Set to a value attenuated from.
[0150] FIG. 13 shows an example of a frame for exemplifying the operation of an improved audio decoder. Suppose frames 1, 3 and 4 are good (ie received) frames and frames 2, 5-8 are lost frames. If the current lost frame is the first lost frame after a good frame, the decoder g<sub>p</sub>Set to an arbitrarily large number (eg 0.95) for all subframes of lost frames. With reference to Figure 13, this applies to lost frames 2 and 5. 1st lost frame 5 g<sub>p</sub>Is gradually attenuated, g of other lost frames 6-8<sub>P</sub>To set. Thus, for example, g<sub>p</sub>When is set to 0.95 with lost frame 5, g<sub>p</sub>Can be set to 0.9 for lost frame 6, 0.85 for lost frame 7, and 0.8 for lost frame 8. g<sub>c</sub>For, the decoder averages g from previously received frames<sub>p</sub>Calculate and this average g<sub>p</sub>If exceeds a certain threshold, g<sub>c</sub>Is set to zero for all subframes of lost frames. Average g<sub>p</sub>If does not exceed a certain threshold, the decoder uses the same configuration technique for aperiodic-like signals as described above.<sub>c</sub>To set.
[0151] After the decoder estimates the lost parameters in the lost frame (eg LSF, pitch lag, gain, classification, etc.) and synthesizes the resulting voice, the decoder uses extrapolation technology to synthesize the lost frame voice. The energy can be matched with the energy of the previously received frame. This further improves the accuracy of reproducing the original sound even if the frame is lost.
【0152】<u style="single">Seed for generating fixed codebook excitation</u>To save bandwidth, the voice encoder does not have to send a fixed codebook excitation to the decoder during periods of background noise or silence. Instead, both encoders and decoders can use Gauss time series generators to randomly generate local excitation values. Both the encoder and the decoder are configured to generate the same random excitation values on the same order. As a result, it is not necessary to send the excitation values from the encoder to the decoder because the decoder can locally generate the same random excitation values that the encoder generated for a given noise frame. To generate a random excitation value, the Gauss time series generator uses the initial seed to generate a first random excitation value, and then the generator updates the seed to the new value. The generator then uses the updated seed to generate the next random excitation value, updating the seed to yet another value. Figure 14 shows how the Gauss time series generator in the voice encoder uses a seed to generate a random excitation value, and then updates this seed to generate the next random excitation value. Here is a hypothetical example of a frame to illustrate. Suppose frames 0 and 4 contain audio signals and frames 2, 3 and 5 contain silence or background noise. When the first noise frame (ie, frame 2) is found, the encoder uses the initial seed (called "seed 1") to generate a random excitation value for use as a fixed codebook excitation for this frame. For each sample in this frame, change the seed to generate a new fixed codebook excitation. Thus, if the frame was sampled 160 times, the seed would change 160 times. Therefore, by the time the next noise frame (noise frame 3) is encountered, the encoder uses a second and a different seed (ie, seed 2) to generate a random excitation value for this frame. Technically, the seed changes with each sample in the first frame, so the first support in the second frame The seed for the sample is not the "second" seed, but for convenience the seed for the first sample in the second frame is referred to here as seed 2. For noise frame 4, the encoder uses a third seed (different from the first and second seeds). To generate a random excitation value for noise frame 6, the Gauss time series generator may start over from seed 1 or proceed with seed 4, depending on the implementation of the voice communication system. By configuring the encoders and decoders to update the seeds in the same way, the encoders and decoders can generate the same seeds and thus the same random excitation values on the same order. However, in the prior art voice communication system, the loss of the frame destroys this simultaneity between the encoder and the decoder.
[0153] FIG. 15 illustrates the hypothetical case shown in FIG. 14 from the perspective of a decoder. Suppose that noise frame 2 is lost and frames 1 and 3 are received by the decoder. Since the noise frame 2 is missing, the decoder assumes it is of the same type as the previous frame 1 (ie the audio frame). Having made a false assumption about the lost noise frame 2, the decoder considers it the first noise frame, even though the noise frame 3 is actually the second noise frame encountered. Since the seed is updated for each sample of every noise frame encountered, the decoder incorrectly uses seed 1 to generate a random excitation value for noise frame 3, even though seed 2 should be used. Thus, the encoder and decoder lose simultaneity as a result of frame loss. Since frame 2 is a noise frame, it is not important for the decoder to use seed 1 while the encoder uses seed 2, because the result is noise that is different from the original noise. The same applies to frame 3. However, if the later received frame contains audio, the seed value error will have a significant effect on this. For example, pay attention to the voice frame 4. Update the adaptive codebook buffer in frame 3 with continuous use of locally generated Gaussian excitation based on seed 2. When frame 4 is processed, the adaptive codebook excitation is extracted from the adaptive codebook buffer of frame 3 based on information such as the pitch lag of frame 4. In some cases frames because the encoder uses seed 3 to update the adaptive codebook buffer in frame 3 and the decoder uses seed 2 (wrong seed) to update the adaptive codebook buffer in frame 3. Differences in updating the adaptive codebook buffer in 3 can cause quality problems within frame 4.
An improved voice communication system constructed in accordance with the present invention uses an initial fixed seed and does not update this seed each time the system encounters a noise frame. Instead, improved encoders and decoders derive seeds for a given frame from the parameters within this frame. For example, spectral information, energy and / or gain information within the current frame can be used to generate seeds for this frame. For example, with the bits representing the spectrum (eg 5 bits b1, b2, b3, b4, b5) and the bits representing the energy (eg 3 bits c1, c2, c3), the strings b1, b2, b3, b4, It can yield b5, c1, c2, c3, and this value is the seed. To give an example in terms of numbers, the seed is represented by 01101011, assuming that the spectrum is represented by 01101 and the energy is represented by 011. Of course, other alternative methods of deriving the seed from the information in the frame are possible and are within the scope of the invention. Thus, in the example of FIG. 15 where noise frame 2 is lost, the decoder can derive a seed for noise frame 3, which is the same seed derived by the encoder. Therefore, the loss of the frame does not destroy the simultaneity of the encoder and the decoder.
Although examples and implementations of the present invention have been shown and described, it is clear that more examples and implementations are within the scope of the invention. Therefore, the present invention should not be limited except that it is limited to the claims and their equivalents.
BRIEF DESCRIPTION OF THE DRAWINGS FIG. 1 is a functional block diagram of a voice communication system having a source encoder and a source decoder.
FIG. 2 is a more detailed functional block diagram of the voice communication system of FIG.
FIG. 3 is a functional block diagram of a voice preprocessor, an exemplary first stage of a source encoder used in an embodiment of the voice communication system of FIG.
FIG. 4 is a functional block diagram showing an exemplary second stage of a source encoder used in an embodiment of the voice communication system of FIG.
FIG. 5 is a functional block diagram showing an exemplary third stage of a source encoder used in an embodiment of the voice communication system of FIG.
FIG. 6 is a functional block diagram showing an exemplary fourth stage of a source encoder used by an embodiment of the voice communication system of FIG. 1 to process aperiodic voice (mode 0).
FIG. 7 is a functional block diagram showing an exemplary fourth stage of a source encoder used by an embodiment of the voice communication system of FIG. 1 to process periodic voice (mode 1).
FIG. 8 is a block diagram of an embodiment of a voice decoder for processing encoded information from a voice encoder constructed according to the present invention.
FIG. 9 is a diagram showing a hypothetical example of a received frame and a lost frame.
FIG. 10 shows a hypothetical example of the minimum spacing between LSFs assigned to each frame in a received frame and a lost frame, as well as in prior art systems and voice communication systems constructed in accordance with the present invention.
FIG. 11 is a diagram illustrating a hypothetical example of how a prior art voice communication system allocates and uses pitch lag and delta pitch lag information for each frame.
FIG. 12 is a diagram illustrating a hypothetical example illustrating how pitch lag and delta pitch lag information is allocated and used for each frame by a voice communication system constructed in accordance with the present invention.
FIG. 13 illustrates a hypothetical example of how an audio decoder constructed in accordance with the present invention allocates adaptive gain parameter information to each frame when there are lost frames.
FIG. 14 illustrates a hypothetical example of how a prior art encoder uses seeds to generate random excitation values for each frame containing silence or background noise.
FIG. 15 How a prior art decoder uses seeds to generate random excitation values for each frame containing silence or background noise in the presence of lost frames, resulting in loss of simultaneity with the encoder. It is a figure which shows the hypothetical example illustrated.
FIG. 16 is a flowchart showing an example of processing aperiodic-like voice according to the present invention.
FIG. 17 is a flowchart showing an example of processing periodic-like voice according to the present invention.
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| JP09120298A | Cites | Japan |
| JP09120299A | Cites | Japan |
| WO97048205A1 | Cites | World Intellectual Property Organization (WIPO) |
| WO99001863A1 | Cites | World Intellectual Property Organization (WIPO) |
| WO00011650A1 | Cites | World Intellectual Property Organization (WIPO) |
93 members in 13 offices
Priority claims9
| Document | Office | Kind | Date |
|---|---|---|---|
| 09617191 | United States of America | – | |
| 61719100 | United States of America | A | |
| 61719100 | United States of America | A | |
| 0101228 | International Bureau of the World Intellectual Property Organization (WIPO) | W | |
| 0101228 | International Bureau of the World Intellectual Property Organization (WIPO) | W | |
| 2000617191 | – | – | – |
| 2001001228 | – | – | – |
| US20000617191 | – | – | – |
| WO2001IB01228 | – | – | – |
Members93
| Document | Office | Kind | |
|---|---|---|---|
| WO0122402A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU7486200A | Australia | A | |
| WO0191112A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU5542201A | Australia | A | |
| WO0207061A2 | World Intellectual Property Organization (WIPO) | A2 | |
| AU6627801A | Australia | A | |
| KR20020033819A | Republic of Korea | A | |
| EP1214706A1 | European Patent Office (EPO) | A1 | |
| TW493161B | Taiwan Province of China | B | |
| WO0207061A3 | World Intellectual Property Organization (WIPO) | A3 | |
| KR20030001523A | Republic of Korea | A | |
| JP2003513296A | Japan | A | |
| EP1301891A2 | European Patent Office (EPO) | A2 | |
| KR20030040358A | Republic of Korea | A | |
| US6574593B1 | United States of America | B1 | |
| BR0014212A | Brazil | A | |
| US6581032B1 | United States of America | B1 | |
| US6604070B1 | United States of America | B1 | |
| EP1338003A1 | European Patent Office (EPO) | A1 | |
| CN1441950A | China | A | |
| US6636829B1 | United States of America | B1 | |
| CN1451155A | China | A | |
| US2003200092A1 | United States of America | A1 | |
| EP1363273A1 | European Patent Office (EPO) | A1 | |
| CN1468427A | China | A | |
| KR20040005970A | Republic of Korea | A | |
| JP2004504637A | Japan | A | |
| JP2004510174A | Japan | A | |
| US6735567B2 | United States of America | B2 | |
| US6757649B1 | United States of America | B1 | |
| JP2004206132A | Japan | A | |
| CN1516113A | China | A | |
| EP1214706B1 | European Patent Office (EPO) | B1 | |
| AT272885T | Austria | T | |
| ATE272885T1 | Austria | T1 | |
| US6782360B1 | United States of America | B1 | |
| DE60012760D1 | Germany | D1 | |
| AU2001255422B2 | Australia | B2 | |
| BR0110831A | Brazil | A | |
| US2004260545A1 | United States of America | A1 | |
| EP1214706B9 | European Patent Office (EPO) | B9 | |
| KR100488080B1 | Republic of Korea | B1 | |
| KR20050061615A | Republic of Korea | A | |
| CN1212606C | China | C | |
| RU2257556C2 | Russian Federation | C2 | |
| DE60012760T2 | Germany | T2 | |
| EP1577881A2 | European Patent Office (EPO) | A2 | |
| EP1577881A3 | European Patent Office (EPO) | A3 | |
| RU2262748C2 | Russian Federation | C2 | |
| US6959274B1 | United States of America | B1 | |
| US6961698B1 | United States of America | B1 | |
| JP2005338872A | Japan | A | |
| JP2006011464A | Japan | A | |
| CN1722231A | China | A | |
| KR100546444B1 | Republic of Korea | B1 | |
| EP1301891B1 | European Patent Office (EPO) | B1 | |
| AT317571T | Austria | T | |
| ATE317571T1 | Austria | T1 | |
| CN1245706C | China | C | |
| CN1252681C | China | C | |
| DE60117144D1 | Germany | D1 | |
| US7054809B1 | United States of America | B1 | |
| CN1267891C | China | C | |
| EP1338003B1 | European Patent Office (EPO) | B1 | |
| DE60117144T2 | Germany | T2 | |
| AT343199T | Austria | T | |
| ATE343199T1 | Austria | T1 | |
| DE60123999D1 | Germany | D1 | |
| US7191122B1 | United States of America | B1 | |
| US2007136052A1 | United States of America | A1 | |
| KR100742443B1 | Republic of Korea | B1 | |
| US7260522B2 | United States of America | B2 | |
| KR100754085B1 | Republic of Korea | B1 | |
| US2007255559A1 | United States of America | A1 | |
| JP4137634B2This record | Japan | B2 | |
| JP4176349B2 | Japan | B2 | |
| JP4222951B2 | Japan | B2 | |
| US2009043574A1 | United States of America | A1 | |
| EP1363273B1 | European Patent Office (EPO) | B1 | |
| AT427546T | Austria | T | |
| ATE427546T1 | Austria | T1 | |
| DE60138226D1 | Germany | D1 | |
| US2009177464A1 | United States of America | A1 | |
| EP2093756A1 | European Patent Office (EPO) | A1 | |
| ES2325151T3 | Spain | T3 | |
| US7593852B2 | United States of America | B2 | |
| US7660712B2 | United States of America | B2 | |
| EP2093756B1 | European Patent Office (EPO) | B1 | |
| US8620649B2 | United States of America | B2 | |
| US2014119572A1 | United States of America | A1 | |
| BRPI0014212B1 | Brazil | B1 | |
| US10181327B2 | United States of America | B2 | |
| US10204628B2 | United States of America | B2 |
39 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of completion of termEXPY | EXPY | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Written notification of registration of transferJAPANESE INTERMEDIATE CODE: R350R350 | R350 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Request for change of ownership or part of ownershipJAPANESE INTERMEDIATE CODE: R313113S111 | S111 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Written notification of registration of transferJAPANESE INTERMEDIATE CODE: R350R350 | R350 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Written notification for declining of transfer of rightsJAPANESE INTERMEDIATE CODE: R360R360 | R360 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Request for change of ownership or part of ownershipJAPANESE INTERMEDIATE CODE: R313113S111 | S111 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Written permission of extension of timeJAPANESE INTERMEDIATE CODE: A602A602 | A602 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Notification of revocation by ex officioJAPANESE INTERMEDIATE CODE: A971091AA91 | AA91 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 |
Numbers
- Publication
- 4137634
- Publication, DOCDB
- 4137634
- Publication, EPODOC
- JP4137634B
- Application
- 2002512896
- Application, DOCDB
- 2002512896
- Application, EPODOC
- JP20020512896
Titles2
- Japanese
- 紛失フレームを取扱うための音声通信システムおよび方法
- English
- Voice communication systems and methods for handling lost frames
Classification
- CPC, 7
- G10L19/005
- G10L19/08
- G10L19/07
- G10L19/083
- G10L25/90
- G10L2019/0012
- G10L19/04
- IPC, 10
- G10L19 00
- G10L19 06
- G10L13 00
- G10L19 005
- G10L19 04
- H03M7 30
- H03M7 36
- H04B14 04
- H04L1 00
- H04M1 00