Vector quantization method and speech encoding method and apparatus
Abstract
ABSTRACT The processing volume for codebook search for vector quantization is to be diminished. In sending data representing an envelope of spectral components of the harmonics from a spectrum evaluation unit 148 of a sinusoidal analytic encoder 114 to a vector quantizer 116 for vector quantization, the degree of similarity between an input vector and all code vectors stored in the codebook is found by approximation for pre-selecting a smaller plural number of code vectors. From these plural pre-selected code vectors, such a code vector minimizing an error with respect to the input vector is ultimately selected. In this manner, a smaller number of candidate code vectors are pre-selected by pre-selection involving simplified processing and subsequently subjected to ultimate selection with high precision.
Term
No projected expiry on record.
- Priority
- Filed
- Granted
- Today
13 claims: 13 independent, 0 dependent
- 1一種向量量化方法,其中一輸入向量與一碼冊內所儲存的碼向量比較,以輸出一所需碼向量之一索引,該方法包括:一預先選擇步驟,在輸入向量與該碼冊內所儲存的所有碼向量間找出近似程度,俾選擇複數具有高度近似性的碼向量;以及一終極選擇步驟,自該預先選擇步驟所選擇的碼向量中選擇中一相對於輸入向量最小化一誤差的碼向量。
- 2如申請專利範圍第1項之向量量化方法,其中該碼冊由複數碼冊組合而成,且其中碼向量自碼冊逐一選出。
- 3如申請專利範圍第1項之向量量化方法,其中使用一選擇性除以各碼向量之一範數或一加權範數而得之輸入向量與該碼向量之內積或輸入向量與該碼向量之加權內積來作為該近似程度。
- 4如申請專利範圍第1項之向量量化方法,其中該輸入向量係由來自一語音信號位於頻率軸上的參數的向量,且使用該碼向量除以一碼向量之一範數而得之一加權內積作為該近似程度,此碼向量以一權值加權,而此權值對應於一能,朝頻率軸上的參數的低頻範圍集中並朝高頻範圍減低。
- 5如申請專利範圍第1項之向量量化方法,其中使用該碼向量除以一碼向量之一範數而得之一可變加權內積作為該近似程度,此碼向量以一對應於一能量之固定權值加權,而此能量朝頻率軸上的參數的低頻範圍集中並朝高頻範圍減低。
- 6一種語音編碼方法,係一輸入語音信號以預先設定的編碼單位在時間軸上分割,並以預先設定的編碼單元編碼者,其步驟包括:一藉由對來自輸入信號的信號進行正弦分析而找出諧波的頻譜成份;一將諸參數向量量化俾予以編碼的步驟,此參數來自諧波的以編碼單位為基礎的頻譜成份,且作為輸入向量;該向量量化包括:一預先選擇步驟,在輸入向量與該碼冊內所儲存所有碼向量間找出近似程度,俾選擇複數具有高度近似性的碼向量;以及一終極選擇步驟,自該預先選擇步驟所選擇的碼向量選擇一相對於輸入向量最小化一誤差的碼向量。
- 7如申請專利範圍第6項之語音編碼方法,其中該碼冊中複數碼冊組合而成,且碼向量自碼冊逐一選出。
- 8如申請專利範圍第6項之語音編碼方法,其中使用一選擇性除以各碼向量之一範數或一加權範數之輸入向量與該碼向量之內積或輸入向量與該碼向量之一加權內積作為該近似程度。
- 9如申請專利範圍第8項之語音編碼方法,其中使用該碼向量除以一碼向量之一範數而得之一加權內積來對該範數加權,此碼向量以一對應於一能量之權值加權,而此能量朝頻率軸上的參數的低頻範圍集中並朝高頻範圍減低。
- 10一種語音編碼裝置,係一輸入語音信號於時間軸上以預先設定的編碼單位分割並以預定設定的編碼單元編碼者,包括一預測編碼裝置,用來找出輸入語音信號的短期預測剩餘;以及一正弦分析編碼裝置,用來藉由正弦分析編碼處理所找出的短期預測剩餘;該正弦分析編碼裝置具有向量量化裝置,用來藉近似法在一輸入向量與一碼冊內所儲存所有碼向量間找出近似程度,以預先選擇複數具有高度近似性之碼向量,而此輸入向量係一參數,來自以餘弦分析法所得諧波的頻譜成份;該向量量化裝置自預先選擇的碼向量中選擇一相對於輸入向量最小化一誤差之碼向量。
- 11如申請專利範圍第10項之語音編碼裝置,其中該碼冊由複數碼冊組合而成,且碼向量自碼冊逐一選出。
- 12如申請專利範圍第10項之語音編碼裝置,其中使用選擇性除以各碼向量之一範數或一加權範數的輸入向量與該碼向量的內積或輸入向量與該碼向量的加權內積作為該近似程度。
- 13如申請專利範圍第12項之語音編碼裝置,其中使用選擇性除以各碼向量之一範數或一加權範數的輸入向量與該碼向量的一內積或輸入向量與該碼向量的一加權內積來作範數之權值。
Independent claims13
383 paragraphs, as filed
The present invention relates to a vector quantization method, in which an input vector quantity is compared with a code vector stored in a codebook to output an index of the optimal code vector of one of the code vectors. The present invention also relates to a speech coding method and device, in which an input speech signal is divided into a predetermined coding unit such as a block or frame, and coding processing including vector quantization is performed on the basis of the coding unit.
A vector quantization has been known so far. In order to digitize and compress audio or video signals, multiple input data are combined into a vector and used to represent a single code (index).
In this vector quantization, the representative modes of various input vectors are determined in advance by, for example, learning and a given code or value stored in the codebook. Then, by pattern matching, the input vector is compared with the relative pattern (code vector), so that the code with the most approximate and most relevant pattern is output. This approximation and correlation are obtained by calculating the distortion value or error energy between the input vector and the relative code vector, and the higher the value, the smaller the distortion or error.
So far, a variety of coding methods are known to use statistical properties in the time domain and frequency domain under signal compression, and the psychological properties of human hearing. This kind of coding method is roughly divided into time domain coding, frequency domain coding and analysis-combined coding.
Among the high-efficiency examples of speech signals is a sine wave analysis coding, such as harmonic coding, sub-band coding (SBC), linear predictive coding (LPC), discrete cosine transform (DCT), modified DCT (MDCT) or fast four The inner leaf transform (FFT).
In the high-efficiency coding of the speech signal, the above-mentioned vector quantization is used as the parameter of the harmonic spectrum components.
At the same time, if the pattern modulus stored in the code book, that is, the number of code vectors is large, or if the vector quantizer is a multi-segment configuration composed of a complex code book, the code vector retrieval operation used for pattern matching is a multiple Increase and increase the processing volume. In particular, if multiple codebooks are grouped together, a process must be performed to find the approximateness of the multiplier of the code vector number in the codebook, thereby greatly increasing the codebook retrieval processing volume.
An object of the present invention is to provide a vector quantization method, a speech coding method and a speech coding device, which can reduce the codebook retrieval processing volume.
In order to achieve the above object, the present invention provides a vector quantization method, including an approximation method to find the similarity between an input vector to be vector quantized and all code vectors stored in a codebook to preselect a complex with a high degree of similarity. The step of digital vector; and a final step of minimizing the error relative to the input vector from the pre-selected code vectors of the complex number.
By making the final selection after the pre-selection, the few candidate code vectors selected by the pre-selection that can simplify the processing are processed by the final selection that is highly accurate and can reduce the processing volume of the codebook retrieval.
The code book is composed of a complex code book, and a complex code vector representing an optimal combination can be selected from each code book. The degree of approximation can be the inner product of the input vector and the code vector selectively divided by a norm or a weighted norm of each code vector.
The present invention also provides a speech coding method, in which an input speech signal or short-term prediction residual is analyzed by sinusoidal analysis to find out the spectral components of the harmonics, and the parameters are derived from the spectral components of the harmonics based on the coding unit. That is, the input vector is vector quantized for encoding. In vector quantization, the approximate method is used to find the degree of similarity between the input vector and all code vectors stored in the codebook, so as to select a few code vectors with high similarity in advance, and finally select the code vectors from these preselected ones. Select a code vector that minimizes the error relative to the input vector.
The degree of approximation can be a selective weighted inner product of an input vector and a code vector divided by a norm of each code vector or a weighted norm. To weight the norm, a weight can be used, with energy concentrated in the low frequency range and energy reduced in the high frequency range. In this way, the degree of approximation can be obtained by dividing the weighted inner product of the code vector by the weighted code vector norm.
The present invention also relates to a speech coding device capable of realizing the speech coding method.
The preferred embodiments of the present invention are described in detail with reference to the drawings.
Figure 1 shows the basic structure of an encoding device (encoder) for implementing the speech encoding method of the present invention.
The basic concept of the speech signal encoder shown in Figure 1 is that the encoder has a first encoding unit 110, which is used to find the remaining short-term prediction of the input speech signal such as linear prediction (LPC) encoding, so as to perform sinusoidal encoding such as harmonic encoding. Analysis and a second encoding unit 120 are used to encode the input speech signal by waveform encoding with phase reproducibility. The further point is that the first encoding unit 110 and the second encoding unit 120 are used to separately encode the audio of the input signal ( V) Coding of the unvoiced (UV) part of the voice and input signal.
The first encoding unit 110 uses a sinusoidal analysis encoding such as harmonic encoding or multi-band excitation (MBE) encoding to encode, for example, LPC residual encoding. The configuration of the second encoding unit 120 uses vector quantization to find the optimal vector through a closed loop search, and uses, for example, an analysis and synthesis method to perform Code Excited Linear Prediction (CELP).
In the embodiment shown in FIG. 1, the speech signal supplied to an input terminal 101 is sent to the LPC inverse filter 111 of a first encoding unit 110 and the LPC analysis and quantization unit 113 of the first encoding unit 110. The LPC coefficients or so-called α-parameters obtained by the LPC analysis and quantization unit 113 are sent to the LPC inverse filter 111 of the first coding unit. The linear prediction residual (LPC residual) of the input speech signal is extracted from the LPC inverse filter 111. As described later, linear spectrum pairs (LSPs) are extracted from the LPC analysis and quantization unit 113 and sent to an output terminal 102. The LPC remaining from the LPC inverse filtering is sent to a sinusoidal analysis and encoding unit 114. The sine analysis coding unit 114 performs tone detection and frequency amplitude calculation of the spectrum message envelope, and uses the V/UV frequency discrimination unit 115 to perform V/UV frequency discrimination. The frequency amplitude data of the spectrum message envelope from the sine analysis unit 114 is sent to a vector quantization unit 116. The codebook index of the vector quantization output from the vector quantization unit 116 as a spectral message envelope is sent to an output terminal 103 via a switch 117, and the output of the sine analysis coding unit 114 is sent to an output terminal 104 via a switch 118 . One of the V/UV discrimination output of the V/UV discrimination unit 115 is sent to an output terminal 105 and used as a control signal to the switches 117, 118. If the input voice signal is voiced (V), the index and pitch are respectively selected and taken out at the output terminals 103 and 104.
The second encoding unit 120 shown in Figure 1 has a code-excited linear predictive coding (CELP) configuration in this embodiment, and uses an analytical synthesis method to vector quantize the time-domain waveform through a closed-loop search. Here In the comprehensive analysis method, the output of a noise code book 121 is synthesized by a weighted synthesis filter. The generated weighted speech is sent to a subtractor 123. The error between the weighted speech and the speech signal supplied to the input terminal 101 is then taken out through a perceptual weighted filter 125, and the found error is sent to a distance calculation circuit 124 for distance calculation. Calculate and retrieve the vector that minimizes this error with the noise codebook. As mentioned above, this CELP encoding is used to encode the unvoiced speech part. The code book index used as the UV data from the noise code book 121 is taken out from the output terminal 107 via a switch 127. When the result of V/UV discrimination is silent (UV), the switch 127 is turned on.
Fig. 2 shows the basic structure of a speech signal decoder. This decoder is a corresponding device of the speech signal encoder shown in Fig. 1, and is used to implement the speech decoding method of the present invention.
Referring to Fig. 2, a codebook index is provided to an input terminal 202, and the codebook index is a quantized output of linear spectral pairs (LSPs) from the output terminal 102 shown in Fig. 1. The outputs of the output terminals 103, 104, and 105 in Figure 1 are the tones of the message quantized output data. The V/UV discriminator output and index data are respectively supplied to the output terminals 203 to 205. Index data as silent data is supplied from the output terminal 107 shown in FIG. 1 to an input terminal 207.
The quantized output index of the message packet of the input terminal 203 is sent to an inverse vector quantization unit 212 for inverse vector quantization to find the remaining spectral information packets of the LPC, and the remaining LPC is sent to the voiced speech synthesizer 211. The voiced speech synthesizer 211 synthesizes the remaining linear predictive coding (LPC) of the voiced speech part by sine synthesis. The synthesizer 211 is also fed with tones and discriminatory outputs from the input terminals 204 and 205. The remaining LPC of the voiced speech from the voiced speech synthesis unit 211 is sent to an LPC synthesis filter 214. The index data of the UV data from the input terminal is sent to a silent voice synthesis unit 220, where the noise code book is referred to to extract the LPC remaining of the silent part. The remaining LPC is also sent to the LPC synthesis filter 214. In the LPC synthesis filter 214, the LPC residue of the voiced part and the unvoiced part of the LPC residue are processed by LPC synthesis. Alternatively, LPC synthesis may be used to process the total LPC surplus of the voiced part and the LPC surplus of the unvoiced part together. The LSP index data from the input terminal 202 is sent to the LPC parameter regeneration unit 213, where the α parameter of the LPC is taken out and sent to the LPC synthesis filter. At the output terminal 201, the speech signal synthesized by the LPC synthesis filter 214 is taken out.
With reference to Fig. 3, a more detailed structure of the speech signal decoder shown in Fig. 1 will be described. In Figure 3, elements and components similar to those shown in Figure 1 are marked with the same reference numbers.
In the speech signal encoder shown in Figure 3, the speech signal supplied to the input terminal 101 is filtered by a high-frequency filter HPF109 to eliminate a signal in an undesired range, and then supplied to the LPC analysis/quantization unit 113 An LPC (Linear Predictive Coding) analysis circuit 132 and an inverse LPC filter 111.
The LPC analysis circuit 132 of the LPC analysis/quantization unit 113 uses a Hamming window with a 256-bit sample to form a block of input signal waveform block length, and uses an autocorrelation method to find the so-called α-parameter linear prediction parameter. The frame interval as a data output unit is set to approximately 160 samples. If the sample frequency fs is for example 8KHz, a single frame interval is 20msec or 160 samples.
The α-parameters from the LPC analysis circuit 132 are sent to an α-LSP conversion circuit 133 to be converted into line spectrum pair (LSP) parameters. This direct type filter coefficient parameter looking into α- converted into a number such as +, i.e., five pairs of LSP parameters. This conversion is carried out by, for example, the Newton-Raphson method. The reason why α-parameters are converted into LSP parameters is that LSP parameters are superior to α-parameters in interpolation characteristics.
The LSP parameter from the α-LSP conversion circuit is a matrix or vector quantized by the LSP quantizer 134. Before vector quantization, the difference between the frame and the frame can be taken, or the complex number frame can be collected for matrix quantization. In this example, the two frames of the LSP parameters that are each 20msec long and are calculated every 20msec are processed by matrix quantization and vector quantization together.
The quantized output of the quantizer 134, that is, the number of index data quantized by the LSP is taken from a terminal 102, and the quantized LSP is sent to an LSP interpolation circuit 136.
The LSP interpolation circuit 136 interpolates the LSP vector every 20 msec or 40 msec to provide an eighth rate. That is, the LSP vector is updated every 2.5 msec. The reason is that if the remaining waveform is analyzed/synthesized by the harmonic encoding/decoding method, the message envelope of the synthesized waveform provides a very gentle waveform, so that if the LPC coefficient changes rapidly every 20msec, an external noise may be generated. News. That is, if the LPC coefficient is gradually changed every 2.5 msec, the generation of such external noise can be prevented.
In order to use the interpolated LSP vector generated every 2.5 msec to perform inverse filtering of the input speech, an LSP-to-α conversion circuit 137 is used to convert the LSP parameters into α-parameters, which are filtering parameters such as ten-bit direct filtering. An output of the LSP to α conversion circuit 137 is sent to an LPC inverse filter circuit 111, which performs inverse filtering with an α-parameter updated every 2.5 msec to generate a smooth output. One output of the inverse LPC filter 111 is sent to a quadrature conversion circuit 145 such as a -DCT circuit in a sinusoidal analysis coding unit such as a harmonic coding circuit.
The alpha parameter from the LPC analysis circuit 132 of the LPC analysis/synthesis unit 113 is sent to a perceptual weighting filter calculation circuit 139, where data for perceptual weighting is searched. These weighted data are sent to the perceptual weighted vector quantizer 116 of the second encoding unit 120, the perceptual weighted filter 125 and the perceptual weighted synthesis filter 122.
The sinusoidal analysis and coding unit 114 of the harmonic coding circuit analyzes the output of the inverse LPC filter 111 by the harmonic coding method. That is to say, perform tone detection, frequency amplitude AM calculation of individual harmonics, and voice (V)/unvoiced (UV) discrimination, and the frequency amplitude of individual harmonics or the number of message packets that change with the tone can be converted by size conversion. Constant.
In the illustrated example of the sinusoidal analysis encoding unit 114 shown in FIG. 3, a common harmonic encoding is used. In particular, in multi-band excitation (MBE) coding, the voiced part and the unvoiced part appear in each frequency domain or band (in the same block or frame) at the same time point. In other harmonic coding techniques, it uniquely determines whether the voice in a block or a frame is voiced or unvoiced. In the following description, as far as MBE coding is concerned, if all bands are UV, the given frame is judged as UV. For specific technical examples of the above-mentioned MBE analysis and synthesis method, please refer to the Japanese Patent Application No. 4-91442 filed in the name of the authorized person.
The open-loop pitch retrieval unit 141 and the zero-crossing computer 142 of the sine analysis coding unit 114 shown in FIG. 3 feed the input voice signal from the input method 101 and the signal from the HPF 109 respectively. The LPC residue or linear prediction residue from the inverse LPC filter 111 is supplied to the orthogonal transform circuit 145 of the sinusoidal analysis coding unit 114. The open-loop search unit takes out the LPC remaining of the input signal and performs a rougher tone search through the open-loop search. As described later, the extracted coarse pitch data is supplied to a fine pitch searching unit 146 through a closed-loop search. The normalized autocorrelation maximum value obtained from the open-loop tone retrieval unit 141 is normalized by normalizing the autocorrelation maximum value of the remaining LPC together with the coarse tone data, and the normalized autocorrelation maximum value is taken out together with the coarse tone data and sent to the V/UV frequency discrimination unit 115.
The orthogonal transform circuit 145 performs orthogonal transform such as discrete Fourier transform, and converts the LPC residue on the time axis into spectral frequency amplitude data on the frequency axis. The output of the orthogonal transform circuit 145 is sent to the fine pitch retrieval unit 146, and a spectrum estimation unit 148 is configured to calculate the value of the spectrum frequency amplitude or message envelope.
The coarser pitch data extracted by the open-loop pitch retrieval unit 141 and the frequency domain data obtained by the orthogonal transform unit 145 with DFT are fed to the fine pitch retrieval unit 146. The fine pitch retrieval unit 146 changes the pitch data by ± several samples at a rate of 0.2 to 0.5 centered on the coarse pitch value data, so as to finally reach the fine pitch data value with the best decimal point (floating value). The analysis and synthesis method is used as a fine retrieval technique to select a tone, so that the power spectrum is closest to the power spectrum of the original sound. The tone data from the closed-loop fine search unit 146 is sent to an output terminal 104 via a switch 118.
In the spectrum estimation unit 148, the amplitude of each harmonic and the spectrum information as the sum of the harmonics are estimated based on the amplitude and pitch of the remaining orthogonal transform as the LPC, and sent to the fine pitch search unit 146, V/UV detection The frequency unit 115 and the perceptual weight vector quantization unit 116.
The V/UV discrimination unit 115 discriminates the V/UV of the picture based on the output of the orthogonal transform circuit 145, the optimal pitch from the fine pitch retrieval unit 146, and the standardized autogenous correlation r(p) from the open-loop pitch retrieval unit 141 ) The maximum value and the zero-crossing calculation value from the zero-crossing computer 142. In addition, the boundary position used for MBE and V/UV discrimination based on band can also be used as a condition for V/UV discrimination. The frequency discriminating output of the V/UV frequency discriminating unit 115 is taken out at the output terminal 105.
The output unit of the spectrum estimation unit 148 or the input unit of the vector quantization unit 116 is provided with a data number conversion unit (a unit for performing a sampling rate conversion). Given that the number of bands separated on the frequency axis is different from the number of data and tones, the data number conversion unit is used to set the amplitude data AM of a message envelope to a constant value. That is, if the effective waveband reaches 3400KHz, the effective waveband can be divided into 8 to 63 wavebands according to the pitch. The mMX+1 number of the amplitude data AM obtained from the band and the band changes from 8 to 63 in the range. In this way, the data number conversion unit changes the amplitude data of a variable number of mMX+1 into predetermined M number data such as 44 data.
The predetermined M-number amplitude data or envelope data such as 44 from the data number conversion unit and provided to one of the output units of the spectrum estimation unit 148 or one of the input units of the vector quantization unit 116 together take the predetermined number of data such as 44 data as one. The unit is processed by the vector quantization unit 116 performing weighted vector quantization. This weight is supplied by the output of the perceptual weight filter calculation circuit 139. The message packet index from the vector quantizer 116 is fetched from an output terminal 103 by a switch 117. Before the weighted vector quantization, it is better to use an appropriate leakage coefficient in a vector composed of a predetermined number of data to obtain the difference between the frame and the frame.
The second encoding unit 120 is described here. The second encoding unit 120 has a so-called CELP encoding structure, which is particularly used to partially encode the input speech signal. In the CELP coding structure used in the silent part of the input speech signal, a noise output corresponding to the LPC residue of the silent sound as the representative output value of the noise code book or the so-called random code book 121 is passed through a gain control circuit 126 is sent to a perceptual synthesis filter 122. The weighted synthesis filter 122 synthesizes the input noise LPC by LPC synthesis, and sends the generated weighted silent signal to the subtractor 123. The divider 123 feeds a signal from the input terminal 101 via a high-frequency filter (HPF) 109, and is perceptually weighted by a perceptual weighting filter 125. The subtractor looks for the difference or error between the signals from the synthesis filter 122. At the same time, a zero input response of the perceptual weighted synthesis filter is subtracted from the perceptual weighted filter output 125 in advance. This error is fed to a distance calculation circuit 124 for distance calculation. Searching in the noise codebook will minimize the error to the representative vector value. The above is an overview of vector quantization of time-domain waveforms by means of closed-loop retrieval and comprehensive analysis.
The waveform index of the code book is retrieved from the noise code book 121 and the gain index of the code book is retrieved from the gain circuit 126 as data taken from the unvoiced (UV) part of the second encoder 120 using the CELP encoding structure. The waveform index, that is, the UV data from the code book 121 is sent to an output terminal 107s via a switch 127s, and the gain index, that is, the UV data of the gain circuit is sent to an output terminal 107g via a switch 127g.
The switches 127s, 127g, 117 and 118 are turned on and off according to the V/UV decision result from the V/UV frequency discrimination unit 115. In particular, when the V/UV discrimination result of the voice signal of the transmitted picture indicates that it is voiced (V), turn on the switches 117 and 118. If the voice signal of the transmitted picture indicates that the voice signal is silent (UV), that is Turn on the switch for 127s, 127g.
Figure 4 shows a more detailed structure of a speech signal decoder shown in Figure 2. In Figure 4, the same numbers are used to indicate the corresponding parts in Figure 2.
In FIG. 4, the vector quantization output of the LSPs corresponding to the output terminal 102 shown in FIGS. 1 and 3, that is, the codebook index, is provided to an input terminal 202.
The LSP index is sent to the LSP inverse vector quantizer 231 for the LPC parameter regeneration unit 213, and the inverse vector is quantized into a line of spectrum pair (LSP) data, which is then supplied to the LSP for interpolation. The generated interpolation data is converted into α parameters by an LSP to α conversion circuit 234, 235 and sent to the LPC synthesis filter 214. The LSP interpolation circuit 232 and the LSP to α conversion circuit 235 are used for voiced (V) sounds, and the LSP interpolation circuit 233 and the LSP to α conversion circuit 235 are used for unvoiced (UV) sounds. The LPC synthesis filter 214 is composed of the LPC synthesis filter 236 of the voiced speech part and the LPC synthesis filter 237 of the unvoiced speech part. That is, the LPC coefficient interpolation is performed independently for the voiced voice part and the unvoiced voice part to prevent the transition part from the voiced voice part to the unvoiced voice part, or vice versa. The undesirable effects of the interpolation of LSPs.
The code index data corresponding to the weighted vector quantized spectrum information envelope Am is supplied to the input terminal 203 in FIG. 4, and this weighted vector quantized spectrum envelope Am corresponds to the output of the terminal 103 of the encoder shown in FIGS. 1 and 3. The tone data is supplied to an input terminal 204 from the terminal 104 shown in FIGS. 1 and 3, and the V/UV discrimination data is supplied to an input terminal 205 from the terminal 105 shown in FIGS. 1 and 3.
The vector quantization index data of the spectrum information envelope Am from the input terminal 203 is supplied to an inverse vector quantizer 212 for inverse vector quantization, where a conversion opposite to the data number conversion is performed. The generated spectrum envelope data is sent to a sinusoidal synthesis circuit 215.
If the difference between the frames is found before the spectral vector quantization during encoding, that is, the inverse vector quantization decodes the difference between the frames to generate spectral envelope data.
The sinusoidal synthesis circuit 215 is fed with the tone from the input terminal 204 and the V/UV frequency discrimination data from the input terminal 205. The remaining LPC data corresponding to the output of the LPC anti-filter 111 shown in FIGS. 1 and 3 is taken out from the sinusoidal synthesis circuit 215 and sent to an adder 218. The specific technology of sine synthesis is disclosed in, for example, Japanese Patent Application Nos. 4-91442 and 6-198451 filed by the licensee.
The envelope of the vector quantizer 212 and the tone and V/UV discrimination data from the input terminals 204 and 205 are supplied to a noise synthesis circuit 216 configured to add noise to the voiced part (V). An output of the noise synthesis circuit 216 is sent to an adder 218 via a weighted superimposition circuit 217. Specifically, if sine wave synthesis is used to generate excitation as the input of LPC synthesis filter sent to voiced sound, it will produce a deep feeling like a man's voice in the bass sound, and the sound will change sharply between the voiced sound and the unvoiced sound. In this way, an unnatural sense of hearing is generated. In view of this, noise is added to the voiced part of the remaining signal of the LPC. This noise takes into account the relevant parameters of the speech coded data. These parameters are input to the LPC synthesis filter input of the voiced speech, which are related to the excitation such as the pitch, the amplitude of the spectral message envelope, the maximum amplitude in the picture or the remaining signal level Wait.
The sum output of the adder 218 is sent to a voiced voice synthesis filter 236 for LPC synthesis filtering, where LPC synthesis is performed to form time waveform data, and then filtered by a post-filter 238v for voiced voice, and sent To adder 239.
The waveform index and gain index, that is, the UV data from the output terminals 107s and 107g in FIG. 3, are respectively supplied to the input terminals 207s and 207g shown in FIG. 4, and then supplied to the silent speech synthesis unit 220. The waveform index from the input terminal 207s is sent to the noise code book 221 of the silent speech synthesis unit 220, and the gain index from the input terminal 207g is sent to the gain circuit 222. The representative value output read from the noise code book 21 is composed of a noise signal corresponding to the LPC residue of the silent speech. This becomes a preset gain amplitude in the gain circuit 222 and is sent to a windowing circuit 223 to open the window and flatten the boundary of the voiced speech part.
One output of the windowing circuit 223 is sent to a synthesis filter 237 for unvoiced (UV) speech of the LPC synthesis filter 214. The data sent to the synthesis filter 237 is processed by LPC synthesis into time waveform data for the silent part. The time waveform data of the silent part is filtered by a post filter for the silent part 238u before being sent to an adder 239.
In the adder 239, the time waveform signal from the post-filtering for the voiced speech 238v and the time waveform data from the post-filtering 238u for the unvoiced speech for the silent speech are added to each other, and the sum of the data obtained is Take it out at the output terminal 201.
Depending on the required sound quality, the above-mentioned speech signal encoder can output data with different bit speeds. That is, the output data can be output at a variable bit rate.
Specifically, the bit rate of the output data can be switched between a low bit rate and a high bit rate. For example, if the low bit rate is 2 kbps and the high bit rate is 6 kbps, the output data is data with a bit rate at the bit rate of the fifth picture.
The tone data from the output terminal 104 is always output at a bit rate of 8 bits/20msec for voiced speech, and the V/UV discriminator output from the output terminal 105 is output at 1 bit/20msec at any time . The LSP quantization index output from the output terminal 105 is switched between 32 bits/40msec and 48 bits/40msec. On the other hand, the index is switched between 15 bits/20msec and 87 bits/20msec while the output terminal 103 outputs voiced speech (V). The index for the silent part (UV) output from the output terminals 107s and 107g is switched between 11 bits/10msec and 23 bits/5msec. The output data for unvoiced sound (UV) is 40 bits/20msec per 2kbps and 120kbps/20msec per 6kbps. On the other hand, the output data for voiced sound (V) is 39 bits/20msec per 2kbps and 117kbps/20msec per 6kbps.
The index for LSP quantization, the index for voiced speech (V) and the index for unvoiced speech (UV) related to the configuration of relevant parts will be described later.
Here, parameters 6 and 7 illustrate the matrix quantization and vector quantization in the LSP quantizer 134.
The α-parameters from the LPC analysis circuit 132 are sent to the α-LSP circuit 133 to be converted into LSP parameters. If the P-bit LPC analysis is executed in the LPC analysis circuit 132, the α-parameter is calculated. These P α-parameters are converted into LSP parameters and kept in a buffer 610.
The buffer 610 outputs the second frame of the LSP parameter. The two frames of the LSP parameters are the -by-first matrix quantizer 620<sub>1</sub>And a second matrix quantizer 620<sub>2</sub>The matrix quantizer 620 constitutes matrix quantization. These two frames of LSP parameters are in the first matrix quantizer 620<sub>1</sub>In matrix quantization, the obtained quantization error is further matrix quantized in the matrix quantizer. Matrix quantization uses correlation between the time axis and the frequency axis.
Will come from the matrix quantizer 620<sub>2</sub>The quantization error input for the two frames is input by a first vector quantizer 640<sub>1</sub>With a second vector quantizer 640<sub>2</sub>The vector quantization unit 640. First vector quantizer 640<sub>1</sub>It consists of two vector quantization units 650 and 660, and the second vector quantizer 640<sub>2</sub>It is composed of the two-vector quantization unit 670.680. The obtained quantized error vector is further used by the second vector quantizer 640<sub>2</sub>The vector quantization unit 670, 680 further vector quantization. The vector quantization described above uses correlation on the frequency axis.
The matrix quantization unit 620 that performs the aforementioned matrix quantization includes at least one first matrix quantizer 620<sub>1</sub>, Used to perform the first matrix quantization step, and a second matrix quantizer 620<sub>2</sub>, Used to perform the second matrix quantization step to quantize the quantization error matrix generated by the first matrix quantization. The vector quantization unit 640 that performs the above-mentioned vector quantization includes at least one first vector quantizer 640<sub>1</sub>, Used to perform a first vector quantization step, and a second vector quantizer 640<sub>2</sub>, Used to perform a second matrix quantization step to quantize the quantization error matrix generated by the first vector quantization.
Here is a detailed description of matrix quantization and vector quantization.
The LSP parameters for the two frames stored in the buffer 600, that is, a 10×2 matrix is sent to the first matrix quantizer 620<sub>1</sub>. First matrix quantizer 620<sub>1</sub>The LSP parameters for the two frames are sent to a weighted distance calculation unit 623 via the LSP parameter adder 621 to find the minimum weighted distance.
During the codebook search, the first matrix quantizer 620<sub>1</sub>The resulting distortion size d<sub>MQ1</sub>Expressed by equation (2):<maths><img file="TW360859B_D0001.tif" /></maths>Where X<sub>1</sub>Is the LSP parameter, X<sub>1</sub>'Is a quantized value, and t and i are P-size numbers.
The weight limits on the frequency axis and time axis that are not considered are expressed in equation (2):<maths><img file="TW360859B_D0002.tif" /></maths>Regardless of t, x(t, 0)=0, x(t, p+1)=π.
The weight w of equation (2) is also used for matrix quantization and vector quantization on the downstream side.
The calculated weighted distance is sent to a matrix quantizer MQ<sub>1</sub>622 is matrix quantized. The 8-bit index output by the matrix quantization is sent to a signal switch 690. The quantized value obtained by matrix quantization is subtracted from the LSP parameter used in the two frames in the buffer 610 in an adder 621. The weighted distance calculation unit 623 calculates the weighted distance of every two frames for matrix quantization in the matrix quantization unit 622. And choose the quantization value that minimizes the weighted distance. One output of the adder 621 is sent to the second matrix quantizer 620<sub>2</sub>One of the adders 631.
Similar to the first matrix quantizer 620<sub>1</sub>, The second matrix quantizer 620<sub>2</sub>Perform matrix quantization. An output of the adder 621 is sent to a weighted distance calculation unit 633 via the adder 631, where the minimum weighted distance is calculated.
During the codebook search, the second matrix quantizer 620<sub>2</sub>The desired distortion size d<sub>MQ2</sub>Expressed by equation (3):<maths><img file="TW360859B_D0003.tif" /></maths>
The weighted distance is sent to a matrix calculation unit (MQ<sub>2</sub>) 632 is matrix quantized. An 8-bit index output by matrix quantization is sent to a signal switch 690. The weighted distance calculation unit 633 then uses the output of the adder 631 to calculate the weighted distance. One of the adders 631 outputs a frame and then a frame is sent to the first vector quantizer 640<sub>1</sub>The adders 651,661.
First vector quantizer 640<sub>1</sub>Vector quantization is performed one frame after another. One output of the adder 631 is sent to each weighted distance calculation unit 653, 663 through the adders 651, 661 for calculating the minimum weighted distance, one frame by one frame.
Quantization error X<sub>2</sub>And quantization error X<sub>2</sub>The difference between'is a (10×2) matrix. If this difference is X<sub>2</sub>-X<sub>2</sub>'=[<u style="single">x</u><sub>3-1</sub>,<u style="single">x</u><sub>3-2</sub>] Is represented by the first vector quantizer 640<sub>1</sub>The resulting misalignment size d<sub>VQ1</sub>, D<sub>VQ2</sub>Expressed by equations (4) and (5):<maths><img file="TW360859B_D0004.tif" /></maths>
<maths><img file="TW360859B_D0005.tif" /></maths>
The weighted distance is sent to a vector quantization unit 652 and a vector quantization unit VQ<sub>2</sub>For vector quantization. Each 8-bit index output by this vector quantization is sent to the signal switch 690. The quantized value is subtracted from the input two-frame quantization error vector by the adders 651 and 661. Next, the weighted distance calculation units 653 and 663 use the output of the adders 651 and 661 to calculate the weighted distance to select the quantized value that minimizes the weighted distance. The outputs of the adders 651 and 661 are sent to the second vector quantizer 640<sub>2</sub>The adder 671,681.
During the codebook search, the second vector quantizer 640<sub>2</sub>The distortion size d obtained by the vector quantizer 672, 682<sub>VQ3</sub>, D<sub>VQ4</sub>Relative to<u style="single">x</u><sub>4-1</sub>=x<sub>3-1</sub>-<u style="single">x</u><sub>3-1</sub>' <u style="single">x</u><sub>4-2</sub>=x<sub>3-2</sub>-<u style="single">x</u><sub>3-2</sub>'
Expressed by equations (6) and (7):<maths><img file="TW360859B_D0006.tif" /></maths>
The weighted distance is sent to the vector quantizer (VQ<sub>3</sub>)672 and vector quantizer (VQ<sub>4</sub>)682 for vector quantization. The 8-bit output index from the vector quantization is removed by the adders 671, 681 from the input quantization error vector of the two frames. Next, the weighted distance calculation units 673 and 683 use the output of the adders 671 and 681 to calculate the weighted distance to select the quantized value that minimizes the weighted distance.
During the codebook learning period, the General Lloyd algorithm is used to learn according to the magnitude of individual distortion.
The magnitude of distortion may be different during the codebook search period and during the learning period.
The 8-bit index data from the matrix quantization units 622, 632 and the vector quantization units 652, 662, 672, and 682 are converted by the signal conversion switch 690 and output at the output terminal 691.
Specifically, in the case of low bit rate, take out the first matrix quantizer 620 that performs the first matrix quantization step, and the second matrix quantizer 620 that performs the second matrix quantization step<sub>2</sub>And the first vector quantizer 640 that performs the first vector quantization step<sub>1</sub>In the case of high bit rate, the low bit rate output is added to the second vector quantizer 640 which performs the second vector quantization step<sub>2</sub>On the output, and take out the sum.
Thus, 2kbps and 6kbps are respectively a 32-bit/40msec index and a 48-bit/40msec index.
The matrix quantization unit 620 and the vector quantization unit 640 perform weighting that is consistent with the parameter characteristics representing the LPC coefficients and is restricted to the frequency axis and/or the time axis.
First, the weighting that is limited to the frequency axis and consistent with the LSP parameter characteristics will be explained. If the number of digits (orders) p=10, the LSP parameter X(i) is divided into L<sub>1</sub>={X(i)1<img file="TW360859B_D0007.tif" />i<img file="TW360859B_D0008.tif" />2}
L<sub>2</sub>={X(i)3<img file="TW360859B_D0009.tif" />i<img file="TW360859B_D0010.tif" />6}
L<sub>3</sub>={X(i)7<img file="TW360859B_D0011.tif" />i<img file="TW360859B_D0012.tif" />10} Three sets of high, medium and low ranges. If L<sub>1</sub>, L<sub>2</sub>With L<sub>3</sub>The group weights are 1/4, 1/2, and 1/4, and the weights limited to the frequency axis are expressed in equations (8), (9) and (10):<maths><img file="TW360859B_D0013.tif" /></maths>
The weighting of individual LSP parameters is only performed in each group, and this weighting is restricted by the weighting of each group.
Looking along the time axis, the sum of individual frames must be 1, so that the limits in the time axis are based on frames. The weights limited only to the direction of the time axis are expressed in equation (11):<maths><img file="TW360859B_D0014.tif" /></maths>Among them, 1i10 and 0t1.
Therefore, the equation (11) is not limited to the weighting in the frequency axis direction between the two frames with the number of frames t=0 and t=1. This is only limited to the weighting in the time axis direction between the two frames processed by matrix quantization.
During the learning period, the total number of frames with total T used as learning data are weighted according to equation (12):<maths><img file="TW360859B_D0015.tif" /></maths>Among them, 1i10 and 0tT.
It is explained that the weighting is limited to the direction of the frequency axis and the direction of the time axis. If the number of bits P=10, the LSP parameter x(i, t) is divided into L<sub>1</sub>={x(i, t)1<img file="TW360859B_D0016.tif" />i<img file="TW360859B_D0017.tif" />2, 0<img file="TW360859B_D0018.tif" />t<img file="TW360859B_D0019.tif" />1}
L<sub>2</sub>={x(i, t)3<img file="TW360859B_D0020.tif" />i<img file="TW360859B_D0021.tif" />6, 0<img file="TW360859B_D0022.tif" />t<img file="TW360859B_D0023.tif" />1}
L<sub>3</sub>={x(i, t)7<img file="TW360859B_D0024.tif" />i<img file="TW360859B_D0025.tif" />10, 0<img file="TW360859B_D0026.tif" />t<img file="TW360859B_D0027.tif" />1} High, medium and low three groups of ranges. If each group L<sub>1</sub>, L<sub>2</sub>With L<sub>3</sub>The weights are 1/4, 1/2 and 1/4, and the weighting is limited to the frequency axis, which is expressed in equations (13), (14) and (15):<maths><img file="TW360859B_D0028.tif" /></maths>
According to equations (13) to (15), every three frames in the frequency axis direction and the two frames processed by matrix quantization in the time axis direction are weighted and restricted. This is valid during the codebook search and the learning period.
During learning, all frames of all data are weighted. The LSP parameter x(i, t) is divided into: L<sub>1</sub>={x(i, t)1<img file="TW360859B_D0029.tif" />i<img file="TW360859B_D0030.tif" />2, 0<img file="TW360859B_D0031.tif" />t<img file="TW360859B_D0032.tif" />T}
L<sub>2</sub>={x(i, t)3<img file="TW360859B_D0033.tif" />i<img file="TW360859B_D0034.tif" />6, 0<img file="TW360859B_D0035.tif" />t<img file="TW360859B_D0036.tif" />T}
L<sub>3</sub>={x(i, t)7<img file="TW360859B_D0037.tif" />i<img file="TW360859B_D0038.tif" />10, 0<img file="TW360859B_D0039.tif" />t<img file="TW360859B_D0040.tif" />T} high, medium and low range. If each group L<sub>1</sub>, L<sub>2</sub>With L<sub>3</sub>The weights are 1/4, 1/2 and 1/4 respectively, limited to each group L in the frequency direction and time direction<sub>1</sub>, L<sub>2</sub>With L<sub>3</sub>The weight is expressed by equations (16), (17) and (18):<maths><img file="TW360859B_D0041.tif" /></maths>
By equations (16) to (18), the three ranges in the frequency axis direction and all the frames in the time axis direction can be weighted.
In addition, the matrix quantization unit 620 and the vector quantization unit 640 perform weighting according to the change range of the LSP parameter. In the V to UV or UV to V transition region representing a small number of frames in all voice frames, the LSP parameter changes significantly due to the frequency difference between the consonant and the vowel. Therefore, the weighting shown in equation (19) can be multiplied by W'(i, t) performing the weighting to emphasize the transition area.
<maths><img file="TW360859B_D0042.tif" /></maths>
The following equation (20):<maths><img file="TW360859B_D0043.tif" /></maths>
Can be used instead of equation (19).
In this way, the LSP quantization unit 134 performs second-order matrix quantization and second-order vector quantization to make the number of bits of the output index variable.
Fig. 8 shows the basic structure of the vector quantization unit 116, and Fig. 9 shows a more detailed structure of the vector quantization unit 116 shown in Fig. 8. The schematic structure of the weighted vector quantization of the spectrum information envelope Am in the vector quantization unit 116 is described here.
First, in the speech signal encoding device shown in Fig. 3, an icon configuration for data number conversion is used to provide a constant data number of the amplitude of the spectrum information envelope on an output side or vector of the spectrum estimation unit 148 The input terminal of the quantization unit 116.
It is conceivable that there are many methods for this kind of data conversion. In this embodiment, the dummy data inserted from the last data in a block to the first value in the block, or a preset such as repeating the last data or the first data in a block Suppose data is added to the amplitude data of a block of an effective band on the frequency axis to increase the number of data to N<sub>F</sub>, The number of amplitude data equal to the Os times such as 8 times has been found to oversample the Os tuples such as 8-tuples for the limited wave bandwidth type. Insert this ((mMx+1)×Os) amplitude data to expand to a larger N such as 2408<sub>M</sub>number. This N<sub>M</sub>The data is auxiliary sampled to be converted into the above-mentioned preset data M number such as 44 data. In fact, only the data needed to formulate the fundamental M data is calculated by supersampling and linear interpolation, without finding all the above N<sub>M</sub>data.
The vector quantization unit 116 used for weighted vector quantization in Figure 8 includes at least a first vector quantization unit 500 for performing the first vector quantization step, and a second vector quantization unit 510 for performing the second vector quantization Step: quantize the first vector quantization period with the quantization error vector generated by the first vector quantization unit. The first vector quantization unit 500 is a so-called first-stage vector quantization unit, and the second vector quantization unit 510 is a so-called second-stage vector quantization unit.
Output vector of one of the spectrum estimation unit 148<u style="single">x</u>, That is, the message packet data with a predetermined segment M is input into one of the input terminals 501 of the first vector quantization unit 500. This output vector<u style="single">x</u>The vector quantization unit 502 performs quantization by the weight vector. In this way, a waveform index output by the vector quantization unit 502 is output at an output terminal 503, and a quantized value<u style="single">x</u><sub>0</sub>'Is output at an output terminal 504 and sent to the adders 503 and 513. Adder 505 self-source vector<u style="single">x</u>Subtract quantized value<u style="single">x</u><sub>0</sub>'To give a multi-order quantified error vector<u style="single">y</u>。
Quantization error vector<u style="single">y</u>It is sent to one of the vector quantization units 511 of the second vector quantization unit 510. In Figure 8, the second vector quantization unit 511 is composed of a complex vector quantizer or a second vector quantizer 511<sub>1</sub>,511<sub>2</sub>composition. Divide the size of this quantization error vector Y, so that the two vector quantizer 511<sub>1</sub>,511<sub>2</sub>Medium-weighted vector quantization is quantified. These vector quantizers 511<sub>1</sub>,511<sub>2</sub>The output waveform index is at output 512<sub>1</sub>,512<sub>2</sub>Output, while the quantized value<u style="single">y</u><sub>1</sub>',<u style="single">y</u><sub>2</sub>'Is connected to the dimension direction and sent to an adder 513. The adder 513 converts the quantized value<u style="single">y</u><sub>1</sub>',<u style="single">y</u><sub>2</sub>'Add to quantized value<u style="single">x</u><sub>0</sub>'To generate a quantified value<u style="single">x</u><sub>1</sub>'And output at an output terminal 514.
In this way, in the case of low bit rate, the output of the first vector quantization step performed by the first vector quantization unit 500 is taken, and in the case of high bit rate, the output of the first vector quantization step and the output of the second vector quantization unit are output 510 Perform the output of the second quantization step.
Specifically, as shown in FIG. 9, the vector quantizer 502 of the first vector quantization unit 500 in the vector quantization section 116 has a two-stage structure of 44-size L ordinal numbers.
That is, the sum of the 44-size vector quantization codebook and the 32-size codebook and a gain g<sub>i</sub>Multiply and use it as a 44-size spectrum message envelope vector<u style="single">x</u>The quantized value. As shown in Figure 8, this two code book is CB<sub>0</sub>With CB<sub>1</sub>, And the output vector is<u style="single">s</u><sub>1i</sub>,<u style="single">s</u><sub>1j</sub>, Where 0i and j31. On the other hand, the gain codebook CB<sub>g</sub>One output is g<sub>1</sub>, Where 0131, g<sub>1</sub>It is vectorless. A final output<u style="single">x</u><sub>0</sub>'Is g<sub>1</sub>(<u style="single">s</u><sub>1i</sub>+<u style="single">s</u><sub>1j</sub>)。
Obtained from the above-mentioned LPC remaining MBE analysis and converted into a spectrum information of a preset size, Am as<u style="single">x</u>. The decisive thing is<u style="single">x</u>How to effectively quantify.
The quantization error energy E is defined by E=W(H<u style="single">x</u>-Hg1(<u style="single">s</u><sub>0i</sub>+<u style="single">s</u><sub>1j</sub>)}∥<sup>2</sup>=WH{<u style="single">x</u>-{<u style="single">x</u>-g<sub>1</sub>(<u style="single">s</u><sub>0i</sub>+<u style="single">s</u><sub>1j</sub>)}∥<sup>2</sup>.........(21), where H refers to the characteristics on the frequency axis of the LPC synthesis filter and W refers to a matrix used to weight the representative characteristics of the perceptual weight on the frequency axis.
If the α parameter of the LPC analysis result of the current frame is α<sub>1</sub>(1iP), the value of L corresponding to a point such as 44 is sampled from the frequency response of equation (22):<maths><img file="TW360859B_D0044.tif" /></maths>
For calculation, Os is next to a string of 1, α<sub>1</sub>, Α<sub>2</sub>,......Α<sub>p</sub>Fill in to get a string of 1, α<sub>1</sub>, Α<sub>2</sub>,......Α<sub>p</sub>, 0, 0, ........., 0, such as 256 points of data. Then, borrow 256 points FFL to calculate the points related to a value from 0 to π (r<sub>e</sub><sup>2</sup>+im<sup>2</sup>)<sup>1/2</sup>, And find the reciprocal of this result. These reciprocals are auxiliary sampled to L points, such as 44 points, and form a matrix with L points as diagonal elements.
<maths><img file="TW360859B_D0045.tif" /></maths>
A perceptual weighting matrix W is represented by equation (23):<maths><img file="TW360859B_D0046.tif" /></maths>Where α<sub>1</sub>Is the result of LPC analysis, and λ<sub>a</sub>, Λ<sub>b</sub>Is a constant, λ=0.4 and λ=0.9.
The matrix W can be calculated from the frequency response of the above equation. For example, FFT is in the range of 0 to π, α1λb, α2λ1b<sup>2</sup>,............Αpλb<sup>p</sup>, 0, 0, ... ... 256 point data of 0 is executed to obtain (r<sub>2</sub><sup>2</sup>[i]+Im<sup>2</sup>[i])<sup>1/2</sup>, Where 0i128. The frequency division of the denominator should be 256-point FFT, and the range from 0 to π is complexed to 1, α1λa, α2λ1a<sup>2</sup>,............Αpλa<sup>p</sup>, 0, 0, ......... 0 is calculated at 128 points (r<sub>2</sub><sup>'2</sup>[i]+Im<sup>'2</sup>[i])<sup>1/2</sup>, Where 0i128. The frequency response of equation (23) can be obtained by<maths><img file="TW360859B_D0047.tif" /></maths>Find out, where 0i128.
This is obtained by the following method for each correlation point of the 44 size vector, for example. For more accuracy, linear interpolation should be used. However, in the following examples, the closest point is used instead.
That is, ω[i]=ω0[nint{128i/L}], where 1iL.
In the equation, nint(X) is a function that returns a value to the nearest x. As for H, use a similar method to find h(1), h(2),...h(L), that is<maths><img file="TW360859B_D0048.tif" /></maths>
Another example is to find H(z)W(z) first, and then find the frequency response to reduce the FFT time. That is, equation (25):<maths><img file="TW360859B_D0049.tif" /></maths>The denominator of is expanded to<maths><img file="TW360859B_D0050.tif" /></maths>For example, 256 points of data use a string of 1, β<sub>1</sub>, Β<sub>2</sub>,..., β<sub>2P</sub>, 0, 0, ........., 0 to generate. Then take 256 points FF T, the frequency response of the amplitude is<maths><img file="TW360859B_D0051.tif" /></maths>Among them, 0i128. therefore,<maths><img file="TW360859B_D0052.tif" /></maths>Among them, 0i128. This is calculated for each corresponding point of the L size vector. If the number of FFT points is small, linear interpolation should be used. Only the closest value is determined by:<maths><img file="TW360859B_D0053.tif" /></maths>Find out. Among them, 1iL. If the matrix with these values as diagonal elements is W',<maths><img file="TW360859B_D0054.tif" /></maths>
Equation (26) is the same matrix as the above equation (24). Alternatively, H(exp(jω))W(exp(jω)) can be directly calculated from equation (25) in terms of ωiπ for use in ωh[i]. Alternatively, the pulse wave response of equation (25) with an appropriate length, such as 40 points, can be obtained, and FFT can be performed to obtain the frequency response of the used wave amplitude.
<maths><img file="TW360859B_D0055.tif" /></maths>
Using this matrix, that is, the frequency characteristics of the weighted synthesis filter, and rewriting equation (21), we get: E=W'(<u style="single">x</u>-g<sub>1</sub>(<u style="single">s</u><sub>0i</sub>+<u style="single">s</u><sub>1j</sub>))∥………(27)
The method of learning the waveform code book and gain code book is explained here. CB<sub>0</sub>Choose a code vector<u style="single">s</u><sub>0C</sub>To minimize the expected distortion of all frames k. If there are M such frames, and if<maths><img file="TW360859B_D0056.tif" /></maths>Minimize is enough. ω<sub>k</sub>',<u style="single">x</u><sub>k</sub>, G<sub>k</sub>and<u style="single">s</u><sub>1k</sub>Respectively refer to the weight of the k-th frame, the input to the k-th frame, the gain of the k-th frame and the output of the CB1 codebook of the k-th frame.
To minimize equation (28),<maths><img file="TW360859B_D0057.tif" /></maths>therefore,<maths><img file="TW360859B_D0058.tif" /></maths>Good<maths><img file="TW360859B_D0059.tif" /></maths>Where 0 refers to an inverse matrix and W<sub>k</sub><sup>'T</sup>Refers to a W<sub>K</sub>'The transpose matrix. Secondly consider gain optimization.
The expected value of distortion related to the k-th frame of the codeword gc of the selected gain is expressed by the following formula:<maths><img file="TW360859B_D0060.tif" /></maths>untie<maths><img file="TW360859B_D0061.tif" /></maths>Get it<maths><img file="TW360859B_D0062.tif" /></maths>and<maths><img file="TW360859B_D0063.tif" /></maths>
The above equations (31) and (32) give the waveform<u style="single">s</u><sub>0i</sub>,<u style="single">s</u><sub>1i</sub>Optimize the centroid conditions and give 0i31, 0j31 and 0131 this gain g<sub>1</sub>, Which is to optimize the decoding output. At the same time, s<sub>1i</sub>Can be the same as<u style="single">s</u><sub>0i</sub>The way to get.
Secondly, consider the optimal coding conditions, that is, the closest conditions. Give input every time<u style="single">x</u>And the weighted W', that is, when the frame is processed one by one, the distortion is calculated, even if E=W'(X-g1(<u style="single">s</u><sub>1i</sub>+<u style="single">s</u><sub>1j</sub>))∥<sup>2</sup>Minimized<u style="single">s</u><sub>01</sub>and<u style="single">s</u><sub>1i</sub>。
Basically, E is at g1 (0131),<u style="single">s</u><sub>0i</sub>(0i31) and<u style="single">s</u><sub>0j</sub>(0j31) all the combinations of the loop method, that is to find under 32×32×32=32768, to find the minimum value of E<u style="single">s</u><sub>0i</sub>,<u style="single">s</u><sub>1i</sub>Group. However, since this requires huge calculations, this embodiment searches for waveforms and gains in sequence. At the same time, circular search is used<u style="single">s</u><sub>0i</sub>,<u style="single">s</u><sub>1i</sub>combination. There are 32×32=1024<u style="single">s</u><sub>0i</sub>,<u style="single">s</u><sub>1i</sub>combination. In the following description, for simplicity,<u style="single">s</u><sub>1i</sub>+<u style="single">s</u><sub>1j</sub>by<u style="single">s</u><sub>m</sub>Express.
The above equation (27) becomes E=W'(<u style="single">x</u>-glsm)<sup>2</sup>. For further simplification,<u style="single">x</u><sub>w</sub>=W'<u style="single">x</u>and<u style="single">s</u><sub>w</sub>=W'<u style="single">s</u><sub>m</sub>, And got<maths><img file="TW360859B_D0064.tif" /></maths>
Therefore, if g1 is sufficiently correct, the search can be carried out in the following two steps: (1) Search<maths><img file="TW360859B_D0065.tif" /></maths>Minimized<u style="single">s</u><sub>w</sub>, And (2) search the closest<maths><img file="TW360859B_D0066.tif" /></maths>G<sub>1</sub>. If the above is rewritten using the original expression, (1)'Search<maths><img file="TW360859B_D0067.tif" /></maths>Minimized set<u style="single">s</u><sub>0i</sub>and<u style="single">s</u><sub>1i</sub>, And (2)'Search the closest<maths><img file="TW360859B_D0068.tif" /></maths>G<sub>1</sub>。
The above equation (35) represents an optimal coding condition (the closest condition).
The processing volume of vector quantization for codebook retrieval is described here.
K's<u style="single">s</u><sub>0i</sub>and<u style="single">s</u><sub>1i</sub>Size and L<sub>0</sub>With L<sub>1</sub>Code book CB0, CB1 size relationship is 0i<L<sub>0</sub>, 0j<L<sub>1</sub>, Addition, processing of the sum product and the square of the numerator, the volume is 1, and the sum product processing volume of the denominator is 1, the processing volume of equation (35) (1)' is approximately: numerator: L<sub>0</sub>. L<sub>1</sub>. (K. (H1)+1)
Denominator: L<sub>0</sub>. L<sub>1</sub>. (K. (H1))
Comparison scale: L<sub>0</sub>. L<sub>1</sub>And got L<sub>0</sub>. L<sub>1</sub>. The sum of (4K+2). If L<sub>0</sub>=L<sub>1</sub>=32, and K=44, the processing volume is 182272 ordinal.
In this way, all the (i) processing of equation (35) is not performed, but each vector is selected in advance<u style="single">s</u><sub>0i</sub>and<u style="single">s</u><sub>1i</sub>The P number. Since (or allowable) and gain input are not considered, look for (1)' of equation (35), so that the numerator value of (2)' of equation (35) is always positive. That is, to maximize (1)' of equation (35) and include<u style="single">x</u><sup>t</sup>W<sup>t</sup>W(<u style="single">s</u><sub>0i</sub>+<u style="single">s</u><sub>1i</sub>) Of the pole vector.
Here is a method as an illustrative example of this pre-selection method: (Procedure 1) Selection<u style="single">s</u><sub>0i</sub>P of<sub>0</sub>Number, calculated from the upper ordinal side, making<u style="single">x</u><sup>t</sup>W<sup>t</sup>W<u style="single">s</u><sub>0i</sub>Maximize; (procedure 2) select<u style="single">s</u><sub>1i</sub>P of<sub>1</sub>Number, calculated from the upper ordinal side, making<u style="single">x</u><sup>t</sup>W<sup>t</sup>W<u style="single">s</u><sub>1i</sub>Maximize; and (procedure 3) estimate<u style="single">s</u><sub>0i</sub>P of<sub>0</sub>Number and<u style="single">s</u><sub>1i</sub>P<sub>1</sub>(1)' equation of equation (35) for all combinations of numbers. exist<maths><img file="TW360859B_D0069.tif" /></maths>, That is, in the estimation of the square root of the equation (1) of equation (35), if no matter i or j, the denominator is<u style="single">s</u><sub>0i</sub>+<u style="single">s</u><sub>1i</sub>The weighted norm of is essentially a constant, and the above is valid. In fact, the scale of the denominator of equation (a1) is not constant. Next, the pre-selection method that takes this into consideration will be explained.
Here is an explanation of the effect of reducing the processing volume when the denominator of equation (a1) is assumed to be constant. Since L<sub>0</sub>. The processing volume of K needs to be used for the retrieval of (Program 1) and the processing volume of (L0-1)+(L0-2)+...+(L0-P0)=P0L0-P0(1+P0)/2 It needs to be used for scale comparison, so the total processing volume is L0(K+P0)-P0(1+P0)/2. Procedure 2 also requires the same processing volume. Add them together, the pre-selected processing volume is: L0(K+P0)+L1(K+P1)-P0(1+P0)/2-P1(1+P1)/2
Next, the ultimate selection program of program 3 will be explained, the numerator related to (1)' processing of equation (35): P0. P1. (1+K+1)
Denominator: P0. P1. K. (1+1)
Scale comparison: P0. In the case of P1, P0 is obtained. The total number of P1(3K+3).
For example, if P0=P1+6, L0=L1=32, and K=44, the final selected processing volume and the pre-selected volume are 4860 and 3158, respectively, resulting in an ordinal total of 8018. If the pre-selected number is increased to 10 and P0=P1=0, the final selected processing volume is 13,500, and the pre-selected volume is 3346, resulting in a total number of 16846.
If the number of vectors selected in advance is 10 per code book, the processing volume is 16846/182272, which is approximately one-tenth of the previous volume, compared with the calculation of 182272, which is not omitted.
At the same time, the value of the denominator of the equation (1)' of equation (35) is a non-constant value, and is changed according to the selected code vector. Here is an explanation of the pre-selection method that takes the approximate value of this norm into consideration for a certain procedure.
In order to find the maximum value of equation (a1), that is, the square root of (1)' equation of equation (35), because<maths><img file="TW360859B_D0070.tif" /></maths>Therefore, the left side of equation (a2) can be maximized. In this way, the left side is expanded to<maths><img file="TW360859B_D0071.tif" /></maths>Maximize the first and second terms.
Since the numerator of the first term of equation (a3) is just<u style="single">s</u><sub>0i</sub>Function, so the first term is relative to<u style="single">s</u><sub>0i</sub>maximize. On the other hand, since the numerator of the second term of equation (a3) is only<u style="single">s</u><sub>1j</sub>Function of, so the second term is relative to<u style="single">s</u><sub>1j</sub>maximize. That is, this method is embodied in<maths><img file="TW360859B_D0072.tif" /></maths>Including (program 1): self-maximizing equation (a4) vector upper ordinal selection<u style="single">s</u><sub>0i</sub>The number of Q0; (Procedure 2): Self-maximizing equation (a5) vector's upper ordinal selection<u style="single">s</u><sub>1j</sub>The number of Q1; and (Procedure 3): Estimate (1)' of equation (35) in the selected<u style="single">s</u><sub>0i</sub>All combinations of Q0 number and selected Q1 number.
At the same time, W'=WH/<u style="single">x</u>, W and H are input vectors<u style="single">x</u>Function, and W is of course the input vector<u style="single">x</u>The function.
Therefore, W should have been calculated from an input vector<u style="single">x</u>To another vector to calculate the denominator of equations (a4) and (a5). However, it is not advisable to excessively deplete the processing volume used for pre-selection. Therefore, these equal denominators use the typical or representative value of W', and each<u style="single">s</u><sub>0i</sub>and<u style="single">s</u><sub>1j</sub>Calculate and compare with<u style="single">s</u><sub>0i</sub>and<u style="single">s</u><sub>1j</sub>The values of are stored in the table together. At the same time, since the division in the actual retrieval processing means the loading in the processing, the values of equations (a6) and (a7) are:<maths><img file="TW360859B_D0073.tif" /></maths>Store it up. In the above equation, W* is obtained from the following equation (a8):<maths><img file="TW360859B_D0074.tif" /></maths>Where W<sub>K</sub>'W for a frame', the V/UV of this frame is audible after being observed, so that<maths><img file="TW360859B_D0075.tif" /></maths>
Figure 10 shows that W* is represented by the following equation (a10):<maths><img file="TW360859B_D0076.tif" /></maths>In this case, the specific examples of each of W[0] to W[43].
As for the numerators of equations (a4) and (a5), find W'and use it in an input vector<u style="single">x</u>To another vector. The reason is that at any speed,<u style="single">s</u><sub>0i</sub>,<u style="single">s</u><sub>1j</sub>and<u style="single">x</u>The inner product of must be calculated, so if<u style="single">x</u><sup>t</sup>W<sup>t</sup>W'is calculated, the processing volume will only slightly increase.
When roughly estimating the processing volume required by the preselected method, the processing volume of L0(K+1) is required for the retrieval of program 1, and Q0 is required for the numerical comparison. L0-Q0(1+Q0)/2 processing volume. The above procedure 2 also needs similar treatment. Add these processing volumes, the pre-selected processing volume is L0(K+Q0+1)+L1(K+Q1+1)-Q0(1+Q0)/2-Q1(1+Q1)/2
As far as the ultimate selection process of program 3 is concerned, the numerator: Q0. Q1. (1+K+1)
Denominator: Q0. Q1. K. (1+1)
Comparison value: Q0. Q1
The total is Q0. Q1(3K+3)
For example, if Q0=Q1=6, L0=L1=32, and K=44, the final selected processing volume and the pre-selected processing volume are 4860 and 3222, respectively, and the total is 8082 (8 ordinal values). If the number of pre-selected vectors is increased to 10, for example, Q0=Q1=10, the final selected processing volume and the pre-selected processing volume are 13500 and 3410, respectively, for a total of 16910 (8 ordinal values).
The ordinal value of these calculation results is the same as the processing volume of approximately 8018 or P0=P1=10 in the case of P0=P1=6 without normalization (that is, without dividing by the weighted norm).
For example, if the vector number of the individual codebook is set to P, the processing volume will be reduced by 16910/182272, where 182272 is the processing volume without omission. In this way, the processing volume is reduced to no more than one-tenth of the original processing volume.
Based on the specific example of the SNR (S/N ratio) in the preselected situation and the segment SNR of 20msec segment, and using the analyzed and synthesized speech without the above preselection as a reference, the SNR is 16.8dB and the segment SNR is 18.7 dB, and compared to the SNR without standardization and P0=P1=6, the SNR is 14.8dB and the segment SNR is 17.5dB, and the SNR with the same number of vectors as preselected and weighted and normalized is 17.8dB. The segment SNR is 19.8dB. That is, by using the right and standardized operations without standardized operations, the SNR and segment SNR increase by 2 to 3 dB.
Using the conditions (centroid condition) of equations (31) and (32) and the condition of equation (35), codebooks (CB0, CB1, and CBg) can be simultaneously used by the so-called general Loyd formula (GLA).
In this embodiment, a user input is used<u style="single">x</u>W'divided by the norm of is used as W'. That is, in equations (31), (32) and (35), W'/<u style="single">x</u>Replace W'.
Alternatively, the weight Wused for perceptual weighting when vector quantization is performed by the vector quantizer 116 is defined by the above equation (26). However, the weight W'that takes the temporary mask into consideration can also be obtained by calculating the current weight W', which has already taken the past W'into consideration.
At time n, wh(1), wh(2), ........., wh(L) of the above equation (26) are calculated as whn(1), whn(2),........., whn(L )To represent.
If the weekly value is taken into consideration, the weight at time N is defined as An(i), and 1iL, An(i)=λA<sub>n-1</sub>(i)+(1-λ)whn(i), (whn(i)<img file="TW360859B_D0077.tif" />A<sub>n-1</sub>(i))=whn(i), (whn(i)>A<sub>n-1</sub>(i)) Where λ can be set to, for example, λ=0.2. Among the An(i) found to be 1iL, a matrix with this An(i) as a diagonal element can be used as the above-mentioned weighting.
In this way, the waveform index value obtained by weighted vector quantization<u style="single">s</u><sub>0i</sub>,<u style="single">s</u><sub>1j</sub>They are respectively output at output terminals 520 and 522, and the gain index g1 is output at an output terminal 521. And, the quantized value<u style="single">x</u><sub>0</sub>'Is output at the output terminal 504 and sent to the adder 505.
Adder self-spectrum message envelope vector<u style="single">x</u>Subtract the quantized value to obtain a quantized error vector<u style="single">y</u>. Specifically, this quantized error vector<u style="single">y</u>Is sent to the vector quantization unit 511, and the vector quantizer 511<sub>1</sub>To 511<sub>8</sub>The size is divided and quantized by weighted vector quantization. The second vector quantization unit 510 uses a larger number of bits than the first vector quantization unit 500. Therefore, the memory capacity of the code book and the processing volume (complexity) of the code book retrieval are greatly increased. In this way, it is impossible to perform vector quantization with the same size of 44 as that of the first vector quantization unit 500. Therefore, the vector quantization unit 511 of the second vector quantization unit 510 is composed of a complex vector quantizer, and the input quantized value is divided into complex low-sized vectors in size to perform weighted vector quantization.
Used in vector quantizer 511<sub>1</sub>To 511<sub>8</sub>Quantized value<u style="single">y</u><sub>0</sub>to<u style="single">y</u><sub>7</sub>, The relationship between the size and the number of bits is shown in Figure 11.
Auto vector quantizer 511<sub>1</sub>To 511<sub>8</sub>Output index value Id<sub>vq0</sub>To Id<sub>vq7</sub>At output 523<sub>1</sub>To 523<sub>8</sub>Output. The bit sum of these index data is 72.
If the vector quantizer 511 is connected due to the dimension direction<sub>1</sub>To 511<sub>8</sub>Output quantized value<u style="single">y</u><sub>0</sub>'to<u style="single">y</u><sub>7</sub>'The value obtained is 1, the quantized value<u style="single">y</u>'and<u style="single">x</u><sub>0</sub>'That is, a quantized value is obtained by adding together by the adder 513<u style="single">x</u><sub>1</sub>'. Therefore, the quantized value is<u style="single">x</u><sub>1</sub>'=<u style="single">x</u><sub>0</sub>'+<u style="single">y</u>'=<u style="single">x</u>-<u style="single">y</u>+<u style="single">y</u>'To represent. That is, the final quantization error vector is<u style="single">y</u>'-<u style="single">y</u>。
If the quantized value from the second vector quantizer 510<u style="single">x</u>'To be decoded, the speech signal decoding device does not need the quantized value from the first quantization unit 500. Only the index data from the first quantization unit 500 and the second quantization unit 510 are needed.
The learning method and codebook search in the vector quantization unit 511 will be described later.
As for the learning method, the quantized error vector<u style="single">y</u>Use the weight W'to divide into eight low-dimensional vectors as shown in Figure 11<u style="single">y</u><sub>0</sub>to<u style="single">y</u><sub>7</sub>. If the weight W'is a matrix with 44 auxiliary sampling values as diagonal elements:<maths><img file="TW360859B_D0078.tif" /></maths>The weight W'is divided into the following eight matrices:<maths><img file="TW360859B_D0079.tif" /></maths><maths><img file="TW360859B_D0080.tif" /></maths>
So divided into low-dimensional<u style="single">y</u>The term with W'is Y respectively<sub><u style="single">i</u></sub>With W<sub>i</sub>', where 1i8.
The distortion value E is defined as E=W<sub>i</sub>'(<u style="single">y</u><sub>i</sub>-<u style="single">s</u>)∥<sup>2</sup>………(37)
Codebook vector<u style="single">s</u>for<u style="single">y</u><sub>i</sub>Quantify the results. Find the code vector of the codebook that maximizes the distortion value E.
In the codebook study, use the general-purpose Lorde formula (GLA) for further weighting. First, the optimal centroid condition for learning is explained. If there is an M input vector<u style="single">y</u>, Select the code vector s as the optimized quantization result, and the training data is<u style="single">y</u><sub>k</sub>, The distortion expected value J is obtained from equation (38), this equation minimizes the distortion center when weighted relative to all frames k:<maths><img file="TW360859B_D0081.tif" /></maths>untie<maths><img file="TW360859B_D0082.tif" /></maths>Get<maths><img file="TW360859B_D0083.tif" /></maths>Take the transpose value of both sides to get<maths><img file="TW360859B_D0084.tif" /></maths>therefore<maths><img file="TW360859B_D0085.tif" /></maths>
In the above equation (39),<u style="single">s</u>Is an optimized representative vector and represents an optimized centroid condition.
In terms of optimizing coding conditions, you can search to make W<sub>i</sub>'(<u style="single">y</u><sub>i</sub>-<u style="single">s</u>)∥<sup>2</sup>Minimize<u style="single">s</u>. W during the search<sub>i</sub>'No need to communicate with W during the study period<sub>i</sub>'Same, and can be an unweighted matrix:<maths><img file="TW360859B_D0086.tif" /></maths>
By configuring the vector quantization unit 116 in the speech signal encoder with a two-stage vector quantization unit, the number of output index bits can be made variable.
At the same time, the number of data composed of the harmonic spectrum obtained by the spectrum information package estimation unit 148 changes with the pitch, so that, for example, when the effective frequency band is 3400KHz, the number of data ranges from 8 to 63. Including the vector of such data together to form a block<u style="single">v</u>Department of variable size vector. In the above embodiment, before vector quantization, the size is converted into a preset data number, such as 44 size input vector<u style="single">x</u>. This variable/fixed dimensionality conversion means the above-mentioned data number conversion, and can be implemented using the above-mentioned supersampling and linear interpolation.
If the error processing is performed on the vector<u style="single">x</u>, Converted into a fixed-size codebook search to minimize this error, that is, there is no need to select a vector with a variable size relative to the original<u style="single">v</u>The code vector that minimizes this error.
In this embodiment, when selecting a code vector of a fixed dimension, a complex code vector is temporarily selected, and finally the ultimate optimized variable-size code vector is selected from these temporarily selected complex code vectors. At the same time, only variable size selection processing can be performed, and fixed size temporary selection is not performed.
Figure 12 shows the schematic structure of the original variable size optimization vector selection. A variable data number of the spectrum information packet obtained by the evaluation unit 148 is sealed, that is, a variable size vector<u style="single">v</u>Input to the input terminal 541. Like the above-mentioned data number conversion circuit, a variable/fixed size conversion circuit 542 inputs this variable size to the vector<u style="single">v</u>Input fixed size vector<u style="single">x</u>(For example, a 44-size vector composed of 44 data), which is sent to a terminal 501. Fixed size input vector<u style="single">x</u>The fixed-size code vector read from a fixed-size code book 530 is sent to a fixed-size selection circuit 535 for a selection-selection operation or code book search, and the code book 530 is selected from the code book 530 to minimize the weighting error or distortion. The code vector.
In the embodiment of FIG. 12, the fixed-size code vector obtained from the fixed-size codebook 530 is converted by a fixed/variable size conversion circuit 544 of the same variable size as the original size. The converted size code vector is sent to a variable vector conversion circuit 545 to calculate the code vector and the input vector<u style="single">v</u>The weighted distortion between the time and the selection process or the codebook search is performed, and the code vector that will reduce the distortion to the minimum is selected from the codebook 530.
That is, the fixed size selection circuit 535 minimizes the weighted distortion by temporarily selecting a number of code vectors as candidate code vectors, and performs weighted distortion calculations on these candidate code vectors in the variable size conversion circuit 545, so as to finally select the candidate code vectors. The code vector that minimizes distortion.
Here is a brief description of the application range of vector quantization using temporary selection and ultimate selection. This vector quantization can not only use the size conversion on the spectrum composition of the harmonics in the harmonic encoding, but also apply the weighted vector quantization of the variable-size harmonics. The remaining harmonic encodings of the LPC, such as this licensees early Japanese public license The multi-band excitation (MBE) encoding disclosed in Application 4-91422, or the remaining MBE encoding of LPC, can also be applied to vector quantization of variable-size input using a fixed-size codebook.
In terms of temporary selection, if the code book includes a waveform code book and a gain code book and the gain is determined by the variable distortion calculation, you can select a part of the multi-segment quantizer configuration or only search the waveform code book for temporary selection . Alternatively, the above-mentioned pre-selection may be used for temporary selection. Specifically, a fixed-size vector can be obtained by approximation (approximation of weighted distortion)<u style="single">x</u>Similarity with all code vectors stored in this codebook. In this case, the aforementioned pre-selection can be used to perform temporary fixed-size selection, and the final selection can be performed on the pre-selected candidate code vector to select the one that minimizes the weighted distortion of the variable size. Alternatively, before making the final selection, not only can a pre-selection be performed, but also a high-precision distortion calculation can be performed for precise selection.
Specific examples of vector quantization using temporary selection and ultimate selection are explained with reference to the diagrams.
In FIG. 12, the code book 530 is composed of a waveform code book 531 and a gain code book 532. The waveform code book 531 is composed of two code books CB0 and CB1. The output code vectors of these waveform code books CB0 and CB1 are marked with<u style="single">s</u><sub>0</sub>,<u style="single">s</u><sub>1</sub>, And the gain of the gain circuit 533 determined by the gain code book 532 is marked with g. Variable size input vector from an input 541<u style="single">v</u>A variable/fixed size circuit 542 is processed by size conversion (herein referred to as D1), and then used as a fixed size vector via terminal 501<u style="single">x</u>It is supplied to a subtractor 536 of a selection circuit 535, where the vector is obtained from the fixed-size code vector read from the code book 530<u style="single">x</u>The difference is weighted by a weighting circuit 536 to provide an error minimizing circuit 538. The weighting circuit 537 uses a weight value W'. The fixed size code vector read from the code book 530 is processed by the variable/fixed size selection circuit 544 by size conversion (referred to as D2), and is supplied to a selector 546 of a variable size selection circuit 545, where, Self-variable size input vector<u style="single">v</u>The difference of the code vector is obtained, weighted by a weighting circuit, and then supplied to an error minimizing circuit 548. The weighting circuit 537 uses a weight W<sub>v</sub>。
The error of the error minimizing circuit 538, 548 means the aforementioned distortion or distortion value. A decrease in error or distortion corresponds to an increase in similarity or correlation.
In essence, as explained with reference to equation (27), the selection circuit 535 that performs a fixed-size temporary selection will minimize the distortion value represented by equation (b1): E<sub>1</sub>=W'(<u style="single">x</u>-g(<u style="single">s</u><sub>0</sub>+<u style="single">s</u><sub>1</sub>))∥<sup>2</sup>.........(B1)
Please note that the weight W in the weighting circuit 537 is based on W'=WH/<u style="single">x</u>...(b2), where H refers to a matrix with the frequency response characteristic of LPC synthesis filter as the diagonal element, and W refers to a matrix with the frequency response characteristic of perceptually weighted filter as the diagonal element.
The first search will minimize the distortion value E of equation (b1)<sub>1</sub>of<u style="single">s</u><sub>0</sub>,<u style="single">s</u><sub>1</sub>, G. It should be noted that by temporarily selecting from the fixed size, starting from the upper order side, reduce the distortion value E<sub>1</sub>Sequence, find the L group<u style="single">s</u><sub>0</sub>,<u style="single">s</u><sub>1</sub>With g. Then in group L<u style="single">s</u><sub>0</sub>,<u style="single">s</u><sub>1</sub>Make the ultimate choice with g, minimize E2=W<sub>v</sub>(<u style="single">v</u>-D<sub>2</sub>g(<u style="single">s</u><sub>0</sub>+<u style="single">s</u><sub>1</sub>))∥<sup>2</sup>.........(B3) is used as the optimized code vector.
The search and study of equation (b1) can be explained by referring to equation (27) and the following equation.
Here is an explanation of the centroid condition for codebook learning according to equation (b3).
For code book CB0, which is one of wave code book 531 in code book 530, select a code vector from it<u style="single">s</u><sub>0</sub>, To minimize an expected value of distortion related to all frames k.
If these frames are M, they can be minimized<maths><img file="TW360859B_D0087.tif" /></maths>
To minimize equation (b4), solve equation (b5):<maths><img file="TW360859B_D0088.tif" /></maths>
In this equation (b6), O<sup>1</sup>Refers to the inverse matrix, and W<sub>VK</sub><sup>T</sup>Refers to W<sub>VK</sub>The transpose matrix. This equation (b6) represents the waveform vector<u style="single">s</u><sub>0</sub>Optimized centroid conditions.
Perform code vectoring on the code book CB1 of the other wave code book 531 in the code book 530 in the same way as above.<u style="single">s</u><sub>1</sub>Selection, so its description is omitted for conciseness.
Now consider the centroid condition of the gain g from the gain code book 532 in the code book 530.
The distortion of the k-th frame is expected to be worth from equation (7), and the codeword is selected from these k-th frames:<maths><img file="TW360859B_D0089.tif" /></maths>
To minimize equation (b7), solve the following equation (b8):<maths><img file="TW360859B_D0090.tif" /></maths>In order to<maths><img file="TW360859B_D0091.tif" /></maths>
This equation (b9) represents the centroid condition of gain.
Next consider the closest condition of equation (b3).
Retrieved due to equation (b3)<u style="single">s</u><sub>0</sub>,<u style="single">s</u><sub>1</sub>, The number of g groups is temporarily limited to L by a fixed size, so equation (b3) directly refers to L groups<u style="single">s</u><sub>0</sub>,<u style="single">s</u><sub>1</sub>Calculate with g, and choose the distortion E<sub>2</sub>, That is, optimized code vector, minimized<u style="single">s</u><sub>0</sub>,<u style="single">s</u><sub>1</sub>With group g.
Here is a description of a sequential search when the L used for temporary selection is extremely large, or if the temporary selection is not performed, it is directly selected in the variable size<u style="single">s</u><sub>0</sub>,<u style="single">s</u><sub>1</sub>With g, it is regarded as an effective waveform and gain method.
If the value of i, j and 1 are added to the equation (b3)<u style="single">s</u><sub>0</sub>,<u style="single">s</u><sub>1</sub>And g, and rewrite this form of (b3), you can get: E2=W<sub>v</sub>(<u style="single">v</u>-D<sub>2</sub>g<sub>1</sub>(<u style="single">s</u><sub>0i</sub>+<u style="single">s</u><sub>1j</sub>))∥<sup>2</sup>.........(B10)
Although minimizing g in equation (b10),<u style="single">s</u><sub>0i</sub>,<u style="single">s</u><sub>1j</sub>It can be searched in a circular manner, but if 01<32, 0i<32 and 0j<32, the above equation (b10) must be 32<sup>3</sup>=32768 type of calculation, which will lead to huge processing. Here is an explanation of the method of searching the waveform and gain in sequence.
Gain g<sub>1</sub>Waveform code vector<u style="single">s</u><sub>0i</sub>,<u style="single">s</u><sub>1j</sub>It was decided later. set up<u style="single">s</u><sub>0i</sub>+<u style="single">s</u><sub>1j</sub>=<u style="single">s</u><sub>m</sub>, Equation (b10) can be obtained by E<sub>2</sub>==W<sub>v</sub>(<u style="single">v</u>-D<sub>2</sub>g<sub>1</sub><u style="single">s</u><sub>m</sub>)∥<sup>2</sup>.........(B11) to represent.
If set<u style="single">v</u><sub>w</sub>=W<sub>v</sub><u style="single">v</u>,<u style="single">s</u><sub>w</sub>=W<sub>v</sub>D<sub>2</sub><u style="single">s</u><sub>m</sub>, Equation (b11) becomes<maths><img file="TW360859B_D0092.tif" /></maths>
Therefore, if g<sub>1</sub>Can be sufficiently precise, that is, search maximization<maths><img file="TW360859B_D0093.tif" /></maths>of<u style="single">s</u><sub>w</sub>Closest to<maths><img file="TW360859B_D0094.tif" /></maths>G<sub>1</sub>。
By replacing the original variables to rewrite equations (b13) and (b14), the following equations (b15) and (b16) are obtained.
Search maximization<maths><img file="TW360859B_D0095.tif" /></maths>of<u style="single">s</u><sub>0i</sub>,<u style="single">s</u><sub>1j</sub>Group and closest<maths><img file="TW360859B_D0096.tif" /></maths>G<sub>1</sub>。
Using equations (b6) and (b9) waveforms and gain centroid conditions and equations (b15) and (b16) optimized coding conditions (closest conditions), the code book (CB0, CB1, CBg) can be at the same time Simultaneous learning with the general-purpose Lowe-De-German program (GLA).
As mentioned above, compared to the method using equations (27), etc., especially equations (31), (32) and (35), the above equations (b6), (b9), b(15) are used, The learning method of b(16) is in the original input vector<u style="single">v</u>It is extremely superior in terms of minimizing distortion after converting to a variable size vector.
However, due to the extremely complicated processing of equations (b6) and (b9), especially (b6), the centroid condition can be used. This centroid condition comes from the optimization equation (b27), that is, (b1), only Use the closest conditions of equations (b15) and (b16).
It is also advisable to use the method described in equation (27) during the codebook study, and only use the method using equations (b15) and (b16) during retrieval. It is also possible to perform temporary selection in a fixed size by referring to the method described in equation (27), etc., and directly evaluate equation (b3) only for the selected plural number (L) group, for searching.
In any case, by performing a search by distortion evaluation according to equation (b3) after temporary selection or selection in a circular manner, learning or code vector search with less distortion can be performed.
I would like to briefly explain that it is best to calculate the distortion with the original input vector<u style="single">v</u>The same variable size reason.
If the minimum distortion of the fixed size is the same as that of the variable size, that is, there is no need to minimize the distortion of the variable size. However, since the size conversion D2 performed by the fixed/variable size conversion circuit 544 is not an orthogonal matrix, the two minimizations are inconsistent with each other. Accordingly, if the distortion is minimized to a fixed size, this minimization does not necessarily mean that the variable size of the distortion is minimized, so that if the obtained variable size vector is to be optimized, it is necessary to optimize the variable size of the distortion.
Figure 13 shows an example in which the gain when the code book is divided into a waveform code book and a gain code book is a variable gain, and the distortion is optimized to a variable size.
Specifically, the fixed-size code vector read from the code book 531 is sent to the fixed/variable conversion circuit to be converted into a variable-size vector, and then it is sent to the gain control circuit 533. The selection circuit 545 can select the optimal gain for the code vector in the gain circuit 533, which is based on the variable size code vector and the input vector from the gain control circuit 533<u style="single">v</u>It is processed by fixed/variable size conversion. Its structure and operation are the same as the embodiment shown in Figure 12 in other directions.
Turning to the waveform code book 531, in the variable size selection in the selection circuit, a single code vector can be selected, and the variable size selection can be performed only in terms of gain.
By multiplying the code vector converted by the fixed/variable size conversion circuit 544 with the gain, the fixed/variable size conversion method of the fixed/variable size conversion method of multiplying the code vector and the gain shown in Fig. 12 can be used. The optimal gain is selected under the effect of the size conversion.
A further specific example of the vector quantization combination of temporary selection of fixed size and final selection of variable size is described.
In the following specific example, the fixed-size first code vector read from the first code book is converted into a variable size of the input vector, and the fixed-size second code vector read from the second code book is added To the first code vector of variable size processed by the above-mentioned fixed/variable size conversion. From the sum of the code vectors obtained by the addition, the optimal code vector that minimizes the error in the input vector is selected from at least the second code book.
In the example in Figure 14, the fixed-size first code vector read from the first code book CB0<u style="single">s</u><sub>0</sub>It is sent to the fixed/variable size conversion circuit 544 to be converted into a variable size, which is equal to the input vector of the terminal 541<u style="single">v</u>Variable size. The fixed-size second code vector read from the second code book CB1 is sent to an adder 549 to be added to the variable-size code vector read from the fixed/variable size conversion circuit 544. The code vector and result of the adder are sent to the selection circuit, where the sum vector from the adder 549 is selected, and the input vector is selected to minimize<u style="single">v</u>The optimal code vector of the error. The second code book CB1 is applied to a range from the harmonic lower side of the input vector to the input vector of the code book CB1. The gain circuit 533 of the gain g is only provided between the first code book CB0 and the fixed/variable size conversion circuit 544. Since the structure is the same as that shown in Fig. 12 in other respects, the same parts are marked with the same reference numbers, and corresponding descriptions are omitted for brevity.
In this way, the remaining code vectors in the fixed size from the code book CB1 are added and read from the code book CB0 and converted into a variable size code vector, thereby adding them to the fixed size code vector from the code book CB1 Subtract the distortion caused by the fixed/variable size conversion.
A distortion E calculated by the selection circuit 545 shown in Fig. 14<sub>3</sub>Obtained by the following formula: E<sub>3</sub>=W<sub>v</sub>(<u style="single">v</u>-(D<sub>2</sub>g<u style="single">s</u><sub>0</sub>+<u style="single">s</u><sub>1</sub>))∥<sup>2</sup>.........(B17)
In the example of FIG. 15, the gain circuit 533 is arranged on the output side of the adder 549. In this way, the addition result of the code vector read from the first code book CB0 and converted by the fixed/variable size conversion circuit 544 and the code vector read from the code book CB1 is multiplied by the gain g. Since the gain of multiplying the code vector from CB0 is very similar to the gain of multiplying the code vector from codebook CB1, the common gain is used in the correction part (quantization of quantization error). The distortion calculated by the selection circuit 545 shown in Fig. 15 is obtained by the following formula: E<sub>4</sub>=W<sub>v</sub>(<u style="single">v</u>-g(D<sub>2</sub>g<u style="single">s</u><sub>0</sub>+<u style="single">s</u><sub>1</sub>))∥<sup>2</sup>.........(B18)
This example is the same as that shown in Fig. 14 in other ways, so the explanation is omitted for the sake of brevity.
In the example of Figure 16, not only a gain circuit 535A with a gain g is provided on the output side of the first code book CB0 shown in Figure 14, but also a gain circuit 533B with a gain g is provided on the second code book. The output side of CB1. The distortion calculated by the selection circuit 545 in Fig. 16 is equal to the distortion E shown in equation (b18)<sub>4</sub>. Since the configuration of the example in Fig. 16 is otherwise the same as that of the example in Fig. 14, corresponding descriptions are omitted for the sake of brevity.
Figure 17 shows an example n, where the first code book in Figure 14 is composed of two waveform code books CB0 and CB1. Code vectors from these waveform codebooks<u style="single">s</u><sub>0</sub>,<u style="single">s</u><sub>1</sub>The addition and the result of the sum is multiplied by the gain g by the gain circuit 533 before being sent to the fixed/variable size conversion circuit 544. The variable size code vector from the fixed/variable size conversion circuit 544 and the code vector from the second code book CB2<u style="single">s</u><sub>2</sub>It is added by an adder 549 before being sent to the selection circuit 545. The desired distortion E of the selection circuit 545 in Fig. 17<sub>5</sub>Obtained by the following formula: E<sub>5</sub>=W<sub>v</sub>(<u style="single">v</u>-(gD<sub>2</sub>(<u style="single">s</u><sub>0</sub>+<u style="single">s</u><sub>1</sub>)+s<sub>2</sub>))∥<sup>2</sup>.........(B19)
The configuration of the example in Fig. 17 is the same as that of the example in Fig. 14, so the corresponding description is omitted for the sake of brevity.
Here is an explanation of the search method in equation (b18).
An example of the first search method includes the search minimization E<sub>4</sub>'=W'(<u style="single">x</u>-g<sub>1</sub><u style="single">s</u><sub>0i</sub>))∥<sup>2</sup>.........(B20)<u style="single">s</u><sub>0i</sub>, G<sub>1</sub>, And then retrieve to minimize E<sub>4</sub>=W<sub>v</sub>(<u style="single">v</u>-g<sub>1</sub>(D<sub>2</sub><u style="single">s</u><sub>0i</sub>+<u style="single">s</u><sub>0i</sub>))∥<sup>2</sup>.........(B21)
In another instance, the retrieval is maximized<maths><img file="TW360859B_D0097.tif" /></maths>of<u style="single">s</u><sub>0i</sub>To maximize retrieval<maths><img file="TW360859B_D0098.tif" /></maths>of<u style="single">s</u><sub>1j</sub>And retrieve the closest<maths><img file="TW360859B_D0099.tif" /></maths>G<sub>1</sub>。
In the third retrieval method, the retrieval minimizes E<sub>4</sub>'=W(<u style="single">x</u>-g<sub>1</sub><u style="single">s</u><sub>0i</sub>)∥<sup>2</sup>.........(B25)<u style="single">s</u><sub>0i</sub>With g<sub>1</sub>To maximize retrieval<maths><img file="TW360859B_D0100.tif" /></maths>of<u style="single">s</u><sub>1j</sub>And at the end retrieve the closest<maths><img file="TW360859B_D0101.tif" /></maths>G<sub>1</sub>。
Next, the centroid condition of the equation (b20) of the first retrieval method will be explained. Borrowing code vector<u style="single">s</u><sub>0i</sub>Centroid<u style="single">s</u><sub>0c</sub>,will<maths><img file="TW360859B_D0102.tif" /></maths>minimize. To minimize, solve<maths><img file="TW360859B_D0103.tif" /></maths>In order to get<maths><img file="TW360859B_D0104.tif" /></maths>Similarly, in order to gain the centroid g of g<sub>c</sub>, From the above equation (b20), the solution<maths><img file="TW360859B_D0105.tif" /></maths>In order to get<maths><img file="TW360859B_D0106.tif" /></maths>
On the other hand, is a vector<u style="single">s</u><sub>1j</sub>Of the centroid, the solution is the centroid condition of the first retrieval method<maths><img file="TW360859B_D0107.tif" /></maths>and<maths><img file="TW360859B_D0108.tif" /></maths>And got<maths><img file="TW360859B_D0109.tif" /></maths>From equation (b21), find the vector<u style="single">s</u><sub>0i</sub>Heart shape<u style="single">s</u><sub>0c</sub>And got<maths><img file="TW360859B_D0110.tif" /></maths>and<maths><img file="TW360859B_D0111.tif" /></maths>Similarly, the centroid g of the gain g<sub>c</sub>It can be obtained by the following formula:<maths><img file="TW360859B_D0112.tif" /></maths>
Calculate the code vector by the above equation (b20)<u style="single">s</u><sub>0i</sub>Centroid method and centroid g for calculating gain<sub>c</sub>The method is represented by equation (b33). As far as the centroid is calculated using equation (b21), the vector<u style="single">s</u><sub>1j</sub>Centroid<u style="single">s</u><sub>1c</sub>,vector<u style="single">s</u><sub>1j</sub>Centroid<u style="single">s</u><sub>1c</sub>Centroid g with gain g<sub>c</sub>They are represented by equations (b36), (b39) and (b40) respectively.
When actually learning the codebook with LGA, you can use equations (b30), (b36) and (b40) to learn at the same time<u style="single">s</u><sub>0</sub>,<u style="single">s</u><sub>1</sub>, G's method. Please note that the above equations (b22), (b23) and (b24) can be used in the search method (closest condition). In addition, various combinations of centroid conditions shown in equations (b30), (b33), (b36), (b39) or (b40) can be selectively used.
The method of searching the distortion value corresponding to the equation (b17) in Fig. 14 will be explained. In this case, retrievable will minimize E<sub>3</sub>'=W'(<u style="single">x</u>-g<sub>1</sub><u style="single">s</u><sub>0i</sub>))∥<sup>2</sup>.........(B41)<u style="single">s</u><sub>0i</sub>, G<sub>1</sub>, And then search will minimize E<sub>3</sub>=W<sub>v</sub>(<u style="single">x</u>-g<sub>1</sub>(D<sub>2</sub><u style="single">s</u><sub>0i</sub>+<u style="single">s</u><sub>1j</sub>))∥<sup>2</sup>.........(B42)<u style="single">s</u><sub>1j</sub>。
In the above equation (b41), select all g<sub>1</sub>and<u style="single">s</u><sub>0i</sub>The group is not practical, so the retrieval one minimizes<maths><img file="TW360859B_D0113.tif" /></maths>The number of L gains,<maths><img file="TW360859B_D0114.tif" /></maths>And minimize<i>E</i><sub>3</sub>=∥<i>W</i><sub><i>v</i></sub>(<i><u style="single">v</u>-</i>(<i>D</i><sub>2</sub><i>g<u style="single">s</u></i><sub>0<i>i</i></sub>+<i><u style="single">s</u></i><sub>1<i>j</i></sub>))∥<sup>2</sup>.........(B45)<u style="single">s</u><sub>1j</sub>。
Secondly, the centroid conditions are derived from equations (b41) and (b42). In this case, the procedure differs depending on the equation used.
First, if equation (b41) is used, and the code vector<u style="single">s</u><sub>0i</sub>The centroid is<u style="single">s</u><sub>0c</sub>, Which is to minimize<maths><img file="TW360859B_D0115.tif" /></maths>And got<maths><img file="TW360859B_D0116.tif" /></maths>
Similarly, just the centroid g<sub>c</sub>In terms of equation (b43), use equation (b41) to find<maths><img file="TW360859B_D0117.tif" /></maths>
If vector<u style="single">s</u><sub>1j</sub>Heart shape<u style="single">s</u><sub>1c</sub>Use equation (b42) to find, that is, the solution<maths><img file="TW360859B_D0118.tif" /></maths>and<maths><img file="TW360859B_D0119.tif" /></maths>And got<maths><img file="TW360859B_D0120.tif" /></maths>
Similarly, the code vector<u style="single">s</u><sub>0i</sub>Heart shape<u style="single">s</u><sub>0c</sub>Centroid g with gain g<sub>c</sub>, Can be obtained by equation b(42).
<maths><img file="TW360859B_D0121.tif" /></maths>
<maths><img file="TW360859B_D0122.tif" /></maths>
At the same time, you can use the above equations (b47), (b48) or (b51), or use the above equations (b51), (b52) or (b55) to perform codebook learning by GLA.
The second encoding unit 120 using the CELP encoding configuration of the present invention has a multi-stage vector quantization processing section (the two-stage encoding section 120 in the embodiment in FIG. 18<sub>1</sub>With 120<sub>2</sub>). If the transmission bit rate can be switched between, for example, 2kbps and 6kbps, the configuration in Figure 18 is designed to match the transmission bit rate of 6kbps, and the waveform and gain are changed between 23bit/5msec and 15bit/5msec Index output. The processing flow of the configuration in Figure 18 is shown in Figure 19.
Please refer to Fig. 18. A first encoding unit 300 in Fig. 18 corresponds to the first encoding unit 113 in Fig. 3, and an LPC analysis circuit 302 in Fig. 18 corresponds to the LPC analysis circuit 132 shown in Fig. 3. An LSP parameter quantization circuit 303 corresponds to the configuration from α to LSP conversion circuit 133 to LSP to α conversion circuit 137 in Figure 3, and the perceptual weighting filter 304 in Figure 18 corresponds to the perceptual weighting filter calculation in Figure 3. Circuit 139 and perceptual weighting filter 125. Therefore, in Figure 18, an output that is the same as the output of the LSP-to-α conversion circuit 137 of the first encoding unit 113 in Figure 3 is supplied to a terminal 305, and an output that is the same as the perceptual weighted filter calculation circuit 139 in Figure 3 The same output as the output of is supplied to a terminal 307, and an output identical to the output of the perceptual weighting filter 125 in FIG. 3 is supplied to the terminal 306. However, unlike the perceptual filter 125, the perceptual weighted filter 304 in Figure 18 uses the input voice data and pre-quantized α parameters instead of using the output of the LSP-α conversion circuit 137 to generate the same value as the perceptual weighted filter 125 in Figure 3. The output signal is the same perceptually weighted signal.
The second coding unit 120 shown in the second section of Fig. 18<sub>1</sub>With 120<sub>2</sub>Among them, the subtractors 313 and 323 correspond to the subtractor 123 in FIG. 3, and the distance calculation circuits 314 and 324 correspond to the distance calculation circuit 124 in FIG. 3. In addition, the gain circuits 311 and 321 correspond to the gain circuit 126 in FIG. 3, and the random code books 310 and 320 and the code books 315 and 325 correspond to the noise code book 121 in FIG.
In the structure of Fig. 18, the LPC analysis circuit 302 in step S1 in Fig. 19 will supply the input voice data from a terminal 301<u style="single">x</u>Divide the frame into the above-mentioned frame for LPC analysis to obtain an α parameter. The LSP parameter quantization circuit 303 converts the α parameter from the LPC analysis circuit 302 into LSP parameters to quantify the LSP parameters. The quantized LSP parameters are inserted and converted into α-parameters. The LSP parameter quantization circuit 303 self-converts the α-parameter from the quantized LSP parameter, that is, the quantized LSP parameter generates an LPC synthesis filter function 1/H(z), passes through the terminal 305, and combines the generated LPC synthesis filter function 1/H (z) Send to the first section of the second encoding unit 120<sub>1</sub>One is perceptually weighted synthesis filter 312. Perceptually weighted filter 304 is free as shown in Figure 19 S<sub>2</sub>The input speech data and the pre-quantized α-parameters shown in the step generate the same perceptual weighted signal as that output by the perceptual weighting filter 125 shown in FIG. 3. That is, first, the LPC synthesis filter function W(z) is generated from the pre-quantized α-parameters. The resulting filter function W(z) is applied to the input speech data<u style="single">x</u>To produce<u style="single">x</u><sub>w</sub>, Send this perceptual weighted signal to the first section of the second encoding unit 120 via the terminal 306<sub>1</sub>The subtractor 313.
In the first paragraph, the second encoding unit 120<sub>1</sub>Among them, the representative output value of the random code book 310 output by the 9-bit waveform index is sent to the gain circuit 311, which combines the representative output from the random code book 310 and the gain code book 315 from the 6-bit gain index output. Multiply the gain (scalar). The representative output value multiplied by the gain circuit 311 and the gain is sent to the perceptually weighted synthesis filter at 1/A(z)=(1/H(z))*W(z). As shown in step S3 in FIG. 19, the weighted synthesis filter 312 sends the 1/A(z) zero input reaction output to the subtractor 313. The subtractor 313 reacts to the zero input of the perceptually weighted synthesis filter 312 and outputs the perceptually weighted signal from the perceptually weighted filter 304<u style="single">x</u><sub>w</sub>Perform subtraction and take out the difference or error as a reference vector. In the first paragraph, the second encoding unit 120<sub>1</sub>During the search, as shown in step S4 in Figure 19, this reference vector is sent to the distance calculation circuit 314, where the distance is calculated and the waveform vector that minimizes the energy E of the quantization error is searched<u style="single">s</u>And gain g. Here, 1/A(z) is in the zero state. That is, if the waveform vector synthesized in Z state 1/A(z) in the codebook<u style="single">s</u>for<u style="single">s</u><sub>syn</sub>, That is, search to minimize the waveform vector of equation (40)<u style="single">s</u>And gain g:<maths><img file="TW360859B_D0123.tif" /></maths>
Minimize quantization error energy E<u style="single">s</u>And g solid can be fully searched, but the following methods can be used to reduce the amount of calculation.
The first method is used to retrieve the minimum energy E<sub>3</sub>Wave vector<u style="single">s</u>It is defined by the following equation (41):<maths><img file="TW360859B_D0124.tif" /></maths>
According to the first method<u style="single">s</u>, The ideal gain is expressed as equation (42):<maths><img file="TW360859B_D0125.tif" /></maths>
Therefore, according to the second method, search to minimize the g of equation (43): Eg=(g<sub>ref</sub>-g)<sup>2</sup>………(43)
Since E is a quadratic function of g, minimizing g of Eg minimizes E.
According to the s and g obtained by the first and second methods, the quantized error vector<u style="single">e</u>It can be calculated by the following equation (44):<u style="single">e</u>=<u style="single">r</u>-g<u style="single">s</u><sub>syn</sub>………(44)
This is the same as in the first stage, quantized as the second stage second coding unit 120<sub>2</sub>The reference.
That is, the signals supplied to the terminals 305 and 307 are directly from the first stage of the second encoding unit 102<sub>1</sub>The perceptually weighted synthesis filter 312 is supplied to the second stage of the second encoding unit 120<sub>2</sub>The perceptually weighted synthesis filter 322. The first section of the second coding unit 120<sub>1</sub>Quantization error vector<u style="single">e</u>Supplied to the second segment of the second encoding unit 120<sub>2</sub>One of the subtractors.
Step S5 in Figure 19 is performed similarly to the second encoding unit 120 in the second stage of the first step.<sub>2</sub>Processing. That is, a representative output value of the random code book 320 from the 5-bit waveform index output is sent to the gain circuit 321, where the representative output value of the code book 320 and the gain code book from the 3-bit gain index output Multiply the gain of 325. One output of the weighted synthesis filter 322 is sent to the subtractor, where the output of the perceptual weighted synthesis filter 322 and the first-stage quantization error vector are found<u style="single">e</u>The difference between. This difference is sent to a distance calculation circuit 324 to calculate the distance to search for the waveform vector that minimizes the energy E of the quantization error<u style="single">s</u>With gain g
Waveform index output of random code book 310, the first segment of the second encoding unit 120<sub>1</sub>The gain index output of the gain code book 315, the index output of the random code book 320 and the second segment of the second encoding unit 120<sub>2</sub>The index output of the gain code book 325 is sent to an index output switching circuit 330. If 23 bits are output from the second encoding unit 120, that is, the first segment and the second segment of the second encoding unit 120 are totaled<sub>1</sub>,120<sub>2</sub>The gain code book 315, 325 and the random code book 310, 320 index data, and output it. If the output is 15 bits, the first segment of the second encoding unit 120 is output<sub>1</sub>Index data of the random code book 310 and the gain code book 315.
Then, as shown in step S6, the filtering state is updated to calculate the zero input response output.
In this embodiment, the second segment second encoding unit 120<sub>2</sub>The number of index bits is 5 for the waveform vector and 3 for the gain. If the appropriate waveform and gain do not appear in the code book in this case, the quantization error may increase rather than decrease.
Although O can be used for gain to prevent this problem from occurring, only three bits are used for gain. If one of them is set to 0, the quantizer performance will be significantly deteriorated. For this purpose, we provide all-O vectors to waveform vectors with a larger number of bits. The above search is performed without the all-O vector, and if the quantization error has increased in the end, the all-O vector is selected. The gain is random. In this way, the quantization error can be prevented from being in the second stage of the second coding unit 120<sub>2</sub>Increase.
For the two-stage configuration, please refer to Figure 18 for the description above, but the number of stages can be greater than 2. In this case, if the vector quantization of the first segment of the closed loop search is completed, that is, the quantization error of the (N-1) segment is used as the reference input to perform the quantization of the Nth segment under 2N, and the quantization of the Nth segment The quantization error is used as the reference input for the (N+1) section.
It can be seen from Figures 18 and 19 that by using a multi-segment vector quantizer in the second coding unit, compared to using direct vector quantization with the same number of bits or using a conjugate codebook, it increases the number of calculations. In particular, in CELP encoding, it is important to reduce the number of retrieval operations. In this CELP encoding, vector quantization of the time-axis waveform is performed, and the closed loop retrieval used is performed by a comprehensive analysis method. In addition, by using two segments of the second encoding unit 120<sub>1</sub>,120<sub>2</sub>The output of the two is converted between the first-stage second coding unit 120 without the second-stage second coding unit, and the number of bits can be converted smoothly. If the first segment and the second segment of the second encoding unit 120<sub>1</sub>,120<sub>2</sub>Combine and output, the decoder can smoothly cooperate with the configuration by selecting one of the index outputs. That is, using a pair operating at 2 kbps to decode parameters encoded at, for example, 6 kbps, the decoder can smoothly cooperate with the configuration. In addition, if the second segment of the second encoding unit 120<sub>2</sub>The waveform codebook contains O vector, that is, compared to the case where O is added to the gain, it can prevent the increase of quantization error, and there is less performance degradation.
The code vector (waveform vector) of the random code book can be generated, for example, by the following method.
For example, the code vector of the random code book can be generated by cutting the so-called Gaussian noise. Specifically, the codebook can be generated by generating Gaussian noise, cutting the Gaussian noise with an appropriate threshold, and standardizing the cut Gaussian noise.
However, there are various types of voices. For example, Gaussian noise can handle nearly noisy consonants, such as "sa, shi, su, se, and so", but cannot handle sharply elevated consonants, such as "pa, pi, pu, pe, and po".
According to the present invention, Gaussian noise is applied to certain code vectors, and the remaining part is processed by learning, so that both consonants with sharp raised consonants and consonants that are nearly noisy can be processed. For example, if the threshold is increased, the obtained vector will have several larger peaks, and if the threshold is decreased, the code vector will approximate Gaussian noise. In this way, by increasing the change of the cutting threshold, it is possible to deal with consonants with sharp elevations, such as "pa, pi, pu, pe, and po", or consonants that are nearly noisy, such as "sa, shi, su, se, so", in vain and increase clarity. Figure 20 shows Gaussian noise and cut noise with solid lines and dashed lines, respectively. Figures 20A and 20B show that the cutting threshold is equal to 1.0, that is, the noise with a larger threshold, and the cutting threshold is equal to 0.4, that is, the noise with a smaller threshold. It can be seen from Figures 20A and 20B that if a larger threshold is selected, a vector with several larger peaks can be obtained, and if a smaller threshold is selected, the noise is close to the Gaussian noise itself.
In order to achieve this, a preliminary codebook is prepared by cutting Gaussian noise and an appropriate number of non-learning code vectors are set. The non-learning code vectors are selected according to the increasing order of the difference value to deal with consonants that are close to noise such as "sa, shi, su, se and so". The vectors obtained by learning use the LBG program to learn. The encoding performed under the closest conditions uses a fixed code vector and a code vector derived from learning. Under the waveform condition, only the code vector to be learned is updated. In this way, the code vector to be learned can handle sharply elevated consonants, such as "pa, pi, pu, pe, and po".
For these code vectors, an optimal gain can be learned by general learning.
Figure 21 shows the processing flow of forming a codebook by cutting Gaussian noise.
In Fig. 21, in step S10, the number of learning times n is set to n=0 to start. Error D<sub>0</sub>=, set the maximum number of learning n<sub>max</sub>, And set a threshold ε with learning termination conditions.
In the second step S11, a preliminary codebook is generated by cutting the Gaussian noise. In step 12, some code vectors are fixed as non-learning code vectors.
In the next step S13, the codebook is used for encoding. Calculate the error in step S14. In step S15, it is judged whether (D<sub>n-1</sub>-D<sub>n</sub>/D<sub>n</sub><ε, or n=n<sub>max</sub>). If the result is "Yes", the procedure is terminated. If the result is "No", the procedure goes to S16.
In step S16, the code vector not used for encoding is processed. In the next step S17, the code book is updated. In step 18, increase the number of learning n times before returning to step 13.
A specific example of a voice/unvoiced (V/UV) frequency discrimination unit 115 is described in FIG. 3.
The V/UV discrimination unit 115 performs V/UV discrimination of a frame. The main body of this frame is the output of one of the orthogonal transform circuits 145, and an optimal pitch from the high-precision audio search list, from the spectrum evaluation unit 148 The spectrum amplitude data is a maximum normalized autocorrelation value r(p) from the open-loop pitch retrieval unit 141 and a zero-crossing calculation value from the zero-crossing computer 412. The boundary position of the band base result similar to the V/UV judgment used by MBE is also used as a condition for the main body of the frame.
Here is an explanation of the V/UV discrimination conditions for MBE using the band-based V/UV discrimination results.
The parameter or amplitude representing the value of the m-th harmonic under MBE<maths><img file="TW360859B_D0126.tif" /></maths>To represent.
In this equation, S(j) is the spectrum obtained from the remaining DFF of the LPC, and E(j) is the spectrum of the basic signal, specifically 256-point Hamming window, a<sub>m</sub>, B<sub>m</sub>It is the low and high values corresponding to the frequency of the m-th band and represented by an index j, and this m-th band corresponds to the m-th harmonic. To perform band-based V/UV discrimination, a noise-to-signal ratio (NSR) is used. M-th band<maths><img file="TW360859B_D0127.tif" /></maths>
If the NSR value is greater than the repeatedly set threshold, such as 0.3, that is, if the error is large, it can be determined that the approximation between S(j) and AmE(j) is not good. That is, the E(j) excitation signal is not suitable as a base. In this way, the main body of the band is judged to be silent (UV). If on the contrary, it can be judged that the approximation is excellent, it is judged to be sound (V).
It should be noted that the NSR of individual bands (harmonics) represents the approximation between harmonics. The sum of gain weighted harmonics of NSR is NSR<sub>all</sub>=(Σ<sub>m</sub>A<sub>m</sub>NSR<sub>m</sub>)/(Σ<sub>m</sub>A<sub>m</sub>) Defined as NSR<sub>all</sub>。
The basic algorithm used in V/UV frequency discrimination is based on this spectrum approximation NSR<sub>all</sub>It depends on whether it is greater than or less than a certain threshold. Here, the threshold is set to Th<sub>NSR</sub>= 0.3. This basic algorithm is related to the maximum remaining autocorrelation of LPC, and the frame of the frame is related to the zero-crossing. For NSR<sub>all</sub><Th<sub>NSR</sub>In terms of the algorithm, if the algorithm is applicable or if the algorithm is not applicable, the main body of the frame becomes V and UV respectively.
A specific algorithm is as follows: in NSR<sub>all</sub><Th<sub>NSR</sub>In this case, if numZero XP<24, frmPow>340 and r0>0.32, the frame body is V; in NSR<sub>all</sub>TH<sub>NSR</sub>In this case, if numZero XP>30, frmPow<900 and r0>0.23, the main body of the frame is UV; among them, the individual variables are defined as follows: numZero XP: the number of zero crossings per frame
frmPow: frame screen
r0: maximum autocorrelation
Perform V/UV frequency discrimination with reference to the algorithm representing the above-mentioned set of specific algorithms.
Here is a more detailed description of the structure of the basic part shown in Figure 4 and the operation of the speech signal decoder.
In the inverse vector quantizer 212 of the spectral information envelope, an inverse vector quantizer configuration corresponding to the vector quantizer of the speech encoder is used.
For example, if the configuration shown in Figure 12 is used for vector quantization, the code vector is read out by the party making the decision.<u style="single">s</u><sub>0</sub>,<u style="single">s</u><sub>1</sub>, And from the waveform code book CB0, CB1 and gain code book DB<sub>g</sub>Read the gain g and take it out as a g(<u style="single">s</u><sub>0</sub>+<u style="single">s</u><sub>1</sub>) A vector of fixed size is converted into a variable size vector corresponding to the vector size of the original harmonic spectrum (fixed/variable size conversion).
As shown in Figures 14 to 17, the encoder has a configuration of a vector quantizer that adds fixed-size code vectors to variable-size code vectors, read from the variable-size code book (CB0 code book in Figure 14) The code vector is converted in a fixed/variable size method and added to the fixed size code vector read from the fixed size code book (code book CB1 in Figure 14). This fixed size corresponds to the magnitude of the low range of the harmonics. . Then take out the sum.
As mentioned above, the LPC synthesis filter 214 in FIG. 4 is divided into a voiced speech (V) synthesis filter 236 and an unvoiced speech (UV) filter 237. If LSPs continue to be inserted every 20 samples, that is, every 2.5msec, without separate synthesis filtering, and no V/UV distinction, LSPs with completely different properties are inserted at the transition from V to UV or UV to V. As a result, the LPC of UV and V are used as the surplus of V and UV respectively, so that weird sounds will be produced. In order to avoid such adverse effects, the LPC synthesis filter is divided into V and UV, and the LPC coefficient interpolation is independently used for V and UV.
The coefficient interpolation method of the LPC filters 236 and 237 in this case will be explained. Specifically, as shown in Figure 22, the LSP interpolation is transformed according to the V/UV state.
Take 10 ordinal LPC analysis as an example. The equally spaced LSP in Figure 22 is an LSP equivalent to the α-parameter used in the smoothing band filter characteristics and the gain is consistent, that is, α<sub>0</sub>=1, α<sub>1</sub>=α<sub>2</sub>=.........=α<sub>10</sub>=0, and 0α10.
This type of 10-order LPC analysis, that is, the 10-order LSP is a kind of LSP equivalent to the full flat wave band spectrum as shown in Figure 23. The LSPs are arranged at 11 equidistant positions between 0 and π at equal intervals. In this case, the entire band gain of the synthesis filter has a minimum flux characteristic at the moment.
Figure 24 schematically illustrates how the gain is changed. Specifically, Figure 15 shows 1/H<sub>uv(z)</sub>Gain and 1/H<sub>v(z)</sub>How does the gain change during the transition from the unvoiced (UV) part to the voiced (V) part.
As far as the unit of interpolation is concerned, in 1/H<sub>v(z)</sub>In the case of the coefficient, it is 2.5msec (20 samples), and at 1/H<sub>uv(z)</sub>In this case, the bit rate relative to 2kbps is 10msec (80 samples), and the bit rate relative to 6kbps is 5msec (40 samples). In terms of UV, since the second encoding unit 120 uses the analysis and synthesis method to perform waveform matching, the LSPs adjacent to the V portion can be interpolated instead of the LSPs at equal intervals. It should be noted that in the encoding of the UV part in the second encoding part 120, the zero input response is set to 0 by clearing the internal state of the synthesis filter 122 weighted with 1/A(z) in the transition part from V to UV.
The output of the LPC synthesis filters 236, 237 are sent to the post filters 238u, 238v that are set individually and independently. The intensity and frequency response of the LPC synthesis filter are set to different values for V and UV, and the intensity and frequency response of the post-filtering are set to have different values for V and UV.
Here is an explanation of the windowing of the boundary between the V and UV parts of the LPC residual signal, that is, the excitation of the LPC synthesis filter input. The window is opened by the sine synthesis circuit 223 of the voiced speech synthesis unit 211 and the windowing circuit 223 of the silent speech synthesis unit 220 shown in FIG. 4. The method for the synthesis of part V of the incentive is detailed in the Japanese Patent Application No. 4-91422 filed by the licensee, and the rapid synthesis method of the incentive is detailed in the Japanese Patent Application No. 4-91422 filed by the licensee. No. 6-198451. In the illustrated embodiment, this fast synthesis method is used to generate the V part excitation generated by this fast synthesis method.
In the voiced (V) part, sine synthesis is performed by interpolation using adjacent frames. As shown in Figure 25, all waveforms between the nth and (n+1)th frames can be generated. However, as shown in Figure 25, the signal part on both sides of the V and UV parts of the (n+1)th frame and (n+2)th frame, or on both sides of the UV part and the V part For part of the signal part, the UV part only encodes or decodes ±80 samples (the total number of 160 samples is equal to one frame interval). As a result, as shown in Figure 26, the windowing is performed outside the midpoint CN between adjacent frames on the V side, while on the UV side, it is performed at the midpoint CN, so as to coincide with the boundary portion. The opposite procedure is used for the transition from UV to V. The window on the V side can also be indicated by a dotted line as shown in Figure 26.
Here is an explanation of the noise synthesis and noise addition in the voice (V) part. These operations are performed by noise synthesis circuit 216, overlap and add circuit 217, and adder 218 in Figure 4 by adding noise to the remaining voiced part of the LPC. Part of the incentive parameters are taken into consideration.
That is, the above parameters enumerate one by one with the pitch delay Pch, the spectral amplitude Am[i] of the voiced sound, in the frame A<sub>max</sub>The maximum spectral amplitude in and the remaining signal level Lev. The pitch delay Pch is the number of samples relative to a preset sampling frequency fs in a pitch period, for example, fs=8kHz, and i in the spectral amplitude Am[i] is an integer, so that when 0<i<1, fs The number of harmonics in the /2 band is equal to I=Pch/2.
The processing of the noise synthesis circuit 216 is performed in the same manner as the silent voice synthesis such as multi-band coding (MBE). Figure 27 shows a specific embodiment of the noise synthesis circuit.
That is, referring to Fig. 27, a white noise 401 outputs Gaussian noise, and this Gaussian noise is then processed by a short-term Fourier transform (STFT) by an STFT processor 402 to generate a signal on the frequency axis. The power spectrum of the noise. The Gaussian noise is a time-domain white noise signal waveform, which is windowed by an appropriate windowing function such as a Hamming window with a preset length of, for example, 256 samples. The power spectrum from the STFT processor 402 is sent to a multiplier 403 for amplitude processing, and then multiplied by the output of the noise amplitude control circuit 410. One output of the amplifier 403 is sent to an inverse STFT (ISTFT) processor 404, which uses the phase of the original white noise as the phase for conversion into a time domain signal to ISTFT. An output of the ISTFT processor 404 is sent to a weighted overlap and add circuit 217.
In the embodiment of FIG. 27, the time-domain noise is generated from the white noise generator 401, and is processed by orthogonal transformation such as STFT to generate the frequency-domain noise. Alternatively, the frequency domain noise can also be directly generated by the noise generator. By directly generating frequency-domain noise, orthogonal transformation processing operations such as those used for STFT or ISTFT can be omitted.
Specifically, a method can be used to generate any number within a ±x range, and the generated number can be treated as the real and imaginary part of the FFT spectrum, or a method can be used to generate from 0 to a maximum number (max ), to treat it as the amplitude of the FFT spectrum, and generate any number from -π to +π, and treat it as the bit phase of the FFT spectrum.
In this way, the STFT processor 402 in FIG. 27 can be omitted to simplify the structure or reduce the processing volume.
The noise amplitude control circuit 410 has a basic structure as shown in Fig. 28, and the multiplier 403 controls the multiplication coefficient according to the spectral amplitude Am[i] of the voiced (V) sound to find the synthesized noise amplitude Am noise [i], the vocal (V) sound is supplied from the quantizer 212 of the spectral information envelope shown in FIG. 4 through a terminal 411. That is, in Figure 28, the output of the optimal noise mixing value calculation circuit 416 inputted with the spectral amplitude Am[i] and the pitch delay Pch is weighted by a noise weighting circuit 417, and the generated output is sent to a noise weighting circuit 417. The multiplier multiplies a spectrum amplitude Am[i] to generate a noise amplitude-Am-noise[i].
The first specific embodiment of noise synthesis and addition is described, in which the noise amplitude-Am-noise[i] becomes the second of the above four parameters, that is, a function of the pitch delay Pch and the spectral amplitude Am[i].
For this function f<sub>1</sub>(Pch, Am[i]): f<sub>1</sub>(Pch, Am[i])=0, where 0<i<noise_b×I, f<sub>1</sub>(Pch, Am[i])=Am[i]×noise_mix, where noise_b×Ii<I, and noise_mix=K×Pch/2.0.
It should be noted that the maximum value of noise_mixing is the cut noise_mixing_max. Take K=0.02, noise_mix_max=0.3, and noise_b=0.7 as an example, where noise_b is a constant, which determines which part of the entire band to add this noise. In this embodiment, the added noise is a frequency range higher than 70%, that is, if fs=8kHz, the noise is added in a range from 400×0.7=2800kHz or even 4000kHz.
Here is a description of the second specific embodiment of noise synthesis and addition, in which the noise amplitude Am-noise [i] is three of the four parameters, namely the pitch delay Pch, the spectral amplitude Am[i] and the maximum spectral amplitude Amax Function f<sub>2</sub>(Pch, Am[i], Amax).
For this function f<sub>2</sub>(Pch, Am[i], Amax): f<sub>2</sub>(Pch, Am[i], Amax)=0, where 0<i<noise_b×I, f<sub>1</sub>(Pch, Am[i], Amax)=Am[i]×noise_mix, where noise_b×Ii<I, and noise_mix=K×Pch/2.0.
It should be noted that the maximum value of noise_mixing is noise-mixing_max, and in one example, K=0.02, noise_mixing-max=0.3 and noise b=0.7.
If Am[i]×Noise_Mix>Amax×C×Noise_Mix, f<sub>2</sub>(Pch, Am[i], Amax)=Amax×C×noise_mixing, where the constant C is set to 0.3 (C=0.3). Since the conditional equation can be used to avoid excessive levels, the above-mentioned K and noise_mix max value can be further increased, and if the high range level is higher, the noise level can be further increased.
In the third specific embodiment of noise synthesis and addition, the above-mentioned noise amplitude Am-noise[i] can be a function of all the above-mentioned four parameters, that is, f<sub>3</sub>(Pch, Am[i], Amax, Lev).
Function f<sub>3</sub>(Pch, Am[i], Am[max], Lev) The specific example is basically similar to the above function f<sub>2</sub>(Pch, Am[i], Amax). The remaining signal level Lev is the spectral amplitude Am[i] or the average square (RMS) of the signal level obtained on the time axis. The difference from the second specific embodiment is that the value of K and noise_mix_max is set as a function of Lev. That is, if Lev is smaller or larger, the values of K and noise_mix_max are set to larger and smaller values, respectively. Alternatively, the value of Lev can be set to be inversely proportional to the value of K and noise_mix_max.
Here is an explanation of post filtering.
Figure 29 shows a post filter, which can be used as the post filter 238u, 238v of the Figure 4 embodiment. The spectrum shaping filter 440, which forms the main part of the post-filtering, is composed of a formant emphasis filter 441 and a high-range emphasis filter 442. An output of the spectrum shaping filter 440 is sent to a gain adjustment circuit 443 for correcting the gain change caused by the spectrum shaping. The gain G of the gain adjustment circuit 443 is determined by comparing an input with an output Y of the spectrum shaping filter 440 with a gain control circuit 445 to calculate a gain change and calculate a correction value.
If the denominator Hv(z) and Huv(z) coefficients of the LPC synthesis filter, that is, the -parameter, use α<sub>i</sub>To express, the characteristic PF(z) of the spectrum forming filter 440 can be expressed as:<maths><img file="TW360859B_D0128.tif" /></maths>
The fractional part of this equation represents the characteristics of formant emphasis filtering, and (1-kz<sup>-1</sup>The part) represents the characteristics of high-range emphasis filtering. β, γ and k are constants, such as β=0.6, γ=0.8 and k=0.3.
The gain of the gain adjustment circuit 443 is derived from:<maths><img file="TW360859B_D0129.tif" /></maths>
In the above equation, x(i) and y(i) represent the input and output of the spectral filter 440, respectively.
It should be noted that, as shown in Figure 30, the coefficient update period of the spectrum shaping filter 440 is just like the α-parameter, that is, the update period of the LPC synthesis filter coefficient is 20 samples or 2.5 msec. However, the update period of the gain G of the gain adjustment circuit 443 is 160 samples or 20msec.
By setting the coefficient update period of the spectrum shaping filter 443 to be greater than the update period of the coefficients of the spectrum shaping filter 440 as a post-filter, adverse effects caused by gain adjustments and declines in other ways can be prevented.
That is, in the general post-filtering, the coefficient update period of the spectrum shaping filter is set equal to the gain update period, and as shown in Figure 30, if the gain update period is selected to be 20 samples and 2.5msec, the gain value changes at one pitch It becomes uniform during the cycle, resulting in tick noise. In this embodiment, by setting the gain conversion period to be longer, for example, equal to one frame or 160 samples, or 20 msec, a sudden change in the gain value can be avoided. Conversely, if the spectral shaping filter coefficient is 160 samples or 20msec, it cannot produce smooth changes in the filter characteristics, which will cause adverse effects on the synthesized waveform. However, by setting the filter coefficient update cycle to 20 samples or a shorter value of 2.5msec, post-filtering can be more effective.
By performing gain interface processing between adjacent frames, the filter coefficient and gain of the previous frame and the current frame are multiplied by W(i)=i/20 (0i<20) and 1-Wi, where 0 i<20, it will gradually become stable, and the products will be added together. Figure 31 shows the gain G of the previous frame<sub>1</sub>How to match the gain G of the current frame<sub>1</sub>Merge. Specifically, the gain and filter coefficients of the previous frame are gradually reduced, while the gain and filtering of the current frame are gradually increased. The filtering internal state of the current frame and the previous frame start from the same state at the time point T in Fig. 31, that is, from the last state of the previous frame.
The above-mentioned signal encoding and signal decoding device can be used as a voice codebook in a portable communication terminal or a portable telephone as shown in Figures 32 and 33.
Figure 32 shows a transmission side of a portable terminal. This terminal uses the voice coding unit 160 as shown in Figures 1 and 3. The voice signal collected by a microphone 161 is amplified by an amplifier 162 and sent to the voice encoding unit 160 shown in FIGS. 1 and 3. The speech encoding unit 160 performs encoding as described in conjunction with FIGS. 1 and 3. The output signal of the output terminal in FIGS. 1 and 2, that is, the output signal of the speech encoding unit 160 is sent to a transmission channel encoding unit 164, which performs channel encoding on the supplied signal. The output signal of the transmission channel encoding unit 164 is sent to a modulation circuit 165 for modulation, and then supplied to an antenna 168 via a digital/analog (D/A) converter 165 and an RF amplifier 167.
Figure 33 shows a receiving side of a portable terminal using the speech decoding unit shown in Figures 2 and 4. The voice signal received by the antenna 261 in Fig. 33 is sent to a demodulation circuit 264 via an analog/digital (A/D converter 263), and the demodulated signal is sent to a transmission channel decoding unit 265 therefrom. The decoding unit An output signal of 265 is supplied to a speech decoding unit 260 as shown in Figures 2 and 4. The speech decoding unit 260 decodes the signal in the same manner as shown in Figures 2 and 4. In Figures 2 and 4 The output of the output terminal 201, which is a signal of the speech decoding unit 260, is sent to a digital/analog (DA) converter 266. An analog speech signal from the D/A converter 266 is sent to a speaker 268.
The present invention is not limited to the above-mentioned embodiment. For example, the structure of the speech analysis side (encoder) in Figures 1 and 3 or the structure of the speech synthesis side (decoder) in Figures 2 and 4 are hardware as described above, but a digital signal processor can be used, for example. (DSP) is implemented as a software program. The synthesis filter 236, 237 or the post-filter 238, 238u on the decoding side can be designed as a single LPC synthesis filter or a single post-filter, and is not divided into being used for voiced speech or unvoiced speech. The present invention is not limited to transmission or recording/reproduction, but can be applied to various purposes, such as pitch conversion, voice conversion, computer speech synthesis or noise suppression.
<p>109High Pass Filter HPF</p><p>110First coding unit</p><p>111LPC Inverse Filter</p><p>113LPC analysis/quantification unit</p><p>114Sine analysis coding unit</p><p>115V/UV frequency discrimination unit</p><p>116Vector quantization unit</p><p>120Second coding unit</p><p>121Noise Code Book</p><p>122Perceptually weighted synthesis filter</p><p>123Subtractor</p><p>124Distance calculation circuit</p><p>125Perceptually weighted filter</p><p>132LPC analysis circuit</p><p>133α-LSP conversion circuit</p><p>134LSP quantizer</p><p>136LSP interpolation circuit</p><p>137α conversion circuit</p><p>141Open loop pitch search unit</p><p>142Zero Crossover Calculator</p><p>145Orthogonal Transformation Circuit</p><p>146Fine Pitch Retrieval Unit</p><p>148Spectral Evaluation Unit</p><p>211Voice Synthesizer</p><p>212Inverse vector quantization unit</p><p>213LPC parameter regeneration unit</p><p>214LPC synthesis filter</p><p>216Noise synthesis circuit</p><p>217weighted superposition circuit</p><p>218Adder</p><p>231Inverse Vector Quantizer</p><p>232,233LSP interpolation circuit</p><p>236,237LPC synthesis filter</p><p>610Buffer</p><p>620Matrix quantizer</p><p>621LSP parameter adder</p><p>622Matrix quantizer MQ<sub>1</sub></p><p>623Weighted distance calculation unit</p><p>631Adder</p><p>633Weighted distance calculation unit</p><p>640Vector quantization unit</p><p>650, 660, 670, 680vector quantizer part</p><p>651,661,671,681Adder</p><p>652,662vector quantization unit</p><p>653,663Weighted distance calculation unit</p><p>672,682Vector quantizer</p><p>673,683Weighted distance calculation unit</p>
Figure 1 is a block diagram showing the basic structure of a speech signal encoding device (encoder) implementing the encoding method of the present invention.
Fig. 2 is a block diagram showing the basic structure of a speech signal decoding device (decoder) which embodies the decoding method of the present invention.
Figure 3 is a block diagram showing a more detailed structure of the speech signal encoder shown in Figure 1.
Figure 4 is a block diagram showing a more detailed structure of the speech signal decoder shown in Figure 2.
Figure 5 is a graph showing the bit speed of the speech signal encoding device.
Figure 6 is a block diagram showing a more detailed structure of the LSP quantizer.
Figure 7 is a block diagram showing a basic structure of the LSP quantizer.
Figure 8 is a block diagram showing a more detailed structure of the vector quantizer.
Figure 9 is a block diagram showing a more detailed structure of the vector quantizer.
Figure 10 is a figure showing a specific example of the weight W(i) used for weighting.
Figure 11 is a graph showing the relationship between the quantization value, the size number, and the number of bits.
Figure 12 is a circuit block diagram showing the schematic structure of a vector quantizer used for variable size codebook retrieval.
Figure 13 is a circuit block diagram showing the schematic structure of a vector quantizer used for variable size codebook retrieval.
Figure 14 is a circuit block diagram showing a first schematic structure of a vector quantizer using a code book for variable size and a code book for fixed size.
Figure 15 is a circuit block diagram showing a second schematic structure of a vector quantizer using a code book for variable size and a code book for fixed size.
Figure 16 is a circuit block diagram showing a third schematic structure of a vector quantizer using a code book for variable size and a code book for fixed size.
Figure 17 is a circuit block diagram showing a fourth schematic structure of a vector quantizer using a code book for variable size and a code book for fixed size.
Figure 18 is a circuit block diagram showing the specific structure of a CULP encoding section (second encoder) of the speech encoding device of the present invention.
Figure 19 is a flowchart showing the processing flow of the configuration shown in Figure 16.
Figure 20 shows the state of Gaussian noise and noise cut at different thresholds.
Figure 21 is a flow chart, shown in time. The flow chart of generating a waveform code book by learning.
Figure 22 is a diagram showing the state of LSP switching according to U/UV conversion.
Figure 23 shows a 10-order linear spectrum pair based on the α-parameter. This α-parameter is obtained by the 10-order LPC analysis method.
Figure 24 shows the gain change state from a silent (UV) frame to a vocal (V) frame.
Figure 25 shows the insert operation for waveform or spectral elements synthesized between frame and frame.
Figure 26 shows the overlap between the vocal (V) frame and the unvoiced (UV) frame.
Figure 27 shows the noise addition processing at the moment of voiced speech synthesis.
Figure 28 shows an example of amplifying calculation of noise added at the moment of voiced speech synthesis.
Figure 29 shows a schematic structure of post-filtering.
Figure 30 shows the filter coefficient update cycle and gain update cycle of post-filtering.
Figure 31 shows the merging process of the frame boundary between the post-filtering gain and the filter coefficient.
Figure 32 is a block diagram showing the transmission side structure of a portable terminal using the speech signal encoding device of the present invention.
Figure 33 is a block diagram showing the structure of the receiving side of a portable terminal using the speech signal encoding device of the present invention.
Vector quantization method and speech coding method and device
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US8200500B2 | Cited by | United States of America | Applicant |
| US7805313B2 | Cited by | United States of America | Applicant |
| US7903824B2 | Cited by | United States of America | Applicant |
| US7720230B2 | Cited by | United States of America | Applicant |
| US8204261B2 | Cited by | United States of America | Applicant |
| US8238562B2 | Cited by | United States of America | Applicant |
| US8340306B2 | Cited by | United States of America | Applicant |
| US7761304B2 | Cited by | United States of America | Applicant |
| US7941320B2 | Cited by | United States of America | Applicant |
| US7693721B2 | Cited by | United States of America | Applicant |
| US7644003B2 | Cited by | United States of America | Applicant |
15 members in 9 offices
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 251614 | Japan | – | |
| 25161496 | Japan | A |
Members15
| Document | Office | Kind | |
|---|---|---|---|
| EP0831457A2 | European Patent Office (EPO) | A2 | |
| ID18313A | Indonesia | A | |
| JPH1097298A | Japan | A | |
| KR19980024885A | Republic of Korea | A | |
| CN1188957A | China | A | |
| EP0831457A3 | European Patent Office (EPO) | A3 | |
| TW360859BThis record | Taiwan Province of China | B | |
| US6611800B1 | United States of America | B1 | |
| EP0831457B1 | European Patent Office (EPO) | B1 | |
| DE69726525D1 | Germany | D1 | |
| CN1145142C | China | C | |
| DE69726525T2 | Germany | T2 | |
| JP3707153B2 | Japan | B2 | |
| MY120520A | Malaysia | A | |
| KR100543982B1 | Republic of Korea | B1 |
1 legal event, as the office reported them to INPADOC
Events
| Event | Code | |
|---|---|---|
| Annulment or lapse of patent due to non-payment of feesLapsedMM4A | MM4A |
Numbers
- Publication
- 360859
- Application
- 86113292
Titles4
- Chinese
- 向量量化方法及語音編碼方法及裝置
- English
- Vector quantization method and speech encoding method and apparatus
- Unlabeled
- 向量量化方法及語音編碼方法及裝置
- Unlabeled
- Vector quantization method and speech coding method and device
Classification
- CPC, 4
- H03M7/3082
- G10L19/038
- G10L19/0208
- G10L2019/0013
- IPC, 8
- G10L19 038
- G10L15 02
- G10L19 04
- G10L19 08
- G10L19 16
- G10L19 20
- G10L25 00
- H03M7 30