Method and apparatus for CELP coding an audio signal while distinguishing speech periods and non-speech periods
Abstract
This record has no abstract on file.
Term
Term ended
Expired 23 August 2015, 11.1 years ago.
- Priority and filed
- Granted
- Expired
- Today
4 claims: 4 independent, 0 dependent
- 1An autocorrelation analysis means for obtaining information on an autocorrelation matrix from an input acoustic signal. Vocal tract prediction coefficient analysis means for obtaining vocal tract prediction coefficient from the analysis results of the above autocorrelation analysis means, and Predictive gain coefficient analysis means for obtaining the predicted gain coefficient from the above vocal tract prediction coefficient The non-audio signal section of the input acoustic signal is detected from the input acoustic signal, the voice path prediction coefficient, and the predicted gain coefficient, and the autocorrelation information in the non-audio signal section is used to obtain the autocorrelation information in the non-audio signal section. An autocorrelation adjusting means for changing and adjusting the weighted composite value of the correlation information and the above autocorrelation information in the past non-audio signal section, A vocal tract prediction coefficient compensating means for obtaining a post-compensated vocal tract prediction coefficient that compensates for the vocal tract prediction coefficient in the non-voice signal section from the above-adjusted autocorrelation information. A code excitation linear prediction coding apparatus including a coding means for coding a code excitation linear prediction coding of an input acoustic signal using the post-compensation vocal tract prediction coefficient and an adaptive excitation signal. 【請求項1】 入力音響信号から自己相関マトリクスの情報を求める自己相関分析手段と、 上記自己相関分析手段の分析結果から声道予測係数を求める声道予測係数分析手段と、 上記声道予測係数から予測利得係数を求める予測利得係数分析手段と、 上記入力音響信号と上記声道予測係数と上記予測利得係数とから入力音響信号の非音声信号区間を検出し、この非音声信号区間における上記自己相関の情報を、この非音声信号区間における上記自己相関の情報と過去の非音声信号区間における上記自己相関の情報との重み付け合成値に変更調節する自己相関調節手段と、 上記調節後の自己相関の情報から非音声信号区間における声道予測係数を補償した補償後声道予測係数を得る声道予測係数補償手段と、 上記補償後声道予測係数と適応励振信号とを使用して入力音響信号をコード励振線形予測符号化する符号化手段とを備えたことを特徴とするコード励振線形予測符号化装置。
- 2An autocorrelation analysis means for obtaining autocorrelation information from an input acoustic signal, Vocal tract prediction coefficient analysis means for obtaining vocal tract prediction coefficient from the analysis results of the above autocorrelation analysis means, and Predictive gain coefficient analysis means for obtaining the predicted gain coefficient from the above vocal tract prediction coefficient The LSP coefficient is obtained from the voice tract prediction coefficient, and the non-audio signal section of the input acoustic signal is detected from the input acoustic signal, the voice tract prediction coefficient, and the predicted gain coefficient, and the LSP coefficient in this non-audio signal section is detected. To a weighted composite value of the LSP coefficient in the non-audio signal section and the LSP coefficient in the past non-audio signal section. A vocal tract prediction coefficient compensating means for obtaining a post-compensated vocal tract prediction coefficient that compensates for the vocal tract prediction coefficient in the non-voice signal section from the adjusted LSP coefficient. A code excitation linear prediction coding apparatus including a coding means for coding a code excitation linear prediction coding of an input acoustic signal using the post-compensation vocal tract prediction coefficient and an adaptive excitation signal. 【請求項2】 入力音響信号から自己相関の情報を求める自己相関分析手段と、 上記自己相関分析手段の分析結果から声道予測係数を求める声道予測係数分析手段と、 上記声道予測係数から予測利得係数を求める予測利得係数分析手段と、 上記声道予測係数からLSP係数を求めると共に、上記入力音響信号と上記声道予測係数と上記予測利得係数とから入力音響信号の非音声信号区間を検出し、この非音声信号区間における上記LSP係数を、この非音声信号区間における上記LSP係数と過去の非音声信号区間における上記LSP係数との重み付け合成値に変更調節するLSP係数調節手段と、 上記調節後のLSP係数から非音声信号区間における声道予測係数を補償した補償後声道予測係数を得る声道予測係数補償手段と、 上記補償後声道予測係数と適応励振信号とを使用して入力音響信号をコード励振線形予測符号化する符号化手段とを備えたことを特徴とするコード励振線形予測符号化装置。
- 3An autocorrelation analysis means for obtaining autocorrelation information from an input acoustic signal, Vocal tract prediction coefficient analysis means for obtaining vocal tract prediction coefficient from the analysis results of the above autocorrelation analysis means, and Predictive gain coefficient analysis means for obtaining the predicted gain coefficient from the above vocal tract prediction coefficient A non-voice signal section is detected from the input acoustic signal, the predicted gain coefficient, and the vocal tract prediction coefficient, and the vocal tract prediction coefficient in this non-voice signal section is used as the vocal tract prediction coefficient in this non-voice signal section. Vocal tract coefficient adjusting means for obtaining the adjusted vocal tract prediction coefficient by changing and adjusting to a weighted composite value with the above vocal tract prediction coefficient in the past non-audio signal section. A code excitation linear prediction coding apparatus including a coding means for coding a code excitation linear prediction coding of an input acoustic signal using the adjusted vocal tract prediction coefficient and an adaptive excitation signal. 【請求項3】 入力音響信号から自己相関の情報を求める自己相関分析手段と、 上記自己相関分析手段の分析結果から声道予測係数を求める声道予測係数分析手段と、 上記声道予測係数から予測利得係数を求める予測利得係数分析手段と、 上記入力音響信号と上記予測利得係数と上記声道予測係数とから非音声信号区間を検出し、この非音声信号区間における上記声道予測係数を、この非音声信号区間における上記声道予測係数と過去の非音声信号区間における上記声道予測係数との重み付け合成値に変更調節して、調節後の声道予測係数を得る声道係数調節手段と、 上記調節後の声道予測係数と適応励振信号とを使用して入力音響信号をコード励振線形予測符号化する符号化手段とを備えたことを特徴とするコード励振線形予測符号化装置。
- 4An autocorrelation analysis means for obtaining autocorrelation information from an input acoustic signal, Vocal tract prediction coefficient analysis means for obtaining vocal tract prediction coefficient from the analysis results of the above autocorrelation analysis means, and Predictive gain coefficient analysis means for obtaining the predicted gain coefficient from the above vocal tract prediction coefficient A non-voice signal section is detected for each band pass processing signal from the band pass processing signal obtained by band pass processing from the input acoustic signal and the predicted gain coefficient, and the non-voice for each band pass processing signal is detected. A filter coefficient for noise removal is generated according to the detection result of the signal section, and noise removal is performed using the filter coefficient generated for the input acoustic signal to generate a target for generating a composite audio signal. Noise removal means to generate signals and Synthetic voice generation means for generating the synthetic voice signal using the vocal tract prediction coefficient, and A code excitation linear prediction coding apparatus including a coding means for code excitation linear prediction coding of an input acoustic signal using the vocal tract prediction coefficient and the target signal. 【請求項4】 入力音響信号から自己相関の情報を求める自己相関分析手段と、 上記自己相関分析手段の分析結果から声道予測係数を求める声道予測係数分析手段と、 上記声道予測係数から予測利得係数を求める予測利得係数分析手段と、 上記入力音響信号から帯域通過処理して得た帯域通過処理信号と、上記予測利得係数とから、各帯域通過処理信号毎に非音声信号区間を検出し、この各帯域通過処理信号毎の非音声信号区間の検出結果に応じてノイズ除去のためのフィルタ係数を生成し、上記入力音響信号に対して生成された上記フィルタ係数を使用してノイズ除去を行って合成音声信号の生成のためのターゲット信号を生成するノイズ除去手段と、 上記声道予測係数を使用して上記合成音声信号を生成する合成音声生成手段と、 上記声道予測係数と上記ターゲット信号とを使用して入力音響信号をコード励振線形予測符号化する符号化手段とを備えたことを特徴とするコード励振線形予測符号化装置。
Independent claims4
105 paragraphs, as filed
Description: TECHNICAL FIELD [Detailed description of the invention]
【0001】
INDUSTRIAL APPLICABILITY The present invention relates to a code-excited linear predictive coding (CELP) device, and particularly to consider the influence of an acoustic signal in a non-audio signal section.
【0002】
PROBLEM TO BE SOLVED: To encode / decode voice, the voice section and other silence / noise sections are processed in the same manner. As a voice coding method, for example, it is disclosed in the following documents.
References: Proc. IEEE ICSSP, 1990, pp. 461-464, "VECTOR SUM EXCITED LINEAR PREDICTION (VSELP) SPEECH CODING AT 8 kbps", Gerson and Jazz.
This document describes the VSELP method, which is currently defined as the standard for voice coding methods for North American digital cellular. A similar method is adopted for the voice coding method for digital cellular in Japan.
【0005】
However, the configuration of the CELP-based encoder emphasizes the coding characteristics of the voice section, and when noise is encoded / decoded, the synthesized sound is an unnatural sound. It was jarring.
The noise section of the synthesized sound encoded and decoded by the CELP system encoder is obtained from the LPC analysis (linear predictive analysis) that the code book used as the excitation source is optimized for speech. Since the spectral estimation error to be obtained is different for each frame, the sound becomes unnatural sound far from the noise before coding, which causes deterioration of call quality.
From the above, code excitation linear prediction that can reduce the influence of acoustic signals (noise, rotating sound, vibration sound, etc.) on the coded output, especially in the non-voice signal section, and perform good voice reproduction. The provision of a coding device is requested.
【0008】
[Means for Solving the Problems] Therefore, the code excitation linear prediction coding device of the invention of claim 1 obtains autocorrelation information (for example, autocorrelation matrix or autocorrelation coefficient) from an input acoustic signal. "Autocorrelation analysis means", "Voiceway prediction coefficient analysis means" for obtaining the voiceway prediction coefficient from the analysis results of the autocorrelation analysis means, and "Prediction gain coefficient analysis means" for obtaining the prediction gain coefficient from the voiceway prediction coefficient. Then, the non-audio signal section of the input acoustic signal is detected from the input acoustic signal, the voice path prediction coefficient, and the predicted gain coefficient, and the autocorrelation information in the non-audio signal section is obtained in the non-audio signal section. "Autocorrelation adjusting means" that changes and adjusts the weighted composite value of the autocorrelation information and the autocorrelation information in the past non-voice signal section, and the voice in the non-voice signal section from the adjusted autocorrelation information. Code excitation linear prediction coding of the input acoustic signal using the "voiceway prediction coefficient compensating means" that obtains the post-compensation voiceway prediction coefficient that compensates for the way prediction coefficient, and the above-mentioned post-compensation voiceway prediction coefficient and adaptive excitation signal. It is provided with a "coding means" to solve the above-mentioned problems.
The vocal tract prediction coefficient of the vocal tract prediction coefficient analysis means can be obtained by, for example, LPC (linear analysis coding). The predicted gain coefficient can also be obtained as, for example, the reflection coefficient of the vocal tract. By the above autocorrelation adjusting means, for example, the autocorrelation information adjusted so as to reduce the noise is obtained by combining the autocorrelation information of the section determined to be noise in the past and the autocorrelation information of the current frame. Can be done.
By obtaining the vocal tract prediction coefficient in the non-voice signal section from this adjusted autocorrelation information, the vocal tract prediction coefficient after compensation for the acoustic signal (for example, noise) in the non-voice signal section is compensated. To get. By using the adaptive excitation signal of the adaptive codebook to code-encourage the input acoustic signal with this post-compensated vocal tract prediction coefficient, the coded output in the non-voice signal section can be reduced in noise. It can be suitable.
Further, the code excitation linear predictive coding device of the invention of claim 2 has a "self-correlation analysis means" that obtains autocorrelation information from an input acoustic signal, and a voice path prediction from the analysis results of the autocorrelation analysis means. A "voice path prediction coefficient analysis means" for obtaining a coefficient, a "prediction gain coefficient analysis means" for obtaining a prediction gain coefficient from the voice path prediction coefficient, and an LSP (line spectrum pair) coefficient obtained from the voice path prediction coefficient, as well as The non-audio signal section of the input acoustic signal is detected from the input acoustic signal, the voice path prediction coefficient, and the predicted gain coefficient, and the LSP coefficient in the non-audio signal section is used as the LSP coefficient in the non-audio signal section. The "LSP coefficient adjusting means" that changes and adjusts to the weighted composite value with the LSP coefficient in the past non-voice signal section, and the compensated voice path that compensates the voice path prediction coefficient in the non-voice signal section from the adjusted LSP coefficient. It is equipped with a "voiceway prediction coefficient compensating means" for obtaining a prediction coefficient and a "coding means" for code-exciting linear prediction coding of an input acoustic signal using the above-compensated voiceway prediction coefficient and an adaptive excitation signal. , The above-mentioned problem is solved.
In order to reduce the influence of the acoustic signal in the non-audio signal section, the voice path prediction coefficient is converted into the LSP coefficient, and at the stage of this LSP coefficient, the LSP coefficient of the past frame is also referred to and the LSP coefficient is used. By adjusting, it becomes easier to obtain the LSP coefficient that suppresses the spectral fluctuation with respect to the spectral fluctuation in the audio signal section, and finally the LSP coefficient is converted to the voice path prediction coefficient to obtain the adaptive excitation signal of the adaptive code book. By using the code excitation linear predictive coding of the input acoustic signal, it is possible to make the coded output particularly in the non-voice signal section more suitable for noise reduction than in the past.
Further, the code excitation linear prediction coding apparatus according to the third aspect of the present invention includes an "autocorrelation analysis means" for obtaining an autocorrelation matrix or an autocorrelation coefficient from an input acoustic signal, and an analysis result of the autocorrelation analysis means. "Voiceway prediction coefficient analysis means" for obtaining the voiceway prediction coefficient from, "Prediction gain coefficient analysis means" for obtaining the prediction gain coefficient from the voiceway prediction coefficient, the input acoustic signal, the prediction gain coefficient, and the voiceway. The non-voice signal section is detected from the prediction coefficient, and the voice tract prediction coefficient in this non-voice signal section is used as the voice tract prediction coefficient in this non-voice signal section and the voice tract prediction coefficient in the past non-voice signal section. The input acoustic signal is coded using the "voiceway coefficient adjusting means" that obtains the adjusted voiceway prediction coefficient by changing and adjusting to the weighted composite value of, and the above-adjusted voiceway prediction coefficient and the adaptive excitation signal. It is provided with a "coding means" for excitation linear prediction coding to solve the above-mentioned problems.
With such a configuration, the vocal tract prediction coefficient in the non-voice signal section is directly obtained by using the vocal tract prediction coefficient in the past non-voice signal section, and the non-voice signal section can be calculated with a very small amount of calculation. It can be encoded to reduce the influence of the acoustic signal.
Furthermore, the code excitation linear predictive coding device of the invention of claim 4 has a "self-correlation analysis means" that obtains autocorrelation information from an input acoustic signal, and a voice path from the analysis results of the autocorrelation analysis means. A "voiceway prediction coefficient analysis means" for obtaining a prediction coefficient, a "prediction gain coefficient analysis means" for obtaining a prediction gain coefficient from the voiceway prediction coefficient, and a band passage processing signal obtained by band passage processing from the input acoustic signal. And, from the above predicted gain coefficient, a non-audio signal section is detected for each band passing processing signal, and a filter coefficient for noise removal is determined according to the detection result of the non-audio signal section for each band passing processing signal. A "noise removing means" that generates a target signal for generating a composite audio signal by performing noise removal using the filter coefficient generated for the input acoustic signal, and the voice path prediction coefficient. "Synthetic voice generation means" that generates the synthetic voice signal using the above, and "coding means" that code-excited linear predictive coding of the input acoustic signal using the voice path prediction coefficient and the target signal. To solve the above-mentioned problems.
The noise removing means is composed of a filter for removing noise in a non-voice signal section from an input acoustic signal, and the filter coefficient of this filter is a voice tract prediction coefficient, a prediction gain coefficient, or the like. By obtaining using the band passage processing signal, it is possible to obtain a target signal from which noise has been removed. Therefore, by using the target signal after removing the noise and performing code excitation linear predictive coding, it is possible to obtain a coded output in which the influence of noise in the non-audio signal section is removed.
【0017】
BEST MODE FOR CARRYING OUT THE INVENTION Next, preferred embodiments of the present invention will be described with reference to the drawings. In the embodiment of the present invention, first (1), the composite filter coefficient is adjusted based on the frame voice / noise determination using the autocorrelation matrix, the LSP coefficient, or the direct prediction coefficient, and the noise section is used. It is configured to reduce unnatural noise.
(2) Further, the target signal for selecting the optimum code vector is filtered based on the subframe voice / noise determination to reduce the noise.
"First Embodiment": (A) Specifically, in the noise section abnormal noise suppression type CELP type voice encoder, the input signal is separated into voice and noise in frame units, and the autocorrelation matrix of the current noise section frame and the continuous pre-noise section frame are separated. A new autocorrelation matrix is calculated in combination with the autocorrelation matrix, LPC analysis is performed using the new autocorrelation matrix, the composite filter coefficient is obtained, quantized, and transmitted to the decoder, and the optimum is optimized using the above composite filter coefficient. The code book vector is searched.
Next, a detailed configuration for realizing the configuration (A) described above will be described. FIG. 1 is a functional configuration diagram of a CELP coding device. In FIG. 1, the CELP coding apparatus includes a frame power calculation unit 101, an autocorrelation matrix calculation unit 102, an LPC analysis unit 103, a synthesis filter 104, an adaptive codebook 105, a noise codebook 106, and a gain. Adder with codebook 107, weighted distance calculation unit 108, LSP quantizer 109, voice / noise determination unit 110, autocorrelation matrix adjustment unit 111, prediction gain calculation unit 112, and multipliers 113 and 114. It is composed of a device 115, a subtractor 116, a quantizer 117, and a multiplexing unit 130.
The characteristic part of the configuration of FIG. 1 is the vocal tract coefficient by the autocorrelation matrix calculation unit 102, the voice / noise determination unit 110, the autocorrelation matrix adjustment unit 111, and the LPC analysis unit 103. This is to be corrected to eliminate the cause of reproducing a jarring sound by CELP coding of a noise portion other than the conventional voice section.
When the input audio digital signal (audio vector signal) S is given as the original signal vector that is collected in frame units and input as a vector, the frame power calculation unit 101 obtains the frame power and obtains the frame power signal. It is given to the multiplexing unit 130 as P. When the input audio digital signal S in frame units is given, the "autocorrelation matrix calculation unit 102" obtains the autocorrelation matrix R for obtaining the vocal tract coefficient, and the LPC analysis unit 103 and the autocorrelation matrix adjustment unit 111. Give to and.
When the "LPC analysis unit 103" obtains the voiceway prediction coefficient a from the autocorrelation matrix R and gives it to the prediction gain calculation unit 112, and when the autocorrelation matrix Ra from the autocorrelation matrix adjustment unit 111 is given, The optimum voice tract prediction coefficient aa obtained by modifying the voice tract prediction coefficient a by the autocorrelation matrix Ra is obtained and given to the synthesis filter 104 and the LSP quantizer 109.
The predicted gain calculation unit 112 converts the voice tract prediction coefficient a into a reflection coefficient, obtains a predicted gain from this reflection coefficient, and gives this to the voice / noise determination unit 110 as a predicted gain signal pg. The "voice / noise determination unit 110" is given a pitch coefficient signal ptch from the adaptive codebook 105, and further, the input voice digital signal S in frame units, the voice path prediction coefficient a, and the prediction gain signal. From pg, it is determined whether the signal S of the frame is an audio signal or a noise signal other than the audio signal, and the audio / noise determination signal v is given to the autocorrelation matrix adjusting unit 111.
The "autocorrelation matrix adjusting unit 111" is a particularly important functional unit, and is a process performed only when it is determined to be noise. The autocorrelation matrix R, the voice / noise determination signal v, and the like. Given the voice path prediction coefficient a, a new autocorrelation matrix Ra is obtained by combining the "autocorrelation matrix of the section determined to be noise in the past" and the "autocorrelation matrix of the current noise frame". It is given to the LPC analysis unit 103.
The "adaptive codebook 105" is provided with a plurality of periodic adaptive excitation vectors in advance, and each of these adaptive excitation vectors is assigned an index number Ip. , The adaptive excitation vector signal ea is output by the optimum index number Ip specified by the weighted distance calculation unit 108, and is given to the multiplier 113 and the pitch signal ptch (the normalized mutual of the input voice signal S and the optimum adaptive excitation vector signal ea). (Correlation signal) is output and given to the voice / noise determination unit 110. Further, the adaptive excitation vector signal inside the adaptive codebook 105 is updated by the optimum excitation vector signal exOP from the output excitation vector signal ex from the adder 115.
The "noise code book 106" is provided with a plurality of noise-induced excitation vector signals in advance, and each of these noise-induced excitation vector signals is given an index number Is. Then, the noise excitation vector signal es is output by the optimum index number Is specified by the weighted distance calculation unit 108 and given to the multiplier 114.
The gain codebook 107 stores in advance gain codes (gain: gain) for the adaptive excitation vector signal and the noisy excitation vector, and an index number Ig is assigned to each of these gain codes. The gain code signal ga is output for the adaptive excitation vector signal and given to the multiplier 113 by the optimum index number Ig specified by the weighted distance calculation unit 108, and the gain is obtained for the noisy excitation vector signal. The code signal gs is output and given to the multiplier 114.
The "multiplier 113 on the adaptive codebook 105 side" multiplies the adaptive excitation vector signal ea with the gain code signal ga to obtain an adaptive excitation vector signal having an optimum gain (magnitude), and is an adder. Give to 115. The "multiplier 114 on the noise code book 106 side" multiplies the noise excitation vector signal es and the gain code signal gs to obtain a noise excitation vector signal having an optimum gain (magnitude) and gives it to the adder 115. .. The "adder 115" adds the adaptive excitation vector signal with the optimum gain and the noise excitation vector signal with the optimum gain to give the excitation vector signal ex to the synthesis filter 104, and at the same time. The optimum excitation vector signal exOP having a relationship in which the sum of squares E calculated by the weighted distance calculation unit 108 is minimized is fed back to the adaptive codebook 105 to be updated and stored.
The synthetic filter 104 can be configured by an IIR (Infinite Impulse Response) type digital filter circuit, and has the optimum voice path prediction coefficient aa after the above modification and an excitation vector (excitation) from the adder 115. A synthetic voice vector signal Sw (synthetic voice signal) is generated from the signal) ex and given to the subtractor 116. That is, the excitation vector signal ex is filtered as the filter (tap) coefficient of the optimum vocal tract prediction coefficient aa after the modification with respect to the IIR type digital filter to obtain the synthesized speech vector signal Sw. The subtractor 116 subtracts the input voice digital signal S and the synthetic voice vector signal Sw, and gives the subtraction result to the weighted distance calculation unit 108 as an error vector signal e.
When the error vector signal e from the subtractor 116 is given, the weighting distance calculation unit 108 performs frequency conversion of the error vector signal e and weights the error vector signal e. The sum of squares of the weighted vector signals is obtained, and the optimum index numbers Ia and Is corresponding to the optimum adaptive excitation vector signal, noise excitation vector signal, and gain code signal so that the vector signal E obtained by this sum of squares is minimized. , Ig is obtained and given to the adaptive codebook 105, the noise codebook 106, and the gain codebook 107.
The gain code quantizer 117 quantizes the gain code signals ga and gs and gives them to the multiplexing unit 130 as a gain code quantization signal. The LSP quantizer 109 LSP quantizes the vocal tract prediction coefficient aa optimally modified by the noise removal process, and gives the vocal tract prediction coefficient quantization signal <aa> to the multiplexing unit 130.
The "multiplexing unit 130" includes the above-mentioned frame power signal P, a gain code quantization signal, a voiceway prediction coefficient quantization signal <aa>, an index number Ip for selecting an adaptive excitation vector, and a gain. The index number Ig for code selection and the index number Is for noise excitation vector selection are multiplexed, and the multiplexed data obtained by this multiplexing is output as the encoded data of the CELP coding apparatus.
(Operation): The input audio digital signal S given to the input terminal 100 is obtained by the frame power calculation unit 101 for power in frame units, and is given to the multiplexing unit 130 as a frame power signal P. At the same time, the audio digital signal S is given to the autocorrelation matrix unit 102 to obtain the autocorrelation matrix R, and the autocorrelation matrix R is further given to the autocorrelation matrix adjusting unit 111. Further, the input voice vector signal S is also given to the voice / noise determination unit 110, where it is determined whether the input voice digital signal S is voice or noise other than voice. Is determined by using other bitch signals, voice path prediction coefficient a, predicted gain signal pg, and the like.
From the autocorrelation matrix R obtained by the autocorrelation matrix calculation unit 102, the vocal tract prediction coefficient a is obtained by the LPC analysis unit 103, and the vocal tract prediction coefficient a is used by the prediction gain calculation unit 112 to obtain the prediction gain signal pg. Is obtained and given to the voice / noise determination unit 110 together with the vocal tract prediction coefficient a. The input voice digital signal S hits the voice signal or hits the noise in the voice / noise determination unit 110 using the pitch signal pch, the voice path prediction coefficient a, the prediction gain signal pg, and the input digital signal S given from the adaptive codebook 105. Whether or not it is determined, and the voice / noise determination signal v is given to the autocorrelation matrix adjusting unit 111 and given to the autocorrelation matrix adjusting unit 111.
From the autocorrelation matrix R, the voice path prediction coefficient a, and the voice / noise determination signal v, the autocorrelation matrix of the section previously determined to be noise by the autocorrelation matrix adjustment unit 111 and the auto of the current frame. A new autocorrelation matrix Ra is obtained by combining with the correlation matrix. As a result, the autocorrelation matrix for the noise part that caused the harshness is optimally corrected.
A new autocorrelation matrix Ra is given to the LPC analysis unit 103, where a new optimal vocal tract prediction coefficient aa is obtained and given to the synthetic filter 104. A new optimum vocal tract prediction coefficient aa is given as a filter coefficient for the IIR type digital filter, and the excitation vector signal ex is filtered by the synthetic filter 104 to obtain the synthetic speech vector signal Sw.
The difference between the synthetic voice vector signal Sw and the input voice digital signal S is obtained by a subtractor, and this difference signal is given to the weighted distance calculation unit 108 as an error vector signal e. This error vector signal e corresponds to the optimum adaptive excitation vector signal, noise excitation vector signal, and gain code signal such that the squared sum vector signal E is minimized by frequency conversion and further weighting by the weighted distance calculation unit 108. The optimum index numbers Ia, Is, and Ig are obtained. These optimum index numbers Ia, Is, and Ig are given to the multiplexing unit 130, and the adaptive codebook 105, the noise codebook 106, and the noise codebook 106 are used to obtain the optimum excitation vectors ea, es and the gain code signals ga and gs. It is given to the gain code book 107.
The adaptive excitation vector signal ea read by the optimum index number Ia is multiplied by the gain code signal ga read by the index number Ig and given to the adder 115, and is read by the index number Is. The output noisy excitation vector signal es is also multiplied by the gain code signal gs read by the index number Ig and given to the adder 115. The two multiplied signals are added by the adder 115, and the excitation vector signal ex is given to the synthetic filter 104, where the synthetic voice vector signal Sw is obtained.
In this way, the synthetic voice vector signal using the adaptive code book 105, the noise code book 106, and the gain code book 107 until the error between the synthetic voice vector signal Sw and the input voice digital signal S disappears. Sw is generated, and in sections other than voice, the vocal tract prediction coefficient aa is optimally modified to generate a synthetic voice vector signal Sw.
The frame power signal P obtained by the above operation, the gain code quantization signal, the voice path prediction coefficient quantization signal <aa>, the index number Ip for adaptive excitation vector selection, and the gain code selection. The index number Ig for the purpose and the index number Is for the noise excitation vector selection are multiplexed every moment and output as encoded data.
(Details of voice / noise determination unit 110): The voice / noise determination unit 110 detects a noise section using a frame pattern, analysis parameters, and the like. Therefore, first, (1) the analysis parameters are converted into the reflection coefficient r [i] (i = 1, ..., Np, Np = filter order). Here, r [i] is -1.0 <r [i] <1.0 And.
Further, (2) the predicted gain RS using the reflection coefficient r [i] is determined. RS = Π (1.0-r [i]<sup>2</sup>)... (1) Can be represented by. Here, i = 1 to Np.
The reflection coefficient r [0] indicates the slope of the spectrum of the analysis frame signal, and it can be said that the closer | r [0] | is to 0, the flatter the spectrum. Generally, the noise spectrum has a smaller slope than the voice spectrum. Further, the predicted gain RS is a value close to 0 in the sounded section and a value close to 1.0 in the silent / noisy section.
Further, in applications such as mobile phone devices to which a CELP coding device is applied, since the distance between the human mouth, which is a voice source, and the microphone, which is a signal input unit, is short, the frame power is set to voice. It is large in the section and small in the silence (noise section).
Therefore, in determining voice / noise, D = Power · | r [0] | / RS ... (2) Is determined by Dth (threshold value) with respect to this value, and if D> Dth, it can be determined as voice, and if D <Dth, it can be determined as noise.
(Details of Autocorrelation Matrix Adjustment Unit 111): Next, when the adjustment of the autocorrelation matrix R in the above-mentioned autocorrelation matrix adjustment unit 111 is determined to be noise in a series of several m frames in the past. To do. When the autocorrelation matrix of the current frame is R [0] and the noise section autocorrelation matrix before n frames is R [n], the autocorrelation matrix Radj of the noise section after adjustment is Radj = Σ (Wi · R [i])... (3) i = 0 to m-1, ΣWi = 1.0, Wi W<sub>i + 1</sub> 0 Can be represented by.
The autocorrelation matrix adjusting unit 111 performs processing corresponding to such a calculation. The autocorrelation matrix Radj obtained by this process is given to the LPC analysis unit 103.
(Effect of the first embodiment): According to the first embodiment described above, when an input signal other than voice is encoded by a CELP-based coding device, the input signal is divided into frame units. Therefore, it is affected by vocal tract analysis (spectral analysis), and the analysis result differs from the actual one. Further, since the degree of difference in the analysis result varies from frame to frame, the signal after coding / decoding not only differs from the spectrum of the original voice, but also makes the sound jarring. By combining the autocorrelation matrix for spectrum estimation with that of past noise frames, it is possible to suppress the degree of difference in the analysis results between frames and prevent the generation of harsh synthetic sounds. In addition, since human hearing is more sensitive to noise in the fluctuating portion than in the stationary noise section, it is possible to suppress spectral fluctuations between noise frames.
"Second Embodiment": (B) Further, in the configuration of (A) described above, the noise interval composition filter coefficient is converted into the LSP (line spectrum pair: Line Spectrum Pair) coefficient, the spectrum characteristic of the composition filter is obtained, and the composition filter spectrum characteristic and the past. A new LSP coefficient that suppresses spectral fluctuations is obtained by collating with the spectral characteristics of the noise interval composite filter of, and after converting the new LSP coefficient into a composite filter coefficient, it is quantized and transmitted to the decoder, and the composite filter coefficient is used. It is configured to search for the optimum codebook vector.
Next, a detailed configuration for realizing the configuration (B) described above will be described. FIG. 2 is a functional configuration diagram of the CELP coding device. In FIG. 2, the configuration different from FIG. 1 described above is a portion particularly surrounded by a dotted line. That is, "in the portion surrounded by this dotted line", the autocorrelation matrix calculation unit 102, the LPC analysis unit 103A, the voice / noise determination unit 110A, the prediction gain calculation unit 112, and the vocal tract coefficient / LSP conversion unit. It includes 119, an LSP / vocal tract coefficient conversion unit 120, and an LSP coefficient adjustment unit 121.
Except for the portion surrounded by the dotted line, the configuration is substantially the same, and the same operation is performed. Therefore, "The vocal tract coefficient is corrected centering on the portion surrounded by the dotted line, and the conventional voice period is not used. It is stated that the cause of reproducing the jarring sound is eliminated by CELP coding of the noise part of time. "
Therefore, the vocal tract coefficient / LSP conversion unit 119 converts the vocal tract prediction coefficient a into the LSP coefficient l and gives it to the LSP coefficient adjustment unit 121. The "LSP coefficient adjusting unit 121" adjusts the LSP coefficient l from the voice / noise determination signal v from the voice / noise determination unit 110 and the LSP coefficient l from the vocal tract coefficient / LSP conversion unit 119 to adjust the noise. The effect is reduced and the adjusted LSP coefficient la is given to the LSP / vocal tract coefficient converter 120.
The LSP / vocal tract coefficient conversion unit 120 converts the adjusted LSP coefficient la from the LSP coefficient adjustment unit 121 into the optimum vocal tract prediction coefficient aa, and digitally filters the coefficient to the composite filter 104. Give as.
(Details of the LSP coefficient adjusting unit 121): The above-mentioned adjustment of the LSP coefficient is performed when it is determined to be noise in a series of several m frames in the past. Here, the current frame LSP coefficient is LSP-0 [i], the noise interval LSP coefficient before n frames is LSP-n [i], and the adjusted LSP coefficient is i = 1, ..., Np = filter order. When LSP<sub>adj</sub>[I] = ΣW<sub>k</sub> LSP-k [i]... (4) Here, k = 0 to m-1, ΣW<sub>k</sub>= 1.0, i = 0 to Np-1, W<sub>k</sub> W<sub>k + 1</sub>It can be represented by 0.
The LSP coefficient is a coefficient in the cosine region. The process corresponding to such a calculation is performed. The LSP coefficient la obtained by this process is given to the LSP / vocal tract coefficient conversion unit 120.
(Operation): The operation until the optimum vocal tract prediction coefficient aa is obtained will be described, and the generation of the optimum excitation vector signal ex by the codebook will be described because it is the same as that of the first embodiment described above. Is omitted. Therefore, first, the input audio digital signal S is given to the autocorrelation matrix unit 102, and the autocorrelation matrix R is obtained. This autocorrelation matrix R is given to the LPC analysis unit 103A, and the vocal tract prediction coefficient a is obtained. The vocal tract prediction coefficient a is given to the prediction gain calculation unit 112, the vocal tract coefficient / LSP conversion unit 119, and the voice / noise determination unit 110.
As a result, the predicted gain calculation unit 112 obtains the predicted gain signal pg and gives it to the voice / noise determination unit 110. The vocal tract coefficient / LSP conversion unit 119 obtains the LSP coefficient l from the vocal tract prediction coefficient a and gives it to the LSP coefficient adjusting unit 121. On the other hand, when the vocal tract prediction coefficient a, the input voice vector signal A, the pitch signal ptch, and the prediction gain signal pg are given, the voice / noise determination signal v is output and given to the LSP coefficient adjusting unit 121. Be done. The LSP coefficient adjusting unit 121 adjusts the LSP coefficient l to reduce the influence of noise, and the adjusted LSP coefficient la is given to the LSP / vocal tract coefficient converting unit 120. The LSP / vocal tract coefficient conversion unit 120 converts the LSP coefficient la into the optimum vocal tract prediction coefficient aa and gives it to the synthetic filter 104.
With this configuration, the vocal tract prediction coefficient of the noise section is optimally corrected as compared with the conventional case, and the coded signal that is the source of the jarring sound is not generated.
(Effect of Second Embodiment): According to the second embodiment described above, by adjusting the LSP coefficient which is directly related to the spectrum, it is the same as that of the first embodiment described above. In addition to being able to obtain such an effect, it is possible to reduce the amount of calculation because it is not necessary to perform the LPC analysis twice.
"Third Embodiment": (C) Furthermore, in the configuration of (A) described above, the new composite filter coefficient in the current noise section is directly calculated by interpolating the noise section composite filter coefficient with the past noise section composite filter coefficient, and the new composite filter coefficient is obtained. It is configured to be quantized and sent to the decoder to search for the optimal codebook vector using the new composite filter coefficients.
Next, a detailed configuration for realizing the configuration (C) described above will be described. FIG. 3 is a functional configuration diagram of the CELP coding device. In FIG. 3, the configuration different from FIG. 1 described above is a portion particularly surrounded by a dotted line. That is, the portion surrounded by the dotted line includes an autocorrelation matrix calculation unit 102, an LPC analysis unit 103A, a voice / noise determination unit 110, a prediction gain calculation unit 112, and a vocal tract coefficient adjustment unit 126. ing. The "vocal tract coefficient adjusting unit 126" reduces the influence of noise by reducing the vocal tract prediction coefficient from the voice tract prediction coefficient a from the LPC analysis unit 103A and the voice / noise determination signal v from the voice / noise determination unit 110. It is adjusted so as to give the optimum vocal tract prediction coefficient aa to the synthetic filter 104. That is, the new vocal tract prediction coefficient aa is directly obtained by combining the vocal tract prediction coefficient a with the vocal tract prediction coefficient of the past noise section.
Specifically, the adjustment of the vocal tract prediction coefficient is performed when it is determined that noise is generated in a series of several m frames in the past. Then, when the current frame composition filter coefficient is a-0 [i] and the noise interval composition filter coefficient before n frames is an [i], i = 1, ..., Np: Np = filter order. The adjusted filter coefficient is a<sub>adj</sub>[I] = ΣW<sub>k</sub> (Ak) [i]... (5) Here, ΣW<sub>k</sub>= 1.0, W<sub>k</sub> W<sub>k + 1</sub>It can be represented by 0, k = 0 to m-1, and i = 0 to Np-1.
At this time, it is necessary to confirm the stability of the filter using the adjusted coefficient, and it is preferable to control so that the adjustment is not performed when it is determined to be unstable.
(Operation): The operation until the optimum vocal tract prediction coefficient aa is obtained will be described, and the generation of the optimum excitation vector signal ex by the codebook will be described because it is the same as that of the first embodiment described above. Is omitted. Therefore, first, the input voice vector signal S is given to the autocorrelation matrix unit 102, and the autocorrelation matrix R is obtained. This autocorrelation matrix R is given to the LPC analysis unit 103A, and the vocal tract prediction coefficient a is obtained. The vocal tract prediction coefficient a is given to the prediction gain calculation unit 112, the vocal tract coefficient adjustment unit 126, and the voice / noise determination unit 110.
In the prediction gain calculation unit 112, the prediction gain coefficient pg is obtained from the vocal tract prediction coefficient a and is given to the voice / noise determination unit 110. In the voice / noise determination unit 110 to which the input voice digital signal S, the prediction gain coefficient pg, the vocal tract prediction coefficient a, and the pitch signal ptch are given, the voice / noise section is determined and the voice / noise is determined. The determination signal v is obtained and given to the vocal tract coefficient adjusting unit 126. From the voice / noise determination signal v and the vocal tract prediction coefficient a, the optimum vocal tract prediction coefficient aa adjusted so that the influence of noise can be reduced by the vocal tract coefficient adjusting unit 126 is obtained and given to the synthetic filter 104. ..
[0067] With this configuration, the vocal tract prediction coefficient of the noise section is optimally corrected as compared with the conventional case, and the coded signal that is the source of the jarring sound is not generated.
(Effect of Third Embodiment): According to the third embodiment described above, by directly combining the vocal tract coefficient with the vocal tract coefficient of the past noise section, the first embodiment described above is performed. The same effect as that of the above form can be obtained, and the amount of calculation can be reduced because the filter coefficient is directly calculated.
"Fourth Embodiment": (D) Voice / noise determination is performed for each subframe, the noise reduction amount and noise reduction method are determined based on this determination, the target signal vector is calculated according to the determined noise reduction method, and the target signal vector is used. A noise-reducing CELP-based voice encoder is configured so as to search for the optimum code book vector.
Next, a detailed configuration for realizing the configuration (D) described above will be described. FIG. 4 is a functional configuration diagram of the CELP coding device. In FIG. 4, a configuration different from that in FIG. 1 described above is a portion surrounded by a dotted line. That is, the portion surrounded by the dotted line includes a voice / noise determination unit 110B, a noise reduction filter 122, a prediction gain calculation unit 112, a filter bank 124, and a filter control unit 125.
The filter bank 124 is composed of bandpass filters a to n, each of which has a different pass band, and the bandpass filter a outputs a passband signal Sbp1 with respect to the input audio digital signal S. Then, ..., The bandpass filter n outputs the passband signal SbpN with respect to the input audio digital signal S and gives it to the audio / noise determination unit 110B. With such a filter bank configuration, noise in the blocking band is reduced, a passband signal with an increased SN ratio is output, and the voice / noise determination unit 110B determines the voice section / noise section for each passband. Can be easily done.
The prediction gain calculation unit 112 obtains the prediction gain coefficient pg from the vocal tract prediction coefficient a from the LPC analysis unit 103A and gives it to the voice / noise determination unit 110B. The voice / noise determination unit 110B calculates each band noise evaluation function from the passband signals Sbp1 to SbpN from the filter bank 124, the pitch signal ptch, and the predicted gain coefficient pg, and voice / noise for each band. The determination signals v1 to vN are output and given to the filter control unit 125.
The filter control unit 125 adjusts the noise reduction filter coefficient according to the voiced / unvoiced / noise determination for each band from the voice / noise determination signals v1 to vN from the voice / noise determination unit 110B. The adjusted noise reduction filter coefficient nc is given to the noise reduction filter 122. The noise reduction filter 122 is composed of an IIR type or FIR type digital filter, sets a noise reduction filter coefficient nc from the filter control unit 125, and optimally processes the input audio digital signal S by this filter coefficient to reduce noise. The target signal t is output and given to the subtractor 116.
(Operation): The operation until the target signal t is obtained will be described, and the generation of the optimum excitation vector signal ex by the codebook will be the same as that of the first embodiment described above, and thus the description thereof will be omitted. Therefore, first, the input audio digital signal S is given to the autocorrelation matrix unit 102, and the autocorrelation matrix R is obtained. This autocorrelation matrix R is given to the LPC analysis unit 103A, and the vocal tract prediction coefficient a is obtained. The vocal tract prediction coefficient a is given to the prediction gain calculation unit 112 and the composite filter 104, and the prediction gain coefficient pg is obtained by the prediction gain calculation unit 112 and given to the voice / noise determination unit 110B.
On the other hand, the input audio digital signal S is given to the filter bank 124, where the bandpass filters Sbp1 to SbpN are output by the bandpass filters a to n. These band-passing signals Sbp1 to SbpN, pitch signals ptch, and predicted gain coefficient pg are given to the voice / noise determination unit 110B, and voice / noise determination signals v1 to vN for each band are obtained. Using these voice / noise determination signals v1 to vN, the filter control unit 125 adjusts the noise reduction filter coefficient, and the noise reduction filter coefficient nc is given to the noise reduction filter 122.
With this noise reduction filter coefficient nc, the noise reduction filter 122 is optimally set as the filter coefficient of the digital filter so that noise reduction can be performed optimally. With this setting, the noise reduction filter 122 performs filter processing on the input audio digital signal S to obtain the target signal t. The difference e between the target signal t and the synthetic speech signal Sw from the synthetic filter 104 is obtained by the subtractor 116, and the weighted distance calculation unit 108 searches for the optimum index based on this error signal e.
With this configuration, the noise in the noise section is reduced as compared with the conventional case, and the coded signal that is the source of the jarring sound is not generated.
(Effect of the Fourth Embodiment): According to the above-mentioned fourth embodiment, the degree of discomfort is higher than that when only the background noise in the voice section is heard in terms of human hearing. Few. Therefore, by distinguishing the voice section at the time of coding and changing the noise reduction method between the noise section and the voice section, it is possible to improve the audible sound quality without performing complicated processing in the voice section.
Further, by performing noise reduction only on the target signal of the CELP coding device, it is possible to perform noise reduction in subframe units, and it is possible to reduce the sound when the voice / noise determination is incorrect. The influence of the spectrum distortion due to noise reduction can be reduced as well as the influence of the influence.
(Other Embodiments): (1) In the above embodiment, a pulse code book is further provided, and a pulsed excitation vector is used as a waveform code vector to generate a synthetic speech vector. It is also preferable to configure it as follows.
(2) Further, although it has been described that the composite filter 104 of FIG. 1 described above is composed of an IIR type digital filter, other than the FIR (Finite Impulse Response) type digital filter and the IIR type. It is also preferable to use a composite digital filter with an FIR type.
(3) Further, in the above-mentioned CELP coding apparatus, it is also preferable to provide a statistical code book for CELP coding. Such a configuration and a method of creating a statistical code book can also be realized by, for example, the configuration and a method of creating the statistical code book shown in Japanese Patent Application Laid-Open No. 6-130995.
(4) Further, in the above-described embodiment, the CELP encoding device has been described in detail. However, the configuration of the decoding device is described, for example, in the reference document: JP-A-5-165497, "Code Excitation". It can also be decoded with the configuration shown in "Linear Predictive Coder and Decoder".
(5) Further, although the application to the CELP coding apparatus is shown in the above-described embodiment, it can also be applied to the VS (vector sum) ELP coding apparatus. It can also be applied to LD (low delay) -CELP, CS (conjugated structure) -CELP, PSI (pitch synchronization noise) -CELP.
(6) Further, the CELP coding device of the above-described embodiment is effective when applied to a mobile phone or the like, and such a configuration is described in, for example, Japanese Patent Application Laid-Open No. 6-130998 "Compressed voice decoding". It is also effective when applied to the TDMA transmitting device and receiving device shown in "Chemical device". It is also preferable to apply the present invention to a VSELP TDMA transmitter.
(7) Furthermore, as the noise reduction filter 122 of FIG. 4 described above, it is also preferable to apply a Kalman filter in addition to the realization of the IIR type, FIR type, and IIR / FIR composite type digital filter. This Kalman filter is applicable if signal and noise statistics are given, and has the effect of being able to perform optimum operation even when the signal and noise statistics are given so as to change over time.
【0087】
According to the invention of claim 1 as described above, an autocorrelation analysis means for obtaining autocorrelation information from an input acoustic signal and a voice for obtaining an autocorrelation prediction coefficient from the analysis results of the autocorrelation analysis means. The non-voice signal section of the input acoustic signal is detected from the road prediction coefficient analysis means, the prediction gain coefficient analysis means for obtaining the prediction gain coefficient from the voice path prediction coefficient, and the input acoustic signal, the voice path prediction coefficient, and the prediction gain coefficient. An autocorrelation adjusting means for changing and adjusting the autocorrelation information in this non-voice signal section to a weighted composite value of the autocorrelation information in this non-voice signal section and the autocorrelation information in the past non-voice signal section, and adjustment. Using a voiceway prediction coefficient compensating means to obtain a compensated post-voiceway prediction coefficient that compensates for the voiceway prediction coefficient in the non-voice signal section from the later autocorrelation information, and a post-compensated voiceway prediction coefficient and an adaptive excitation signal. By providing a coding means for code excitation linear predictive coding of the input acoustic signal, it is possible to adjust the autocorrelation information of the acoustic signal in the non-audio signal section and reduce the influence of the acoustic signal in the non-audio signal section. Can be done.
Further, according to the invention of claim 2, the autocorrelation analysis means for obtaining the autocorrelation information from the input acoustic signal and the voiceway prediction coefficient analysis means for obtaining the voiceway prediction coefficient from the analysis result of the autocorrelation analysis means. The predicted gain coefficient analysis means for obtaining the predicted gain coefficient from the voice tract prediction coefficient, the LSP coefficient from the voice tract prediction coefficient, and the non-voice of the input acoustic signal from the input acoustic signal, the voice tract prediction coefficient, and the predicted gain coefficient. LSP coefficient adjusting means for detecting a signal section and changing and adjusting the LSP coefficient in this non-voice signal section to a weighted composite value of the LSP coefficient in this non-voice signal section and the LSP coefficient in the past non-voice signal section, and adjustment Input acoustics using a voiceway prediction coefficient compensating means for obtaining a compensated post-voiceway prediction coefficient that compensates for the voiceway prediction coefficient in the non-voice signal section from the later LSP coefficient, and a post-compensation voiceway prediction coefficient and an adaptive excitation signal. By providing a coding means for coding the signal with code excitation linear prediction coding, it is possible to reduce the influence of the acoustic signal in the non-voice signal section because the spectral fluctuation in the non-voice signal section is suppressed at the stage of the LSP coefficient. it can.
Further, according to the invention of claim 3, the autocorrelation analysis means for obtaining the autocorrelation information from the input acoustic signal and the voiceway prediction coefficient analysis means for obtaining the voiceway prediction coefficient from the analysis result of the autocorrelation analysis means. The non-voice signal section is detected from the input acoustic signal, the predicted gain coefficient, and the voice path prediction coefficient, and the voice path prediction in this non-voice signal section is performed by the prediction gain coefficient analysis means for obtaining the predicted gain coefficient from the voice path prediction coefficient. Voiceway coefficient adjusting means for obtaining the adjusted voiceway prediction coefficient by changing and adjusting the coefficient to a weighted composite value of the voiceway prediction coefficient in this non-voice signal section and the voiceway prediction coefficient in the past non-voice signal section. And by providing a coding means to code excitation linear prediction coding of the input acoustic signal using the adjusted voice tract prediction coefficient and the adaptive excitation signal, the non-voice signal section directly from the voice tract prediction coefficient. Since the voice path prediction coefficient of can be adjusted, the influence on the coded output of the acoustic signal in the non-voice signal section can be reduced while the amount of calculation is very small.
Furthermore, according to the invention of claim 4, the autocorrelation analysis means for obtaining the autocorrelation information from the input acoustic signal and the voice tract prediction coefficient analysis for obtaining the voice tract prediction coefficient from the analysis result of the autocorrelation analysis means. From the means, the predicted gain coefficient analysis means for obtaining the predicted gain coefficient from the voice tract prediction coefficient, the band passing processing signal obtained by band passing processing from the input acoustic signal, and the predicted gain coefficient, for each band passing processing signal. A non-voice signal section is detected, a filter coefficient for noise removal is generated according to the detection result of the non-voice signal section for each band passage processing signal, and the filter coefficient generated for the input acoustic signal is used. A noise removing means for generating a target signal for generating a synthetic voice signal by performing noise removal, a synthetic voice generating means for generating the synthetic voice signal using the voice tract prediction coefficient, and a voice tract prediction coefficient. By providing a coding means for code excitation linear predictive coding of the input acoustic signal using the target signal and the target signal, a noise-removed target signal can be obtained, so that the coded output of the non-voice signal section can be obtained. Can be prevented from being affected by noise.
Therefore, a code excitation linear predictive coding device capable of reducing the influence on the coded output of an acoustic signal (noise, rotating sound, vibration sound, etc.) in a non-voice signal section and performing good voice reproduction is provided. It can be realized.
[Simple explanation of drawings]
FIG. 1 is a functional configuration diagram of a CELP coding apparatus according to the first embodiment of the present invention.
FIG. 2 is a functional configuration diagram of a CELP coding apparatus according to a second embodiment of the present invention.
FIG. 3 is a functional configuration diagram of a CELP coding apparatus according to a third embodiment of the present invention.
FIG. 4 is a functional configuration diagram of a CELP coding apparatus according to a fourth embodiment of the present invention.
[Explanation of symbols] 100 ... Input terminal, 101 ... Frame power calculation unit, 102 ... Autocorrelation matrix calculation unit, 103 ... LPC analysis unit, 104 ... Synthesis filter, 105 ... Adaptive codebook, 106 ... Noise codebook, 107 ... Gain codebook, 108 ... Weighted distance calculation unit, 109 ... LSP quantizer, 110 ... Voice / noise determination unit, 111 ... Autocorrelation matrix adjustment unit, 112 ... Prediction gain calculation unit, 113, 114 ... Multiplier, 115, 116 ... Adder, 117 ... Quantizer.
Continuation of front page (58) Surveyed field (Int.Cl.<sup>7</sup>, DB name) G10L 11/00 G10L 19/12
2 priority claims, no other members on record
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 21451795 | Japan | A | |
| JP19950214517 | – | – | – |
7 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| First payment of annual fees (during grant procedure)A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)A01 | A01 |
Numbers
- Publication, DOCDB
- 3522012
- Publication, EPODOC
- JP3522012B
- Application
- 21451795
- Application, DOCDB
- 21451795
- Application, EPODOC
- JP19950214517
Titles
- English
- Code excitation linear-predictive-coding equipment
Classification
- CPC, 3
- G10L19/12
- G10L19/09
- G10L25/78
- IPC, 11
- G10L19 00
- G10L19 038
- G10L19 04
- G10L19 12
- G10L19 125
- G10L19 22
- G10L25 00
- G10L25 78
- G10L25 84
- G10L25 90
- H03M7 30