Method and device to discriminate voice and noise, method and device to reduce noise, voice and noise discriminating program, noise reducing program, and recording medium for program
Abstract
Problem to be solved.To provide a method and a device to discriminate voice and noise with which it is discriminated whether voice and noise that includes non-stationary noise exist in input signals employing one input or not and whether noise including the non-stationary noise is reduced using the discrimination result and to provide a method and a device to reduce noise, a voice and noise discriminating program, a noise reducing program and a recording medium for the programs.
Solution.The noise reducing device includes a plurality featured values computing section 11 which computes a plurality of featured values for voice and noise mixed signals, a featured values analysis section 13 which analyzes information related to voice and noise using the plurality of featured values and input voice and noise mixed signals, a reduction variable computing section 14 which computes reduction variables corresponding to the plurality of noise reduction processes using the analyzed information and the input voice and noise mixed signals and a plurality noise reducing section 15 which on the other hand reduces noise in the input voice and noise mixed signals by a plurality of noise reducing processes employing the computed reduction variables.
Copyright (C)2006,JPO&NCIPI
Term
No projected expiry on record.
- Priority
- Filed
- Published
- Today
20 claims: 8 independent, 12 dependent
- 1For a voice noise mixed signal in which a target voice signal and an unnecessary noise signal are mixed, a plurality of feature quantities are calculated, and information on voice noise is obtained using the plurality of feature quantities and the input voice noise mixed signal. Voice noise discrimination method to be analyzed. 目的となる音声信号と不要な雑音信号の混在する音声雑音混在信号に対して、複数の特徴量を計算し、複数の特徴量と入力される音声雑音混在信号とを用いて音声雑音に関する情報を分析する音声雑音判別方法。
- 3For a voice noise mixed signal in which a target voice signal and an unnecessary noise signal are mixed, a plurality of feature quantities are calculated, and information on voice and noise is obtained using the voice noise mixed signal input with the plurality of feature quantities. We analyzed and calculated the reduction variables corresponding to multiple noise reduction processes using the analyzed information and the input voice noise mixed signal, while using the calculated reduction variables for the input voice noise mixed signal. A noise reduction method characterized in that noise is reduced by a plurality of noise reduction processes, and mixed voice noise signals with reduced noise are added and output. 目的となる音声信号と不要な雑音信号の混在する音声雑音混在信号に対して、複数の特徴量を計算し、複数の特徴量と入力される音声雑音混在信号を用いて音声および雑音に関する情報を分析し、分析した情報と入力される音声雑音混在信号とを用いて複数の雑音低減処理に対応した低減変数を計算し、一方、入力される音声雑音混在信号を計算された低減変数を用いた複数の雑音低減処理で雑音を低減し、雑音を低減した音声雑音混在信号を足し合わせて出力することを特徴とする雑音低減方法。
- 5For a voice noise mixed signal in which a target voice signal and an unnecessary noise signal are mixed, a multiple feature amount calculation unit that calculates a plurality of feature amounts and a voice noise mixed signal that is input with a plurality of feature amounts are used. A voice noise discrimination device including a feature amount analysis unit that analyzes information related to voice and noise . 目的となる音声信号と不要な雑音信号の混在する音声雑音混在信号に対して、複数の特徴量を計算する複数特徴量計算部と、複数の特徴量と入力される音声雑音混在信号とを用いて音声および雑音に関する情報を分析する特徴量分析部を具備する音声雑音判別装置。
- 7For a voice noise mixed signal in which a target voice signal and an unnecessary noise signal are mixed, a multiple feature amount calculation unit that calculates a plurality of feature amounts, and a plurality of feature amounts and an input voice noise mixed signal are used. A feature quantity analysis unit that analyzes information related to voice and noise, and a reduction variable calculation unit that calculates reduction variables corresponding to multiple noise reduction processes using the analyzed information and input voice noise mixed signals, while inputting. A noise reduction device including a plurality of noise reduction units that reduce noise by a plurality of noise reduction processes using a calculated reduction variable for the voice noise mixed signal. 目的となる音声信号と不要な雑音信号の混在する音声雑音混在信号に対して、複数の特徴量を計算する複数特徴量計算部と、複数の特徴量および入力される音声雑音混在信号を用いて音声および雑音に関する情報を分析する特徴量分析部と、分析した情報および入力される音声雑音混在信号を用いて複数の雑音低減処理に対応した低減変数を計算する低減変数計算部と、一方、入力される音声雑音混在信号を計算された低減変数を用いて複数の雑音低減処理で雑音を低減する複数雑音低減部とを具備する雑音低減装置。
- 13For a mixed voice noise signal, the power, cepstrum, and frequency characteristics are calculated, the voice section is determined using the calculated power and kepstram, and the calculated cepstrum, frequency characteristics, voice section information, and input voice are used. A voice noise discrimination program that gives a command to a computer to determine a sudden noise section using a noise mixture signal. 音声雑音混在信号に対して、パワー、ケプストラム、周波数特性を計算し、計算されたパワー、ケプストラムを用いて音声区間を判定し、計算されたケプストラム、周波数特性、音声区間の情報、入力される音声雑音混在信号を用いて突発性雑音区間を判定する指令をコンピュータに対してする音声雑音判別プログラム。
- 15Power, Kepstram, and frequency characteristics are calculated for the mixed voice noise signal, and the voice section is determined using the calculated power and Kepstram, and the calculated Kepstram, frequency characteristics, voice section information, and input voice noise are used. The sudden noise section is determined using the mixed signal, and the signal suppression gain and the sound immediately before the sudden noise is present at the position where the sudden noise is present are used to determine the signal suppression gain using the information of the analyzed and judged voice section and the sudden noise section. Calculates the variables for periodic waveform insertion that repeatedly inserts the periodic waveform of, while suppressing the signal of the input voice noise mixed signal using the calculated signal suppression gain, and for the calculated periodic waveform insertion. A noise reduction processing program that inserts a periodic waveform using variables and issues a command to the computer to add and output a mixed voice noise signal with reduced noise. 音声雑音混在信号に対してパワー、ケプストラム、周波数特性を計算し、計算されたパワー、ケプストラムを用いで音声区間を判定し、計算されたケプストラム、周波数特性、音声区間の情報、入力される音声雑音混在信号を用いて突発性雑音区間を判定し、分析判定した音声区間および突発性雑音区間の情報を用いて信号抑圧ゲインと、突発性雑音の存在する位置に突発性雑音の存在する直前の音声の周期波形を繰り返し挿入する周期波形挿入に対する変数とを計算し、一方、入力される音声雑音混在信号の信号を計算された信号抑圧ゲインを用いて抑圧し、また、計算された周期波形挿入に対する変数を用いて周期波形を挿入し、雑音を低減した音声雑音混在信号を足し合わせて出力する指令をコンピュータに対してする雑音低減処理プログラム。
- 17For a voice noise mixed signal in which a target voice signal and an unnecessary noise signal are mixed, a plurality of feature quantities are calculated, and information on voice and noise is obtained using the voice noise mixed signal input with the plurality of feature quantities. It is analyzed and the signal suppression gain is calculated using the information of the voice section and the sudden noise section, while the signal of the input voice noise mixed signal is suppressed and calculated using the calculated signal suppression gain. Using the information of the voice section and the sudden noise section, the voice average spectrum is calculated in the case of the voice section, and in the case of the section where the sudden noise is mixed in the voice, the voice average spectrum is suppressed. A noise reduction processing program that gives a command to the computer to add up the noise-reduced signals and output them. 目的となる音声信号と不要な雑音信号の混在する音声雑音混在信号に対して、複数の特徴量を計算し、複数の特徴量と入力される音声雑音混在信号を用いて音声および雑音に関する情報を分析し、音声区間、突発性雑音区間の情報を用いて信号抑圧ゲインを計算し、一方、入力される音声雑音混在信号の信号を計算された信号抑圧ゲインを用いて抑圧し、また、計算された音声区間、突発性雑音区間の情報を用いて、音声区間の場合には音声平均スペクトルを計算し、音声に突発性雑音が混在している区間の場合には、音声平均スペクトルまで抑圧し、各々雑音を低減した信号を足し合わせて出力すべき指令を、コンピュータに対してする雑音低減処理プログラム。
- 19For a voice noise mixed signal in which a target voice signal and an unnecessary noise signal are mixed, the power, kepstram, and frequency characteristics are calculated, the voice section is determined using the calculated power and kepstram, and the calculated kepstram is calculated. , Frequency characteristics, voice section information, and input voice noise mixed signal are used to determine the sudden noise section, and the calculated voice section and sudden noise section information are used to calculate the signal suppression gain, while , The signal of the input voice noise mixed signal is suppressed using the calculated signal suppression gain, and the voice average spectrum in the case of the voice section using the calculated voice section and sudden noise section information. In the case of a section where sudden noise is mixed in the voice, the noise reduction gives a command to the computer to suppress the voice average spectrum and add the signals with reduced noise to output. Processing program. 目的となる音声信号と不要な雑音信号の混在する音声雑音混在信号に対して、パワー、ケプストラム、周波数特性を計算し、計算されたパワー、ケプストラムを用いて音声区間を判定し、計算されたケプストラム、周波数特性、音声区間の情報、入力される音声雑音混在信号を用いて突発性雑音区間を判定し、計算された音声区間、突発性雑音区間の情報を用いて信号抑圧ゲインを計算し、一方、入力される音声雑音混在信号の信号を計算された信号抑圧ゲインを用いて抑圧し、また、計算された音声区間、突発性雑音区間の情報を用いて、音声区間の場合には音声平均スペクトルを計算し、音声に突発性雑音が混在している区間の場合には、音声平均スペクトルまで抑圧し、各々雑音を低減した信号を足し合わせて出力すべき指令を、コンピュータに対してする雑音低減処理プログラム。
Independent claims8
27 paragraphs, as filed
The present invention relates to a voice noise discrimination method and device, a noise reduction method and device, a voice noise discrimination program, a noise reduction program, and a recording medium of the program.
As a conventional example 1 of a noise reduction device, for example, there is a noise reduction device for stationary noise (Patent Document 1). reference). This will be briefly explained with reference to FIG. The audio signal S (n), which is the target signal, and the unnecessary ambient noise N (n) such as air conditioning are input as the input signal X (n) = S (n) + N (n). Here, n is an integer value that represents the time representation of the signal as discrete time. This input signal X (n) is converted into a frequency domain signal X (ω) by the frequency domain conversion unit 61, for example, by discrete Fourier transform every short time. ω represents the frequency. The power spectrum Pavx (ω) of the frequency domain signal X (ω) is calculated by the input signal power spectrum calculation unit 62, and the noise power spectrum Pavn (ω) in the frequency domain signal X (ω) is calculated by the noise power spectrum estimation unit 63. ) Is estimated. The loss calculation unit 64 calculates the loss value L (ω) using Pavx (ω) and Pavn (ω), and transfers the loss value to the loss insertion unit 65. In the loss insertion unit 65, the output Y (ω) with reduced noise is calculated by Y (ω) = L (ω) × X (ω) using the loss value L (ω) calculated by the loss calculation unit 64. Is output. The output Y (ω) is converted into the time domain by the time domain conversion unit 66, and the noise-reduced signal Y (n) is output.
According to Conventional Example 1, constant noise can be reduced. However, it is difficult to reduce non-stationary noise because the noise power fluctuates greatly with time. As a conventional example 2, there is a noise reduction device that discriminates a noise section including unsteady noise (see Patent Document 2). This will be briefly explained with reference to FIG. The first sound receiver 71 composed of a microphone array device is composed of a microphone array 72 composed of a plurality of microphone elements and a directivity control unit 73. 74 is the second sound receiver, and these two sound receivers are installed in the same place. A typical example of a microphone array device having a directivity control function is a sound receiver called an adaptive array. Adaptive arrays provide directivity with low sensitivity in the direction of the noise source. As a result, the fluctuation of the SN ratio can be kept small with respect to the position of the noise source and the movement of the speaker. That is, the first receiver output outputs a signal having a larger SN ratio than the second receiver output.
The voice on which noise is superimposed is received by the microphone array 72. The output signal of this microphone array 72 is input to the directivity control unit 73, and the first signal x<sub>1 </sub>Occurs. On the other hand, the output of one microphone element that constitutes the microphone array 72 is x.<sub>2 </sub>And. At this time, as a result of the directivity control by the directivity control unit 73, x<sub>1 </sub>The signal-to-noise ratio in is x<sub>2 </sub>It is larger than the SN ratio in. Next, in the short-time power calculation units 75 and 76, x, respectively.<sub>1 </sub>And x2 short power P<sub>1 </sub>And P<sub>2 </sub>Is calculated and output. The voice section detection unit 77 can detect the voice section by obtaining the difference in power between the two signals.
According to the conventional example 2, the voice section can be detected even for the voice on which the non-stationary noise is superimposed. In addition, noise can be reduced by forming a directivity with low sensitivity in the direction of the noise source using a microphone array. However, it is necessary to install multiple microphones.<patcit num="1"><text>Japanese Unexamined Patent Publication No. 9-258792</text></patcit><patcit num="2"><text>Patent No. 2913105</text></patcit>
<p> A noise reduction device for stationary noise as in Conventional Example 1 can reduce stationary noise superimposed on voice with one input, but cannot reduce non-stationary noise with large temporal fluctuation. A noise reduction device using a plurality of microphones as in Conventional Example 2 can detect a voice section regardless of unsteady noise by directivity control. Further, noise can be reduced by forming a directivity characteristic having low sensitivity in the direction of the noise source. However, it is necessary to install multiple microphones. A general communication and sound collecting device has one microphone, and a noise reducing device with one input is desired in order to avoid an increase in hardware scale and processing calculation amount.</p><p> The present invention discriminates whether or not voice and noise including non-stationary noise are present in an input signal with one input, and reduces noise including non-stationary noise by using the discriminated result. Voice noise Provided are a discrimination method and device, a noise reduction method and device, a voice noise discrimination program, a noise reduction program, and a recording medium for the program.</p>
<p> Claim 1: For a voice noise mixed signal in which a target voice signal and an unnecessary noise signal are mixed, a plurality of feature quantities are calculated, and the voice is voiced using the plurality of feature quantities and the input voice noise mixed signal. A voice noise discrimination method for analyzing information about noise was constructed. Claim 2: In the voice noise discrimination method according to claim 1, the power, cepstrum, and frequency characteristics are calculated for the voice noise mixed signal, and the voice section is determined and calculated using the calculated power and cepstrum. A voice noise discrimination method for determining a sudden noise section was constructed using the input cepstrum, frequency characteristics, voice section information, and input voice noise mixed signal.</p><p> Claim 3: For a voice noise mixed signal in which a target voice signal and an unnecessary noise signal are mixed, a plurality of feature amounts are calculated, and the voice and the input voice noise mixed signal are used with the plurality of feature amounts. Information on noise is analyzed, and reduction variables corresponding to multiple noise reduction processes are calculated using the analyzed information and the input voice noise mixed signal, while the input voice noise mixed signal is calculated reduction. A noise reduction method was constructed in which noise was reduced by a plurality of noise reduction processes using variables, and the noise-reduced mixed voice noise signals were added and output. Claim 4: In the noise reduction method according to claim 3, the power, kepstram, and frequency characteristics are calculated for the voice noise mixed signal, and the voice section is determined and calculated using the calculated power and kepstram. The sudden noise section is determined using the Keptram, frequency characteristics, voice section information, and input voice noise mixed signal, and the calculated voice section and sudden noise section information are used to obtain the signal suppression gain and the sudden noise. Calculates the variables for periodic waveform insertion that repeatedly inserts the periodic waveform of the voice immediately before the presence of sudden noise at the position where the sexual noise exists, while the signal of the input voice noise mixed signal is the calculated signal suppression. A noise reduction method was constructed in which the noise was suppressed by using the gain, the periodic waveform was inserted using the variable for the calculated periodic waveform insertion, and the mixed voice noise signals with reduced noise were added and output.</p><p> Claim 5: A plurality of feature amount calculation unit 11 for calculating a plurality of feature amounts for a voice noise mixed signal in which a target voice signal and an unnecessary noise signal are mixed, and a voice noise input with a plurality of feature amounts. A voice noise discrimination device including a feature analysis unit 13 for analyzing information on voice and noise using mixed signals was configured. Claim 6: In the voice noise discriminating device according to claim 5, a power calculation unit 21 that calculates power for a voice noise mixed signal, a cepstrum calculation unit 22 that calculates cepstrum, and a frequency that calculates frequency characteristics. Using the characteristic calculation unit 23, the voice section determination unit 24 that determines the voice section using the calculated power and cepstrum, the calculated cepstrum, frequency characteristics, voice section information, and the input voice noise mixed signal. A voice noise discrimination device including a sudden noise determination unit 25 for determining a sudden noise section was configured.</p><p> Claim 7: A plurality of feature quantity calculation unit 11 for calculating a plurality of feature quantities for a voice noise mixed signal in which a target voice signal and an unnecessary noise signal are mixed, and a plurality of feature quantities and input voice noise. Feature quantity analysis unit 13 that analyzes information related to voice and noise using mixed signals, and reduction variable calculation that calculates reduction variables corresponding to multiple noise reduction processes using the analyzed information and input voice noise mixed signals. A noise reduction device including a unit 14 and a plurality of noise reduction units 15 that reduce noise by a plurality of noise reduction processes using a calculated reduction variable for the input voice noise mixed signal is configured.</p><p> Claim 8: In the noise reduction device according to claim 7, the power calculation unit 21 that calculates the power for the voice noise mixed signal, the Kepstram calculation unit 22 that calculates the Kepstram, and the frequency characteristic that calculates the frequency characteristic. Suddenly using the calculation unit 23, the voice section determination unit 24 that determines the voice section using the calculated power and the keptram, the calculated kepstram, the frequency characteristics, the voice section information, and the input voice noise mixed signal. The sudden noise determination unit 25 that determines the sexual noise section, the calculated voice section information, the sudden noise section information, the signal suppression gain using the input voice noise mixed signal, and the position where the sudden noise exists. Using the reduction variable calculation unit 14 that calculates the variable for the periodic waveform insertion that repeatedly inserts the periodic waveform of the voice immediately before the presence of sudden noise, and the signal suppression gain that calculates the input voice noise mixed signal. A noise reduction device including a signal suppression unit 26 that suppresses the signal and a periodic waveform insertion unit 27 that inserts a periodic waveform using a variable for the calculated periodic waveform insertion is configured.</p><p> Claim 9: In the noise reduction method according to claim 3, the signal suppression gain is calculated using the information of the voice section and the sudden noise section, while the signal of the input voice noise mixed signal is calculated. Suppressed using the signal suppression gain, and using the calculated information of the voice section and the sudden noise section, the voice average spectrum is calculated in the case of the voice section, and the voice is mixed with the sudden noise. In the case of a section, a noise reduction method was constructed in which the noise average spectrum was suppressed, and the noise-reduced signals were added and output. Claim 10: In the noise reduction method according to claim 1, the power, kepstram, and frequency characteristics are calculated for the voice noise mixed signal, and the voice section is determined and calculated using the calculated power and kepstram. The sudden noise section is determined using the Keptram, frequency characteristics, voice section information, and input voice noise mixed signal, and the signal suppression gain is calculated using the calculated voice section and sudden noise section information. On the other hand, the signal of the input voice noise mixed signal is suppressed by using the calculated signal suppression gain, and the calculated voice section and the information of the sudden noise section are used to suppress the sound in the case of the voice section. A noise reduction method was constructed in which the average spectrum was calculated, and in the case of a section in which sudden noise was mixed in the sound, the sound average spectrum was suppressed, and the signals with reduced noise were added and output.</p><p> 11.: In the noise reduction device according to claim 7, a power calculation unit that calculates power for a voice noise mixed signal, a cepstrum calculation unit that calculates cepstrum, and a frequency characteristic calculation unit that calculates frequency characteristics. And sudden noise using the calculated power, the voice section determination unit that judges the voice section using the sub-cepstrum, the calculated cepstrum, the frequency characteristic, the voice section information, and the input voice noise mixed signal. The sudden noise determination unit that determines the section, the calculated voice section, the information of the sudden noise section, and the reduction variable calculation unit that calculates the signal suppression gain using the input voice noise mixed signal, while being input. The signal suppressor that suppresses the signal of the mixed voice noise signal using the calculated signal suppression gain, and the voice average spectrum in the case of the voice section using the calculated voice section and sudden noise section information. In the case of a section in which sudden noise is mixed in the sound, a noise reduction device including a band-specific suppression unit that suppresses the sound average spectrum is configured.</p><p> Claim 12: In the noise reduction device according to claim 8, a reduction variable for calculating a signal suppression gain and a variable for periodic waveform insertion using information on a voice section and a sudden noise section, and an input voice noise mixed signal. The calculation unit, on the other hand, the signal suppression unit that suppresses the input voice noise mixed signal signal using the calculated signal suppression gain, and the voice using the calculated voice section and sudden noise section information. In the case of a section, the voice average spectrum is calculated, and in the case of a section in which sudden noise is mixed in the voice, a noise reduction device including a band-specific suppression unit that suppresses the voice average spectrum is configured.</p><p> Claim 13: For a mixed voice noise signal, the power, cepstrum, and frequency characteristics are calculated, the voice section is determined using the calculated power and kepstram, and the calculated cepstrum, frequency characteristics, and voice section information, We constructed a voice noise discrimination program that gives a command to the computer to judge the sudden noise section using the input voice noise mixed signal. 14: A recording medium on which the voice noise discrimination program according to claim 13 is recorded is configured. Claim 15: Power, Keptram, and frequency characteristics are calculated for a mixed voice noise signal, the voice section is determined using the calculated power and Kepstram, and the calculated Kepstram, frequency characteristics, and voice section information and input are used. The sudden noise section is determined using the mixed voice noise signal, and the signal suppression gain and the presence of the sudden noise at the position where the sudden noise exists are used by using the information of the voice section and the sudden noise section analyzed and judged. The variable for the periodic waveform insertion that repeatedly inserts the periodic waveform of the voice immediately before is calculated, while the signal of the input voice noise mixed signal is suppressed using the calculated signal suppression gain, and is also calculated. A noise reduction processing program was constructed that inserts a periodic waveform using variables for the periodic waveform insertion, and issues a command to the computer to add and output a mixed voice noise signal with reduced noise.</p><p> 16: A recording medium on which the noise reduction processing program according to claim 15 is recorded is configured. Claim 17: For a voice noise mixed signal in which a target voice signal and an unnecessary noise signal are mixed, a plurality of feature amounts are calculated, and the voice and the input voice noise mixed signal are used with the plurality of feature amounts. Information on noise is analyzed, the signal suppression gain is calculated using the information in the voice section and the sudden noise section, while the signal of the input voice noise mixed signal is suppressed using the calculated signal suppression gain. In addition, using the calculated information of the voice section and the sudden noise section, the voice average spectrum is calculated in the case of the voice section, and the voice average spectrum is calculated in the case of the section where the sudden noise is mixed in the voice. A noise reduction processing program was constructed to give a command to the computer to output a command to be output by adding the signals with reduced noise.</p><p> 18: A recording medium on which the noise reduction processing program according to claim 17 is recorded is configured. Claim 19: For a voice noise mixed signal in which a target voice signal and an unnecessary noise signal are mixed, the power, keptrum, and frequency characteristics are calculated, and the voice section is determined using the calculated power and kepstram. The sudden noise section is determined using the calculated Keptram, frequency characteristics, voice section information, and input voice noise mixed signal, and the signal suppression gain is calculated using the calculated voice section and sudden noise section information. On the other hand, the signal of the input voice noise mixed signal is suppressed by using the calculated signal suppression gain, and the calculated voice section and sudden noise section information are used in the case of the voice section. Calculates the voice average spectrum, and in the case of a section where sudden noise is mixed in the voice, suppresses the voice average spectrum and gives a command to the computer to add up the signals with reduced noise and output them. The noise reduction processing program was configured. 20: A recording medium on which the noise reduction processing program according to claim 19 is recorded is configured.</p>
<p> According to the present invention, it is possible to determine whether or not noise including voice and non-stationary noise is present in the input signal by calculating and analyzing a plurality of feature quantities with respect to the input signal. Further, when noise is present, it is possible to reduce noise according to the type of noise by combining a plurality of noise reduction devices using the analysis result. Further, since the present invention realizes the processing by one input, it becomes easy to use it in combination with the existing communication and sound collecting device.</p>
The best mode for carrying out the invention will be described with reference to the figures. FIG. 1 is a block diagram illustrating an embodiment of a noise reduction device. With reference to FIG. 1, first, the target signal and the input signal mixed with unnecessary ambient noise are transferred to the plurality of feature amount calculation unit 11. The plurality of feature amount calculation unit 11 includes a combination of a plurality of feature amount calculation units 12 for calculating the feature amount with respect to the input signal. The plurality of feature amount calculation unit 11 calculates various feature amounts for the input signal, and transfers the plurality of feature amounts to the feature amount analysis unit 13. The feature amount analysis unit 13 estimates the state and characteristics of the input signal using a plurality of feature amounts and input signals, and transfers the estimated input signal state and characteristic information to the reduction variable calculation unit 14. The reduction variable calculation unit 14 is a noise reduction unit of the plurality of noise reduction units 15 so that the noise reduction effect is optimized according to the input signal and the input signal state and characteristic information estimated by the feature quantity analysis unit 13. 16 reduction variables are determined and transferred to each noise reduction unit 16.
On the other hand, the input signal is also transferred to the plurality of noise reduction units 15. The plurality of noise reduction units 15 are composed of a combination of a plurality of noise reduction units 16 that perform noise reduction on the input signal using the reduction variables calculated by the reduction variable calculation unit 14. The plurality of noise reduction units 15 perform various noise reductions on the input signal and output each reduced signal. Each signal output by the plurality of noise reduction units 16 is weighted so that each noise reduction unit 16 works effectively according to the noise existing in the input signal. Furthermore, all are added and standardized, and output as an output signal.
Other embodiments will be described with reference to FIG. In this embodiment, assuming voice communication, it is determined whether the input signal X (n) is voice or non-voice, and whether or not there is sudden noise, and the sudden noise is reduced. Here, n is an integer value that represents the time representation of the signal as discrete time. First, the input signal X (n) is transferred to the plurality of feature amount calculation unit 11. The plurality of feature amount calculation unit 11 includes each feature amount calculation unit of the power calculation unit 21, the cepstrum calculation unit 22, and the frequency characteristic calculation unit 23. The feature amount calculated by each feature amount calculation unit is transferred to the feature amount analysis unit 13. Here, each feature calculation unit calculates features using power, cepstrum, and frequency characteristics, but in addition, autocorrelation function, analysis using wavelet transformation, pattern recognition, linear prediction, zero intersection counting, A feature amount calculation unit that calculates the feature amount of voice and noise by using band filter bank analysis or the like may be used.
The power calculation unit 21 calculates the power level of the input signal and outputs it as a feature amount. Power level is PX (n) = X (n)<sup>2</sup> Is sought after. The time average is, for example, Pavx (n) = (1 / A) Σ<sub>m</sub>γ<sub>m</sub>Calculated as PX (nm). Here, γ<sub>m </sub>Is, for example, γ<sub>m </sub>= (γ)<sup>m</sup> Γ <1, A is (1 / A) Σ<sub>m</sub>γ<sub>m</sub>It is a constant for normalization such that = 1. The power calculation unit 21 transfers Pavx (n) as a feature quantity to the voice section determination unit 24. The cepstrum calculation unit 22 calculates the cepstrum of the input signal and outputs the peak value indicating the periodicity of the signal as a feature amount. Cepstrum is obtained, for example, by the logarithmic inverse Fourier transform of the short-time amplitude spectrum | X (ω) | of the waveform described in "Digital Speech Processing" p.44-47 by Sadaoki Furui. The peak of the high cepstrum high keffrency portion represents the basic period, and the value of this peak is transferred to the voice section determination unit 24 and the sudden noise determination unit 25 as the feature amount C1.
The frequency characteristic calculation unit 23 calculates a value characterized by fluctuations in the high frequency range of the frequency characteristic, and outputs this as a feature amount. Generally, in the frequency characteristics of speech, voiced sounds have a peak in the low frequency range where the fundamental frequency exists. On the other hand, the frequency characteristics of sudden noise are flat. Therefore, it is considered that the fluctuation in the high frequency range is characterized by sudden noise. The processing flow is shown in Fig. 3. First, in S21, the input signal X (n) is divided into frames for each fixed section using a time window. Next, in S22, for example, the frequency domain signal X (ω) is converted by the discrete Fourier transform every short time. Generally, the signal converted to the frequency domain is a complex number, and X (ω) = X<sub>r</sub>(ω) + jX<sub>i</sub>Let it be (ω). Next, in S23, the power spectrum P of the frequency band<sub>X</sub>Find (ω). Power spectrum is P<sub>X</sub>(ω) = (X<sub>r</sub>(ω))<sup>2 </sup>+ (X<sub>i</sub>(ω))<sup>2 </sup>Is calculated by. Next, in S24, the power spectrum is divided into M bands. For example, consider dividing the frequency band up to the Nyquist frequency into equal parts. Next, in S25, the average value of the power spectrum is obtained for each band and used as a representative value for each band. Furthermore, in S26, the weight w so that the influence in the high frequency range becomes large with respect to the representative value for each band.<sub>m </sub>Multiply by (m = 1, ..., M). w<sub>m </sub>For example, w<sub>m </sub>Use the sin function calculated by = sin (π (m-1) / 2 (M-1)). The representative value for each of M bands is used as one feature vector, and V<sub>l </sub>And. The subscript l represents the current processing frame. In S27, the correlation with the feature vector of the immediately preceding frame is taken as the feature amount C2. C2 takes power into account, C2 = (| V<sub>l</sub>|<sup>2 </sup> | V<sub>l-1</sub>|<sup>2 </sup>) / (V<sub>l</sub> V<sub>l-1</sub>). In S28, the feature amount C2 is transferred to the sudden noise determination unit 25.
The feature amount analysis unit 13 estimates the state and characteristics of the input signal using the plurality of feature amounts and input signals sent from the plurality of feature amount calculation units 11, and obtains the estimated input signal state and characteristic information. Transfer to the reduction variable calculation unit 14. Here, the feature amount analysis unit 13 includes a voice section determination unit 24 and a sudden noise determination unit 25. The voice section determination unit 24 determines whether the input signal is voice or non-voice, and transfers the corresponding flag to the reduction variable calculation unit 14. The sudden noise determination unit 25 determines whether or not the sudden noise is present, and transfers the corresponding flags to the reduction variable calculation unit 14.
The voice section determination unit 24 determines whether the input signal is voice or non-voice by using the power level of the input signal. Here, first, the flag F1 is set when the Pavx (n) transferred from the voice section power calculation unit 21 exceeds the threshold value T1. When the flag F1 continues above the threshold value T2, it is presumed that the input signal is audible. If not, it is estimated that there is no sound, and the flag F3 is transferred to the reduction variable calculation unit 14 and the sudden noise determination unit 25. When it is sound and the feature quantity C1 transferred from the cepstrum calculation unit 22 is larger than the threshold value T3, it is estimated to be a periodic voice interval, and the flag F2 is set to the reduction variable calculation unit 14, sudden noise determination. Transfer to part 25. When the feature amount C1 is smaller than the threshold value T3, it is estimated as a non-voice section, and the flag F3 is transferred to the reduction variable calculation unit 14 and the sudden noise determination unit 25.
The processing flow of the sudden noise determination unit 25 will be described with reference to FIG. If the flag transferred from the voice section judgment unit 24 is F3, the sudden noise determination unit 25 is presumed to be non-voice, and therefore does not output anything (S42). If the flag transferred from the voice section determination unit 24 is F2, it is presumed to be voice and processing is performed (S43). The sudden noise determination unit 25 uses the transferred cepstrum feature amount C1 and the frequency characteristic feature amount C2 to determine whether or not sudden noise exists in the processing frame, and if so, from anywhere in the frame. Estimate the presence of sexual noise. The feature amount C1 represents the periodicity of the signal, and it is considered that the smaller the value, the more sudden noise exists. The feature amount C2 represents the fluctuation in the high frequency band of the signal, and it is considered that the larger the value, the more sudden noise exists. Therefore, when C1 is smaller than the threshold value T4 and C2 is larger than the threshold value T5, it is estimated that sudden noise exists in the processing frame, and the flag F4 is set (S45). Next, the absolute value of the original signal of the processing frame is taken, and the position S1 having the largest value is estimated to be the position where the sudden noise exists (S46). It is estimated that S2 before the margin M1 from S1 is the position where the sudden noise starts, and S2 and the flag F4 are transferred to the reduction variable calculation unit 14 (S47). On the other hand, when C1 is larger than the threshold value T4 or C2 is smaller than the threshold value T5, it is estimated that there is no sudden noise in the processing frame and the flag F5 is transferred to the reduction variable calculation unit 14 (S44). ..
The reduction variable calculation unit 14 determines the reduction variables of the plurality of noise reduction units 16 by using the information on the state and characteristics of the input signal and the transferred input signal. Here, the signal suppression gain G of the signal suppression unit 26 and the periodic waveform insertion are used by using the flag of whether the sound is transferred from the feature analysis unit 13 or not, and if it is voice, whether or not sudden noise is present. Determine the number of repetitions R of part 27. Figure 5 shows the processing flow of the reduction variable calculation unit 14. First, when the flag transferred from the voice section determination unit 24 is F3, the signal is presumed to be non-voice, so the signal is completely suppressed. The signal suppression gain G = 0 and the number of repetitions R = 0 are set, and the signals are transferred to the signal suppression unit 26 and the periodic waveform insertion unit 27, respectively (S52). On the other hand, when the flag transferred from the voice section determination unit 24 is F2, the signal is presumed to be voice, so the presence of sudden noise is confirmed (S53). When the flag transferred from the sudden noise determination unit 25 is F5, it is estimated that there is no sudden noise, so the signal is passed as it is. The signal suppression gain G = 1 and the number of repetitions R = 0 are set, and the signals are transferred to the signal suppression unit 26 and the periodic waveform insertion unit 27, respectively (S54). On the other hand, when the flag transferred from the sudden noise determination unit 25 is F4, it is estimated that the sudden noise exists, so the sudden noise is reduced. The signal suppression gain G = 1 and the number of repetitions R = R1 are set, and the signals are transferred to the signal suppression unit 26 and the periodic waveform insertion unit 27, respectively (S55). At the same time, the position S2 at which the sudden noise started, which has been transferred from the sudden noise determination unit 25, is transferred to the periodic waveform insertion unit 27.
The plurality of noise reduction units 15 use the reduction variables transferred from the reduction variable calculation unit 14 to perform noise reduction processing of the input signal in each reduction variable unit. Here, the plurality of noise reduction units 15 include a signal suppression unit 26 and a periodic waveform insertion unit 27. In addition, in each noise reduction unit, a noise reduction device for stationary noise as in Conventional Example 1, a center clip that suppresses a signal at a level below the threshold value, a process that suppresses a band above the threshold value in the frequency domain, and the like are performed. The signal suppression unit 26, which may be used for noise reduction processing, suppresses the entire input signal. For the transferred input signal X (n), the signal suppression gain G transferred from the reduction variable calculation unit 14 is used, and GX (n) is output.
The periodic waveform insertion unit 27 reduces the sudden noise by repeatedly inserting the periodic waveform of the voice immediately before the presence of the sudden noise at the position where the sudden noise exists. For the insertion of this periodic waveform, the method described in "PCM coding method for JT-G711 voice frequency band signal Appendix 1 high-quality low-calculation algorithm for packet loss compensation for standard JT-G711" is used. First, the immediately preceding cycle is detected with S2 transferred from the reduction variable calculation unit 14 as the disappearance start point. OLA (overlap addition) is performed by covering the triangular window with 1/4 cycle from 5/4 cycle before disappearance start point to 1 cycle before disappearance start point and 1/4 cycle immediately before disappearance start point. Then, using one cycle immediately before the disappearance start point, the R-time repeated composite signal transferred from the reduction variable calculation unit 14 is created and inserted into the original signal. In the next αR period, OLA with the original signal is performed. α is a constant that adjusts the OLA interval. In general, sudden noise is rapidly attenuated. Therefore, by reducing the number of repetitions R of periodic waveform insertion and lengthening the OLA section by adjusting α, sudden noise can be reduced and processing with less distortion can be realized. Output the processed waveform.
For the processed signal output from the signal suppression unit 26 and the periodic waveform insertion unit 27, w according to the noise existing in the input signal, respectively.<sub>a </sub>, W<sub>b b </sub>Multiply the weight of. Here, when the input signal is non-voice w<sub>a </sub>= 1, w<sub>b b </sub>= 0, when there is no sudden noise in the voice w<sub>a</sub>= 1, w<sub>b b</sub>= 0, when there is sudden noise in the voice w<sub>a </sub>= p<sub>a </sub>, W<sub>b b </sub>= p<sub>b b </sub>And. p<sub>a </sub>, P<sub>b b </sub>Is a parameter representing the effect of noise reduction processing by inserting a periodic waveform. All the weighted signals are added and standardized, and output as an output signal. The processing of each block of the noise reduction device of the present invention may be performed by a DSP (Digital Signal Processor). Further, it may be made to function by executing a program by a computer. In this case, the program is recorded on a CD-ROM, a floppy (registered trademark) disk, a magnetic disk, or the like, and is taken into the program memory in the computer. The program may be downloaded to this program memory by communication.
FIG. 8 shows the results of computer simulation showing the effect of the example of the noise reduction device. The horizontal axis of the figure represents time, and the vertical axis represents amplitude. Figure 8 (a) shows the waveform of a signal in which the sound of hitting a desk with a pen is superimposed three times as sudden noise on the voice signal. FIG. 8 (b) shows the waveform of the output signal simulated using the computer in the embodiment using the signal of FIG. 8 (a) as an input. As shown in this figure, it can be seen that by using the examples, only the sudden noise superimposed on the voice can be reduced. A further embodiment of the noise reduction device will be described with reference to FIG. The configuration of this embodiment is almost the same as the configuration of the embodiment described with reference to FIGS. The difference is that part 97 is used. The feature amount analysis unit 13 and the signal suppression unit 96 in FIGS. 1 and 2 perform the same operations as the feature amount analysis unit 13 and the signal suppression unit 26 of the embodiments of FIGS. 1 and 2.
The reduction variable calculation unit 98 uses the input signal, the state of the input signal transferred via the multiple feature calculation unit 11 and the feature analysis unit 13, and the characteristic information. Determine 15 reduction variables. Here, the signal suppression gain G of the signal suppression unit 96 is determined by using the flag of whether the information transferred from the feature amount analysis unit 13 is voice or non-voice, and if it is voice, whether or not sudden noise is present. Then, the flag is transferred to the band-specific suppression unit 97. First, when the flag transferred from the voice section determination unit 94 is F3, the signal is presumed to be non-voice, so the signal is completely suppressed. The signal suppression gain G = 0 is transferred to the signal suppression unit 96, and the flag F3 is transferred to the band-specific suppression unit 97. On the other hand, when the flag transferred from the voice section determination unit 94 is F2, the signal is presumed to be voice, so the presence of sudden noise is confirmed. When the flag transferred from the sudden noise determination unit 95 is F5, it is estimated that there is no sudden noise, so the signal is passed as it is. The signal suppression gain G = 1 is transferred to the signal suppression unit 96, and the flag F5 is transferred to the band-specific suppression unit 97. On the other hand, when the flag transferred from the sudden noise determination unit 95 is F4, it is estimated that the sudden noise exists, so the sudden noise is reduced. The signal suppression gain G = 1 is transferred to the signal suppression unit 96, and the flag F4 is transferred to the band-specific suppression unit 97.
The band-specific suppression unit 97 reduces sudden noise by dividing the input signal into bands and suppressing the input signal to the average level of voice for each band. The processing flow of the band-specific suppression unit 97 will be described with reference to FIG. When the flag transferred from the reduction variable calculation unit 98 is F3, the signal is presumed to be non-voice, so the band-specific suppression unit 97 does not operate (S102). When the flag transferred from the reduction variable calculation unit 98 is F5, it is estimated that the signal is in the voice section and there is no sudden noise, so the voice average spectrum is calculated. First, the input signal X (n) is divided into frames at regular intervals using a time window. Next, it is converted into a frequency domain signal X (ω) by, for example, discrete Fourier transform every short time. Generally, the signal converted to the frequency domain is a complex number, and X (ω) = X<sub>r</sub>(ω) + jX<sub>i</sub>Let it be (ω). Next, the frequency band is the power spectrum P<sub>x</sub>Find (ω). Power spectrum is P<sub>x</sub>(ω) = X<sub>r</sub>(ω)<sup>2 </sup>+ X<sub>i</sub>(ω)<sup>2 </sup>Calculated by (S104). Next, the power spectrum is divided into M bands. For example, consider dividing the frequency band up to the Nyquist frequency into 128 bands. Next, the average value P of the power spectrum for each band<sub>xm</sub>Find (k) (k = 1 ~ M) and use it as a representative value for each band (S105). Average value of the obtained power spectrum P<sub>xm</sub>(k) and the stored audio average power spectrum value P<sub>sp</sub>Average P with (k)<sub>sp</sub>(k) = (1-α)<sub>sp</sub>) P<sub>xm (</sub>k) + α<sub>sp</sub> P<sub>sp</sub>Calculate (k) and add a new voice average power spectrum value P<sub>sp</sub>Save as (k) (S106). Here, α<sub>sp</sub>(0 α<sub>sp</sub>1) is the forgetting coefficient, where α<sub>sp</sub>Set to = 0.98. Also, P<sub>sp</sub>For the initial value of (k), the long-term average spectrum of voice calculated in advance is used. When the flag transferred from the reduction variable calculation unit 98 is F4, the signal is in the voice section, and it is estimated that sudden noise exists, so sudden noise suppression processing is performed. Input signal power spectrum P<sub>x</sub>Find (ω) (S107), and for each ω, the audio average power spectrum value P of the corresponding band<sub>sp</sub>Weight w on (k)<sub>Psp</sub>Compare with the one multiplied by. Px (ω)> w<sub>Psp</sub> P<sub>sp</sub>When (k), noise suppression processing output power spectrum P<sub>xout</sub>(ω) = w<sub>Psp</sub> P<sub>sp</sub>Let it be (k) and suppress it. Meanwhile, P<sub>x</sub>(ω) w<sub>Psp</sub> P<sub>sp</sub>At (k), P<sub>xout</sub>(ω) = P<sub>x</sub>Let (ω) be (S108). Noise suppression processing output power spectrum P<sub>xout</sub>The inverse Fourier transform is performed using (ω), the frame is synthesized using the time window, and the time domain signal is output.
For the processed signal output from the signal suppression unit 96 and the band-specific suppression unit 97, w according to the noise existing in the input signal, respectively.<sub>a</sub>, W<sub>b b</sub>Multiply the weight of. Here, when the input signal is non-voice w<sub>a</sub>= 1, w<sub>b b</sub>= 0, when there is no sudden noise in the voice w<sub>a</sub>= 1, w<sub>b b</sub>= 0, when there is sudden noise in the voice w<sub>a</sub>= p<sub>a</sub>, W<sub>b b</sub>= p<sub>b b</sub>And. p<sub>a</sub>, P<sub>b b</sub>Is a parameter that represents the effect of band-specific suppression. All the weighted signals are added and standardized, and output as an output signal.
<figref num="1">The figure explaining the Example.</figref><figref num="2">The figure explaining another embodiment.</figref><figref num="3">The flow chart explaining the process of the frequency characteristic calculation part of an Example.</figref><figref num="4">The flow figure explaining the processing of the sudden noise determination part of an Example.</figref><figref num="5">The flow diagram explaining the derivation method of the reduction variable of an Example.</figref><figref num="6">The figure explaining the conventional example.</figref><figref num="7">The figure explaining another conventional example.</figref><figref num="8">Results of computer simulation showing the effect of the examples.</figref><figref num="9">The figure explaining the further embodiment.</figref><figref num="10">The processing flow diagram explaining the processing of the suppression part by band of a further embodiment.</figref>
Code description
11 Multiple feature calculation unit 12 Feature calculation unit 13 Feature analysis unit 14 Reduction variable calculation unit 15 Multiple noise reduction unit 16 Noise reduction unit 21 Power calculation unit 22 Cepstrum calculation unit '23 Frequency characteristic calculation unit 24 Voice section judgment unit 25 Sudden noise judgment unit 26 Signal suppression unit 27 Periodic waveform insertion unit 61 Frequency domain conversion unit 62 Input signal power spectrum calculation unit 63 Noise power spectrum estimation unit 64 Loss calculation unit 65 Loss insertion unit 66 Time domain conversion unit 71 Sound receiver 72 Microphone array 73 Directional control unit 74 Sound receiver 75, 76 Short-time power calculation unit 77 Voice section detection unit 91 Power calculation unit 92 Cepstrum calculation unit 93 Frequency characteristic calculation unit 94 Voice section judgment unit 95 Sudden noise judgment unit 96 Signal suppression unit 97 Band-specific suppression unit 98 Reduction variable calculation unit
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| JP2011145372A | Cited by | Japan | Search report |
| JP5788873B2 | Cited by | Japan | Examiner |
| JP2008054071A | Cited by | Japan | Search report |
| CN102918592A | Cited by | China | Search report |
| JP2017513046A | Cited by | Japan | Search report |
| JP2011002723A | Cited by | Japan | Examiner |
| WO2010146711A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| JP2019164898A | Cited by | Japan | Search report |
| JPWO2008111462A1 | Cited by | Japan | Search report |
| US2011235812A1 | Cited by | United States of America | Pre-grant |
| US8676571B2 | Cited by | United States of America | Applicant |
| JP2007251917A | Cited by | Japan | Search report |
| WO2011148861A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| CN102804260A | Cited by | China | Search report |
| JP2014021438A | Cited by | Japan | Examiner |
| JP2011203500A | Cited by | Japan | Examiner |
| JP2015158696A | Cited by | Japan | Examiner |
| JPWO2017037830A1 | Cited by | Japan | Search report |
| WO2010146711A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| JP2000163099A | Cites | Japan | Search report |
| JP2000330597A | Cites | Japan | Search report |
| JP2001236085A | Cites | Japan | Search report |
| JP2001344000A | Cites | Japan | Search report |
| JPH04230800A | Cites | Japan | Search report |
| JPH1098348A | Cites | Japan | Search report |
6 priority claims, no other members on record
Priority claims6
| Document | Office | Kind | Date |
|---|---|---|---|
| 2004066212 | Japan | A | |
| 2004066212 | Japan | – | |
| 2005062200 | Japan | A | |
| 2004200466212 | – | – | – |
| JP20040066212 | – | – | – |
| JP20050062200 | – | – | – |
15 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Written notification of registration of transferR350 | R350 | |
| Written request for registration of change of domicileS531 | S531 | |
| First payment of annual fees (during grant procedure)A61 | A61 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelR150 | R150 | |
| Certificate of patent or registration of utility modelR150 | R150 | |
| Written decision to grant a patent or to grant a registration (utility model)A01 | A01 | |
| Written decision to grant a patent or to grant a registration (utility model)A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Written amendmentA521 | A521 | |
| Notification of reasons for refusalA131 | A131 | |
| Report on retrievalA977 | A977 | |
| Written request for application examinationA621 | A621 | |
| Notification of appointment of power of attorneyRD03 | RD03 |
Numbers
- Publication
- 2005292812
- Publication, DOCDB
- 2005292812
- Publication, EPODOC
- JP2005292812
- Application
- 62200
- Application, DOCDB
- 2005062200
- Application, EPODOC
- JP20050062200
Titles2
- Japanese
- 音声雑音判別方法および装置、雑音低減方法および装置、音声雑音判別プログラム、雑音低減プログラム、およびプログラムの記録媒体
- English
- Voice noise discrimination method and device, noise reduction method and device, voice noise discrimination program, noise reduction program, and recording medium of the program.
Classification
- IPC, 6
- G10L15 02
- G10L15 20
- G10L21 02
- G10L25 93
- H04B1 10
- H04R3 00