Signal processing apparatus
Abstract
Problem to be solved.To provide a signal processing apparatus which is capable of enhancing articulation even in the case that frequency bans where signal components are present or sampling frequencies are different between a target signal to be reproduced and peripheral noise because of difference of frequency bands to be limited.
Solution.The signal processing apparatus which changes frequency characteristics of an input signal whose frequency band is limited to a first frequency range includes: peripheral noise extracting means which extracts peripheral noise included in a collected sound signal; information extracting means which extracts frequency characteristic information on a second frequency range from the peripheral noise extracted by the peripheral noise extracting means; frequency characteristic information extending means which extends the frequency characteristic information extracted by the information extracting means, to the first frequency range in a frequency direction; and signal correcting means which changes frequency characteristics of the input signal in accordance with the frequency characteristic information obtained by the frequency characteristic information extending means.

Term
Projected expiry 27 June 2032.
- Priority and filed
- Published
- Today
- Projected expiry
12 claims: 4 independent, 8 dependent
- 1第1の周波数範囲に帯域制限された入力信号に対して周波数特性を変化させる信号処理装置であって、 集音信号に含まれる周囲雑音を抽出する周囲雑音抽出手段と、 前記周囲雑音抽出手段によって抽出された周囲雑音から第2の周波数範囲の周波数特性情報を抽出する情報抽出手段と、 前記情報抽出手段によって抽出された周波数特性情報に対して、前記第1の周波数範囲へ周波数特性情報を周波数方向に拡張する周波数特性情報拡張手段と、 前記周波数特性情報拡張手段によって得られた周波数特性情報に応じて、前記入力信号の周波数特性を変化させる信号補正手段と、 を備えた信号処理装置。
- 2前記情報抽出手段は、前記第2の周波数範囲よりも狭い第3の周波数範囲に制限された周波数特性情報を抽出することを特徴とする請求項1に記載の信号処理装置。
- 3事前に取得した信号の第2の周波数範囲の周波数特性情報と第1の周波数範囲の周波数特性情報とを対応づけて記憶する記憶手段を更に有し、 前記周波数特性情報拡張手段は、前記記憶手段に記憶された第2の周波数範囲の周波数特性情報と第1の周波数範囲の周波数特性情報との対応を用いて周波数特性情報の拡張を行うことを特徴とする請求項1または請求項2に記載の信号処理装置。
- 4前記周波数特性情報拡張手段は、 前記情報抽出手段によって抽出された周波数特性情報から、前記第1の周波数範囲のうち前記第2周波数範囲を除いた第4の周波数範囲における周波数特性情報を推定し、この推定した周波数特性情報と前記情報抽出手段により抽出された前記周波数特性情報とが、前記第2の周波数範囲と前記第4の周波数範囲との境界において連続するように補正を行うことを特徴とする請求項1乃至請求項3のいずれか1項に記載の信号処理装置。
- 5前記情報抽出手段によって抽出される周波数特性情報は、周波数ごとのマスキングレベルであることを特徴とする請求項1乃至請求項4のいずれか1項に記載の信号処理装置。
- 6前記情報抽出手段によって抽出される周波数ごとのマスキングレベルは、多項式で近似表現されることを特徴とする請求項5に記載の信号処理装置。
- 7第1の周波数範囲に帯域制限された入力信号に対して周波数特性を変化させる信号処理装置であって、 集音信号に含まれる周囲雑音を抽出する周囲雑音抽出手段と、 前記周囲雑音抽出手段によって抽出された周囲雑音から前記第1の周波数範囲よりも狭い第2の周波数範囲の周波数特性情報を抽出する情報抽出手段と、 前記情報抽出手段によって抽出された周波数特性情報に対して、前記第1の周波数範囲へ周波数特性情報を周波数方向に拡張する周波数特性情報拡張手段と、 前記周波数特性情報拡張手段によって得られた周波数特性情報に応じて、前記入力信号の周波数特性を変化させる信号補正手段と を備えた信号処理装置。
- 8前記情報抽出手段は、前記第2の周波数範囲よりも狭い第3の周波数範囲に制限された周波数特性情報を抽出することを特徴とする請求項7に記載の信号処理装置。
- 9事前に取得した信号の第2の周波数範囲の周波数特性情報と第1の周波数範囲の周波数特性情報とを対応づけて記憶する記憶手段を更に有し、 前記周波数特性情報拡張手段は、前記記憶手段に記憶された第2の周波数範囲の周波数特性情報と第1の周波数範囲の周波数特性情報との対応を用いて周波数特性情報の拡張を行うことを特徴とする請求項7または請求項8に記載の信号処理装置。
- 10前記周波数特性情報拡張手段は、 前記情報抽出手段によって抽出された周波数特性情報から、前記第1の周波数範囲のうち前記第2周波数範囲を除いた第4の周波数範囲における周波数特性情報を推定し、この推定した周波数特性情報と前記情報抽出手段により抽出された前記周波数特性情報とが、前記第2の周波数範囲と前記第4の周波数範囲との境界において連続するように補正を行うことを特徴とする請求項7乃至請求項9のいずれか1項に記載の信号処理装置。
- 11前記情報抽出手段によって抽出される周波数特性情報は、周波数ごとのマスキングレベルであることを特徴とする請求項7乃至請求項10のいずれか1項に記載の信号処理装置。
- 12前記情報抽出手段によって抽出される周波数ごとのマスキングレベルは、多項式で近似表現されることを特徴とする請求項11に記載の信号処理装置。
Independent claims12
116 paragraphs, as filed
The present invention relates to a signal processing device that improves intelligibility for signals such as voice, music, and audio.
When reproducing a signal such as voice, music, or audio, the clarity of the target signal is reduced due to the influence of ambient noise other than the desired signal such as voice, music, or audio (hereinafter referred to as the target signal). In some cases. Therefore, in order to improve the clarity of the target signal, it is necessary to perform signal processing according to the ambient noise contained in the collected signal. Conventionally, as such a signal processing method, there have been a method using the volume of ambient noise and a method using the frequency characteristic of ambient noise (for example, Patent Document 1).
<p><patcit num="1"><text>Japanese Unexamined Patent Publication No. 2001-188599</text></patcit></p>
<p num="0004"> However, since the restricted frequency band is different between the target signal and the ambient noise, the frequency band in which the signal component exists may be different, or the sampling frequency may be different. In such a case, the conventional signal processing device has a problem that the volume and frequency characteristics of ambient noise cannot be obtained with high accuracy, resulting in deterioration of sound quality and inability to improve intelligibility.</p><p num="0005"> In addition, for target signals such as audio signals and music / audio signals, the surroundings collected by using the conventional technology that expands the band such as using aliasing, using non-linear functions, and using linear predictive analysis are used as they are. Even if the noise band is expanded, there is a problem that the frequency characteristics of ambient noise cannot be estimated with high accuracy.</p><p num="0006"> The present invention has been made to solve the above-mentioned problems, and the frequency band in which the signal component exists is different due to the difference in the limited frequency band between the target signal to be reproduced and the ambient noise, or the sampling frequency. It is an object of the present invention to provide a signal processing device capable of improving clarity even when the frequencies are different.</p>
<p num="0007"> In order to achieve the above object, the present invention is a signal processing device that changes the frequency characteristics of an input signal band-limited in the first frequency range, and extracts ambient noise contained in the sound collection signal. For the ambient noise extracting means, the information extracting means for extracting the frequency characteristic information in the second frequency range from the ambient noise extracted by the ambient noise extracting means, and the frequency characteristic information extracted by the information extracting means. , The frequency characteristic of the input signal is changed according to the frequency characteristic information expanding means for extending the frequency characteristic information to the first frequency range in the frequency direction and the frequency characteristic information obtained by the frequency characteristic information expanding means. It is configured to include a signal correction means.</p>
<p num="0008"> According to the present invention, the intelligibility is different even when the frequency band in which the signal component exists is different or the sampling frequency is different because the limited frequency band is different between the target signal to be reproduced and the ambient noise. It is possible to provide a signal processing device capable of improving the above. </p>
<figref num="1">The circuit block diagram which shows the structure of the communication apparatus which applied the 1st Embodiment of the signal processing apparatus which concerns on this invention.</figref><figref num="2">The circuit block diagram which shows the structure of 1st Example of the signal processing part which concerns on this invention.</figref><figref num="3">The circuit block diagram which shows the structural example of the ambient noise estimation part of the signal processing part shown in FIG.</figref><figref num="4">The circuit block diagram which shows the structural example of the ambient noise information band expansion part of the signal processing part shown in FIG.</figref><figref num="5">FIG. 5 is a processing flow diagram for explaining the operation of the dictionary generation method in the dictionary storage unit of the ambient noise information band expansion unit shown in FIG.</figref><figref num="6">The circuit block diagram which shows the structural example of the signal characteristic correction part of the signal processing part shown in FIG.</figref><figref num="7">FIG. 3 is a circuit block diagram showing a configuration of a communication device and a digital audio player to which the first embodiment of the signal processing device according to the present invention is applied.</figref><figref num="8">The circuit block diagram which shows the structure of the modification 1 of the signal processing part which concerns on this invention.</figref><figref num="9">FIG. 6 is a circuit block diagram showing a configuration example of an ambient noise information band expansion unit of the signal processing unit shown in FIG.</figref><figref num="10">FIG. 5 is a processing flow diagram for explaining the operation of the dictionary generation method in the dictionary storage unit of the ambient noise information band expansion unit shown in FIG.</figref><figref num="11">The figure which shows the example of the wide band masking threshold.</figref><figref num="12">The circuit block diagram which shows the structural example of the signal characteristic correction part of the signal processing part shown in FIG.</figref><figref num="13">The circuit block diagram which shows the structure of the modification 3 of the signal processing part which concerns on this invention.</figref><figref num="14">The circuit block diagram which shows the structural example of the ambient noise information band extension part of the signal processing apparatus shown in FIG.</figref><figref num="15">FIG. 6 is a processing flow diagram for explaining the operation of the dictionary generation method in the dictionary storage unit of the ambient noise information band expansion unit shown in FIG.</figref><figref num="16">The figure which shows the example for demonstrating the operation of the threshold value correction part of the ambient noise information band extension part shown in FIG.</figref><figref num="17">FIG. 6 is a processing flow diagram for explaining the operation of another generation method of the dictionary in the dictionary storage unit of the ambient noise information band expansion unit shown in FIG.</figref><figref num="18">FIG. 6 is a processing flow diagram for explaining the operation of another generation method of the dictionary in the dictionary storage unit of the ambient noise information band expansion unit shown in FIG.</figref><figref num="19">The circuit block diagram which shows the structure of the communication device and the digital audio player which applied the 2nd Example of the signal processing part which concerns on this invention.</figref><figref num="20">The circuit block diagram which shows the structure of the 2nd Example of the signal processing apparatus which concerns on this invention.</figref><figref num="21">FIG. 6 is a circuit block diagram showing a configuration example of an ambient noise estimation unit and an ambient noise suppression processing unit of the signal processing device shown in FIG. 20.</figref>
Hereinafter, embodiments of the present invention will be described with reference to the drawings. (First Example) FIG. 1 shows a configuration of a communication device according to an embodiment of the present invention. The communication device shown in this figure shows a receiving system of a wireless communication device such as a mobile phone, and includes a wireless communication unit 1, a decoder 2, a signal processing unit 3, and a digital analog (D / A). It includes a converter 4, a speaker 5, a microphone 6, an analog-to-digital (A / D) converter 7, a downsampling unit 8, an echo suppression processing unit 9, and an encoder 10. In the present embodiment, the target signal to be reproduced will be described as an audio signal of a far-end speaker included in the received input signal.
The wireless communication unit 1 wirelessly communicates with a wireless base station accommodated in the mobile communication network, and establishes a communication link with a communication partner station through the wireless base station and the mobile communication network to communicate.
The decoder 2 decodes the received data received from the communication partner station by the wireless communication unit 1 every frame (= 20 [ms]), which is a predetermined time unit, and decodes the digital input signal x [n]. ] (N = 0,1, ... 2N-1) is obtained and output to the signal processing unit 3 in frame units. However, this input signal x [n] is a wideband signal whose sampling frequency is fs'[Hz] and the band is limited from fs_wb_low [Hz] to fs_wb_high [Hz]. Here, the relationship between the sound collection signal z [n] described later and the sampling frequency fs [Hz] is fs'= 2fs. Also, when the sampling frequency is fs'[Hz], the data length of one frame is 2N samples. That is, N = 20 [ms] × fs [Hz] ÷ 1000.
The signal processing unit 3 receives an input signal x [n] in units of one frame according to the sound collection signal z [n] (n = 0,1, ... N-1) whose echo is reduced by the echo suppression processing unit 8 described later. ] (N = 0,1,... 2N-1) is subjected to signal correction processing to change the volume or frequency characteristics, and the output signal is converted to y [n] (n = 0,1,... 2N-1). ) Is output to the D / A converter 4 and the downsampling unit 8. A specific configuration example of the signal processing unit 3 will be described in detail later.
The D / A converter 4 converts the signal-corrected output signal y [n] into an analog signal y (t) and outputs the signal to the speaker 5. The speaker 5 outputs an output signal y (t), which is an analog signal, to the acoustic space.
The microphone 6 collects sound, acquires a sound collecting signal z (t) which is an analog signal, and outputs the sound to the A / D converter 7. This analog signal includes a voice signal of a near-end speaker, a noise component caused by other surrounding environment, an output signal y (t), an echo component caused by an acoustic space, and the like. For example, noise components include noise generated by trains, car noise such as cars, and street noise in crowds. In the present embodiment, since the voice signal of the near-end speaker is a necessary signal desired for communication with the communication partner station as a communication device, components other than the voice signal of the near-end speaker are ambient noise. Treat as.
The A / D converter 7 converts the sound collecting signal z (t), which is an analog signal, into a digital signal, and converts the digital sound collecting signal z'[n] (n = 0,1, ... N-1). Then, it is output to the echo suppression processing unit 8 in units of N samples. However, this sound collection signal z [n] is a narrow band signal whose sampling frequency is fs [Hz] and the band is limited from fs_nb_low [Hz] to fs_nb_high [Hz]. However, it is assumed that fs_wb_low fs_nb_low <fs_nb_high <fs / 2 fs_wb_high <fs' / 2 is satisfied.
The downsampling unit 8 downsamples the output signal y [n] from the sampling frequency fs'[Hz] to fs [Hz], and limits the band from fs_nb_low [Hz] to fs_nb_high [Hz] to y'[. n] (n = 0,1, ... N-1) is output to the echo suppression processing unit 9.
The echo suppression processing unit 9 uses the downsampled output signal y'[n] (n = 0,1, ... N-1) to collect the sound collection signal z'[n] (n = 0,1, ...). ... N-1) is processed to reduce the echo component, and the echo-reduced signal is set as z [n] (n = 0,1, ... N-1) to the signal processing unit 3 and the encoder 10. Output. Here, the echo suppression processing unit 9 may be carried out by the existing technique described in, for example, Japanese Patent Application Laid-Open No. 4047867, Japanese Patent Application Laid-Open No. 2006-203358, Japanese Patent Application Laid-Open No. 2007-60644, and the like.
The encoder 10 encodes the echo-reduced sound collection signal z [n] (n = 0,1, ... N-1) in the echo suppression processing unit 8 for each N sample and outputs it to the wireless communication unit 1 to wirelessly. It is transmitted to the communication partner station as transmission data by the communication unit 1.
Next, an embodiment of the signal processing unit 3 will be described. In the following explanation, for example, fs = 8000 [Hz], fs'= 16000 [Hz], fs_nb_low = 340 [Hz], fs_nb_high = 3950 [Hz], fs_wb_low = 50 [Hz], fs_wb_high = 7950 [Hz]. To do. The band-limited frequency band and sampling frequency are not limited to this. Also, here, N = 160.
FIG. 2 shows a configuration example of the signal processing unit 3. The signal processing unit 3 includes an ambient noise estimation unit 31, an ambient noise information band expansion unit 32, and a signal characteristic correction unit 33. These can also be realized by one processor and software recorded on a storage medium (not shown).
The ambient noise estimation unit 31 estimates a signal other than the voice signal of the near-end speaker from the echo-reduced signal in the echo suppression processing unit 8 as ambient noise, and extracts a feature amount that characterizes the ambient noise. Since the sound collection signal z [n] is a narrow band signal, the ambient noise is also a narrow band signal. Therefore, the feature amount that characterizes the ambient noise is referred to as narrowband signal information. The narrow band signal information may be any feature quantity that characterizes ambient noise, such as a power spectrum, an amplitude spectrum or a phase spectrum, a PARCOR coefficient or a reflection coefficient, a line spectral frequency, a cepstrum coefficient, or a cepstrum coefficient.
The ambient noise information band expansion unit 32 estimates the feature amount that characterizes the ambient noise when the ambient noise is expanded to the same frequency band (broadband) as the frequency band of the input signal x [n] by using the narrow band signal information. To do. This feature amount is called broadband signal information.
The signal characteristic correction unit 33 corrects the signal characteristics of the target signal by using the ambient noise information band expansion unit 32.
In this way, even if the ambient noise is a signal in a narrow band, the clarity can be improved by the correction process in the signal characteristic correction unit 33 by estimating the feature amount when the ambient noise is expanded in a wide band.
In the following description, a specific configuration of the signal processing unit 3 will be described. In the following description, the narrowband signal information will be described as the power spectrum of ambient noise, and the wideband signal information will be described as the power value (broadband power value) when the ambient noise is extended to a wideband signal.
FIG. 3 shows a configuration example of the ambient noise estimation unit 31. The ambient noise estimation unit 31 includes a frequency domain conversion unit 311, a power calculation unit 312, an ambient noise section determination unit 313, and a frequency spectrum update unit 314.
The ambient noise estimation unit 31 extracts ambient noise other than the voice signal of the near-end speaker from the echo-reduced sound collection signal z [n] (n = 0,1, ... N-1) in the echo suppression processing unit 8. Estimate the power spectrum of this signal | N [f, w] |<sup>2</sup> Is extracted and output to the ambient noise information band expansion unit 32.
The frequency domain conversion unit 311 receives the sound collection signal z [n] (n = 0,1, ... N-1) of the current frame f. Then, samples for the number of overlapping samples by window hanging are extracted from the sound collection signal one frame before this frame, combined with the input signal of the current frame in the time direction, and appropriately zero-packed to perform the frequency domain. Extract the signals for the samples required for conversion. The overlap, which is the ratio of the shift width of the sound collecting signal z [n] to the data length of the sound collecting signal z [n] in the next frame, may be 50%, but here, as an example, one frame Overlap with the front Assuming that the number of samples is L = 48, 2M = 256 samples are prepared from the sound collection signal L sample one frame before and the N = 160 samples and L samples of the sound collection signal z [n] of the frame. To do. Windowing is performed by multiplying this 2M sample by the window function of a sinusoidal window. Then, the frequency domain conversion is performed on the signal of the windowed 2M sample. The conversion to the frequency domain can be performed by the FFT, for example, with the order of the FFT being 2M. The data length is set to the power of 2 (2M) and the order of the frequency domain conversion is set to the power of 2 (2M) by zero-packing the signal to be subjected to the frequency domain conversion. Not exclusively.
When the sound collection signal z [n] is a real signal, the frequency spectrum Z [f, w] (w = 0,1) is obtained by removing the redundant M = 128 bins from the signal obtained by frequency domain conversion. ,... M-1) is obtained, and this is output. However, ω represents a frequency bin. In the case of a real signal, it is originally the M-1 (= 127) bin that is redundant, and the highest frequency bin w = M (= 128) should be taken into consideration. However, the signal to be converted in the frequency domain here is premised on a digital signal including a band-limited audio signal, and the band limitation does not affect the sound quality without considering the frequency bin w = M in the highest region. Therefore, for the sake of simplification of the explanation thereafter, the description does not consider the frequency bin w = M in the highest region. Of course, the highest frequency bin w = M may be considered. At that time, the highest frequency bin w = M should be treated in the same way as w = M-1 or treated independently.
The window function used for window hanging is not limited to humming windows, but is appropriately changed to other symmetric windows (Hanning windows, Blackman windows, sine wave windows, etc.) or asymmetric windows used in voice coding processing. You can do it. In addition, frequency domain transform is DFT (Discrete Fourier Transform) and separation. Other orthogonal transforms that transform into the frequency domain, such as the Discrete Cosine Transform (DCT), can be substituted.
The power calculation unit 312 is a power spectrum that is the sum of the squares of the real part and the imaginary part in the frequency spectrum Z [f, w] (w = 0,1, ... M-1) output from the frequency domain conversion unit 311 | Z [f, w] |<sup>2</sup> Calculate and output (w = 0,1, ... M-1).
The ambient noise section determination unit 313 uses the sound collection signal z [n] (n = 0,1, ... N-1) and the power spectrum output from the power calculation unit 312 | Z [f, w] |<sup>2</sup> (w = 0,1, ... M-1) and the power spectrum of the ambient noise of each frequency band one frame before output from the frequency spectrum update unit 314 | N [f-1, w] |<sup>2</sup> Is the section (ambient noise section) in which the ambient noise is predominantly included in the sound collection signal z [n], or the voice signal of the near-end speaker and the ambient noise that are not included in the ambient noise are mixed. It determines which of the sections (audio sections) is being used for each frame, and outputs the frame judgment information vad [f] indicating the judgment result for each frame. Here, the frame determination information vad [f] = 0 is set in the ambient noise section, and vad [f] = 1 is set in the audio section. From this point onward, the case where only the component is present or the component is contained in a much larger amount than the other components (when the component is contained in excess of a predetermined threshold value) is expressed as "dominantly contained".
Specifically, the sound collection signal z [n] (n = 0,1,... N-1) and the power spectrum | Z [f, w] |<sup>2</sup> And the power spectrum of ambient noise one frame before | N [f-1, w] |<sup>2</sup> Is used to calculate multiple features and output the frame judgment information vad [f]. Here, as multiple features, the first-order autocorrelation coefficient Acorr [f, 1], the maximum autocorrelation coefficient Acorr_max [f], the sum total SN ratio by frequency snr_sum [f], and the SN ratio variance by frequency snr_var [f] ] Will be described as an example.
First, as shown in Eq. (1), the k-th order autocorrelation coefficient Acorr [f, k] (k = 1,... N-1), which is normalized by the power in frame units and takes an absolute value, is calculated. To do.<maths num="1"></maths> At the same time, the first-order autocorrelation coefficient Acorr [f, 1] with k = 1 is also calculated. The first-order autocorrelation coefficient Acorr [f, 1] takes a value from 0 to 1, and the closer it is to 0, the stronger the noise property. That is, it is determined that the smaller the value of the first-order autocorrelation coefficient, the more ambient noise is included in the sound collection signal, and the less audio signals are not included in the ambient noise. Then, as shown in Eq. (2) from the normalized k-th order autocorrelation coefficient Acorr [f, k] (k = 1, ... N-1), the maximum autocorrelation coefficient Acorr [f, k] ] Is calculated and used as the maximum value of the autocorrelation coefficient Acorr_max [f]. The maximum value of the autocorrelation coefficient, Acorr_max [f], takes a value from 0 to 1, and the closer it is to 0, the stronger the noise. That is, it is determined that the smaller the value of the autocorrelation coefficient, the more ambient noise is included in the sound collection signal, and the less audio signals are not included in the ambient noise.<maths num="2"></maths> Next, the power spectrum | Z [f, w] |<sup>2</sup>And ambient noise power spectrum | N [f, w] |<sup>2</sup>With the input of and, the SN ratio of each frequency band, which is the ratio thereof, is calculated by the equation (3) as snr [f, w] (w = 0,1, ... M-1) expressed in dB here.<maths num="3"></maths> Then, the sum of the SN ratio snr [f, w] (w = 0,1, ... M-1) of each frequency band is calculated by the equation (4), and is used as the total SN ratio value snr_sum [f] for each frequency. The total SN ratio value snr_sum [f] by frequency takes a value of 0 or more, and the smaller this value is, the more ambient noise that is a noise component is included in the sound collection signal, and the less audio signals are not included in the ambient noise. Judged.<maths num="4"></maths> Further, the variance of the SN ratio snr [f, w] (w = 0,1, ... M-1) of each frequency band is calculated by the equation (5) and used as the frequency-specific SN ratio variance value snr_var [f]. The frequency-specific signal-to-noise ratio dispersion value snr_var [f] takes a value of 0 or more, and it is judged that the smaller this value is, the more ambient noise is included as a noise component, and the less audio signals are not included in the ambient noise.<maths num="5"></maths> Finally, there are multiple features, such as the first-order autocorrelation coefficient Acorr [f, 1], the maximum autocorrelation coefficient Acorr_max [f], the total SN ratio by frequency snr_sum [f], and the SN ratio variance by frequency. Using the value snr_var [f], each of them is weighted by a predetermined weighting, and the ambient noise ratio type [f] is calculated as the weighted sum of a plurality of features. Here, it is assumed that the smaller the ambient noise degree type [f] is, the more dominant the ambient noise is, and the larger the ambient noise degree type [f] is, the more dominant the audio signal is not included in the ambient noise. Weight w by the learning algorithm etc.<sub>1、</sub>w<sub>2</sub>, W<sub>3</sub>, W<sub>4</sub>(However, w<sub>1</sub> 0, w<sub>2</sub> 0, w<sub>3</sub> 0, w<sub>4</sub> 0) is set, and the calculation is performed by the equation (6). Then, if the ambient noise degree type [f] is larger than the predetermined threshold value THR, vad [f] = 1, and if the ambient noise degree type [f] is equal to or less than the predetermined threshold value THR, vad [f] = 0. ..<maths num="6"></maths> In the above explanation, when obtaining a plurality of feature quantities, it is described that processing is performed for each frequency bin. However, a plurality of adjacent frequency bins by frequency domain conversion are grouped together to form a group, and processing is performed for each group. It doesn't matter. Further, it may be realized by frequency domain conversion such as a band division filter such as a filter bank.
It should be noted that it is not necessary to use all of the plurality of feature quantities described above, or other feature quantities may be additionally used. Further, the codec information output from the wireless communication unit 1 or the decoder 2, for example, the voice detection information indicating whether the voice is voice or not by the silent insertion descriptor (SID) or the voice detector (VAD) or the pseudo background noise. Information on whether or not it has been generated may be used.
The frequency spectrum update unit 314 contains the frame determination information vad [f] output from the ambient noise section determination unit 313 and the power spectrum output from the power calculation unit 312 | Z [f, w] |<sup>2</sup> Using (w = 0,1,... M-1), it is the power spectrum of ambient noise in each frequency band | N [f, w] |<sup>2</sup> Estimate (w = 0,1, ... M-1) and output. For example, the power spectrum of a frame determined to be a section (ambient noise section) in which ambient noise is predominantly included with the frame determination information vad [f] set to 0 | Z [f, w] |<sup>2</sup> Is forgotten on a frame-by-frame basis to calculate the average power spectrum, which is used as the power spectrum of ambient noise in each frequency band | N [f, w] |<sup>2</sup> Output as (w = 0,1, ... M-1). Specifically, the power spectrum of ambient noise in each frequency band | N [f, w] |<sup>2</sup> Is calculated by the power spectrum of the ambient noise in each frequency band one frame before, as shown in Eq. (8) | N [f-1, w] |<sup>2</sup> Perform recursively using. However, the forgetting coefficient α in equation (7)<sub>N</sub>[ω] is a coefficient of 1 or less, preferably about 0.75 to 0.95.<maths num="7"></maths> The ambient noise information band extension 32 is a power spectrum of ambient noise in each frequency band | N [f, w] |<sup>2</sup> Is used to generate the power value of the signal including the frequency band component that exists in the input signal x [n] but does not exist in the sound collecting signal z [n].
FIG. 4 is a diagram showing a configuration example of the ambient noise information band expansion unit 32. The ambient noise information band expansion unit 32 includes a power normalization unit 321, a dictionary storage unit 322, and a wideband power calculation unit 323.
The ambient noise information band expansion unit 32 calculates narrowband feature amount data from narrowband signal information, and models in advance the correspondence between the narrowband feature amount data calculated from the narrowband signal information and the wideband feature amount data. Then, the wideband feature amount data is calculated by using the correspondence between this model and the acquired narrowband feature amount data, and the wideband signal information is generated from the wideband feature amount data. As mentioned above, here the narrowband signal information is the power spectrum of ambient noise. Further, here, it is assumed that the wideband feature data and the wideband signal information are the same, and the wideband signal information is the volume indicated by the wideband power value N_wb_level [f]. A method using a GMM (Gaussian mixture model) is used to model the correspondence between narrow-band feature data and wide-band feature data. Here, the narrowband power value Pow_N [f] and the normalized power spectrum of ambient noise | Nn [f, w] |<sup>2</sup> (w = 0,1,... M-1) is concatenated in the order direction and used as Dnb-order narrowband feature data, and the broadband power value N_wb_level [f] is used as Dwb-order wideband feature data (Dnb =). M + 1, Dwb = 1).
First, in order to calculate the narrow band feature data from the narrow band signal information, the power normalization unit 321 has the power spectrum of the ambient noise output from the ambient noise estimation unit 31 | N [f, w] |<sup>2</sup> (w = 0,1, ... M-1) is input, and narrowband feature data is calculated using the power spectrum of this ambient noise. One of the narrowband feature data is the narrowband power value Pow_N [f], which is the sum of each frequency bin of the power spectrum calculated based on the equation (8).<maths num="8"></maths> As other narrowband feature data, the narrowband power value Pow_N [f] is used and the power spectrum of each frequency bin according to Eq. (9) | N [f, w] |<sup>2</sup>Normalized power spectrum | Nn [f, w] |<sup>2</sup>Is calculated.<maths num="9"></maths> The dictionary storage unit 322 is a mixed number Q (here, Q =) learned by modeling the correspondence between the Dnb-order narrow-band feature data and the Dwb-order wideband feature data based on the ambient noise collected in advance. 64) GMM dictionary λ1<sub>q</sub>= {w<sub>q</sub>, μ<sub>q</sub>, Σ<sub>q</sub>} (Q = 1, ..., Q) is stored. In addition, w<sub>q</sub>Indicates the mixed weight of the qth mixed normal distribution, μ<sub>q</sub>Is the mean vector of the qth mixed normal distribution, Σ<sub>q</sub>Represents the covariance matrix (diagonal covariance matrix or total covariance matrix) of the mixed normal distribution of order q. The average vector μ<sub>q</sub>And the covariance matrix Σ<sub>q</sub>The order, which is the number of components of, is Dnb + Dwb.
Prior dictionary λ1 in the dictionary storage unit 322<sub>q</sub>A flowchart of the learning generation method of is shown in FIG. 5 and will be described.
The signal used to generate the GMM is a wideband signal that is band-limited from fs_wb_low [Hz] to fs_wb_high [Hz] with a sampling frequency fs'[Hz] similar to the input signal x [n]. It is a signal group. It is desirable for this signal group to have many different environments and different volumes. In the following, the signal group of the wideband signal used for GMM generation is collectively referred to as wideband signal data wb [n]. n represents the time (sample).
First, the wideband signal data wb [n] is input, downsampled to the sampling frequency fs [Hz] by a downsampling filter, and the narrowband signal data is band-limited to a narrow band from fs_nb_low [Hz] to fs_nb_high [Hz]. Obtain nb [n] (step S101). In this way, a band-limited signal group is generated in the same manner as the sound collection signal z [n]. Although not shown, when the algorithm delay occurs in the downsampling filter or the band limitation process, the narrow band signal data nb [n] is synchronized with the wide band signal data wb [n].
Next, the narrowband feature data Pnb [f, d] (d = 1, ..., Dnb) is extracted from the narrowband signal data nb [n] in units of frames f (step S102). Narrowband feature data Pnd [f, d] is feature data representing narrowband signal information of a predetermined order. In step S102, first, the frequency domain conversion process is performed from the narrow band signal data nb [n] for each frame in the same manner as the process in the frequency domain conversion unit 311 described above, and the power spectrum of the Mth order narrow band signal data nb [n] is performed. (Step S1021). Next, the power is calculated for each frame from the narrowband signal data nb [n] by the same processing as the processing in the power normalization unit 321 described above, and the primary power value is obtained (step S1022). Then, the normalized power spectrum of the M-th order narrowband signal data nb [n] is obtained from these power spectra and power values (step S1023). Then, the M-th order normalized power spectrum and the first-order power value are connected in the order direction (dimensional direction) in frame units, and the narrow band feature amount data Pnb [f, of order Dnb (= M + 1)) is connected. d] (d = 1, ..., Dnb) is generated (step S1024).
On the other hand, in parallel with the above, the wideband feature data Pwb [f, d] (d = 1, ..., Dwb) is extracted from the wideband signal data wb [n] in units of frames f (step S103). Broadband feature data Pwb [f, d] is feature data representing wideband signal information of a predetermined order. In step S103, first, from the wideband signal data wb [n], the number of FFT points of the processing in the frequency domain conversion unit 311 described above is doubled to 4M points for each frame, and the frequency domain conversion processing is performed in the same manner to obtain a 2Mth-order wideband signal. Obtain the power spectrum of the data wb [n] (step S1031). Next, the power is calculated for each frame from the wideband signal data wb [n] by the same processing as the processing in the power normalization unit 321 described above to obtain the first-order power value. Let this power value be the broadband feature data Pwb [f, d] of order Dwb (= 1) (step S1032).
Next, the narrow-band feature data Pnb [f, d] (d = 1, ..., Dnb) and the wide-band feature data Pwb [f, d] (d = 1, ..., Dwb) are time-synchronized. The two feature data are concatenated in the order direction (dimensional direction) on a frame-by-frame basis to generate the concatenated feature data P [f, d] (d = 1, ..., Dnb + Dwb) of the order Dnb + Dwb. (Step S104).
Then, an initial GMM with a mixture number Q = 1 is generated from the concatenated feature data P [f, d], and the average vector of each GMM is slightly shifted to generate another mixture distribution, thereby doubling the mixture number Q. The process of increasing to 3 and the process of performing the likelihood maximization learning of GMM until convergence by the EM algorithm using the concatenated feature data P [f, d] are alternately repeated, and the mixture number Q (here, Q = 64). ) GMM λ1<sub>q</sub>= {w<sub>q</sub>, μ<sub>q</sub>, Σ<sub>q</sub>} (Q = 1, ..., Q) is generated (step S105). For EM algorithms, DA Reynols and RCRose, "Robust text-independent speaker identification using Gaussian mixture models", IEEE Trans. Speech and Audio Processing, Vol.3, no.1, pp.72-83, Jan.1995. There is a detailed description in the literature. Returning to the description of FIG. The wideband power calculation unit 323 has a narrowband power value Pow_N [f] output from the power normalization unit 321 and a Dnb-order power spectrum in which ambient noise is normalized | Nn [f, w] |<sup>2</sup> (w = 0,1, ... M-1) are concatenated and input as narrowband feature data Pn_nb [f] (d = 1, ..., Dnb). Further, the wideband power calculation unit 323 is charged from the dictionary storage unit 322 to the GMM dictionary λ1.<sub>q</sub>= {w<sub>i</sub>, μ<sub>q</sub>, Σ<sub>q</sub>} (Q = 1, ..., Q) is read, and according to the minimum mean square error (MMSE) estimation, as shown in Eq. (10), soft clustering by multiple normal distribution models and continuous By linear regression, the feature data is converted to feature data corresponding to the wide band with the extended frequency band, and the wide band power value N_wb_level [f], which is the wide band feature data, is calculated from the narrow band feature data Pn_nb [f]. And output. Equation (10) is described as a vector in the dimension (d = 1, ..., Dnb + Dwb) direction. Also, the average vector μ<sub>q</sub>(D = 1, ..., Dnb + Dwb) is in the dimensional direction, μ<sub>q</sub><sup>N</sup>(D = 1,..., Dnb) and μ<sub>q</sub><sup>W</sup>A covariance matrix Σ that is divided into (d = Dnb,..., Dnb + Dwb) and is a (Dn + Dw) × (Dn + Dw) matrix.<sub>q</sub>Is also a Dn × Dn matrix as shown below.<sub>q</sub><sup>NN</sup>And Dn × Dw matrix Σ<sub>q</sub><sup>NW</sup>And Dw × Dn matrix Σ<sub>q</sub><sup>WN</sup>And Dw × Dw matrix Σ<sub>q</sub><sup>WW</sup>Divide into and.<maths num="10"></maths> In the ambient noise information band expansion unit 32, the wideband feature amount data and the wideband signal information are the same. Therefore, in this way, the power spectrum of the ambient noise, which is the narrowband signal information | N [f, w] |<sup>2</sup>From this, the wideband power value N_wb_level [f], which is the wideband signal information, can be obtained.
FIG. 6 shows a configuration example of the signal characteristic correction unit 33. The signal characteristic correction unit 33 includes a frequency domain conversion unit 331, a correction degree determination unit 332, a correction processing unit 333, and a time domain conversion unit 334. The input signal x [n] (n = 0,1,... 2N-1) and the wideband power value N_wb_level [f] are input to the signal characteristic correction unit 33, and the input signal x [n] is included in the sound collection signal. A signal correction process is performed to clarify the sound so that it is not buried in the ambient noise, and the corrected output signal y [n] (n = 0,1,... 2N-1) is output.
The frequency domain conversion unit 331 receives an input signal x [n] (n = 0,1, ...) Instead of the sound collection signal z [n] (n = 0,1, ... N-1) in the frequency domain conversion unit 311. 2N-1) is entered. The frequency domain conversion unit 331 outputs the frequency spectrum X [f, w] of the input signal x [n] by the same processing as the frequency domain conversion unit 311. For example, the frequency domain conversion unit 331 sets the number of samples that overlap with one frame before to be L = 96, and sets the input signal L sample one frame before and the input signal x [n] of the frame to 2N = 320 samples. Prepare 4M = 512 samples from zero packing for L samples. Then, the frequency domain transform by the FFT is performed with the order of the FFT as 4M for the signal obtained by multiplying the 4M sample by the window function by the sinusoidal window.
The wideband power value N_wb_level [f] output from the ambient noise information band expansion unit 32 is input to the correction degree determination unit 332. Then, the correction gain G [f, w] (w = 0,1, ... 2M-1) is calculated and output by the equation (11).<maths num="11"></maths> N of equation (11)<sub>0</sub>Is the reference power value of ambient noise set by measuring the power of ambient noise in a normal usage environment in advance with the same sampling frequency and band limitation as the input signal x [n]. By doing so, even in an environment where the power value of ambient noise is larger than the normal usage environment (that is, an environment containing a lot of ambient noise), the correction gain G [f, w] can be set to be larger by that amount. , The input signal x [n] can be clarified.
In the correction processing unit 333, the frequency spectrum X [f, w] (w = 0,1, ... 2M-1) of the input signal x [n] and the correction gain G [f, w] (w = 0,1,... 2M-1) is input. Then, the frequency spectrum X [f, w] of the input signal x [n] is corrected by the equation (12), and the frequency spectrum Y [f, w] (w = 0) of the output signal y [n] which is the correction result is corrected. , 1,... 2M-1) is output.<maths num="12"></maths> The time domain conversion unit 334 performs time domain conversion (frequency inverse conversion) on the frequency spectrum Y [f, w] (w = 0,1, ... 2M-1) output from the correction processing unit 333. The output signal y [n], which is a corrected signal, is calculated by appropriately performing a process of returning the overlap in consideration of window hanging in the frequency domain conversion unit 331. For example, the frequency spectrum Y [f, w] (w = 0,1, ... 2M-1) in the pair and the input signal x [n] is taken into account that the actual signal frequency spectrum Y [f, After restoring w] to w = 0,1,... 4M-1, perform IFFT (Inverse Fast Fourier Transform) at 4M points, and consider windowing, and output signal y, which is a corrected signal one frame before. The overlap is returned using [n], and the output signal y [n] is calculated.
As described above, even if the frequency band in which the signal component exists or the sampling frequency is different between the reproduced input signal and the sound collection signal, the frequency band of the input signal regarding the volume of the sound collection signal By expanding and estimating the sound collection signal in consideration of, the volume of the sound collection signal can be obtained with high accuracy, and the clarity of the input signal can be improved.
In the above description, the case where the present invention is applied to a communication device has been described, but as shown in FIG. 7A, the present invention can also be applied to a digital audio player. This digital audio player includes a storage unit 11 using a flash memory or an HDD (Hard Disk Drive), and the decoder 2 decodes the music / audio data read from the storage unit 11. At this time, the target signal, which is a desired signal to be decoded and reproduced, is a music / audio signal. The sound collection signal z (t) collected by the microphone 6 of this digital audio player is a noise component caused by the voice of the near-end speaker and the surrounding environment, an output signal y (t), and an echo component caused by the acoustic space. It is composed of such as, and does not include music / audio signals. In this case, unlike the communication device, the voice of the near-end speaker is unnecessary, so all these components including the voice of the near-end speaker are treated as ambient noise.
Further, as shown in FIG. 7B, it is also possible to apply the present invention to a communication device and apply it to a voice band expansion communication device. This voice band expansion communication device is configured to include a decoder 2A and a signal band expansion processing unit 12 between the decoder 2A and the signal processing unit 3. Then, the signal processing unit 3 in this case performs the above-described processing on the input signal x'[n] whose band is expanded.
The processing performed by the signal band expansion processing unit 12 converts a narrow band input signal whose band is limited from fs_nb_low [Hz] to fs_nb_high [Hz] into a wide band signal from fs_wb_low [Hz] to fs_wb_high [Hz]. The process of expanding the band may be carried out by the existing techniques described in, for example, Tokuto 3189614, Tokuto 3243174, and JP-A-9-55778.
(Modification example 1 of the signal processing unit) Next, the narrowband signal information used in the signal processing unit will be described using the power spectrum of ambient noise, and the wideband signal information will be described as an example of the masking threshold value (broadband masking threshold value) when the ambient noise is expanded to a wideband signal. To do.
FIG. 8 shows the configuration. The signal processing unit 30 includes an ambient noise information band expansion unit 34 and a signal characteristic correction unit 35 in place of the ambient noise information band expansion unit 32 and the signal characteristic correction unit 33 used in the signal processing unit 3. Will be done.
FIG. 9 shows a configuration example of the ambient noise information band expansion unit 34. The ambient noise information band expansion unit 34 includes a power normalization unit 321, a dictionary storage unit 342, a wideband power spectrum calculation unit 343, a wideband masking threshold value calculation unit 344, and a power control unit 345.
Similar to the ambient noise information band expansion unit 32, the ambient noise information band expansion unit 34 exists in the input signal x [n] and exists in the sound collection signal z [n] with the power spectrum of the ambient noise as an input. Generates information (broadband signal information) including frequency band components that do not. That is, the ambient noise information band expansion unit 34 calculates the narrow band feature data from the narrow band signal information, and models in advance the correspondence between the narrow band feature data calculated from the narrow band signal information and the wide band feature data. Then, the wideband feature amount data is calculated by using the correspondence between this model and the acquired narrowband feature amount data, and the wideband signal information is generated from the wideband feature amount data. However, the ambient noise information band expansion unit 34 uses a method of using a codebook by vector quantization for modeling the correspondence between the narrow band feature data and the wide band feature data. Here, the normalized power spectrum of ambient noise | Nn [f, w] |<sup>2</sup> Wideband power spectrum normalized by ambient noise using (w = 0,1,... M-1) as Dnb-order narrowband feature data | Nw [f, w] |<sup>2</sup> (w = 0,1,... 2M-1) is used as the Dwb-order broadband feature data (Dnb = M, Dwb = 2M). Specifically, the ambient noise information band expansion unit 34 has a power spectrum of ambient noise | N [f, w] |<sup>2</sup> Power spectrum of ambient noise with (w = 0,1,... M-1) as input | N [f, w] |<sup>2</sup>The power spectrum of the frequency band component that exists in the input signal x [n] but does not exist in the sound collection signal z [n] is generated by frequency band expansion, and the masking threshold for the band-extended power spectrum is generated. Is obtained, and the resulting wideband masking threshold N_wb_th [f, w] (w = 0,1,... 2M-1) is output.
The dictionary storage unit 342 is a dictionary λ2 of a codebook of size Q (here, Q = 64) learned in advance by modeling the correspondence between the Dnb-order narrow-band feature data and the Dwb-order wideband feature data.<sub>q</sub>= {μx<sub>q</sub>, μy<sub>q</sub>} (Q = 1, ..., Q) is stored. In addition, μx<sub>q</sub>Is the centroid vector of narrowband features data in the qth codebook, μy<sub>q</sub>Represents the centroid vector of the broadband feature data in the qth codebook. The order of the code vector in the codebook is the centroid vector μx of the narrowband feature data.<sub>q</sub>And wideband feature data centroid vector μy<sub>q</sub>Dnb + Dwb, which is the sum of the components of.
Prior dictionary λ2 in the dictionary storage unit 342<sub>q</sub>A flowchart of the learning generation method of is shown in FIG. 10 and will be described. In the following description, the above-mentioned dictionary λ1<sub>q</sub>The same numbers are assigned to the same processes as the learning generation method of, and duplicate explanations are omitted as necessary to simplify the explanations.
The signal used to generate the dictionary of the codebook is the same as the input signal x [n], and a wide band signal with a sampling frequency fs'[Hz] limited from fs_wb_low [Hz] to fs_wb_high [Hz] is collected in advance. It is a group of sounded signals. It is desirable for this signal group to have many different environments and different volumes. In the following, the signal group of the wideband signal used for generating the dictionary of the codebook is collectively referred to as wideband signal data wb [n]. Further, n represents a time (sample).
First, the wideband signal data wb [n] is input and downsampled to the sampling frequency fs [Hz] to obtain the narrowband signal data nb [n] (step S101). Then, the narrowband feature data Pnb [f, d] (d = 1, ..., Dnb), which is the feature data representing the narrowband signal information, is extracted from the narrowband signal data nb [n] (step S202). In this step S202, the power spectrum (Mth order) of the narrowband signal data nb [n] is obtained (step S1021), the power value of the narrowband signal data nb [n] is obtained (step S1022), and these powers are obtained. A normalized power spectrum of narrowband signal data nb [n] is obtained from the spectrum and power value (step S1023), and this is used as narrowband feature data Pnb [f, d] (d) of order Dnb (= M). = 1,..., Dnb) Narrowband feature data is extracted by.
On the other hand, the wideband feature data Pwb [f, d] (d = 1, ..., Dwb), which is the feature data representing the wideband signal information, is extracted from the wideband signal data wb [n] (step S203). In this step S203, the power spectrum of the wideband signal data wb [n] is obtained (step S1031), and the power value of the wideband signal data wb [n] is obtained from the wideband signal data wb [n] in frame units (step S2032). ), The normalized power spectrum of the broadband signal data wb [n] is obtained in frame units from these power spectra and power values (step S2033), and this is obtained by the broadband feature data Pwb [of order Dwb (= 2M)]. Broadband feature data is extracted by setting f, d] (d = 1, ..., Dwb).
Next, the narrowband feature data Pnb [f, d] (d = 1, ..., Dnb) and the wideband feature data Pwb [f, d] (d = 1, ..., Dwb) are concatenated to form the order Dnb. + Dwb concatenated feature data P [f, d] (d = 1, ..., Dnb + Dwb) is generated (step S104).
From the above-mentioned concatenated feature data P [f, d] to the size Q (here, Q = 64) codebook dictionary λ2<sub>q</sub>= {μx<sub>q</sub>, μy<sub>q</sub>} (Q = 1, ..., Q) is generated by using a clustering method such as a k-means algorithm or an LBG algorithm (step S205). In step S205, first, the narrowband centroid vector μx<sub>1</sub>Is the average of all narrowband feature data, and the wideband centroid vector μy<sub>1</sub>Is used as the average of all the broadband feature data to generate an initial codebook of size Q = 1 (step S2051). Then, it is determined whether or not the size Q of the codebook has reached a predetermined number (64 in this case) (step S2052). If the size Q of the codebook does not reach the specified number, the codebook λ2<sub>q</sub>Narrowband centroid vector μx in each code vector of<sub>q</sub>And wideband centroid vector μy<sub>q</sub>Is slightly shifted to generate another code vector to double the size Q of the codebook (step S2053). Then, for the concatenated feature data P [f, d] of the order Dnb + Dwb, the codebook λ2<sub>q</sub>Narrowband centroid vector μx in each code vector of<sub>q</sub>The code vector that minimizes the predetermined distance scale (for example, Euclidean distance or Mahalanobis distance) is obtained, and the concatenated feature data P [f, d] is assigned to the corresponding code vector. Then codebook λ2<sub>q</sub>Using the concatenated feature data P [f, d] assigned to each code vector of, a new narrowband centroid vector μx for each code vector<sub>q</sub>And wideband centroid vector μy<sub>q</sub>In search of the codebook λ2<sub>q</sub>Is updated (step S2054). If the size Q of the codebook reaches the specified number, the codebook λ2<sub>q</sub>= {μx<sub>q</sub>, μy<sub>q</sub>} (Q = 1, ..., Q) is output.
The wideband power spectrum calculation unit 343 uses the normalized power spectrum of the ambient noise output from the power normalization unit 321 | Nn [f, w] |<sup>2</sup> Input (w = 0,1, ... M-1) as Dnb next feature data, and enter the codebook dictionary λ2 from the dictionary storage unit 342.<sub>q</sub>= {μx<sub>q</sub>, μy<sub>q</sub>} (Q = 1, ..., Q) is read out, and the wideband power spectrum is based on the correspondence between the Dnb-order narrowband feature data and the Dwb-order wideband feature data | Nw [f, w] |<sup>2</sup> Find (w = 0,1,... 2M-1). Specifically, there are Q narrowband centroid vectors μx<sub>q</sub>From (q = 1,..., Q), the normalized power spectrum of ambient noise | Nn [f, w] |<sup>2</sup> Find the one with the closest distance on a predetermined distance scale with (w = 0,1,... M-1), and find the wide band centroid vector μy in the code vector with the shortest distance.<sub>q</sub>Wideband power spectrum | Nw [f, w] |<sup>2</sup> Let (w = 0,1,... 2M-1).
The wideband masking threshold value calculation unit 344 is a wideband power spectrum output from the wideband power spectrum calculation unit 343 | Nw [f, w] |<sup>2</sup> Using (w = 0,1,... 2M-1) as an input, the wideband masking threshold N_wb_th1 [f, w] (w = 0,1,... 2M-1), which is the masking threshold for ambient noise, is calculated for each frequency component. To do.
Generally, the masking threshold can be calculated by convolving a function called a spreading function into the power spectrum of the signal. That is, the wideband masking threshold value N_wb_th1 [f, w] (w = 0,1, ... 2M-1) for ambient noise is calculated by the equation (13) with the spreading function as the function sprdngf (). Wideband power spectrum of ambient noise | Nw [f, w] |<sup>2</sup>If is less than or equal to the wideband masking threshold N_wb_th1 [f, w], it is masked by the wideband power spectrum of ambient noise in frequency bands other than frequency bin ω. FIG. 11 shows an example of a broadband masking threshold value of ambient noise collected in various environments such as outdoors, where the horizontal axis is frequency [Hz] and the vertical axis is power [dB].<maths num="13"></maths> Here, bark [w] represents a bark value obtained by converting the frequency bin ω into a bark scale, and in the spreading function, it is appropriately converted into a bark scale bark [w]. The Bark scale is a scale that is set finer in the low range and coarser in the high range in consideration of auditory resolution.
Here, the spreading function is used as the function sprdngf (), and the method defined in ISO / IEC 13818-7 is used. For the spreading function, other methods described in the literature such as ITU-R 1387 and 3GPP TS 26.403 may be used. In addition to the Bark scale, a spreading function using a scale obtained from a human pitch perception characteristic such as a Mel scale or an ERB scale or an auditory filter may be appropriately used.
The power control unit 345 has a narrow band power value Pow_N [f] output from the power normalization 321 and a wide band masking threshold N_wb_th1 [f, w] (w = 0,1, ... With 2M-1) as the input, the wideband masking threshold N_wb_th1 [f, w] has the same power at fs_nb_low [Hz] to fs_nb_high [Hz] as the narrowband power value Pow_N [f]. It is controlled by amplifying or attenuating f, w], and this power-controlled N_wb_th1 [f, w] is output as the wideband masking threshold N_wb_th [f, w].
In this way, in the ambient noise information band expansion unit 34, the power spectrum of the ambient noise, which is the narrow band signal information | N [f, w] |<sup>2</sup>From, the wideband masking threshold value N_wb_th [f, w], which is the wideband signal information, is obtained.
FIG. 12 shows a configuration example of the signal characteristic correction unit 35. The signal characteristic correction unit 35 includes a frequency domain conversion unit 331, a power calculation unit 352, a masking threshold value calculation unit 353, a masking determination unit 354, a power smoothing unit 355, a correction degree determination unit 356, and a correction processing unit. It includes 333 and a time domain conversion unit 334.
The signal characteristic correction unit 35 inputs the input signal x [n] (n = 0,1, ... 2N-1) and the wideband masking threshold N_wb_th [f, w], and the input signal x [n] becomes the sound collecting signal. A signal correction process is performed to clarify the sound so that it is not buried in the ambient noise contained in it, and the corrected output signal y [n] (n = 0,1,... 2N-1) is output.
The power calculation unit 352 has two parts, a real part and an imaginary part, in the frequency spectrum X [f, w] (w = 0,1, ... 2M-1) of the input signal x [n] output from the frequency domain conversion unit 331. Power spectrum that is the sum of powers | X [f, w] |<sup>2</sup> Calculate and output (w = 0,1, ... 2M-1).
The masking threshold calculation unit 353 uses the power spectrum of the input signal x [n] output from the power calculation unit 352 | X [f, w] |<sup>2</sup> With (w = 0,1,... 2M-1) as the input and the spreading function as the function sprdngf (), the wideband masking threshold X_th [f, w] (w) of the input signal x [n] in the equation (14) = 0,1,... 2M-1) is calculated and output. The broadband masking threshold X_th [f, w] is the power spectrum of the input signal x [n] | X [f, w] |<sup>2</sup> If is less than or equal to the wideband masking threshold X_th [f, w] of the input signal x [n] For example, the power spectrum of the input signal x [n] in a frequency band other than the frequency bin ω | X [f, w] |<sup>2</sup> Indicates that it is masked by.<maths num="14"></maths> The masking determination unit 354 uses the power spectrum output from the power calculation unit 352 | X [f, w] |<sup>2</sup> (w = 0,1, ... 2M-1) and the wideband masking threshold value X_th [f, w] output from the masking threshold value calculation unit 353 are used as inputs, and each frequency band is masked by the input signal x [n] itself. The masking judgment information X_flag [f, w] (w = 0,1,... 2M-1) indicating whether or not the data is displayed is output. Specifically, the power spectrum | X [f, w] |<sup>2</sup> And the broadband masking threshold X_th [f, w] are compared in magnitude, and the power spectrum | X [f, w] |<sup>2</sup> If is greater than or equal to the broadband masking threshold X_th [f, w], then X_flag [f, w] = 0, assuming that the frequency component is not masked by other frequency components in the input signal x [n]. Also, the power spectrum | X [f, w] |<sup>2</sup> If is less than the broadband masking threshold X_th [f, w], then that frequency component is masked by other frequency components in the input signal x [n] and X_flag [f, w] = 1. The power smoothing unit 355 is a power spectrum output from the power calculation unit 352 | X [f, w] |<sup>2</sup> With (w = 0,1, ... 2M-1) and the masking judgment information X_flag [f, w] output from the masking judgment unit 354 as inputs, the power spectrum | X [f, w] |<sup>2</sup> Is smoothed by the moving average by the triangular window according to the equation (15), and the smoothed power spectrum | X<sub>S</sub>[f, w] |<sup>2</sup> Is output. Note that K is the range for calculating smoothing, and α<sub>X</sub>[j] is a smoothing coefficient such that the coefficient increases as j approaches 0. For example, with K = 3, α<sub>X</sub>[j] is [0.1, 0.2, 0.4, 0.8, 0.4, 0.2, 0.1].<maths num="15"></maths> The correction degree determination unit 356 uses the smoothed power spectrum output from the power smoothing unit 355 | X.<sub>S</sub>[f, w] |<sup>2</sup> (w = 0,1, ... 2M-1), masking judgment information X_flag [f, w] (w = 0,1, ... 2M-1) output from masking judgment unit 354, and ambient noise information band expansion unit 32. The correction gain G [f, w] (w = 0,1,... 2M-1) is calculated by inputting N_wb_th [f, w] (w = 0,1,... 2M-1) output from. And output. The specific calculation of the correction gain G [f, w] is first masked by the masking determination information X_flag [f, w] to other frequency components in the input signal x [n] (X_flag [f, w] = If the frequency band is determined to be 1), set G [f, w] = 1 so that amplification and attenuation by correction are not performed. Then, for the frequency band determined by the masking determination information X_flag [f, w] not to be masked by other frequency components in the input signal x [n] (X_flag [f, w] = 0), the power spectrum | X [ f, w] |<sup>2</sup> And the broadband masking threshold N_wb_th [f, w] of ambient noise are compared. Where the power spectrum | X [f, w] |<sup>2</sup> If is greater than or equal to the wideband masking threshold N_wb_th [f, w] of ambient noise, the frequency component is not masked by other frequency components in the sound collection signal z [n], so G [f, w] = 1 and correction is applied. Avoid amplification. On the other hand, power spectrum | X [f, w] |<sup>2</sup> If is less than the wideband masking threshold N_wb_th [f, w] of ambient noise, it is judged that the ambient noise is masked due to the presence of ambient noise even though it can be perceived if the ambient noise in the sound collection signal z [n] is small. Then, the correction gain G [f, w] is smoothed to the wideband masking threshold N_wb_th [f, w] of the ambient noise as in the equation (16).<sub>S</sub>[f, w] |<sup>2</sup>Calculated based on the ratio with. The function F is a smoothed power spectrum | X.<sub>S</sub>[f, w] |<sup>2</sup>It is a function that amplifies the spectral gradient of the ambient noise so that it is close to parallel to the shape of the wideband masking threshold N_wb_th [f, w] of ambient noise. Here, α and β are positive constants, and γ is either a positive or negative constant. These constants are used to adjust the degree of amplification of the input signal x [n].<maths num="16"></maths><maths num="17"></maths> In the correction degree determination unit 356, the correction gain G [f, w] thus obtained is further smoothed by the moving average by the triangular window according to the equation (22), and the smoothed correction gain G is obtained.<sub>S</sub>[f, w] may be used. Note that K is the range for calculating smoothing, and α<sub>X</sub>[j] is a smoothing coefficient such that the coefficient increases as j approaches 0. For example, with K = 3, α<sub>G</sub>[j] is [0.1, 0.2, 0.4, 0.8, 0.4, 0.2, 0.1].<maths num="18"></maths> As described above, even if the reproduced input signal and the sound collection signal have different frequency bands in which signal components exist or different sampling frequencies, the power spectrum, which is the frequency characteristic of the sound collection signal, is By expanding the band and estimating by adding the frequency band of the input signal, the frequency characteristic of the sound collection signal can be obtained with high accuracy, and the clarity of the input signal can be improved.
When applying this modification to the voice band expansion communication device shown in FIG. 7 (b), the frequency f_limit (f_limit is about 500 to 1200 [Hz]) preset in the signal band expansion processing unit 12, for example. When the low frequency band below f_limit = 1000 [Hz] is extended, that is, when fs_wb_low <fs_nb_low and fs_wb_low <f_limit, the signal characteristic correction unit 35 does not perform signal correction processing for the frequency band below f_limit. To do so. In the low frequency range (frequency below f_limit), the ambient noise varies widely depending on the sound collecting environment and the type of noise component. Therefore, by doing so, the signal band expansion processing unit 12 extends the low frequency band. It is possible to prevent the signal correction process from becoming unstable due to variations in ambient noise.
(Modification 2 of signal processing unit) In this modification, the narrowband signal information used by the signal processing unit 30 shown in FIG. 8 is used as the power spectrum of ambient noise, and the wideband signal information is used as the wideband power spectrum of ambient noise (when the ambient noise is extended to a wideband signal). The case of power spectrum) will be described as an example. In this case, the ambient noise information band expansion unit 34 takes the power spectrum of the ambient noise, which is the narrowband signal information, as an input, calculates the normalized power spectrum of the ambient noise as the narrowband feature amount data, and wideband feature amount data. The normalized wideband power spectrum of the ambient noise is calculated using the correspondence between the narrowband feature data and the wideband feature data modeled in advance, and the ambient signal information is obtained from this wideband feature data. Try to generate a wideband power spectrum of noise. In order to model the correspondence between the narrow-band feature data and the wide-band feature data, the method using GMM shown in FIG. 5 is used. According to this, even if the reproduced input signal and the sound collection signal have different frequency bands in which signal components exist or different sampling frequencies, the power spectrum, which is the frequency characteristic of the sound collection signal, is By expanding the band and estimating by adding the frequency band of the input signal, the frequency characteristic of the sound collection signal can be obtained with high accuracy, and the clarity of the input signal can be improved.
(Modification 3 of signal processing unit) Next, the narrowband signal information used in the signal processing unit will be described using the power spectrum of ambient noise, and the wideband signal information will be described as an example of the masking threshold value (broadband masking threshold value) when the ambient noise is expanded to a wideband signal. To do.
FIG. 13 shows the configuration. The signal processing unit 300 has a configuration in which the ambient noise information band expansion unit 36 is used instead of the ambient noise information band expansion unit 34 used in the signal processing unit 30.
FIG. 14 shows a configuration example of the ambient noise information band expansion unit 36. The ambient noise information band expansion unit 36 includes a power normalization unit 321, a narrow band masking threshold value calculation unit 362, a band control unit 363, a dictionary storage unit 364, a wideband masking threshold value calculation unit 365, and a threshold value correction unit 366. , A power control unit 345 is provided.
Similar to the ambient noise information band expansion unit 34, the ambient noise information band expansion unit 36 inputs information (narrow band signal information) in the frequency band component of the sound collection signal z [n] to the input signal x [n]. Generates information (broadband signal information) including frequency band components that exist and do not exist in the sound collection signal z [n]. That is, the ambient noise information band expansion unit 36 calculates the narrow band feature data from the narrow band signal information, models the correspondence between the narrow band feature data and the wide band feature data in advance, and acquires this model. The wideband feature amount data is calculated by using the correspondence with the narrowband feature amount data, and the wideband signal information is generated from the wideband feature amount data. At this time, the ambient noise information band expansion unit 36 uses a method of using a codebook by vector quantization for modeling the correspondence between the narrow band feature data and the wide band feature data. Here, the band-controlled narrow-band masking threshold of ambient noise N_th [f, w] (w = 0,1, ... M)<sub>C</sub>It is used as the Dnb-order narrow-band feature data of -1), and is used as the Dwb-th-order wide-band feature data of the ambient noise broadband masking threshold N_wb_th1 [f, w] (w = 0,1, ... 2M-1) (-1). Dnb = M<sub>C</sub>, Dwb = 2M). Specifically, the ambient noise information band expansion unit 36 has a power spectrum of ambient noise | N [f, w] |<sup>2</sup> Using (w = 0,1, ... M-1) as an input, the masking threshold of ambient noise is obtained, the masking threshold is band-limited, and the band-limited masking threshold exists in the input signal x [n]. A frequency band component that does not exist in the sound collection signal z [n] is generated by expanding the frequency band, and the wide band masking threshold N_wb_th [f, w] (w = 0,1, ... 2M-1) is output.
The narrowband masking threshold value calculation unit 362 uses the normalized power spectrum of ambient noise output from the power normalization unit 321 | Nn [f, w] |<sup>2</sup>With (w = 0,1, ... M-1) as the input, the narrowband masking threshold N_th1 [f, w] (w = 0,1, ... M-1), which is the masking threshold for ambient noise, is set for each frequency component. calculate. Similar to the above-mentioned wideband masking threshold value calculation unit 344, the data length of 2M is replaced with M, and the narrowband masking threshold value N_th1 [f, w] (w = 0,1, ... M-1) of ambient noise is It is calculated by the formula (19) with the spreading function as the function sprdngf (). The narrowband masking threshold N_th1 [f, w] is the normalized power spectrum of ambient noise | Nn [f, w] |<sup>2</sup>If is less than or equal to the narrowband masking threshold N_th1 [f, w], it indicates that it is masked by the normalized power spectrum of ambient noise in frequency bands other than frequency bin ω.<maths num="19"></maths> The band control unit 363 receives the narrow band masking threshold N_th1 [f, w] (w = 0,1, ... M-1) of the ambient noise output from the narrow band masking threshold calculation unit 362 as an input, and controls the lower limit of the band. It controls to use only the signal information of the frequency band from the frequency limit_low [Hz] to the upper limit frequency limit_high [Hz] for band control, and outputs the band-controlled narrow band masking threshold N_th [f, w]. However, fs_nb_low limit_low <limit_high fs_nb_high <fs / 2. For example, when limit_low = 1000 [Hz] and limit_high = 3400 [Hz], if these frequency bands are converted to frequency bin ω by equation (24) and considered, the narrow band masking threshold N_th1 [f, w] (w) Only w = 32,33, ... 108 out of = 0,1, ... M-1) should be used. M<sub>C</sub>Is the number of arrays of N_th [f, w], and the bandwidth-controlled narrowband masking threshold N_th [f, w] (w = 0,1,... M)<sub>C</sub>-1) substitutes the narrowband masking threshold N_th1 [f, w] (w = 32,... 108) itself. In this case M<sub>C</sub>= 108-32 + 1 = 77.
As shown in FIG. 11, it can be seen that in the low frequency range, the dispersion / variation of the masking threshold value of ambient noise is large depending on the sound collecting environment and the type of noise component. Since the main component of ambient noise is the noise component, the dispersion / variation of the narrow band masking threshold N_th1 [f, w] also increases in the low frequency range. Therefore, in order to obtain the wideband masking threshold with high accuracy by using a method using a codebook by vector quantization for modeling the correspondence between narrowband feature data and wideband feature data, the variance and variation are large and low. Bandwidth is controlled so that the region is not used. That is, here, it is desirable to set the lower limit frequency limit_low [Hz] for band control to the lower limit of the frequency band so that the variance / variation of the narrow band masking threshold value is smaller than a predetermined value. By doing so, the wideband masking threshold value can be obtained with high accuracy, and the intelligibility of the input signal can be improved.
Further, the masking threshold value is calculated by taking into account not only the power spectrum of the frequency band but also the power spectrum of the surrounding frequency band. Therefore, the masking threshold cannot be calculated accurately in the vicinity of the frequency band in which the band of the original signal for which the masking threshold is obtained is limited. That is, it is desirable to set the upper limit frequency limit_high [Hz] for band control to the upper limit of the frequency band in which the masking threshold can be accurately obtained even when the band limit is taken into consideration. By doing so, the wideband masking threshold value can be obtained with high accuracy, and the intelligibility of the input signal can be improved.
The dictionary storage unit 364 is a pre-learned size Q (here, Q = 64) codebook dictionary λ3 that models the correspondence between the Dnb-order narrowband feature data and the Dwb-order wideband feature data.<sub>q</sub>= {μx<sub>q</sub>, μy<sub>q</sub>} (Q = 1, ..., Q) is stored. In addition, μx<sub>q</sub>Is the centroid vector of narrowband features data in the qth codebook, μy<sub>q</sub>Represents the centroid vector of the broadband feature data in the qth codebook. The order of the code vector in the codebook is the centroid vector μx of the narrowband signal information.<sub>q</sub>And wideband signal information centroid vector μy<sub>q</sub>Dnb + Dwb, which is the sum of the components of.
Prior dictionary λ3 in the dictionary storage unit 364<sub>q</sub>A flowchart of one method of the learning generation method of is shown in FIG. 15 and will be described. In the following description, the dictionary λ2 in the above-described modification 1<sub>q</sub>The same numbers are assigned to the same processes as the learning generation method of, and duplicate explanations are omitted as necessary to simplify the explanations.
First, the wideband signal data wb [n] is input and downsampled to the sampling frequency fs [Hz] to obtain the narrowband signal data nb [n] (step S101). Then, the narrowband feature data Pnb [f, d] (d = 1, ..., Dnb), which is the feature data representing the narrowband signal information, is extracted from the narrowband signal data nb [n] (step S202). In this step S202, the power spectrum (Mth order) of the narrowband signal data nb [n] is obtained (step S1021), the power value of the narrowband signal data nb [n] is obtained (step S1022), and these powers are obtained. The normalized power spectrum of the narrowband signal data nb [n] is obtained from the spectrum and the power value (step S1023), and the masking threshold of the narrowband signal data nb [n] is calculated in the same manner as in Eq. (23). (Step S3024). Then, the masking threshold value of the narrow band signal data nb [n] is band-controlled in the same manner as the processing by the band control unit 363 (step S3025). This is the order Dnb (= M)<sub>C</sub>) Narrowband feature data Pnb [f, d] (d = 1, ..., Dnb) to extract narrowband feature data.
On the other hand, the wideband feature data Pwb [f, d] (d = 1, ..., Dwb), which is the feature data representing the wideband signal information, is extracted from the wideband signal data wb [n] (step S303). In this step S303, the power spectrum (2Mth order) of the wideband signal data wb [n] is obtained (step S1031), and the power value of the wideband signal data wb [n] is obtained from the wideband signal data wb [n] (step S1031). S2032), the normalized power spectrum of the broadband signal data wb [n] is obtained in frame units from these power spectra and power values (step S2033), and the order of the equation (23) is changed from M to 2M in the same manner. The masking threshold of the broadband signal data wb [n] is calculated (step S3034). Wideband feature data is extracted by setting this as the wideband feature data Pwb [f, d] (d = 1, ..., Dwb) of order Dwb (= 2M).
Next, the narrowband feature data Pnb [f, d] (d = 1, ..., Dnb) and the wideband feature data Pwb [f, d] (d = 1, ..., Dwb) are concatenated to form the order Dnb. + Dwb concatenated feature data P [f, d] (d = 1, ..., Dnb + Dwb) is generated (step S104).
Then, from the concatenated feature data P [f, d], the narrowband centroid vector μx in each code vector of the codebook<sub>q</sub>And wideband centroid vector μy<sub>q</sub>Is obtained, and a codebook of size Q (Q = 64 in this case) is generated by using a clustering method such as a k-means algorithm or an LBG algorithm (step S205). Wideband centroid vector μy in each code vector in the codebook<sub>q</sub>The masking threshold of the broadband signal data wb [n] is expressed by the approximate polynomial coefficient, and the approximate polynomial coefficient is expressed by the broadband centroid vector μ'y.<sub>q</sub>Stored in the dictionary as, dictionary λ3<sub>q</sub>= {μx<sub>q</sub>, μ'y<sub>q</sub>} (Q = 1, ..., Q) is generated (step S307). Approximate polynomial coefficient m<sub>p</sub>What is (p = 0, ..., P)? Here, the vertical axis is the power value X [dB] and the horizontal axis is the frequency Y [Hz], and the masking threshold is set to a predetermined order (here, as shown in equation (20)). It is the coefficient of the polynomial approximated by the polynomial of (P, for example, P = 6), and will be referred to as such thereafter.<maths num="20"></maths> By expressing the masking threshold value as an approximate polynomial coefficient and storing it as a dictionary in this way, it is possible to reduce the amount of memory required for storing the dictionary as compared with storing the masking threshold value as a dictionary. Since the number of arrays of is reduced, the amount of processing when using the dictionary can be reduced.
The wideband masking threshold value calculation unit 365 is a band-controlled narrowband masking threshold value N_th [f, w] (w = 0,1, ... M) output from the band control unit 363.<sub>C</sub>-1) is input as Dnb next feature data, and the dictionary λ3 of the codebook is input from the dictionary storage unit 364.<sub>q</sub>= {μx<sub>q</sub>, μ'y<sub>q</sub>} (Q = 1, ..., Q) is read, and the wideband masking threshold N_wb_th1 [f, w] (w = 0) of ambient noise is obtained from the correspondence between the Dnb-order narrowband feature data and the Dwb-order wideband feature data. , 1,... 2M-1) is calculated. Specifically, there are Q narrowband centroid vectors μx<sub>q</sub>From (q = 1, ..., Q), the band-controlled narrowband masking threshold N_th [f, w] (w = 0,1, ... M)<sub>C</sub>Find the closest distance to -1) and the predetermined distance scale, and find the wide band centroid vector μ'y in the code vector with the shortest distance.<sub>q</sub>Is set as it is as an approximate polynomial coefficient of the wideband masking threshold value, and the wideband masking threshold value N_wb_th1 [f, w] (w = 0,1, ... 2M-1) is calculated in the same manner as in Eq. (20).
The threshold value correction unit 366 is from the narrow band masking threshold value N_th1 [f, w] (w = 0,1, ... M-1) of the ambient noise output from the narrow band masking threshold value calculation unit 362 and the wide band masking threshold value calculation unit 365. Using the wideband masking threshold N_wb_th1 [f, w] (w = 0,1,... 2M-1) of the output ambient noise as an input, discontinuity or differential discontinuity near the boundary band in the narrow band and wide band can be obtained. Corrected to eliminate, and the corrected wideband masking threshold N_wb_th2 [f, w] Output (w = 0,1, ... 2M-1). In FIG. 16A, a discontinuity occurs between the narrow band masking threshold N_th [f, w] and the wide band masking threshold N_wb_th1 [f, w] at frequencies around the boundary band fs / 2 [Hz]. An example of the wideband masking threshold N_wb_th2 [f, w] corrected to be eliminated is shown. FIG. 16B shows discontinuity and differential discontinuity between the narrow band masking threshold N_th [f, w] and the wide band masking threshold N_wb_th1 [f, w] at frequencies around the boundary band fs / 2 [Hz]. An example of the wideband masking threshold N_wb_th2 [f, w] is shown, in which both of the above occur and are corrected to eliminate them. In both figures, the solid line is the narrowband masking threshold N_th [f, w], the broken line is the wideband masking threshold N_wb_th2 [f, w], and the thick solid line is the corrected wideband masking threshold N_wb_th2 [f, w]. Represent. However, adjust_low [Hz] <fs / 2 <adjust_high [Hz]. Where adjust_low is the frequency bin ω<sub>L</sub>Frequency bin ω above the frequency corresponding to -1<sub>L</sub>Is less than the frequency corresponding to, and adjust_high is the frequency bin ω<sub>H</sub>Frequency bin ω above the frequency corresponding to<sub>H</sub>It is assumed that the frequency is less than the frequency corresponding to +1. For example, when fs = 8000 [Hz], adjust_low = 3600 [Hz] and adjust_high = 4400 [Hz]. Specifically, when a discontinuity or differential discontinuity is detected at least in the frequency around the boundary band fs / 2 [Hz], the boundary band is equal to or more than adjust_low [Hz] and less than or equal to adjust_high [Hz]. About the vicinity, frequency bin ω<sub>L</sub>, Ω<sub>L</sub>+1 ..., ω<sub>L</sub>+ S and ω<sub>H</sub>, Ω<sub>H</sub>-1 ..., ω<sub>H</sub>Frequency bin ω using the broadband masking threshold N_wb_th1 [f, w] in S<sub>L</sub>+ S + 1 to ω<sub>H</sub>The wideband masking threshold up to S-1 is simulated by the (2S-1) order function, and the corrected wideband masking threshold value N_wb_th2 [f, w] is obtained by performing spline interpolation. Here, spline interpolation may be performed by setting a function that simulates passing through the midpoint between the narrow band masking threshold N_th1 [f, M-1] and the wide band masking threshold N_wb_th1 [f, M].
By correcting the wideband masking threshold value in the threshold value correction unit 366 in this way, the discontinuity or differential discontinuity in the wideband masking threshold value is eliminated, and the discontinuity in the frequency direction is eliminated even in the signal correction, so that there is no sense of discomfort. The signal can be corrected and a high degree of clarity can be obtained.
As described above, even if the frequency band in which the signal component exists or the sampling frequency is different between the reproduced input signal and the sound collection signal, the frequency of the input signal is about the masking threshold of the sound collection signal. By expanding the band in consideration of the band and estimating, the masking threshold of the sound collection signal can be obtained with high accuracy, and the clarity of the input signal can be improved.
(Modification example 4 of the signal processing unit) Prior dictionary λ3 in the dictionary storage unit 364 of the signal processing unit 300<sub>q</sub>A flowchart of another method of the learning generation method of is shown in FIG. 17 and will be described. Here, the dictionary λ3 is derived only from the wideband signal data wb [n] without generating the narrowband signal data nb [n].<sub>q</sub>The method of learning and generating is described. In the following description, the dictionary λ3 in the above-described modification 2<sub>q</sub>The same numbers are assigned to the same processes as the learning generation method of, and duplicate explanations are omitted as necessary to simplify the explanations.
First, in step S303, the wideband feature data Pwb [f, d] (d = 1, ..., Dwb), which is the feature data (here, the masking threshold) representing the wideband signal information, is extracted from the wideband signal data wb [n]. To do. Using only this wideband feature data Pwb [f, d] (d = 1, ..., Dwb), a codebook of size Q is created in step S205. Then, the broadband centroid vector μy in each code vector of the codebook.<sub>q</sub>For the wideband masking threshold of the wideband signal data wb [n], only the wideband masking threshold of the frequency band from the lower limit frequency limit_low [Hz] for bandwidth control to the upper limit frequency limit_high [Hz] for bandwidth control is used. (Step S3025). As a result, a narrowband masking threshold that is band-controlled in a narrow band is obtained, and this is calculated as the narrowband centroid vector μx in each code vector of the codebook.<sub>q</sub>(Q = 1, ..., Q) (step S306). Then, in step S307, the broadband centroid vector μ'y, which is an approximate polynomial coefficient of the masking threshold of the broadband signal data wb [n].<sub>q</sub>Store it in the dictionary together with the dictionary λ3<sub>q</sub>= {μx<sub>q</sub>, μ'y<sub>q</sub>} Is generated.
In the method of FIG. 15 for clustering using narrow-band feature data in combination, the narrow-band feature data includes an error near the boundary band between the narrow band and the wide band. In this way, by clustering using only the wideband feature data and band-limited the wideband centroid vector to obtain the narrowband centroid vector, clustering is performed using only the wideband feature data, which is ideal data. Therefore, clustering can be performed with higher accuracy than the method shown in FIG.
(Modification 5 of signal processing unit) Prior dictionary λ3 in the dictionary storage unit 364 of the signal processing unit 300<sub>q</sub>A flowchart of another method of the learning generation method of is shown in FIG. 18 and will be described. In the following description, the dictionary λ3 in the above-described modification 2<sub>q</sub>The same numbers are assigned to the same processes as the learning generation method of, and duplicate explanations are omitted as necessary to simplify the explanations.
After creating the size Q codebook in step S205, the narrowband centroid vector μx in each code vector of the codebook.<sub>q</sub>The masking threshold of the narrowband signal data nb [n] is expressed by an approximate polynomial as in Eq. (20), and the approximate polynomial coefficient is expressed by the narrowband centroid vector μ'x.<sub>q</sub>(Q = 1, ..., Q) (step S306A). Then, in step S307, the broadband centroid vector μ'y, which is an approximate polynomial coefficient of the masking threshold of the broadband signal data wb [n].<sub>q</sub>Store it in the dictionary together with the dictionary λ3<sub>q</sub>= {μ'x<sub>q</sub>, μ'y<sub>q</sub>} Is generated.
On the other hand, in this method, in the wideband masking threshold value calculation unit 365, the band-controlled narrowband masking threshold value N_th [f, w] (w = 0,1, ... M) output from the band control unit 363 is used.<sub>C</sub>-1) is input as Dnb next feature data, and the dictionary λ3 of the codebook is input from the dictionary storage unit 364.<sub>q</sub>= {μ'x<sub>q</sub>, μ'y<sub>q</sub>} (Q = 1, ..., Q) is read, and the wideband masking threshold N_wb_th1 [f, w] (w = 0) of ambient noise is obtained from the correspondence between the Dnb-order narrowband feature data and the Dwb-order wideband feature data. , 1,... 2M-1) is calculated. Specifically, there are Q narrowband centroid vectors μ'x<sub>q</sub>Band-controlled narrowband masking threshold N_th [f, w] (w = 0,1, ... M) from the approximate polynomial of (q = 1, ..., Q)<sub>C</sub>The wideband centroid vector μ'y in the code vector with the closest distance is obtained by substituting the one with the closest distance from -1) and the predetermined distance scale into the approximate polynomial.<sub>q</sub>Is set as it is as an approximate polynomial coefficient of the wideband masking threshold value, and the wideband masking threshold value N_wb_th1 [f, w] (w = 0,1, ... 2M-1) is calculated in the same manner as in Eq. (20).
In this way, by expressing the narrow band masking threshold value as an approximate polynomial coefficient and storing it as a dictionary, the dictionary can be stored even when compared with the method in FIG. 15 rather than storing the masking threshold value as a dictionary. The amount of memory required for the dictionary can be reduced, and the number of arrays in the dictionary can be reduced, so that the amount of processing when using the dictionary can be reduced.
(Second Example) FIG. 19A shows the configuration of the communication device according to the second embodiment of the present invention.
The communication device shown in this figure shows a receiving system of a wireless communication device such as a mobile phone, and includes a wireless communication unit 1, a decoder 2, a signal processing unit 3A, and a digital analog (D / A). It includes a converter 4, a speaker 5, a microphone 6, an analog-to-digital (A / D) converter 7, a downsampling unit 8, an echo suppression processing unit 9, and an encoder 10.
Similar to the first embodiment, the present invention can be applied not only to the communication device as shown in FIG. 19 (a) but also to the digital audio player shown in FIG. 19 (b). It can also be applied to the voice band expansion communication device shown in FIG. 19 (c).
Next, the signal processing unit 3A will be described. FIG. 20 shows the configuration. The signal processing unit 3A is configured by adding an ambient noise suppression processing unit 37 to the signal processing unit 3 described in the first embodiment. In the following description, the same configurations as those in the above-described embodiment will be numbered the same, and duplicate description will be omitted as necessary.
FIG. 21 shows a configuration example of the ambient noise suppression processing unit 37. The ambient noise suppression processing unit 37 includes a suppression gain calculation unit 371, a spectrum suppression unit 372, a power calculation unit 373, and a time domain conversion unit 374.
The ambient noise suppression processing unit 37 uses the power spectrum of the ambient noise output from the ambient noise estimation unit 31, the power spectrum of the sound collection signal z [n], and the frequency spectrum of the sound collection signal z [n] to collect sound. The noise component which is the ambient noise contained in the signal z [n] is suppressed, and the signal s [n] in which the noise component which is the ambient noise is suppressed is output to the encoder 10. The encoder 10 encodes the signal s [n] output from the ambient noise suppression processing unit 37 and outputs it to the wireless communication unit 1.
The suppression gain calculation unit 371 uses the power spectrum of the sound collection signal z [n] output from the power calculation unit 312 | Z [f, w] |<sup>2</sup> (w = 0,1, ... M-1) and the power spectrum of ambient noise output from the frequency spectrum updater 314 | N [f, w] |<sup>2</sup> (w = 0,1, ... M-1) and the power spectrum of the suppressed signal one frame before output from the power calculation unit 373 | S [f-1, w] |<sup>2</sup> Using (w = 0,1, ... M-1), the suppression gain G [f, w] (w = 0,1, ... M-1) of each frequency band is output. For example, the suppression gain G [f, w] is calculated by the following algorithm or a combination thereof. That is, the Spectral Subtraction method (SF Boll, "Suppression of acoustic noise in speech using spectral subtraction", IEEE Trans. Acoustics, Speech, and Signal Processing, vol.ASSP-29, pp. 113-120 (1979).), Wiener Filter Method (JS Lim, AV Oppenheim, "Enhancement and bandwidth compression of noisy speech", Proc. IEEE Vol.67, No.12, pp.1586-1604, Dec.1979.) And Maximum Likelihood Method (RJ McAulay, ML Malpass, "Speech enhancement using a soft-decision noise suppression filter", IEEE Trans. On Acoustics, Speech, and Signal Processing, vol.ASSP-28, no.2, pp.137-145, Apr.1980.). Here, it is assumed that the suppression gain G [f, w] is calculated by using the Wiener filter method as an example.
The spectrum suppression unit 372 includes the frequency spectrum Z [f, w] of the sound collection signal z [n] output from the frequency domain conversion unit 311 and the suppression gain G [f, w] output from the suppression gain calculation unit 371. With the input of and, the frequency spectrum Z [f, w] of the sound collecting signal z [n] is the amplitude spectrum of the sound collecting signal z [n] | Z [f, w] | (w = 0,1,... M- 1) and phase spectrum θ<sub>Z</sub>Divide into [f, w] (w = 0,1,... M-1) and multiply the amplitude spectrum | Z [f, w] | of the sound collection signal z [n] by the suppression gain G [f, w]. Suppresses the noise component, which is ambient noise, and sets the amplitude spectrum of the suppressed signal | S [f-1, w] | to the phase spectrum.<sub>θZ</sub>Phase spectrum θ of the signal obtained by suppressing [f, w] as it is<sub>S</sub>As [f, w], the frequency spectrum S [f, w] (w = 0,1, ... M-1) of the suppressed signal is calculated.
The power calculation unit 373 uses the power spectrum of the suppressed signal output from the spectral suppression unit 372 from the frequency spectrum S [f, w] (w = 0,1, ... M-1) of the suppressed signal | S [f, w] |<sup>2</sup> Calculate and output (w = 0,1, ... M-1).
The time domain transform unit 374 inputs the frequency spectrum S [f, w] (w = 0,1, ... M-1) of the suppressed signal output from the spectrum suppression unit 372, and sets the frequency domain as the time domain. (For example, IFFT) is performed, and the suppression processing is performed by appropriately adding the suppression processing signal s [n] one frame before in consideration of the overlap due to window hanging in the frequency domain conversion unit 311. Calculate the signal s [n] (n = 0,1,... N-1) in the time domain.
As described above, by using the ambient noise suppression processing together with the ambient noise estimation processing, the input signal is clarified while suppressing the increase in the processing amount, and at the same time, the ambient noise component in the sound collection signal is suppressed to collect high-quality sound. A sound signal can be obtained.
The present invention is not limited to the above-described embodiment as it is, and at the implementation stage, the components can be modified and embodied within a range that does not deviate from the gist thereof. Further, various inventions can be formed by appropriately combining a plurality of components disclosed in the above-described embodiment. Further, for example, a configuration in which some components are deleted from all the components shown in the embodiment is also conceivable. Furthermore, the components described in different embodiments may be combined as appropriate.
For example, the sampling frequency of the input signal (or target signal) is not limited to twice the sampling frequency of the sound collection signal (or ambient noise), and may be an integral multiple or a non-integer multiple. Further, the sampling frequency of the input signal (or target signal) is equal to the sampling frequency of the sound collecting signal (or ambient noise), and the range of the frequency band limitation of the input signal (or target signal) and the sound collecting signal (or surroundings). It does not matter even if the range of the frequency band limitation of noise) is different. Further, the range of the frequency band limitation of the input signal (or the target signal) does not have to include the range of the frequency band limitation of the sound collecting signal (or ambient noise). Furthermore, the frequency band limiting range of the input signal (or target signal) does not have to be adjacent to the frequency band limiting range of the sound collecting signal (or ambient noise).
Further, even if the input signal is a stereo signal instead of a monaural signal, for example, the L (left) channel and the R (right) channel may be subjected to signal processing by the signal processing unit 3, respectively, or a sum signal (L channel and R) may be applied. The same effect can be obtained by performing the above signal processing on each of the channel signal sum) and the difference signal (difference between the L channel signal and the R channel signal). Of course, even if it is a multi-channel signal, the same effect can be obtained by performing the above signal processing on each channel signal in the same manner, for example.
In addition, it goes without saying that various modifications can be made in the same manner without departing from the gist of the present invention.
1 ... wireless communication unit, 2,2A ... decoder, 3,30,300,3A ... signal processing unit, 4 ... digital analog (D / A) converter, 5 ... speaker, 6 ... microphone, 7 ... analog digital (A / D) Converter, 8 ... Downsampling unit, 9 ... Echo suppression processing unit, 10 ... Encoder, 11 ... Storage unit, 12 ... Signal band expansion processing unit, 31 ... Ambient noise estimation unit, 32, 34, 36 ... Ambient noise information band expansion unit, 33, 35 ... Signal characteristic correction unit, 37 ... Ambient noise suppression processing unit, 311, 331 ... Frequency region conversion unit, 312, 352, 373 ... Power calculation unit, 313 ... Ambient noise section determination Unit, 314 ... Frequency spectrum update unit, 321 ... Power normalization unit, 322, 342, 364 ... Dictionary storage unit, 323 ... Wideband power calculation unit, 332,356 ... Correction degree determination unit, 333 ... Correction processing unit, 334 374 ... Time region conversion unit, 343 ... Wideband power spectrum calculation unit, 344,365 ... Wideband masking threshold calculation unit, 345 ... Power control unit, 353 ... Masking threshold calculation unit, 354 ... Masking determination unit, 355 ... Power smoothing unit , 362 ... Narrow band masking threshold calculation unit, 363 ... Band control unit, 366 ... Threshold correction unit, 371 ... Suppression gain calculation unit, 372 ... Spectrum suppression unit.
42 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41 Sheet 42
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| JP2016028229A | Cited by | Japan | Search report |
| JP2015118307A | Cited by | Japan | Search report |
| JP2019219419A | Cited by | Japan | Search report |
| US10127910B2 | Cited by | United States of America | Applicant |
| JP2015118307A | Cited by | Japan | Search report |
| CN111402917A | Cited by | China | Search report |
| WO2015093013A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| JP2000181496A | Cites | Japan | Examiner |
| JP2000181497A | Cites | Japan | Examiner |
| JP2000200099A | Cites | Japan | Examiner |
| JP2000305599A | Cites | Japan | Examiner |
| JP2001188599A | Cites | Japan | Examiner |
| JP2004289614A | Cites | Japan | Examiner |
| JP2007067858A | Cites | Japan | Examiner |
| JP2007171954A | Cites | Japan | Examiner |
| JP2007518291A | Cites | Japan | Examiner |
| JP2009069856A | Cites | Japan | Examiner |
| JPH11298990A | Cites | Japan | Examiner |
2 members in 1 office
Members2
| Document | Office | Kind | |
|---|---|---|---|
| JP2012181561AThis record | Japan | A | |
| JP5443547B2 | Japan | B2 |
11 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Transfer to examiner for re-examination before appeal (zenchi)AppealJAPANESE INTERMEDIATE CODE: A911A911 | A911 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Decision of refusalJAPANESE INTERMEDIATE CODE: A02A02 | A02 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 2012181561
- Application
- 144135
Titles2
- Japanese
- 信号処理装置
- English
- Signal processor
Classification
- IPC, 5
- G10L21 02
- G10L21 0364
- G10L21 0388
- G10L21 057
- G10L21 04