Sound processing device and method, and imaging apparatus
Abstract
Problem to be solved.To suppress a reverberation sound caused by an acoustic resistive element while reducing a wind noise by using the acoustic resistive element, and to provide a high quality sound.
Solution.A sound processing device has first and second microphones, and an acoustic resistive element is provide to the second microphone so as to cover the microphone. A high frequency component of an output signal of the first microphone is obtained by a high pass filter, and a low-frequency component of an output signal of the second microphone is obtained by a low pass filter. The output signal from the high pass filter and the output signal from the low pass filter are added and outputted. Here, an adaptive filter is provided between the second microphone and the low pass filter. By performing estimation learning of a filter coefficient so that a difference between the output signal from the first microphone and the output signal from the second microphone becomes the minimum, a reverberation component in the output signal from the second microphone, which is generated in a closed space between the acoustic resistive element and the second microphone, is suppressed.

Term
No projected expiry on record.
- Priority and filed
- Published
- Today
11 claims: 2 independent, 9 dependent
- 1The first and second microphones, an acoustic resistor provided to cover the second microphone in order to block the movement of air from the outside of the device to the second microphone, and the first microphone. A high-pass filter that passes only the high-frequency component of the output signal of the second microphone, a low-pass filter that passes only the low-frequency component of the output signal of the second microphone, and an output signal of the high-pass filter and the low-pass. An adder that adds and outputs the output signal of the pass filter is provided between the second microphone and the low frequency pass filter, and the output signal of the first microphone and the output of the second microphone are provided. By estimating and learning the filter coefficient so that the difference from the signal is minimized, the reverberation generated in the closed space between the acoustic resistor and the second microphone in the output signal of the second microphone A voice processing device characterized by having an adaptive filter that suppresses components. 第1及び第2のマイクロホンと、 装置外部から前記第2のマイクロホンへの空気の移動を遮断するために、前記第2のマイクロホンを覆うように設けられた音響抵抗体と、 前記第1のマイクロホンの出力信号の高周波成分のみを通過させる高域通過フィルタと、 前記第2のマイクロホンの出力信号の低周波成分のみを通過させる低域通過フィルタと、 前記高域通過フィルタの出力信号と前記低域通過フィルタの出力信号とを加算して出力する加算器と、 前記第2のマイクロホンと前記低域通過フィルタとの間に設けられ、前記第1のマイクロホンの出力信号と前記第2のマイクロホンの出力信号との差が最小になるようフィルタ係数を推定学習することで、前記第2のマイクロホンの出力信号のうちの、前記音響抵抗体と前記第2のマイクロホンとの間の閉空間において発生する残響成分を抑圧する適応フィルタと、 を有することを特徴とする音声処理装置。
- 10In a sound processing device including the first and second microphones and an acoustic resistor provided so as to cover the second microphone in order to block the movement of air from the outside of the device to the second microphone. In the voice processing method, only the high frequency component of the output signal of the first microphone is passed by the high frequency pass filter, and only the low frequency component of the output signal of the second microphone is passed by the low frequency pass filter. The step of adding and outputting the output signal of the high-pass filter and the output signal of the low-pass filter, and the output signal of the first microphone and the second By estimating and learning the filter coefficient so that the difference from the output signal of the microphone is minimized, in the closed space between the acoustic resistor and the second microphone in the output signal of the second microphone. A sound processing method characterized by having a step of suppressing a reverberant component generated. 第1及び第2のマイクロホンと、装置外部から前記第2のマイクロホンへの空気の移動を遮断するために、前記第2のマイクロホンを覆うように設けられた音響抵抗体とを備える音声処理装置における音声処理方法であって、 高域通過フィルタにより、前記第1のマイクロホンの出力信号の高周波成分のみを通過させるステップと、 低域通過フィルタにより、前記第2のマイクロホンの出力信号の低周波成分のみを通過させるステップと、 前記高域通過フィルタの出力信号と前記低域通過フィルタの出力信号とを加算して出力するステップと、 適応フィルタにより、前記第1のマイクロホンの出力信号と前記第2のマイクロホンの出力信号との差が最小になるようフィルタ係数を推定学習することで、前記第2のマイクロホンの出力信号のうちの、前記音響抵抗体と前記第2のマイクロホンとの間の閉空間において発生する残響成分を抑圧するステップと、 を有することを特徴とする音声処理方法。
Independent claims2
119 paragraphs, as filed
The present invention relates to a voice processing technique for reducing wind noise mixed in during recording.
The voice processing device is desired to faithfully record voice under various environments. In outdoor photography, the generation of wind noise (hereinafter referred to as "wind noise") is particularly remarkable. Many mechanical / electrical treatments have been proposed to suppress wind noise. For example, Patent Document 1 discloses a method of suppressing wind noise by attaching a wind noise reducing body (hereinafter referred to as "acoustic resistor") to a sound collecting portion of a housing of an imaging device with an adhesive tape.
<p><patcit num="1"><text>Japanese Unexamined Patent Publication No. 2006-211302</text></patcit></p>
<p> However, in the prior art disclosed in the above-mentioned patent document, it is considered that reverberation is generated inside the sound collecting portion depending on the material of the acoustic resistor, and the sound quality is deteriorated. Therefore, an object of the present invention is to use an acoustic resistor to reduce wind noise and suppress reverberation sound generated by the acoustic resistor to provide high-quality sound.</p>
<p> According to one aspect of the present invention, the first and second microphones and an acoustic signal provided to cover the second microphone in order to block the movement of air from the outside of the device to the second microphone. A resistor, a high-pass filter that passes only the high-frequency component of the output signal of the first microphone, a low-pass filter that passes only the low-frequency component of the output signal of the second microphone, and the high-pass filter. An adder that adds and outputs the output signal of the pass filter and the output signal of the low frequency pass filter, and the output of the first microphone provided between the second microphone and the low frequency pass filter. By estimating and learning the filter coefficient so that the difference between the signal and the output signal of the second microphone is minimized, the acoustic resistor and the second microphone among the output signals of the second microphone can be used. Provided is a voice processing device characterized by having an adaptive filter that suppresses a reverberation component generated in a closed space between them.</p>
<p> According to the present invention, it is possible to provide a recording device in which wind noise is reduced and reverberation is suppressed by an acoustic resistor.</p>
<figref num="1">The figure which shows the structure of the recording apparatus in embodiment.</figref><figref num="2">A perspective view and a cross-sectional view of the image pickup apparatus.</figref><figref num="3">The figure which shows the example of the frequency characteristic of a microphone.</figref><figref num="4">The figure explaining the mounting structure of a microphone.</figref><figref num="5">The figure which shows the structure of the reverberation suppressor.</figref><figref num="6">The figure which shows the operation of the wind detector according to the wind noise.</figref><figref num="7">The figure which shows the structure and operation of a synthesizer.</figref><figref num="8">The figure which shows the example which applied the prior art.</figref><figref num="9">The figure which shows the operation sequence of a switch, a variable filter, and a variable gain.</figref><figref num="10">The figure explaining the wind noise processing in the absence of HPF.</figref><figref num="11">The figure explaining the wind noise processing when there is HPF.</figref><figref num="12">The figure which shows the example of another voice processing apparatus.</figref><figref num="13">The perspective view of the image pickup apparatus in the 2nd Example.</figref><figref num="14">The figure which shows the structure of the voice processing apparatus in 2nd Example.</figref><figref num="15">The figure which shows the structure of the voice processing apparatus in 3rd Example.</figref><figref num="16">The figure which shows the structure of the voice processing apparatus in 4th Example.</figref><figref num="17">The figure explaining the positional relationship between the subject sound and the microphone in the 4th Example.</figref>
Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. Throughout the drawings, the same components are given the same reference numbers.
(Example 1) Hereinafter, an image pickup apparatus including a recording apparatus and a recording apparatus according to the first embodiment of the present invention will be described with reference to FIGS. 1 to 11.
FIG. 1 is a block diagram showing a configuration of a recording device in this embodiment. 2 (a) and 2 (b) are perspective views and cross-sectional views of an image pickup device (camera) provided with the recording device of FIG. 1, respectively. 1 is an image pickup device, 2 is a lens mounted on the image pickup device 1, 3 is a housing of the image pickup device 1, 4 is an optical axis of the lens, 5 is a photographing optical system, and 6 is an image sensor. In addition, 30 is a release button and 31 is an operation button. The image pickup apparatus 1 is provided with a first microphone 7a and a second microphone 7b. 32a and 32b are openings provided in the housing 3 for the microphones 7a and 7b, respectively. An acoustic resistor 41 is attached to the opening 32b. As will be described later, the acoustic resistor 41 can have an uneven thickness structure in the housing 3 or can be configured by a separate component. The image pickup apparatus 1 can record sound at the same time as acquiring an image by using microphones 7a and 7b.
The operation of shooting a moving image by the image pickup device 1 will be described. By pressing the live view button (not shown) prior to shooting a moving image, the image of the image sensor 6 is displayed in real time on the display device provided in the image sensor 1. The image pickup device 1 obtains subject information from the image sensor 6 at a set frame rate in synchronization with the operation of the moving image shooting button, and obtains audio information from the microphones 7a and 7b, and synchronizes these to obtain (not shown). Record in memory. Shooting ends in synchronization with the operation of the movie shooting button.
The configuration of the voice processing device 51 will be described with reference to FIG. 52 is a variable high-pass filter (HPF). Reference numeral 53 denotes a reverberation suppressor, in which, for example, a reverberation suppressor adaptive filter is used. 54a and 54b are the first analog-to-digital converter (ADC) that digitizes the output signal of the microphone, 55 is the first delay device (DL), and 56a and 56b are HPFs for cutting the DC component. 61 is an automatic level correction unit (ALC). In ALC61, 62a and 62b are variable gains for level adjustment, and 63 is a level controller. The 71 is a synthesizer that synthesizes the signal of the first microphone 7a and the signal of the second microphone 7b. In the synthesizer 71, 72 is a low-pass filter (LPF), 73 is a variable HPF, 74 is a variable gain, and 75 is an adder. 81 is a wind-detector. In the wind detector 81, 82a and 82b are bandpass filters (BPF), 83 is a diff, 84 is a second analog-to-digital converter (ADC), 85 is a second delayer, and 86 is a level detector. .. 87 is a switch that controls the reverberation suppressor 53, 88 is a switch that controls the synthesizer 71, and 89 is a mode switching operation unit.
In FIGS. 1 and 2, the housing 3 is provided with openings 32a and 32b for microphones. Here, the opening 32b is provided with an acoustic resistor 41 that covers the second microphone 7b so as to block the movement of air from the outside of the device to the second microphone 7b. On the other hand, the opening 32a is not provided with such an acoustic resistor so that the first microphone 7a can faithfully acquire the subject sound. The acoustic resistor 41 is provided in close contact with the housing 3. The movement of air here is assumed to be the movement of air due to the wind. For example, it is possible to use a material such as porous PTFE, which allows air to move in a slower time than the movement of air by wind and does not allow wind to pass through, as an acoustic resistor.
The voice processing device 51 processes the signal from the first microphone 7a with the HPF52, and then performs analog / digital conversion (A / D conversion) with the ADC 54a. In addition, the output of ADC 54a is delayed by an appropriate amount by the first delayer 55. On the other hand, the voice processing device 51 performs A / D conversion of the signal from the second microphone 7b by the ADC 54b, and then suppresses the reverberation by the reverberation suppressor 53. The operation of the reverberation suppressor 53 and the method of giving a delay in the first delay device 55 will be described later.
The outputs of the first delayer 55 and ADC 54b are processed by HPF56a and 56b for cutting DC components, respectively. Since HPF56a and 56b are intended to remove the offset of the analog part, it is preferable that the components below the audible range can be removed from the DC. Therefore, the cutoff frequency of HPF56a and 56b is set to, for example, about 10Hz.
The outputs of HPF56a and 56b are input to ALC61, and the gain is adjusted by the variable gains 62a and 62b, respectively. At this time, the gains of the variable gains 62a and 62b are controlled in conjunction with each other so that the two signal levels are the same. The level adjuster 63 obtains outputs of variable gains 62a and 62b, and appropriately adjusts the level so that saturation does not occur and the dynamic range can be effectively utilized. At this time, the level adjuster 63 adjusts the level so that the larger of the outputs of the variable gains 62a and 62b is not saturated.
The outputs of the variable gains 62a and 62b are input to the synthesizer 71. The output of the variable gain 62a is sent to the adder 75 after passing through the HPF 73. On the other hand, the output of the variable gain 62b is sent to the adder 75 via the LPF 72 and the variable gain 74. The output synthesized by the adder 75 is output as the sound after wind noise processing.
The output of the first microphone 7a and the output of the reverberation suppressor 53 are input to the BPF 82a and 82b of the wind detector 81, respectively. The purpose of BPF82a and 82b is to pass the subject sound through the range in which the subject sound can be faithfully acquired by the second microphone 7b. Therefore, the pass band is set to, for example, about 30 Hz to 1 kHz. However, the upper limit frequency can be changed depending on the structure of the acoustic resistor 41 and the like. Details will be described later together with the frequency characteristics of the second microphone 7b.
The output of BPF82a is A / D converted by the second ADC 84 and then sent to the second delay device 85. How to give the delay in the second delay device 85 will be described later together with the operation of the reverberation suppressor 53.
The differencer 83 calculates the difference between the output of the second delayer 85 and the output of the BPF 82b, and sends the result to the level detector 86. The operation of the level detector 86 will be described later. The level detector 86 determines the strength of the wind and controls the switch 87 to switch the feedback to the reverberation suppressor 53. The detection result of the level detector 86 is also used to control the switch 88 that controls the synthesizer 71. When the mode switching operation unit 89 is set to OFF by the user, the switch 88 operates so as to always select the process when there is no wind, which will be described later. On the other hand, when the mode switching operation unit 89 is set to Auto by the user, the switch 88 switches the cutoff frequency and variable gain of the HPF52 and HPF73 according to the wind strength determined by the level detector 86. Acts to change 74. The details of this process will be described later.
The effects, desirable characteristics, and reduction of wind noise of the acoustic resistor 41 will be described with reference to FIGS. 1, 3, and 4. FIG. 3 is a diagram schematically showing the frequency characteristics of the microphone, in which the horizontal axis shows the frequency and the vertical axis shows the gain. In FIG. 3, (a) shows the subject sound acquisition characteristic of the first microphone 7a, and (b) shows the subject sound acquisition characteristic of the second microphone 7b. (c) shows the wind noise acquisition characteristic of the first microphone 7a, and (d) shows the wind noise acquisition characteristic of the second microphone 7b. (e) shows the subject sound acquisition characteristic of the output of the synthesizer 71, and (f) shows the wind noise acquisition characteristic of the output of the synthesizer 71. In addition, in order to clarify the difference in characteristics between the first microphone 7a and the second microphone 7b, the characteristics of the first microphone 7a are shown by broken lines in (b) and (d). In FIG. 3, f0 shows the structural cutoff frequency due to the acoustic resistor 41, and f1 shows the cutoff frequencies of LPF72 and HPF73 in the synthesizer 71 shown in FIG.
As shown in FIG. 3A, it is desirable that the subject sound acquisition characteristic of the first microphone 7a is flat in the audible range. This makes it possible to faithfully acquire the subject sound. As shown in FIG. 3 (b), the second microphone 7b has different characteristics because the acoustic resistor 41 is provided so as to block the movement of air from the subject. At frequencies lower than the cutoff frequency of the acoustic resistor 41, the audio signal is passed relatively faithfully. This is because the acoustic resistor 41 is vibrated by the sound which is a dense wave of air, and the acoustic resistor 41 vibrates the air inside the device in the same manner. On the other hand, at a frequency higher than the cutoff frequency of the acoustic resistor 41, the audio signal is cut off. This is a state in which the acoustic resistor 41 is vibrated by the sound which is a sparse and dense wave of air, but cannot move because the sparse and dense reverses faster than the acoustic resistor 41 vibrates. Thus, the acoustic resistor 41 acts as a structural LPF. The frequency f0 at which structural cut begins is called the cutoff frequency of the acoustic resistor 41.
It is known that the power of wind noise is concentrated in the low frequency range. For example, as shown in Fig. 3 (c), the power of wind noise in the first microphone 7a often has the characteristic of rising from about 1 kHz toward low frequencies. Even if the shape is not as shown in Fig. 3 (c), the wind noise is dominated by low-frequency (500 Hz or less) components. As shown in Fig. 3 (d), the second microphone 7b has less lift of low-frequency components due to wind noise. A large pressure difference is likely to occur in the vicinity of the first microphone 7a due to turbulent flow. On the other hand, since the second microphone 7b is provided with the acoustic resistor 41 so as to block the movement of air from the subject, a large pressure difference due to turbulence or the like does not occur. This is the reason why the output of the second microphone 7b has less lift of low frequency components due to wind noise.
Consider processing these signals with synthesizer 71. As described with reference to FIG. 1, the signal of the first microphone 7a is processed by HPF73. This corresponds to cutting out the portion shown in 91 in FIG. 3 (a) and the portion shown in 93 in FIG. 3 (c). The signal of the second microphone 7b is processed by LPF72. This corresponds to cutting out the part shown by 92 in FIG. 3 (b) and the part shown in 94 in FIG. 3 (d). After passing through the adder 75, the subject sound characteristics are as shown in FIG. 3 (e), and the wind noise characteristics are as shown in FIG. 3 (f). The parts shown by 91a, 92a, 93a, and 94a in FIGS. 3 (e) and 3 (f) are the parts shown by 91, 92, 93, and 94, respectively, which are dominant. The reason for saying "dominant" is that the other is not necessarily zero due to the characteristics of LPF72 and HPF73. As is clear from FIGS. 3 (e) and 3 (f), the subject sound characteristic of the output of the synthesizer 71 is flat in the audible range, and the wind noise characteristic is the characteristic of the microphone provided with the acoustic resistor 41. ing.
Figure 4 shows an example of the microphone mounting structure. In FIG. 4, 33a and 33b are holding elastic bodies of the first microphone 7a and the second microphone 7b, respectively. Reference numeral 34 denotes a sleeve for holding the second microphone 7b and the acoustic resistor 41.
FIG. 4A shows an example in which the acoustic resistor 41 is attached to the outside of the housing 3. In the example of FIG. 4A, since the acoustic resistor 41 may be attached after assembling the device, the assembling property can be improved.
FIG. 4B shows an example in which the acoustic resistor 41 is attached to the inside of the housing 3. In the example of FIG. 4B, the acoustic resistor 41 is not exposed to the outside of the housing 3, which is excellent in terms of aesthetics.
FIG. 4C shows an example in which a part of the housing 3 also functions as the acoustic resistor 41. In the example of FIG. 4 (c), a part of the housing 3 serving as the acoustic resistor 41 is thinned so as to be vibrated by sound waves. In the example of FIG. 4 (c), the number of parts is reduced, and the acoustic resistor 41 does not need to be attached to the housing 3, which is excellent in terms of aesthetics. However, in the example of FIG. 4 (c), since the housing 3 and the acoustic resistor 41 are integrated, the degree of freedom in design is generally reduced. (The strength of the housing 3 may be limited by the thickness of the portion forming the acoustic resistor 41, which makes it difficult to achieve both.)
FIG. 4D shows an example in which the second microphone 7b and the acoustic resistor 41 are held by a sleeve 34 having sufficiently high rigidity. It is desirable that the sleeve 34 has a first-order resonance frequency at a frequency sufficiently higher than the frequency band desired to be acquired by the second microphone 7b (meaning that the resonance frequency of the sleeve 34 is higher than f0 in FIG. 3). In the example of FIG. 4 (d), the acoustic resistor 41 is attached to the sleeve 34 with high rigidity, so that it is not affected by the unnecessary resonance of the mounting structure and is in the pass band (at a frequency lower than f0 in FIG. 3). The desired audio signal can be obtained.
Next, the reverberation suppressor 53 will be described with reference to FIGS. 1 and 5. Since the second microphone 7b has a structure covered by the acoustic resistor 41, reverberation may occur in the closed space. In this embodiment, a reverberation suppressor 53 is provided to suppress such reverberation.
The specific configuration of the reverberation suppressor 53 is shown in Fig. 5. The reverberation suppressor 53 is composed of an adaptive filter. This adaptive filter is the difference between the output of the diffifier 83, which represents the magnitude of wind noise, that is, the output signal of the first microphone 7a and the output signal of the second microphone 7b, as will be specifically described below. Estimate and learn the filter coefficient so that As a result, the reverberation component generated in the closed space between the acoustic resistor 41 and the second microphone 7b in the output signal of the second microphone 7b is suppressed. By using such an adaptive filter, it is possible to appropriately process the change in the gripping state of the camera by the user and the change in the reverberation generation state due to the temperature change.
The principle of reverberation suppression will be briefly explained. Let s be the subject sound, g1 be the subject sound acquisition characteristic of the first microphone 7a, g2 be the subject sound acquisition characteristic of the second microphone 7b, and r be the effect of reverberation. g1 and g2 are equivalent to the inverse Fourier transform of the characteristics in the frequency space shown in FIG. The signal x1 of the first microphone 7a and the signal x2 of the second microphone 7b obtained in an environment where the second microphone 7b has reverberation are given as in Eq. (1).
<maths num="1"><img file="JP2012129652A_D0001.tif" /></maths>
However, in Eq. (1), * is an operator indicating convolution. As described in FIG. 3, at frequencies lower than f0, similar subject sounds can be obtained with the first microphone 7a and the second microphone 7b. Furthermore, as shown in FIG. 1, only the components in the appropriate band are extracted by BPF82a and 82b. That is, the band passed by the BPF is the audible range, which is a frequency lower than f0 in FIG. Due to the characteristics of human hearing, the sensitivity is extremely low in the band below 50Hz. For details, refer to the A characteristic curve. Therefore, the BPF82a and 82b may be designed to pass, for example, 30 Hz to 1 kHz. If BPF82a and 82b are BPF and the signals after passing through BPF are x1_BPF and x2_BPF, the following equation holds.
<maths num="2"><img file="JP2012129652A_D0002.tif" /></maths>
g1 g2 and g1 * BPF = g2 * BPF are equivalent to being able to obtain similar subject sounds with the first microphone 7a and the second microphone 7b at frequencies lower than f0. As is clear from Eq. (2), the inputs of the differencer 83 in FIG. 1 are equal if there is no reverberation effect r. The effect of reverberation can be reduced by operating the adaptive filter with x1_BPF = d as the desired response and x2_BPF = u as the input from equation (2).
When the filter of the reverberation suppressor 53 is expressed by h, the adaptive filter output y is given by the following equation.
<maths num="3"><img file="JP2012129652A_D0003.tif" /></maths>
However, in Eq. (3), n indicates that it is the signal of the nth sample, M is the filter order of the reverberation suppressor 53, and the subscript of h is the value of the filter h of the nth sample. Shown. The input u may use x2_BPF.
Furthermore, since the desired response can use x1_BPF for d, the error signal e is expressed as follows.
<maths num="4"><img file="JP2012129652A_D0004.tif" /></maths>
Various adaptive algorithms have been proposed, but here, as an example, the update formula for h in the LMS algorithm is shown below.
<maths num="5"><img file="JP2012129652A_D0005.tif" /></maths>
However, in Eq. (5), μ is a step size parameter. According to the above, after giving an appropriate initial h, u approaches d by updating h using Eq. (5). That is, the influence of r is reduced and becomes close to x1_BPF = x2_BPF. At this time, | h * r | = 1 holds in the pass band of BPF. However, in an environment where wind noise is dominant, the update of Eq. (5) is not performed correctly, so the switch 87 stops the estimation learning of the adaptive filter. The control sequence of the switch 87 will be described later together with the operation of the wind detector 81.
As described above, the reverberation suppressor 53 suppresses the reverberation. On the other hand, as is clear from FIG. 5, in the reverberation suppressor 53, the signal is delayed according to the order of the adaptive filter. To compensate for these, in FIG. 1, a first delay device 55 and a second delay device 85 are provided. Typically, a delay of half the filter order (= M / 2) of the reverberation suppressor 53 may be given (if M is an odd number, a nearby value may be used). At this time, for example, by setting h (M / 2) = 1 and initializing all other h to 0, the adaptive algorithm can be operated with the state without reverberation as the initial value. When an appropriate initial value for suppressing reverberation is stored in the memory, h may be initialized with that value before starting the operation. For example, it is conceivable to set the initial value as follows. The filter coefficient can be estimated to some extent based on the design values such as the dimensions around the microphones 7a and 7b and the material of the structural member. Therefore, the filter coefficient obtained from the design value may be set as the initial value. Further, the filter coefficient when the power of the recording device is turned off may be stored in the memory and set as the initial value at the next startup of the recording device. Further, in the production process of the recording device, the filter coefficient may be calculated by generating a predetermined reference sound and stored in the memory, and this may be set as the initial value at the time of starting the recording device.
Next, the operation of ALC61 will be described. ALC is provided to effectively utilize the dynamic range while suppressing the saturation of the audio signal. Since the power of the audio signal fluctuates greatly with respect to the time axis, it is necessary to adjust the level appropriately. The level adjuster 63 provided on the ALC61 monitors the output from the variable gains 62a and 62b.
First, the attack operation will be described. When it is determined that the signal with the higher level exceeds a predetermined level, the gain is lowered by a predetermined step. This operation is repeated at a predetermined cycle. This operation is called an attack operation. It is possible to prevent saturation by the attack operation.
Next, the recovery operation will be described. When the signal with the higher level does not exceed the predetermined level for a predetermined time, the gain is increased by a predetermined step. This operation is repeated at a predetermined cycle. This operation is called a recovery operation. The recovery operation makes it possible to obtain sound in a quiet environment.
The variable gains 62a and 62b in the ALC61 are operating in tandem. That is, when the variable gain 62a is attacked and the gain is lowered, the gain of the variable gain 62b is also lowered by the same amount. By performing such an operation, the level difference between the signal channels is eliminated, and the discomfort is reduced when the signals between the channels are mixed by the synthesizer 71 later.
Next, the wind detector 81 will be described. Let w1 be the wind noise picked up by the first microphone 7a, and w2 be the wind noise picked up by the second microphone 7b. As explained in Fig. 3, the power of wind noise is concentrated in the low frequency range, so it is not blocked by BPF82a and 82b. Therefore, w1-w2 is obtained as the output of the differencer 83. It is assumed that the effect of reverberation described above can be ignored. Even in the actual environment, the effect of reverberation is sufficiently small compared to wind noise and is negligible.
The level detector 86 performs LPF processing appropriately after calculating the absolute value of the output of the diffifier 83. The cutoff frequency of the LPF may be determined by the stability of the wind detector and the detection speed, but it may be about 0.5 Hz. Since the LPF integrates the signals in the cutoff band and passes the signals in the passing band as they are, the result is the same effect as the integration operation + HPF. Therefore, if the absolute value calculation maintains a high level for a certain period of time (which changes depending on the cutoff frequency described above), a large output is obtained. In other words, it is equivalent to monitoring Σ | w1-w2 | for an appropriate period of time.
Figure 6 shows an example of the output signal of the wind detector 81 due to the difference in wind strength. 6 (a), (b), and (c) are diagrams showing the signals obtained by the first microphone 7a and the second microphone 7b, where the horizontal axis shows the time and the vertical axis shows the signal level. .. In FIGS. 6A, 6B, and 6C, the signal level +1 indicates the level at which the positive signal is saturated. In addition, FIG. 6 (a) shows signals in a state of no wind, FIG. 6 (b) shows signals in a state of weak wind, and FIG. 6 (c) shows signals in a state of strong wind. It can be seen that the signal level of the first microphone 7a increases according to the strength of the wind, and that wind noise is generated. On the other hand, it can be seen that the signal level of the second microphone 7b is not so high as that of the signal level of the first microphone 7a. It is shown that the wind noise is reduced by the effect of the acoustic resistor 41.
At this time, the result of performing the above-mentioned treatment of the wind detector 81 is shown in FIG. 6 (d). The horizontal axis of FIG. 6 (d) shows the same time as in FIGS. 6 (a), (b), and (c), and the vertical axis shows the output of the wind detector. The BPF82a and 82b have a pass band of 30 Hz to 1 kHz, and the cutoff frequency of the LPF in the level detector 86 is 0.5 Hz. It can be seen that the output of the wind detector 81 changes around zero when there is no wind, and the value increases according to the strength of the wind. Also, in Fig. 6 (d), the signal near 0 seconds is small because the rise is delayed due to the influence of the LPF in the level detector 86. There is a delay of the degree shown in the rising edge of the signal in Fig. 6 (d) before the wind is detected. Since there is a problem that if the delay is reduced, it is easily affected by the fluctuation of the wind, in this embodiment, the wind is detected with a delay as shown in FIG.
The output of the wind detector 81 is used not only for the switch 87 of the reverberation suppressor 53 described above, but also for the switching of the HPF 52 described later and the switching of the synthesis process in the synthesizer 71.
Next, the operation of the synthesizer 71 will be described with reference to FIGS. 1 and 7. In FIG. 1, it has been described that the cutoff frequency and the variable gain 74 of the HPF73 are changed based on the output of the wind detector 81, but a specific change method will be described with reference to FIG.
FIGS. 7 (a) and 7 (c) show configuration examples of the synthesizer 71, respectively. 7 (b) and 7 (d) are diagrams showing a method of changing the variable portion of FIGS. 7 (a) and 7 (c), respectively.
First, the configuration of FIG. 7A will be described. The synthesizer 71 shown in FIG. 7 (a) has the same configuration as that shown in FIG. In FIG. 7 (a), the cutoff frequency of LPF72 is fixed, for example, 1 kHz. In FIG. 7 (b), the upper row schematically shows the gain of the variable gain 74, and the lower row schematically shows the cutoff frequency of the HPF 73. The horizontal axis of Fig. 7 (b) is common to the two graphs, and Wn1, Wn2, and Wn3 are values indicating the magnitude of wind noise, indicating that the wind noise is stronger in this order.
As shown in Fig. 7 (b), if the wind noise is smaller than the predetermined value Wn1, it is assumed that wind processing is not required, and the gain of the variable gain 74 is set to 0 and the cutoff frequency of the HPF73 is set to 50 Hz. As a result, by passing through the circuit shown in Fig. 7 (a), the signal from the second microphone 7b is completely cut off, and the frequency higher than 50Hz, which is the cutoff frequency of HPF73, is dominant in the sound. The signal of) can be obtained only from the first microphone 7a. Since it is not necessary to use the signal of the second microphone 7b provided with the acoustic resistor 41, it is considered that the sound of the subject is faithfully obtained.
The time when the wind noise exceeds the level of Wn1 and is between Wn1 and Wn2 will be described. At this time, the value of the variable gain 74 gradually increases, and the cutoff frequency of the HPF73 gradually rises. By performing the above-mentioned control, the ratio of the signal from the second microphone 7b provided with the acoustic resistor 41 is gradually increased in the low frequency audio signal. Wind noise has a large effect on the signal from the first microphone 7a, but the wind noise is reduced by increasing the cutoff frequency of HPF73.
The time when the wind noise exceeds the level of Wn2 and is between Wn2 and Wn3 will be described. At this time, the value of the variable gain 74 is fixed at 1, and the cutoff frequency of the HPF73 gradually rises. By performing the above-mentioned control, the sound existing between the cutoff frequency of LPF72 and the cutoff frequency of HPF73 is lost, but the wind noise can be further reduced. If the cutoff frequency of HPF73 is raised excessively, the deterioration of the subject sound will become too great, so I try not to raise it above the appropriate cutoff frequency. In the example of Fig. 6 (b), when the magnitude of wind noise exceeds Wn3, the cutoff frequency of HPF73 is fixed at 2kHz and does not change any more.
The configuration of FIG. 7 (c), which is another example, will be described. The synthesizer 71 shown in FIG. 7 (c) is provided with a variable LPF76 instead of the fixed LPF72 and the variable gain 74. In FIG. 7 (d), the upper row schematically shows the cutoff frequency of the variable LPF76, and the lower row schematically shows the cutoff frequency of the HPF73. The horizontal axis of FIG. 7 (d) is common to the two graphs, and Wn1, Wn2, and Wn3 are values indicating the magnitude of wind noise, indicating that the wind noise is stronger in this order.
As shown in FIG. 7 (d), if the wind noise is smaller than the predetermined value Wn1, it is considered that wind processing is not required, and the cutoff frequencies of the variable LPF76 and HPF73 are set to 50 Hz. As a result, by passing through the circuit shown in Fig. 7 (c), the signal from the second microphone 7b is almost completely cut off, and the frequency higher than 50Hz, which is the cutoff frequency of HPF73, dominates the sound. The signal of) can be obtained only from the first microphone 7a. Since it is not necessary to use the signal of the second microphone 7b provided with the acoustic resistor 41, it is considered that the sound of the subject is faithfully obtained.
The time when the wind noise exceeds the level of Wn1 and is between Wn1 and Wn2 will be described. At this time, the cutoff frequencies of the variable LPF76 and HPF73 are gradually raised while remaining in agreement. By performing the above-mentioned control, the low-frequency audio signal gradually uses the signal from the second microphone 7b provided with the acoustic resistor 41. Wind noise has a large effect on the signal from the first microphone 7a, but the wind noise is reduced by increasing the cutoff frequency of HPF73.
The time when the wind noise exceeds the level of Wn2 and is between Wn2 and Wn3 will be described. At this time, the cutoff frequency of the variable LPF76 is fixed at 1 kHz, and the cutoff frequency of the HPF73 is further raised. By performing the above-mentioned control, the sound existing between the cutoff frequency of LPF72 and the cutoff frequency of HPF73 is lost, but the wind noise can be further reduced. If the cutoff frequency of HPF73 is raised excessively, the deterioration of the subject sound will become too great, so I try not to raise it above the appropriate cutoff frequency. In the example of Fig. 7 (d), when the magnitude of wind noise exceeds Wn3, the cutoff frequency of HPF73 is fixed at 2kHz and does not change any more.
In the above description, an example of moving the HPF73 wider than the operation of the variable gain 74 and the variable LPF76 has been described. Obviously, by setting Wn2 = Wn3, it is possible to operate the HPF73 only in the same range as the variable gain 74 and the variable LPF76. If the operation is restricted, the effect of reducing wind noise is reduced, but the subject sound can be faithfully acquired. On the other hand, the magnitude of wind noise generated in the first microphone 7a when the wind blows greatly differs depending on the mounting structure of the microphone and the like. The settings of Wn1, Wn2, and Wn3 may be adjusted by comparing the need to reduce wind noise with the need to faithfully acquire the subject sound.
In the above description, in the example of the synthesizer 71 shown in FIG. 7, the range in which the cutoff frequencies of the variable HPF and LPF are changed is specifically shown. The preferred variable range and filter configuration will be briefly described.
In the synthesizer 71 shown in this embodiment, the voices acquired by the plurality of microphones 7a and 7b are synthesized. In such a process of separating into bands and performing synthesis, it is desirable that the phases of the respective paths match, especially in the frequency band where the signals of a plurality of microphones overlap. This is because when the phases are out of phase due to processing in a plurality of paths, the waveforms may not overlap correctly and may cancel each other out. In order to fully satisfy this, it is convenient that HPF73 and LPF72 are composed of FIR filters of the same order. By using the FIR filter, so-called group delay characteristics can be obtained, and signals can be synthesized without contradiction even when processing is performed for each band. When the cutoff frequency of the FIR filter is very low (to be exact, when the ratio is very small when normalized by the ratio to the sampling frequency), the order is very high to obtain sufficient filter performance. A filter is needed. This is derived from the fact that a large number of samples are required to obtain a wave of the frequency to be blocked / passed. Since the order of the filter cannot be increased infinitely, the lower limit of the variable range of the cutoff frequency is determined from this. Since the LPF and HPF are variable in the configuration of FIG. 7 (c), the order of the variable LPF76 and HPF73 becomes very high when the cutoff frequency is set to be very low. Therefore, as a limit for lowering the frequency, in the example of FIG. 7, 50 Hz is illustrated as a range that does not significantly affect the signal in the audible range. As mentioned above, it is not limited to 50Hz and may be set appropriately according to computer resources. In the example of FIG. 7 (a), since only the HPF is variable, only one high-order filter is required as described above. In terms of reducing the amount of calculation, it is superior to the configuration shown in Fig. 7 (c).
On the other hand, the upper limit of the variable range is limited by the second microphone 7b provided with the acoustic resistor 41. As schematically shown in FIG. 3 (b), the band of the subject that can be acquired by the second microphone 7b is limited to f0 due to the influence of the acoustic resistor 41. Since the subject sound is not obtained in the portion exceeding this, the cutoff frequencies of the variable LPF76 and HPF73 in the example of FIG. 7 should be set lower than this. It is f1 in Fig. 3, and obviously f1 <f0 should be set.
The effects of HPF52, variable operation, etc. will be described with reference to FIGS. 1, 3, 6, and 8 to 11. As explained with reference to FIGS. 3 and 6, wind noise is concentrated in low frequencies, and the influence of the first microphone 7a and the second microphone 7b is significantly different. That is, even if the wind is weak, a large wind noise is generated in the first microphone 7a. Problems associated with this may be the saturation of ADC54a and improper operation of ALC61. Since it is easy to understand the saturation of ADC54a, the explanation is omitted, and the problems associated with ALC61 operation when wind noise is generated will be described.
In the absence of HPF52, large wind noise is generated in the first microphone 7a as shown in FIG. It is assumed that the wind noise becomes dominant even when the wind noise and the subject sound are superimposed. In such an environment, the ALC61 adjusts the level by referring to the wind noise level of the first microphone 7a. After that, when the wind noise is processed by the HPF73 in the synthesizer 71, the level of the audio signal drops significantly. As a result, there is a problem that the output from the adder 75 becomes very small. That is, the signal level becomes inappropriate.
For example, it is conceivable to apply the invention shown in Patent Document 1 in order to solve the above-mentioned problems of ADC saturation and signal level becoming inappropriate. An example of the voice processing device 51 at this time is shown in FIG. In Fig. 8, those having the same function as in Fig. 1 are given the same number. In FIG. 8, variable gains 62a and 62b are provided in front of the ADCs 54a and 54b to avoid ADC saturation. Further, another ALC61b is provided after the wind noise processing by the synthesizer 71, and the variable gain 62c and the level adjuster 63b prevent the signal level after the wind processing from becoming inappropriate.
However, the circuit shown in FIG. 8 also has two problems. One is the increase in circuit scale by performing level ALC operation at two locations. The other is the increase in quantization error due to the gain being raised by the ALC61b located behind the synthesizer 71. That is, the level adjuster 63a adjusts the level with a signal containing wind noise, and the level adjuster 63b adjusts the level with a signal containing no wind noise. If the wind noise reduction effect is large, it will be necessary to greatly increase the gain with the level adjuster 63b. At this time, since the signal has already been digitized, the quantization error increases as the level is adjusted.
The quantization error referred to here will be briefly described. For example, when raising the 12 dB gain with the level adjuster 63b, the operation of shifting the digital signal to the left by 2 bits may be performed, but since there is no information corresponding to the lower 2 bits at that time, fill it with an appropriate value (for example, 0). There is a need. In this case, the lower 2 bits are always 0, so only 4 can be expressed after 0 in decimal. In this way, the signal can be expressed only in a discrete manner, and a quantization error occurs with respect to a natural signal (continuous).
Now consider the HPF52 shown in Figure 1. By setting the cutoff frequency of HPF52 appropriately, the main component of wind noise can be removed. As a result, it is possible to prevent saturation of ADC54a and to make appropriate gain adjustment in ALC61. (At the time of ALC61, the subject sound is not buried in the wind noise, so it is possible to perform ALC operation according to the level of the subject sound.)
An example of the cutoff frequency control sequence in the HPF52 will be described with reference to FIG. FIG. 9A shows the operation sequence of the switch 87, FIG. 9B shows the operation sequence of the HPF52, FIG. 9C shows the operation sequence of the variable gain 74, and FIG. 9D shows the operation sequence of the HPF73. .. In addition, in FIGS. 9 (a) to 9 (d), the horizontal axis is common and indicates the magnitude of wind noise. Wn1, Wn2, and Wn3 are values indicating the magnitude of wind noise, indicating that the wind noise is strong in this order. The operations of FIGS. 9 (c) and 9 (d) are the same as those of FIG. 7 (b), and the description thereof will be omitted.
If the wind noise is smaller than the predetermined value Wn1, it is considered that wind processing is not necessary, and the switch 87 is turned ON to perform the adaptive operation of the reverberation suppressor 53 described above. In addition, the cutoff frequency of HPF52 is set to 0Hz (= through without HPF operation). Since it is not necessary to use the signal of the second microphone 7b provided with the acoustic resistor 41, it is considered that the sound of the subject is faithfully obtained.
When the wind noise exceeds the level of Wn1, it is assumed that the wind noise is generated, and the switch 87 is turned off to stop the adaptive operation of the adaptive filter in the reverberation suppressor 53 described above. By performing such control, inappropriate adaptive operation can be suppressed.
The time between Wn1 and Wn2 will be explained. At this time, the cutoff frequency of HPF52 is gradually raised within a range not exceeding the cutoff frequency of HPF73. By performing the above-mentioned control, it is possible to reduce the wind noise generated by the first microphone 7a. Further, by controlling so as not to exceed the cutoff frequency of HPF73, the cutoff frequency of HPF52 does not have a great influence on the output of HPF73.
The effect of this will be explained. Since the HPF52 is provided in the analog part (stage before the ADC) of the audio processing device 51, it is generally composed of an IIR filter (HPF by RC circuit). At this time, HPF52 cannot satisfy the group delay characteristic. On the other hand, even in the IIR filter, the phase delay is small in the pass band, so that the phase delay does not affect even if the group delay characteristic is not satisfied. By controlling the cutoff frequencies of HPF52 and HPF73 as described above, the influence of the phase delay caused by the IIR filter can be reduced. As described above, in the process of separating into bands and performing synthesis, it is desirable that the phases of the respective paths match, especially in the frequency band where the signals of a plurality of microphones overlap. However, it shows that the effect can be reduced even in situations where this is not observed. Further, as described above, the HPF 52 is provided in the analog section of the voice processing device 51, but if the cutoff frequency is continuously changed in the analog circuit, the circuit scale becomes large. By making the circuit suitable for the control sequence as described in FIG. 9, it can be realized by a simple configuration.
Examples of signals processed by the circuit described above are shown in FIGS. 10 and 11. FIG. 10 shows the case where the HPF52 is not provided, and FIG. 11 shows the case where the HPF52 is provided. The signal of FIG. 10 is a signal processed with respect to FIG. 1 with HPF52 removed. Further, as shown in the figure, the graph shows the gain 62a output, the gain 62b output, the HPF73 output, the LPF72 output, and the adder 75 output in order from the top. The horizontal axis shows time, which is common to all graphs. In the examples of FIGS. 10 and 11, the state in which the subject is speaking from around 2.5 seconds (the human voice is the sound to be picked up) is shown. Further, the signals shown in FIGS. 10 and 11 were processed assuming that the wind noise level was at the level of Wn2 in FIG.
The part before 2.5 seconds is the same as the one shown in Fig. 6, and it is in the state of only wind noise. Focusing only on this part, the gain 62a output of FIGS. 10 and 11 seems to be larger in FIG. This is because the gain is actually increased by ALC61. This is clear when looking at the subject sound after 2.5 seconds.
Focusing on the gain 62b output after 2.5 seconds, it can be seen that the signal in FIG. 10 has a clearly lower signal level than the signal in FIG. This is because the ALC61 adjusted the level for the wind noise generated by the first microphone 7a, so the gain became small, and as a result, the subject sound was acquired very small. On the other hand, the signal in FIG. 11 reduces the wind noise generated by the first microphone 7a by the effect of HPF52, so that the gain of ALC61 is kept higher than that in the state of FIG.
Focusing on the HPF73 output in FIG. 10, it can be seen that the wind noise is considerably reduced by appropriately processing the cutoff frequency of the HPF73. However, since the signal level of the HPF73 is significantly lower than the signal level of the gain 62a output, it can be seen that the signal level of the final adder 75 output is very small.
On the other hand, also in FIG. 11, it can be seen that the wind noise is considerably reduced by appropriately processing the cutoff frequency of HPF73. Furthermore, since the output of LPF72 is kept large, it can be seen that the signal level of the output of the final adder 75 is also kept at a sufficient level.
By arranging the HPF52 closer to the microphone than the ADC and ALC in this way, it is possible to obtain high-quality sound.
An example of another circuit configuration of this embodiment is shown in FIG. FIG. 12 (a) shows an example in which the ALC is arranged in the analog section, and FIG. 12 (b) shows an example in which the ALC 61 is arranged behind the synthesizer 71. Even with such a configuration, the effects shown in this embodiment can be obtained.
As described above, according to the present invention, it is possible to obtain a high-quality sound in which the reverberation sound is suppressed while reducing the wind noise by the acoustic resistor.
(Example 2) Hereinafter, a recording device and an imaging device including the recording device according to the second embodiment of the present invention will be described with reference to FIGS. 13 to 14. In the second embodiment, those having the same operation as the first embodiment are numbered the same.
FIG. 13 is a perspective view of the image pickup apparatus. FIG. 13 is similar to FIG. 2, but with the addition of an opening 32c for the microphone. A microphone 7c (not shown) is provided behind the opening 32c.
FIG. 14 is a diagram illustrating a main part of the voice processing device 51 corresponding to the device shown in FIG. FIG. 14 shows an extension to stereo based on the circuit in which ALC is performed in analog as shown in FIG. 12 (a) in the first embodiment. In addition, the reverberation suppressor 53 and the level detector 86 have been simplified / changed. For the first embodiment, the first microphone 7a is extended to two. Here, the microphone 7a and the microphone 7c are microphones that form the left and right stereo channels, and are designed so that their characteristics are the same. On the other hand, the second microphone 7b is provided with the acoustic resistor 41, and has the same characteristics as those of the first embodiment.
The HPF52b, gain 62c, ADC54c, HPF56c for DC component cut, and HPF73b expanded in FIG. 14 have the same movements as the HPF52, gain 62a, ADC54a, HPF56a for DC component cut, and HPF73 shown in Example 1, respectively. To do. Here, the delayers 55a and 55b whose operation changes, the newly installed phase comparator 57, the adder 58, and the gain 59 will be described.
In a stereo recording device, a stereo feeling is given to the signal by the phase difference of the audio signal. On the other hand, in the arrangement as shown in FIG. 13, the second microphone 7b is arranged between the first microphones 7a and 7c. In such a configuration, when considering the phase difference between the microphone 7a and the microphone 7c, the phase of the signal of the second microphone 7b exists in the middle. For example, when the second microphone 7b is placed exactly in the middle so that the microphone 7a and the microphone 7b and the microphone 7c and the microphone 7b are equidistant, the phase is also exactly in the middle. Therefore, in the circuit of FIG. 14, the phase difference between the microphone 7a and the microphone 7c is calculated, and the corresponding delay is given by the delayers 55a and 55b.
For example, consider the case where the signal of the microphone 7c is delayed more than the signal of the microphone 7a. At this time, as will be described later, the reverberation suppressor is adjusted to match the signal in the middle. When mixing with the signal of the microphone 7a, the phase may be advanced, and when mixing with the signal of the microphone 7c, the phase may be delayed. In the first embodiment, it is sufficient to give a delay of half (= M / 2) of the filter order of the reverberation suppressor 53, but 55a gives a smaller delay and 55b gives a larger delay. You just have to give a delay. The absolute value differs depending on the arrangement of the microphones. For example, as described above, when the second microphone 7b is located between the first microphones 7a and 7c, the phase difference calculated by the phase comparator 57 You can shift each half of. By performing the above-mentioned processing, an audio signal can be obtained without impairing the stereo feeling.
The adder 58 and the gain 59 will be described. The adder 28 adds the signals of the microphone 7a and the microphone 7c. Gain 59 halves the output of adder 58. As a result, the output of the gain 59 is the summing average of the microphone 7a and the microphone 7c. The resulting voice phase is intermediate between the microphone 7a and microphone 7c signals. On the other hand, BPF82a allows only a band of about 30 Hz to 1 kHz to pass as shown in Example 1. Further, the voice processing device 51 has a configuration capable of acquiring voice having a high frequency with respect to the pass band of the BPF. The audio signals that can be acquired at this time are arranged so that phase inversion does not occur between the microphone 7a and the microphone 7c signals. From the above, when observing only the band passed by BPF82a, the phase difference existing between the microphone 7a and the microphone 7c signal is small. From this, it can be considered that the signal levels in the 82a pass band are almost added. Therefore, by halving the output with a gain of 59, it is possible to obtain a signal in which the signal level is almost the same as 7a and 7c and the phase is in the middle. In this embodiment, the reverberation suppressor 53 is operated so as to match the output of the gain 59 described above.
With the above configuration, the present invention can be easily applied even in a device for recording in stereo without impairing the stereo feeling.
In this embodiment, the case of stereo (the case where there are two first microphones for acquiring up to the high frequency range) has been described, but a recording device having more microphones can be easily expanded.
(Example 3) Hereinafter, with reference to FIG. 15, a recording device and an imaging device including the recording device according to the third embodiment of the present invention will be described. In the third embodiment, those having the same operation as the first embodiment are numbered the same.
The perspective view of the image pickup apparatus provided with the recording apparatus according to the third embodiment is the same as that of FIG. 2 of the first embodiment, and thus is omitted. FIG. 15 is a diagram illustrating a main part of the voice processing device 51 in the third embodiment. In FIG. 15, an upsampler 96 that changes the sampling frequency of the audio signal is arranged in front of the LPF72. Further, unlike the first embodiment, different values are set for the sampling frequencies in the ADCs 54a and 54b. The sampling frequency of ADC54b is set to a lower value than the sampling frequency of ADC54a. The sampling frequency of ADC84 is set to the same value as ADC54b.
The ADC54b, ADC84, reverberation suppressor 53, and the newly installed upsampler 96 will be described.
The output of the first microphone 7a is branched and sent to the wind detector 81, passed through the BPF82a, and then A / D converted by the ADC84 at a sampling frequency lower than that of the ADC54a. This sampling frequency is a value within the range in which the band passed by BPF82a can be reproduced, and it is desirable that the sampling frequency is 1 / integer of the sampling frequency of ADC54a. For example, if the pass band of BPF82a is 30Hz to 1kHz and the sampling frequency of ADC54a is 48kHz, set it to 3kHz, which is 1/16 of 48kHz. Then, the output of the ADC 84 is delayed by the delay device 85 and sent to the difference device 83.
On the other hand, the signal of the second microphone 7b is A / D converted in ADC 54b to the same sampling frequency as ADC 84. Then, after the reverberation is suppressed by the reverberation suppressor 53, it branches and is sent to the wind detector 81, passes through the BPF 82b, and then is sent to the diffifier 83. Since the sampling frequency of the filter order M of the reverberation suppressor 53 is suppressed to 1/16 by ADC54b, the same effect as the conventional one can be obtained even if it is 1/16 of the conventional one. It leads to a decrease in. As the filter order M of the reverberation suppressor 53 decreases, the delay amount of the delay device 85 also decreases. Since the operation of the differencer 83 and below is the same as that of the first embodiment, it is omitted.
One of the outputs of the branched reverberation suppressor 53 passes through HPF56b, gain is adjusted by ALC61, and sent to the upsampler 96. In the upsampler 96, the output of variable gain 62b is converted to the same sampling frequency as ADC54a and sent to LPF72. Upsampling may cause aliasing, but LPF72 reduces high frequency components and eliminates aliasing.
The operations of HPF 52 or less and LPF 72 or less in the subsequent stage of the first microphone 7a are the same as those in the first embodiment and are omitted.
With the above configuration, the circuit scale and the amount of calculation can be reduced by downsampling the low frequency components and performing the reverberation suppression processing. Further, by performing upsampling after the reverberation suppression process, it is possible to obtain high-quality sound.
(Example 4) Hereinafter, a recording device and an imaging device including the recording device according to the fourth embodiment of the present invention will be described with reference to FIGS. 16 and 17. In the fourth embodiment, those having the same operation as the first embodiment are numbered the same.
The perspective view of the image pickup apparatus provided with the recording apparatus according to the fourth embodiment is the same as that of FIG. 2 of the first embodiment, and thus is omitted. FIG. 16 is a diagram illustrating a main part of the voice processing device 51 in the third embodiment. Reference numeral 97 in FIG. 16 is a cross-correlation calculator that receives the branched outputs of the BPF 82b and the delay device 85, calculates the cross-correlation value of the two signals, and determines whether or not there are multiple directions of arrival of the sound source. The operation of the cross-correlation calculator 97 will be described later. FIG. 17 schematically shows the positional relationship between the sound source generated by the subject sound and the microphones 7a and 7 and the propagation of the sound. FIG. 17 (a) shows the case where the subject sound propagates from one direction. 17 (b) is a schematic diagram when the subject sound propagates from two directions.
A problem when the subject sound propagates from two directions will be described with reference to FIG. Let s1 be the subject sound emitted from a certain subject O1 and s2 be the subject sound emitted from a direction different from that of the subject O1. Then, the transfer function of the sound propagating from the subject O1 to the microphone 7a is T1a, and the transfer function of the sound propagating to the microphone 7b is T1b. Similarly, the transfer functions of the sound propagating from the subject O2 to the microphones 7a and 7b are T2a and T2b, respectively. When the sound source of the subject sound is unidirectional as shown in FIG. 17 (a), the audio signals x1 and x2 acquired by the microphones 7a and 7b are expressed by the following equations, respectively.
<maths num="6"><img file="JP2012129652A_D0006.tif" /></maths>
There is a delay between the signal x1 of the microphone 7a and the signal x2 of the microphone 7b due to the difference in distance from the microphone 7a and the microphone 7b from the subject sound, but the correlation between the two signals is very large just because there is a time lag. Is expensive. On the other hand, when the subject sound propagates from two directions as shown in FIG. 17B, the audio signals x1 and x2 acquired by the microphones 7a and 7b are expressed by the following equations, respectively.
<maths num="7"><img file="JP2012129652A_D0007.tif" /></maths>
Between the signal x1 of the microphone 7a and the signal x2 of the microphone 7b, a delay occurs depending on the distance between the two microphones 7a and 7b and the two subjects O1 and O2. As the positions of the two subjects O1 and O2 move away from each other, the amount of delay between T1a and T1b and T2a and T2b shifts, so that the correlation between the two signals becomes lower. As a result, there arises a problem that the reverberation suppressor 53 is not updated correctly.
Therefore, in the imaging device provided with the recording device according to the fourth embodiment, the cross-correlation calculator 97 is provided, and when the cross-correlation value of the two signals is lower than the specified value, the learning of the reverberation suppressor is stopped. To solve.
The operation of the cross-correlation calculator 97 will be described. The branched output of BPF82b and delayer 85 is sent to the cross-correlation calculator 97. This is an audio signal in the frequency band of 30 Hz to 1 kHz that has passed through the BPF 82a and BPF 82a of the microphone 7a and the microphone 7b, respectively. Let this signal be x1_BPF and x2_BPF, respectively, and the cross-correlation calculator 97 calculates the cross-correlation value of the two signals as follows. The cross-correlation value R (n) of the two signals in the nth sample when the data length is N is calculated by the following equation.
<maths num="8"><img file="JP2012129652A_D0008.tif" /></maths>
Furthermore, when this is normalized by x1_BPF, it is expressed as the following equation.<maths num="9"><img file="JP2012129652A_D0009.tif" /></maths>
Ideally, Rnorm (n) has a maximum value of 1 when the subject sound is unidirectional. However, when the sound source direction of the subject sound is two or more directions, the cross-correlation of the two signals becomes low, so that Rnorm (n) becomes lower than 1. Therefore, if the obtained normalized cross-correlation value Rnorm (n) is lower than the predetermined value Rn1, it is determined that the sound source direction of the subject sound is two or more directions, and the switch 87 is turned off. Stops the adaptive operation of the reverberation suppressor 53.
Further, in the image pickup apparatus according to the third embodiment, the switch 87 is switched according to the detection result of the level detector 86 as in the first embodiment. That is, when the cross-correlation calculator 97 detects that the cross-correlation value is lower than Rn1, or the level detector 86 detects that the wind noise exceeds the level of Wn1, the switch 87 is turned off. The adaptive operation of the adaptive filter in the reverberation suppressor 53 is stopped.
By performing such control, even when the subject sound propagates from two or more directions, an appropriate adaptive operation can be performed, and high-quality sound can be obtained.
(Other embodiments) The present invention is also realized by executing the following processing. That is, software (program) that realizes the functions of the above-described embodiment is supplied to the system or device via a network or various storage media, and the computer (or CPU, MPU, etc.) of the system or device reads the program. This is the process to be executed. In this case, the program and the storage medium that stores the program constitute the present invention.
9 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| JP2014060523A | Cited by | Japan | Search report |
| JP2020039055A | Cited by | Japan | Search report |
| US10244271B2 | Cited by | United States of America | Applicant |
| WO2016203866A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| US2005041825A1 | Cites | United States of America | Search report |
| JP2006262098A | Cites | Japan | Examiner |
| US2007258597A1 | Cites | United States of America | Search report |
| JP2008060625A | Cites | Japan | Examiner |
| WO2009078105A1 | Cites | World Intellectual Property Organization (WIPO) | Examiner |
| JP2009542057A | Cites | Japan | Examiner |
| US5193117A | Cites | United States of America | Search report |
| JPH03106299A | Cites | Japan | Examiner |
| JPH03219798A | Cites | Japan | Search report |
| JPH0654394A | Cites | Japan | Examiner |
| JPH09218687A | Cites | Japan | Examiner |
| JPH0965482A | Cites | Japan | Examiner |
| JPH11331973A | Cites | Japan | Search report |
2 priority claims, no other members on record
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 2010277419 | Japan | A | |
| JP20100277419 | – | – | – |
1 legal event, as the office reported them to INPADOC
Events
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS |
Numbers
- Publication
- 2012129652
- Publication, DOCDB
- 2012129652
- Publication, EPODOC
- JP2012129652
- Application
- 277419
- Application, DOCDB
- 2010277419
- Application, EPODOC
- JP20100277419
Titles3
- English
- Audio processing equipment and methods and imaging equipment
- English
- SOUND PROCESSING DEVICE AND METHOD, AND IMAGING APPARATUS
- Japanese
- 音声処理装置及び方法並びに撮像装置
Classification
- CPC, 2
- G10L21/0208
- G10L2021/02161
- IPC, 4
- H04R3 00
- H04R1 08
- H04N5 225
- H04N101 00