Vocal fly detector and computer program
Abstract
A VF detecting apparatus capable of highly accurate vocal fry (VF) detection includes: a very-short-term peak detection processing unit framing a speech signal with a first frame of a first frame length and first frame shift amount and detecting each power peak;a short-term periodicity detecting unit framing the speech signal with a second frame of a second frame length longer than the first frame length and a second frame shift amount larger than the first frame length and determining presence/absence of periodicity in each of the resulting frame; a periodicity checking unit for detecting power peaks in those frames determined to have no periodicity, from among the detected power peaks; and a similarity checking unit for detecting, for each of the selected power peaks, neighboring power peaks having high cross-correlation and detecting the section therebetween as the VF section.

Term
Term ended
Expired 31 August 2025, 1.1 years ago.
- Priority and filed
- Granted
- Expired
- Today
4 claims: 1 independent, 3 dependent
- 1発話信号中のボーカル・フライ区間を検出するためのボーカル・フライ検出装置であって、 発話信号を、第1のフレーム長でかつ第1のフレームシフト量の第1のフレームでフレーム化するための第1のフレーム化手段と、 前記第1のフレーム化手段の出力する一連の第1のフレームの各々のパワーのピークを検出するためのパワーピーク検出手段と、 前記発話信号を、前記第1のフレーム長よりも大きな第2のフレーム長で、かつ前記第1のフレームシフト量よりも大きな第2のフレームシフト量の第2のフレームでフレーム化するための第2のフレーム化手段と、 前記第2のフレーム化手段の出力する一連の第2のフレームの各々の内部における周期性の有無を判定するための周期性判定手段と、 前記パワーピーク検出手段により検出されたパワーピークのうちで、前記周期性判定手段により周期性がないと判定された前記第2のフレーム内のパワーピークを選択するためのパワーピーク選択手段と、 前記パワーピーク選択手段により選択されたパワーピークの各々について、当該パワーピークを含む所定区間内の他のパワーピークとの間の相互相関が所定のしきい値よりも大きなパワーピークを探索し、前記発話信号中の、当該パワーピークを含む所定の区間をボーカル・フライ区間として検出するための手段とを含む、ボーカル・フライ検出装置。
- 2前記周期性判定手段は、前記一連の第2のフレームの各々において、当該フレーム内での最大パワーピークの、当該フレーム内の所定の遅延範囲内での自己相関値の関数としてフレーム内の周期性の尺度を算出し、当該自己相関値のピークが所定のしきい値関数よりも大きいか否かにしたがって、周期性があるか否かを判定するための手段と、 前記判定するための手段により周期性があると判定された前記第2のフレームのうち、前記周期性の尺度が予め定める定数よりも大きなフレームが所定個数連続している部分以外の前記第2のフレームの前記周期性の尺度の値を、周期性がないと判定される値に補正するための周期性補正手段を含む、請求項1に記載のボーカル・フライ検出装置。
- 3前記発話信号を前記第1のフレーム化手段及び前記第2のフレーム化手段に与えるに先立って、前記発話信号の所定の周波数帯域の成分以外の成分を除波するためのフィルタリング手段をさらに含む、請求項1又は請求項2に記載のボーカル・フライ検出装置。
- 4コンピュータにより実行されると、当該コンピュータを、請求項1~請求項3のいずれかに記載のボーカル・フライ検出装置として動作させる、コンピュータプログラム。
Independent claims4
63 paragraphs, as filed
The present invention relates to a technique for analyzing human voice quality, and more particularly to a VF detection device for detecting a section having a specific voice quality called a vocal fly (hereinafter referred to as "VF") from an utterance signal.
In the dialogue between humans and machines, it is necessary to automatically extract information other than textual information contained in speech (hereinafter referred to as "paralanguage information"). Conventionally, phonological features such as pitch, power, and duration have been used as acoustic features for extracting paralanguage information. However, recent studies have reported that information on voice quality, such as breathability, squeaks, and faintness, depending on the mode of pharyngeal voice source, also plays an important role in the perception of paralanguage information.
The terms VF, creaky voice, glottis fly, pulse register, and laryngealization refer to a relatively discrete series of laryngeal (or glottic) excitations (or short-term pulses). It is used as a representation in the prior art literature. In these voices, the vocal tract is almost completely dampened between successive glottic pulses, usually with a very low fundamental frequency and irregular duration of the glottic cycle. When you hear the VF, the perception is "a fast, continuous tapping sound when you move the rod along the railing", or "a mouth imitation of the engine sound of a motor boat", or "a sound when cooking in a hot frying pan". It is expressed as "sound similar to".
VF is language-dependent, but conveys important paralinguistic information in addition to important linguistic information. In German, VF often occurs near the boundaries of morphemes. In Japanese, in addition to VF occurring in a low-pitched voice, it also occurs in utterances with emotional emphasis, such as a squeaky voice. Rikimi voice conveys paralinguistic information primarily related to emotions or attitudes about surprises, praises, and suffering. A very low fundamental frequency can be seen in the VF utterance part (hereinafter referred to as the "VF segment") in such a voice.
In addition, the VF segment is characterized by its irregularity, which can cause serious errors in the pitch determination algorithm, which plays an important role in the extraction of phonological information. Therefore, knowing where the VF occurs is important not only for extracting paralanguage information, but also for improving pitch determination performance.
The physiological, perceptual, and acoustic attributes of VF have been reported in several research areas. Many of them report qualitative or descriptive material about acoustic features associated with various voice qualities. However, few reports have been made on the evaluation of VF for the purpose of automatic detection.<nplcit num="1"><text>Ishii, CT, "Analysis of Autocorrelation-Based Parameters for Creaky Voice Detection," Proceedings of the 2nd International Conference on Speech Prosody, pp.643-646, 2004. (Ishi, CT, "Analysis of Autocorrelation-based parameters for Creaky Voice Detection," Proc. Of The 2nd International Conference on Speech Prosody: 643-646, 2004.)</text></nplcit>
<p> It has been reported that the fundamental frequency range of VF is consistently lower than 100Hz and the average is around 24-52Hz. Two, and sometimes three, glottic pulses in VF occur at very short intervals, followed by considerable glottal braking.</p><p> Regarding VF, many acoustic analyzes in the time domain, the spectral domain, and the cepstrum domain have been reported. In the usual method, attributes related to periodicity (or harmonicity) are evaluated using a fixed-length short-time analysis frame.</p><p> Using fixed length frames causes problems if the VF segment has a very low fundamental frequency (ie, has a very long interval between pulses). The frame length of a standard (commonly used) analysis frame ranges from 25 ms to 32 ms, but under such conditions, there is often at most one glottic pulse in the analysis frame in the VF segment. Sometimes the frame does not contain any glottic pulses. Without the presence of at least two glottic pulses in the analysis frame, no tuning structure can be found in the spectrum, and it is difficult to generate a peak of correlation that reflects the short-term periodicity between glottic pulses.</p><p> The simplest solution to this is to increase the analysis frame length. In Non-Patent Document 1, periodicity analysis based on autocorrelation is performed using a technique of adaptively changing the frame length. However, such a method can solve only part of the problem. This is because a large analysis frame can contain two glottic pulses with different pulse spacings. In such a case, the tuning structure in the spectrum is disturbed, and the magnitude of the peak of autocorrelation (or cepstrum) is also reduced.</p><p> Therefore, an object of the present invention is to provide a VF detection device that accurately performs VF detection while avoiding problems such as disturbance of the wave-tuning structure in the spectrum and reduction of the peak of autocorrelation.</p><p> Another object of the present invention is to provide a VF detection device that accurately performs VF detection by a method synchronized with the glottis pulse while avoiding the problems of disturbance of the wave tuning structure in the spectrum and reduction of the peak of autocorrelation. is there.</p><p> Yet another object of the present invention is to avoid the problems of disturbance of the wave-tuning structure in the spectrum and reduction of the peak of autocorrelation by using an appropriate analysis frame, and accurately detect VF by a method synchronized with the glottic pulse. It is to provide the VF detection apparatus which performs.</p>
<p> The VF detection device according to the first aspect of the present invention is a device for detecting the VF section in the utterance signal, and the utterance signal has the first frame length and the first frame shift amount. A first framing means for framing with a frame of, a power peak detecting means for detecting the peak of each power of a series of first frames output by the first framing means, and an utterance signal. A second framing means for framing a second frame with a second frame length larger than the first frame length and a second frame shift amount larger than the first frame shift amount. Of the periodicity determining means for determining the presence or absence of periodicity inside each of the series of second frames output by the second framing means, and the power peak detected by the power peak detecting means. , The power peak selection means for selecting the power peak in the second frame determined to be non-periodic by the periodicity determination means, and the power peak selected by the power peak selection means. Search for a power peak whose mutual correlation with other power peaks in a predetermined section including is larger than a predetermined threshold value, and detect a predetermined section including the power peak in the utterance signal as a VF section. Including means for.</p><p> The power peak is detected by the utterance signal framed by the first frame. The presence or absence of periodicity is determined by the utterance signal framed by the second frame. The first frame has a shorter frame length than the second frame, and the frame shift amount is also small. Therefore, a waveform with a low fundamental frequency such as a VF pulse can be detected with higher accuracy than when the second frame is used. On the other hand, since the frame length of the second frame is longer than that of the first frame, it is possible to more accurately determine whether or not there is periodicity in the frame. Among the detected power peaks, those existing in the non-periodic part are likely to be VF pulses. Furthermore, if such a VF pulse candidate shows a high cross-correlation with other adjacent pulses in a predetermined interval, it is more likely that the VF pulse candidate is a VF pulse. By detecting the section including the power peak corresponding to such a VF pulse as the VF section, the VF section can be detected with high accuracy. Since the first and second frames are used for processing, a frame suitable for signal processing can be used, and VF detection can be performed with high accuracy.</p><p> Preferably, the power peak detecting means is larger than the power of any of the other frames in the predetermined section including the frame in the first frame in the series, and the difference is larger than the predetermined first threshold value. Of the power peak candidate detecting means for detecting a large frame as a power peak candidate and the power peak candidate detected by the power peak candidate detecting means, each frame in a section wider than a predetermined section including the frame. It includes means for detecting a frame larger than the power and the maximum value of the difference is larger than a predetermined second threshold value as a power peak.</p><p> More preferably, a section wider than the predetermined section is a period corresponding to 10 milliseconds in the utterance signal.</p><p> More preferably, in each of the second frames in the series, the periodicity determination means within the frame as a function of the autocorrelation value of the maximum power peak within the frame within a predetermined delay range within the frame. It includes means for calculating a measure of periodicity and determining whether or not there is periodicity according to whether or not the peak of the autocorrelation value is larger than a predetermined threshold function.</p><p> As a means for determining, the measure of periodicity may be calculated by multiplying the autocorrelation value for the maximum power peak by a function that becomes a monotonous decrease function for the amount of delay from the maximum power peak in the frame. ..</p><p> Preferably, a predetermined threshold function is obtained by multiplying a predetermined constant greater than 0 and less than 1 by a monotonically decreasing function.</p><p> More preferably, the periodicity determining means further comprises a predetermined number of consecutive frames having a periodicity scale larger than a predetermined constant among the second frames determined to have periodicity by the determining means. It includes a periodicity correction means for correcting the value of the periodicity scale of the second frame other than the present portion to a value determined to be non-periodic.</p><p> More preferably, it further includes filtering means for removing components other than the components of the predetermined frequency band of the utterance signal prior to giving the utterance signal to the first framing means and the second framing means.</p><p> When the computer program according to the second aspect of the present invention is executed by the computer, the computer operates as any of the above-mentioned VF detection devices.</p>
<Summary> In order to solve the problem of frame length, the inventors of the present invention decided to perform a process synchronized with the glottic pulse when no periodicity was found in the fixed-length analysis frame. Therefore, glottic pulse candidates are detected based on the VF attributes of braking and low fundamental frequency. This is based on the phenomenon that the braking that occurs at intervals between long pulses causes the amplitude envelope of the utterance signal, that is, the vertical movement of the local power curve.
Another problem with automatic detection is that many acoustic analyzes analyze the temporal or spectral characteristics of pre-segmented sound utterances with respect to the utterance signal. The practical problem of automatically detecting VFs from the entire utterance, including consonant and non-speech segments, can result in many insertion errors. This is because such segments also usually have the characteristic of aperiodicity. The question, therefore, is how to distinguish between the aperiodicity produced by VF and the reverberation produced by consonants and non-spoken signals in the environment.
With respect to this problem, the present embodiment attempts to solve the problem by evaluating a measure of similarity between successive (or adjacent) glottic pulses. This scale is based on the assumption that the structure of the glottis does not change between the generations of the two glottic pulses, and therefore the glottic response at the two timings will be similar.
<Structure> FIG. 1 shows a block diagram of an automatic dialogue system 100 that employs the vocal fly detection device 122 according to the embodiment of the present invention. With reference to FIG. 1, the automatic dialogue system 100 has a voice recognition device 120 for performing voice recognition for an incoming utterance signal 102 and outputting the voice recognition result 130 as text data, and one of the utterance signals 102. It includes a VF detection device 122 for detecting the VF period and outputting the VF section information 132.
The automatic dialogue system 100 further receives the voice recognition result 130 from the voice recognition device 120 and the VF section information 132 from the VF detection device 122, and paralinguistic information processing using the VF section information 132 and the voice recognition result 130. The response creation device 124 for understanding the speaker's intention and outputting the text information and voice quality information that are appropriate responses by integrating the above, and the response creation device 124 that refers to the response when creating the response. The response creation device 124 instructed the knowledge base 126, which stores the knowledge for creating an appropriate response to the combination of text information and paralinguistic information, and the text information of the response output from the response creation device 124. It includes a voice synthesizer 128 for synthesizing voice with voice quality and outputting it as a voice signal 104. The audio signal 104 is analogized by a circuit (not shown), amplified, and supplied to the speaker.
FIG. 2 shows a block diagram of the VF detection device 122. With reference to FIG. 2, the VF detector 122 includes a bandpass filter 160 for passing only the frequency components of 100 to 1500 Hz, which contains most of the information about the periodicity of the utterance signal 102. Frequency components below 100 Hz are DC components and components that gradually rise and fall, which adversely affect the periodic analysis, and are therefore dewavered by the bandpass filter 160. Further, since the frequency component exceeding 1500 Hz contains a high frequency noise component, this is also dewaved. The pass band of this bandpass filter is selected so that peaks and valleys can be detected from the power curve for each glottis pulse in the VF segment.
The VF detector 122 also uses frames with a frame length of 5 ms and a frame interval of 2.5 ms (referred to herein as "ultra-short-term frames") to localize within the output of the bandpass filter 160. Ultra-short-term peak detection processing unit 162 for detecting peaks of typical power as VF pulse candidates and outputting peak position information 170, and frame lengths of 25 to 32 ms and frame lengths of 10 or 5 ms. The frame used (this is referred to herein as the "short-term frame") is used, and the non-short-term periodicity part that indicates the possibility of VF being present in the output of the bandpass filter 160 is the other part. It includes a short-term periodicity detection unit 164 for detecting and outputting short-term periodicity information 172 separately.
The VF detection device 122 further receives peak position information 170 from the ultra-short-term peak detection processing unit 162 and short-term periodicity information 172 from the short-term periodicity detection unit 164, and among the peaks indicated by the peak position information 170, A periodic inspection unit 166 for selecting a frame including a frame existing in a non-short-term periodicity as a VF frame candidate and outputting it as VF candidate information 176, and a VF candidate information 176 output by the periodic inspection unit 166. And the utterance signal 174 of the frequency component of 100 to 1500 Hz output by the bandpass filter 160, and only the VF candidates having pulses similar to the predetermined range before and after are set as VF, and the VF section indicating the section where VF exists. It includes a similarity inspection unit 168 for outputting information 132.
FIG. 3 shows a block diagram of the ultra-short-term peak detection processing unit 162. With reference to FIG. 3, the ultra-short-term peak detection processing unit 162 includes a framing processing unit 190 for framing the utterance signal 174 of a frequency component of 100 to 1500 Hz output by the bandpass filter 160 by an ultra-short-term frame. An ultra-short-term power calculation unit 192 for calculating and outputting power (this is called "ultra-short-term power") and an ultra-short-term power calculation unit 192 for each of the ultra-short-term frames output by the framing processing unit 190. Of the series of ultra-short-term powers output by, the memory 194 for storing the latest predetermined number of values and the ultra-short-term power stored in the memory 194 are larger than both the ultra-short-term powers of one frame before and after. , And the difference is the predetermined power threshold Pw<sub>TH</sub>A peak comparison unit 196 for estimating a candidate for a VF glottis pulse larger than (for example, 6 to 7 dB) and outputting the peak position as peak position information 170, and a power threshold value Pw used by the peak comparison unit.<sub>TH</sub>Includes a power threshold storage unit 198 for storing.
FIGS. 4 and 5 show the principle of peak detection in the peak comparison unit 196. With reference to FIG. 4, the power is calculated by the ultra-short-term power calculation unit 192 for each of the ultra-short-term frames having a frame length of 5 milliseconds and a frame interval of 2.5 milliseconds, so that power values can be obtained at 2.5-millisecond intervals. Of these power values, those larger than the front and rear power values, such as arrows 210,212,214,216,218, can be peak candidates. Further, in the present embodiment, among these peak candidates, those satisfying the following conditions are set as peak candidates.
With reference to FIG. 5, the value of the power value 232 is the power threshold Pw compared to the power values 230 and 234 of the two frames before and after.<sub>TH</sub>It shall be larger. In the present embodiment, a frame showing this power value in such a case is used as a peak candidate. Like the power value 238, the difference between the power values 236 and 240 of the two front and rear frames is the power threshold Pw.<sub>TH</sub>Those less than are excluded from the peak candidates.
Figures 6 (A) and 6 (B) show the distribution of peak power increase and power decrease in the VF segment and non-VF segment (hereinafter referred to as "NF segment"), respectively, obtained by experiments. The amount of peak rise and fall here refers to the difference between the power value of a certain peak and the power of the frame 4 frames before the peak (that is, the power 10 milliseconds before the peak). According to FIG. 6 (A), it can be seen that a considerably large value is generated in both the amount of increase and the amount of decrease in the power value, reflecting the characteristic that braking occurs in VF. On the other hand, according to FIG. 6B, it can be seen that in the NF segment, the range of 1 to 6 dB is mostly in both the amount of increase and the amount of decrease in the power value.
From this figure, it is not always clear how much value should be selected as the threshold value (power threshold value) for distinguishing between VF and NF. This threshold is selected based on the results of experiments as described later, but a value of 7 dB is used, for example.
The short-term periodicity detection unit 164 shown in FIG. 2 is considered to be in the VF segment among the peak candidates extracted by the ultra-short-term peak detection processing unit 162 for each of the peak candidates thus determined. Has the ability to further select.
With reference to FIG. 7, the short-term periodicity detection unit 164 includes a frame processing unit 250 for framing the output of the bandpass filter 160 with a frame length of 32 milliseconds and a frame interval of 10 milliseconds, and a framing process. Intra-frame periodicity (IFP) by autocorrelation analysis based on the memory 252 for storing the framed utterance signal output by the unit 250 and the utterance signal for each frame stored in the memory 252. The IFP calculation unit 254 for calculating each frame and the IFP value calculated for each frame by the IFP calculation unit 254 are the threshold function IFP with a predetermined periodicity.<sub>TH</sub>If any of the peaks of the IFP value is lower than the threshold function, it is determined that there is no periodicity, and the periodicity determination unit 258 for setting the IFP value of the frame to null and the periodicity Based on the IFP value set by the determination unit 258, only when three or more frames with non-null IFP values are continuous, it is determined as a segment with short-term periodicity, and a short-term cycle indicating whether or not the frame has short-term periodicity. The continuity inspection unit 260 for outputting sexual information 172 and the periodicity threshold function IFP used by the periodicity determination unit 258.<sub>TH</sub>Includes a periodic threshold function storage unit 262 for storing.
The IFP value in the autocorrelation analysis by the IFP calculation unit 254 is defined by the value obtained by normalizing the correlation value of the maximum peak by "frame length / (frame length-delay)". This normalization is for compensating for the characteristic of the autocorrelation function as a monotonous decreasing function that the autocorrelation decreases as the delay amount increases.
In the IFP calculation unit 254, only autocorrelation peaks with a delay amount smaller than 15 milliseconds (corresponding to a fundamental frequency larger than about 66.7 Hz) are analyzed for periodicity. That is, at least two glottic cycles will be included in the analysis frame.
The periodicity determination unit 258 performs the following processing on the autocorrelation peak corresponding to the fundamental frequency larger than 200 Hz. That is, the periodicity of all low harmonics above 66.7 Hz is examined. This process prevents erroneous detection of periodicity due to strong wave tuning around the first formant rather than periodicity due to repetition of the glottic cycle. The low harmonic attributes of the autocorrelation function are shown in FIGS. 8 and 9. FIG. 8 shows the waveform and autocorrelation related to VF containing only one glottic pulse in one frame, and FIG. 9 shows the waveform and autocorrelation related to the ground voice having a high fundamental frequency. These are in the vowel / e / segments extracted from the female speaker's voice. In FIGS. 8 (B) and 9 (B), solid lines 276 and 296 show the threshold function. The threshold function is defined by "predetermined constant x (frame length-delay amount) / (frame length)". As a predetermined constant, a value of 0.5 is used in this embodiment. The threshold function also takes into account the attribute that the autocorrelation function is a monotonous decrease function for delay.
With reference to FIG. 9 (B), in the ground voice segment, for the strong harmonics contained in waveform 290 (FIG. 9 (A)), the peak of the autocorrelation 294 of the low harmonic component is also usually large. The autocorrelation peak 300 for low harmonics above 66.7 Hz (delay less than 15 ms, ie to the left of the dotted line 298) is higher than the threshold function 296.
On the other hand, referring to Fig. 8 (B), for the waveform 270 of the VF segment (Fig. 8 (A)), the autocorrelation function has a strong peak, but the delay is within 15 milliseconds (left side of the dotted line 278). Then, most of the low harmonic components have a value of 280, which is smaller than the threshold function 276, as the value of the autocorrelation function 274. In the present embodiment, the IFP calculation unit 254 has a function of calculating the autocorrelation function of each low harmonic component in this way. The periodicity determination unit 258 inspects the IFP value calculated for each frame by the IFP calculation unit 254, and if any of the peaks is smaller than the value of the threshold function, sets the IFP value of that frame to null. Has a function to do. The continuity inspection unit 260 inspects the IFP value for each frame output by the periodicity determination unit 258, and the frames have short-term periodicity only when at least three frames in which the IFP value is not null are continuous. In other cases, it is judged that there is no short-term periodicity.
Figures 10 (A) and 10 (B) show the distribution of experimentally obtained IFP values for the VF and NF segments, respectively, as white bar graphs. With reference to FIGS. 10 (A) and 10 (B), it can be seen that the number of frames in which the IFP value is null is overwhelmingly large in the VF segment. In FIG. 10, "null_1" is the number of frames in which the IFP value is null due to constraints on the low harmonic component (that is, frames with strong autocorrelation peaks but only weak autocorrelation peaks in low harmonics). Shown, "null_2" indicates the number of frames in which the IFP value is null due to the constraint of aperiodicity (that is, frames without strong autocorrelation peaks).
The periodic inspection unit 166 shown in FIG. 2 receives the peak position information 170 of the VF segment candidate from the ultra-short-term peak detection processing unit 162 and the short-term periodicity information 172 from the short-term periodicity detection unit 164, and the IFP value is set. It has a function of selecting only peak candidates of null frames and giving them to the similarity inspection unit 168 as VF candidate information 176.
FIG. 11 shows a block diagram of the similarity inspection unit 168 shown in FIG. With reference to FIG. 11, the similarity inspection unit 168 is a VF segment that clears the above-mentioned restrictions based on the speech signal 174 of the frequency component of 100 to 1500 Hz and the VF candidate information 176 from the periodicity inspection unit 166. To calculate the inter-pulse similarity (IPS) value calculated as a cross-correlation function between the waveform near each power peak and the waveform near the previous power peak for the power peak candidates of. IPS calculation unit 310 and the threshold IPS determined by experiments as described later.<sub>TH</sub>IPS value for each power peak output from the threshold storage unit 314 of inter-pulse similarity for storing the IPS, the IPS calculation unit 310, and the threshold IPS stored in the threshold storage unit 314.<sub>TH</sub>Compare with and threshold IPS<sub>TH</sub>IPS comparison unit 312 for selecting only power peaks exceeding the above and outputting peak position information, and adjacent (or close within a predetermined search range) based on the peak position information output from IPS comparison unit 312. It includes a VF segment determination unit 316 for merging frames existing between pulses with high IPS values as VF segments and outputting VF section information 132.
As described above, the IPS value calculated by the IPS calculation unit 310 is calculated by the cross-correlation function between the waveform near the power peak to be processed and the waveform near the power peak before that. The frame length for cross-correlation calculation is limited to 15 ms. This is to avoid interference in the similarity calculation due to glottic pulses with irregular intervals.
The cross-correlation is estimated for a width of 5 milliseconds centered on the power peak position, and the maximum value is used as the IPS value. If the IPS value is high, it is considered that there is a high probability that the power peak represents a VF pulse. In the calculation of the IPS value, other power peaks are searched only in the range of 100 milliseconds before the target power peak, and the cross-correlation with the power peak is calculated. A value of 100 milliseconds corresponds to the maximum possible time interval between the excitation pulses of the two glottis. The maximum value of the excitation pulse is a value corresponding to a very low value of 10 Hz as the fundamental frequency.
Figures 10 (A) and 10 (B) show the distribution of experimentally calculated IPS values for the VF segment and NF segment, respectively, as hatched bar graphs. According to Fig. 10 (A), most of the VF segments have large IPS values, and they are concentrated in the range of 0.8 to 0.95. On the other hand, in the NF segment, null_2 has a large value. "Null_2" is set to a null value because the search range is limited to 100 milliseconds, that is, the IPS value is set because there are no other power peaks in the range of 100 milliseconds immediately before the power peak. Indicates what is set to null. On the other hand, in FIG. 10 (A), there is almost no null value of the IPS value.
Also, referring to Fig. 10 (B), the IPS values can be divided into two groups in the NF segment. One is a group with a low IPS value, and the other is a group with a high IPS value. These high IPS values are probably the result of periodicity in the ground voice. Therefore, the IFP value should also be high in this case. Correspondingly, the white bar graph in Fig. 10 (B) shows that there are many NF segments with high IFP values.
<Operation> The automatic dialogue system 100 having the above-described configuration, particularly the VF detection device 122, operates as follows. With reference to FIG. 1, the utterance signal 102 input from the microphone or the like is digitized and given to the voice recognition device 120 and the VF detection device 122. The voice recognition device 120 performs voice recognition processing on the voice signal, and gives the response creation device 124 a voice recognition result 130 composed of text information of a plurality of voice recognition results having a high possibility. On the other hand, the VF detection device 122 performs the operation as described below to identify a frame considered to be a VF segment in the voice signal, and gives the VF section information 132 to the response creation device 124.
The response creation device 124 accesses the knowledge base 126 by using a plurality of candidates included in the voice recognition result 130 given by the voice recognition device 120 and the VF section information 132 given by the VF detection device 122. , Create the most appropriate response as a response from the combination of the speech recognition result candidates and the VF segment. This response consists of the text information of the response and the information that specifies the voice quality of the response voice, and is given to the speech synthesizer 128. The voice synthesizer 128 synthesizes a voice signal 104 for reproducing the specified text information with the specified voice quality and gives it to the speaker.
The operation of the VF detection device 122 will be described below. With reference to FIG. 2, the utterance signal 102 given to the VF detector 122 is given to the bandpass filter 160. The bandpass filter 160 passes only the frequency components of 100 Hz to 1500 Hz of the utterance signal 102 as the utterance signal 174. The utterance signal 174 is given to the ultra-short-term peak detection processing unit 162, the short-term periodicity detection unit 164, and the similarity inspection unit 168.
The ultra-short-term peak detection processing unit 162 detects the peak of power in the ultra-short-term frame by the following processing, and gives it to the periodic inspection unit 166 as peak position information 170. That is, referring to FIG. 3, the framing processing unit 190 frames the utterance signal 174 of the frequency component of 100 to 1500 Hz by an ultra-short-term frame. This ultra-short frame has a frame length of 5 ms and a frame interval of 2.5 ms. The audio signal framed by the ultra-short-term frame is given to the ultra-short-term power calculation unit 192.
The ultra-short-term power calculation unit 192 calculates the ultra-short-term power for each frame, gives the result to the memory 194, and stores it. The memory 194 stores the value of the ultra-short-term power for the latest predetermined number of frames.
The peak comparison unit 196 has a power threshold Pw for each frame as compared with the two frames before and after the frame.<sub>TH</sub>A larger frame is used as a power peak candidate, and peak position information 170 indicating the frame position is output and given to the periodic inspection unit 166.
On the other hand, the short-term periodicity detection unit 164 shown in FIG. 2 detects the periodicity in each frame as follows and gives it to the periodicity inspection unit 166 as short-term periodicity information 172. That is, referring to FIG. 7, the framing processing unit 250 frames the utterance signal with a frame length of 32 milliseconds and a frame interval of 10 milliseconds, and stores it in the memory 252.
The IFP calculation unit 254 calculates the IFP value for each frame stored in the memory 252 and gives it to the periodicity determination unit 258. The periodicity determination unit 258 corrects the IFP value of each frame given by the IFP calculation unit 254 by comparing it with the threshold value function. That is, the periodicity determination unit 258 sets the IFP value of the low harmonic of each frame to null if any of the IFP values of the low harmonics is smaller than the threshold value. The periodicity determination unit 258 gives this IFP value to the continuity inspection unit 260 for each frame.
The continuity inspection unit 260 corrects the IFP value of each frame given by the periodicity determination unit 258 to null if the frames whose values are not null are not continuous for at least three frames. .. The IFP value of each frame after the continuity is inspected by the continuity inspection unit 260 is given to the periodicity inspection unit 166 shown in FIG. 2 as short-term periodicity information 172.
In the periodic inspection unit 166, the IFP value of the frame becomes null due to the short-term periodicity information 172 given by the short-term periodicity detection unit 164 among the peak position information 170 given by the ultra-short-term peak detection processing unit 162. Only the part that is used is used as a candidate for the VF segment, and is given to the similarity inspection unit 168 as VF candidate information 176.
With reference to FIG. 11, the IPS calculation unit 310 of the similarity inspection unit 168 sets the waveform near each power peak and the waveform near the power peak before the power peak candidate specified by the VF candidate information 176. The IPS value between them is calculated and given to the IPS comparison unit 312. The IPS comparison unit 312 contains an IPS value for each power peak calculated by the IPS calculation unit 310 and a threshold IPS stored in the threshold storage unit 314.<sub>TH</sub>Compare with and threshold IPS<sub>TH</sub>Only power peaks that exceed the above are selected, and peak position information is output. This peak position information is given to the VF segment determination unit 316. Based on the peak position information output from the IPS comparison unit 312, the VF segment determination unit 316 sets the frame between adjacent (or adjacent within a predetermined search range) pulses having a high IPS value as the VF segment. Merge and output VF section information 132. This VF section information 132 is given to the response creation device 124 shown in FIG.
<Evaluation of automatic detection> The automatic detection of the VF of the VF detection device 122 according to the above embodiment is performed by the duration of the automatically detected VF segment (VF).<sub>dur</sub>) And the period (VF) that was manually judged and labeled as VF.<sub>dur_human</sub>) Was evaluated. Below, VF<sub>dur</sub>And VF<sub>dur_human</sub>The ratio with is called the VF rate. For the segment labeled as VF, it was determined that it was detected accurately only when the VF rate was greater than 2/3. The number of segments that were not labeled as VF and were determined to be VF by automatic detection (VF)<sub>dur_ins</sub>) Was checked for insertion errors. The detection results and insertion error results were divided into two groups, "detection" and "detection?", Depending on the detection performance or the severity of the insertion error. The "Detected?" Group includes segments detected as "VF" with a VF rate in the range of 1/3 to 2/3, and "VF".<sub>dur_ins</sub>The value of "" is less than 30 milliseconds.
For the various parameters included in the above embodiments, some combinations of values were tested to reduce insertion errors without degrading detection performance. First, the power peak threshold was reset by setting the IPS value to 0.0 and the IFP value to 1.0. This condition corresponds to using only information about power. FIG. 12 shows the detection results when the power threshold value is changed variously. See Figure 12, increasing the power threshold reduces insertion errors (black and shaded areas in the NF group), but also reduces detection rates (black and shaded areas in the VF group). You can see (the hanging part).
Next, the power threshold was fixed at 7 dB and the IPS threshold was set at 0.0. FIG. 13 shows the detection results for various IFP thresholds under this condition. With reference to Figure 13, the detection rate did not change much (VF group), but an IFP threshold of 0.6 could further reduce insertion errors (NF group).
Finally, the power threshold was set to 7 dB and the IFP threshold was set to 0.6, and experiments were conducted on several IPS value thresholds. With reference to Figure 14, setting the IPS value threshold to 0.6 was able to further reduce serious insertion errors (black areas in the "NF" group) and keep the detection rate at a favorable value. We were able to.
For the "R" group (segments in which VF features were not perceived by humans), most of those samples were not detected as VF by automatic detection. However, in the "VF?" Group, some were detected as "VF". Based on these results, it can be said that the VF automatic detection device according to the present embodiment has obtained results that are almost consistent with the results of human perception experiments.
VF for overall detection rate<sub>dur</sub>Total of VF<sub>dur_human</sub>It was calculated by dividing by the total of. For the overall insertion error rate, VF<sub>dur_ins</sub>Total of VF<sub>dur_human</sub>It was calculated by dividing by the total of. For the combination of parameters "power = 7dB, IFP = 0.6, IPS = 0.6", the overall detection rate was 73.3% and the overall insertion error rate was 3.9%. There is room for further improvement in the detection rate of 73.3% by post-processing the detection results. For example, it seems that improvement can be achieved by merging adjacent VF segments. For applications where a slightly higher insertion error rate does not cause any problems, the parameters can be further adjusted to increase the detection rate.
As described above, according to the present embodiment, the vocal fly can be automatically detected by using the combination of the parameters "power, IFP and IPS".
<Realization and operation by computer> The VF detection device 122 and the automatic dialogue system 100 according to this embodiment can be realized by computer hardware, a program executed by the computer hardware, and data stored in the computer hardware. FIG. 15 shows the appearance of the computer system 330, and FIG. 16 shows the internal configuration of the computer system 330.
With reference to FIG. 15, the computer system 330 includes a computer 340 with an FD (flexible disk) drive 352 and a CD-ROM (compact disc read-only memory) drive 350, a keyboard 346, a mouse 348, and a monitor 342. , Includes a microphone 370 and a speaker 372.
With reference to FIG. 16, the computer 340 has a CPU (central processing unit) 356 in addition to the FD drive 352 and the CD-ROM drive 350, and a bus 366 connected to the CPU 356, the FD drive 352 and the CD-ROM drive 350. , A read-only memory (ROM) 358 that stores the boot-up program, etc., a random access memory (RAM) 360 that is connected to the bus 366 and stores program instructions, system programs, work data, etc., and input from the microphone 370. It includes a sound board 368 for digitizing the spoken signal to be output and analogizing the digital audio signal processed by the CPU 356 and giving it to the speaker 372. The computer system 330 may further include a printer (not shown).
Although not shown here, computer 340 may further include a network adapter board that provides connectivity to a local area network (LAN).
The computer program for causing the computer system 330 to operate as the automatic dialogue system 100 and the VF detection device 122 according to the present embodiment is in the CD-ROM 362 or FD 364 inserted in the CD-ROM drive 350 or the FD drive 352. It is stored and then transferred to the hard disk 354. Alternatively, the program may be sent to the computer 340 via a network (not shown) and stored on the hard disk 354. The program is loaded into RAM360 at run time. Programs may be loaded directly into RAM360 from CD-ROM362, from FD364, or over the network.
This program includes a plurality of instructions that cause the computer 340 to operate as the automatic dialogue system 100 and the VF detection device 122 according to this embodiment. Some of the basic functionality required to perform these instructions is provided by an operating system (OS) or third-party program running on the computer 340, or by modules in various toolkits installed on the computer 340. .. Therefore, this program does not necessarily include all the functions necessary to realize the operation as the automatic dialogue system 100 and the VF detection device 122 of this embodiment. This program executes the operation as the above-mentioned automatic dialogue system 100 and VF detection device 122 by calling an appropriate function or "tool" among the instructions in a controlled manner so as to obtain a desired result. It only needs to contain the instructions. The operation of computer system 330 is well known and will not be repeated here.
The embodiments disclosed this time are merely examples, and the present invention is not limited to the above-described embodiments. The scope of the present invention is indicated by each claim of the claims, taking into account the description of the detailed description of the invention, and all changes within the meaning and scope equivalent to the wording described therein. Including.
<figref num="1">It is a block diagram of the automatic dialogue system 100 which adopted the VF detection apparatus 122 which concerns on one Embodiment of this invention.</figref><figref num="2">It is a block diagram of the VF detection device 122 which concerns on one Embodiment of this invention.</figref><figref num="3">It is a block diagram of the ultra-short-term peak detection processing unit 162.</figref><figref num="4">It is a figure which shows the principle of the peak detection in the ultra-short-term peak detection processing part 162.</figref><figref num="5">It is a figure which shows the principle of the peak detection in the ultra-short-term peak detection processing part 162.</figref><figref num="6">It is a graph which shows the result obtained by an experiment about the distribution of the power rise and power fall of a peak in a VF segment and an NF segment.</figref><figref num="7">It is a block diagram of the short-term periodicity detection unit 164.</figref><figref num="8">It is a figure which shows the attribute of the autocorrelation function of a low harmonic when one VF pulse exists in one frame.</figref><figref num="9">It is a figure which shows the attribute of the autocorrelation function of the low harmonic about the ground voice.</figref><figref num="10">It is a graph which shows the distribution of IFP and IPS in the VF and NF segments.</figref><figref num="11">It is a block diagram of the similarity inspection unit 168.</figref><figref num="12">It is a graph which shows the experimental result performed for some power thresholds when IFP threshold = 1 and IPS threshold = 0 are fixed.</figref><figref num="13">It is a graph which shows the experimental result which performed about the threshold value of some IFP when the threshold value of power = 7dB and the threshold value of IPS = 0 was fixed.</figref><figref num="14">It is a graph which shows the experimental result performed for some IPS threshold values when the power threshold value = 7 dB and the IFP threshold value = 0.6 are fixed.</figref><figref num="15">It is a figure which shows the appearance of the computer which realizes the automatic dialogue system 100 and VF detection apparatus 122 which concerns on one Embodiment of this invention.</figref><figref num="16">It is an internal block diagram of the computer shown in FIG.</figref>
Code description
100 automatic dialogue system 102,174 utterance signal 104 audio signal 120 voice recognition device 122 VF detector 124 Response maker 126 Knowledge base 128 Speech synthesizer 130 Speech recognition results 132 VF section information 160 bandpass filter 162 Ultra-short-term peak detection processing unit 164 Short-term periodicity detector 166 Periodic Inspection Department 168 Similarity Inspection Department 170 Peak location information 172 Short-term periodicity information 176 VF candidate information 190,250 Frame processing unit 192 Ultra-short-term power calculation unit 194,252 memory 196 Peak comparison section 254 IFP Calculation Department 258 Periodicity judgment unit 260 Continuity Inspection Department 310 IPS calculation unit 312 IPS comparison section 314 Threshold storage 316 VF segment determination unit
16 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16
Every citation, both ways
| Reference | Relation |
|---|---|
| ドナ・エリクソン,”表現豊かな音声-その生成・知覚と音声合成への応用-”,日本音響学会誌,Vol.61,No.6(2005-06),pp.346-351 | Non-patent |
| 前川喜久雄他,”パラ言語情報の生成と知覚多次元尺度法による布置と音響特徴の関係”,電子情報通信学会技術研究報告,Vol.99,No.74,SP99-10(1999-05),pp.9-16 | Non-patent |
5 members in 3 offices
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 2005250454 | Japan | A | |
| JP20050250454 | – | – | – |
Members5
| Document | Office | Kind | |
|---|---|---|---|
| WO2007026436A1 | World Intellectual Property Organization (WIPO) | A1 | |
| JP2007065226A | Japan | A | |
| US2009089051A1 | United States of America | A1 | |
| JP4736632B2This record | Japan | B2 | |
| US8086449B2 | United States of America | B2 |
19 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 4736632
- Publication, DOCDB
- 4736632
- Publication, EPODOC
- JP4736632B
- Application
- 250454
- Application, DOCDB
- 2005250454
- Application, EPODOC
- JP20050250454
Titles2
- Japanese
- ボーカル・フライ検出装置及びコンピュータプログラム
- English
- Vocal fly detector and computer program
Classification
- CPC, 1
- G10L25/90
- IPC, 8
- G10L15 10
- G10L25 03
- G10L25 21
- G10L25 63
- G10L25 78
- G10L25 84
- G10L25 90
- G10L11 00