Voice recognition device and learning method and learning device to be used in the same device and recording medium on which the same method is programmed and recorded
Abstract
(57) A summary and subject An improvement of the frequency characteristic which a bone conduction microphone has is aimed at, and it uses for speech recognition. It unites and improvement in the recognition performance under noise environment is aimed at. Solution means Using the first code book 110, choose a feature vector and the audio input pattern (bone conduction sound) recorded by first 受音器 101 is outputted, The correction vector memorized by the second code book 121 corresponding to the index is chosen, The feature vector of the sound (air conduction sound) recorded by second 受音器 112 to which the degree of 受音感 is secured by a frequency band larger than the first 受音器 of the above is presumed by adding both the above-mentioned vectors and connecting. Moreover, the speech parameter stored in the voice dictionary 213 as a pattern for reference is given one by one, using the presumed sound as a voice recognition object, and both patterns are compared.
Term
Term ended
Projected expiry passed 24 February 2019, 7.6 years ago.
- Priority and filed
- Published
- Projected expiry
- Today
10 claims: 5 independent, 5 dependent
- 1[Claims] [Claim 1] A means for receiving a sound signal by the first sound receiver and extracting a feature vector from the received sound signal for each frame having a predetermined time length, and a means for temporarily storing the extracted feature vector. A means for storing a representative finite number of characteristic vectors extracted from a sound signal pre-received by the first sound receiver as a first set, and a sound pre-received by the first sound receiver. A representative calculated using the difference between the characteristic vector of the signal and the characteristic vector of the sound signal previously received by the second receiver whose sound receiving sensitivity is ensured in a wider frequency band than the first receiver. A means for storing a finite number of correction vectors as a second set, a means for associating a feature vector belonging to the first set with each correction vector belonging to the second set, and the first set. For each of the feature vectors to which it belongs, a means for calculating the similarity of the feature vectors extracted from the audio signal received by the first sound receiver, and the feature vector having the highest degree of similarity are set in the first set. A means for extracting a correction vector belonging to the second set corresponding to this vector selected from the above, and the extraction for a feature vector extracted from an audio signal received by the first sound receiver. A means for calculating a feature vector generated by adding the corrected correction vectors for each frame, and a means for collating the similarity between the feature vector series stored as a dictionary in advance with respect to the feature vector series. , A sound recognition device including a means for outputting information of a dictionary having the highest degree of similarity from among the collated sounds. 【特許請求の範囲】 【請求項1】 音声信号を第一の受音器により受音し、受音した音声信号から、予め定めた時間長のフレーム毎に特徴ベクトルを抽出する手段と、抽出された特徴ベクトルを一時的に記憶する手段と、第一の受音器で予め受音した音声信号から抽出した代表的な有限個の特徴ベクトルを第一のセットとして記憶する手段と、前記第一の受音器で予め受音した音声信号の特徴ベクトルと前記第一の受音器よりも広い周波数帯域で受音感度が確保される第二の受音器で予め受音した音声信号の特徴ベクトルとの差分を用いて算出した代表的な有限個の補正ベクトルを第二のセットとして記憶する手段と、前記第一のセットに属する特徴ベクトルと前記第二のセットに属する各々の補正ベクトルを対応付ける手段と、前記第一のセットに属する各々の特徴ベクトルに対して、前記第一の受音器で受音した音声信号から抽出された特徴ベクトルの類似度を算出する手段と、類似度の最も高い特徴ベクトルを前記第一のセットの中から選択し、このベクトルに対応する前記第二のセットに属する補正ベクトルを抽出する手段と、前記第一の受音器で受音した音声信号から抽出された特徴ベクトルに対して前記抽出された補正ベクトルを加算して生成される特徴ベクトルをフレーム毎に算出する手段と、この特徴ベクトルの系列に対し、予め辞書として記憶された特徴ベクトル系列との間で類似度を照合する手段と、照合された中から最も類似度の高い辞書の情報を出力する手段とを備えることを特徴とする音声認識装置。
- 4A sound receiving pattern recorded by a first sound receiving device and a second sound receiving device whose sound receiving sensitivity is ensured in a wider frequency band than the first sound receiving device has a predetermined length. The feature amount is calculated by cutting out for each section, and the feature amount extracted through one of the receivers is compared with the code vector stored in the codebook to output the index of the feature vector with the highest degree of similarity. , A codebook learning method of a speech recognition device, characterized in that a correction vector stored in the other codebook corresponding to this index is output. 【請求項4】 第一の受音器と、該第一の受音器より広い周波数帯域で受音感度が確保される第二の受音器とで収録された受音パターンを所定長の区間毎に切り出して特徴量を算出し、一方の受音器を介して抽出された特徴量をコードブックに記憶されたコードベクトルと比較することにより最も類似度の高い特徴ベクトルのインデックスを出力し、このインデックスに対応する他方のコードブックに記憶された補正ベクトルを出力することを特徴とする音声認識装置のコードブック学習方法。
- 7The sound receiving patterns recorded by the first sound receiving device and the second sound receiving device whose sound receiving sensitivity is ensured in a wider frequency band than the first sound receiving device are cut out at intervals of a predetermined time length. By calculating the feature quantity, extracting the feature vector by referring to the code book, and collating the similarity with the code vector stored in advance as a dictionary, the highest degree of similarity is obtained from the collated sounds. In a voice recognition device that outputs dictionary information as a recognition result, a means for receiving a voice signal by the first sound receiver and extracting a feature vector from the received voice signal for each frame having a predetermined time length. A means for temporarily storing the extracted feature vectors, and a means for storing as a first set a finite number of representative feature vectors extracted from the voice signal previously received by the first sound receiver. A typical finite number of correction vectors calculated by using the difference between the feature vector of the voice signal previously received by the first receiver and the feature vector of the voice signal previously received by the second receiver. The sound recognition device is provided with a means for storing the sound as a second set and a means for associating each feature vector belonging to the first set with each correction vector belonging to the second set. Codebook learning device. 【請求項7】 第一の受音器と、該第一の受音器より広い周波数帯域で受音感度が確保される第二の受音器で収録される受音パターンを所定の時間長の区間毎に切り出して特徴量を算出し、コードブックを参照することによって特徴ベクトルを抽出し、予め辞書として記憶されたコードベクトルとの間で類似度を照合することにより、照合された中から最も類似度の高い辞書情報を認識結果として出力する音声認識装置において、音声信号を第一の受音器により受音し、受音した音声信号から、所定の時間長のフレーム毎に特徴ベクトルを抽出する手段と、抽出された特徴ベクトルを一時的に記憶する手段と、前記第一の受音器で予め受音した音声信号から抽出した代表的な有限個の特徴ベクトルを第一のセットとして記憶する手段と、前記第一の受音器で予め受音した音声信号の特徴ベクトルと前記第二の受音器で予め受音した音声信号の特徴ベクトルとの差分を用い算出した代表的な有限個の補正ベクトルを第二のセットとして記憶する手段と、前記第一のセットに属する各々の特徴ベクトルと前記第二のセットに属する各々の補正ベクトルを対応付ける手段とを具備することを特徴とする音声認識装置のコードブック学習装置。
- 9The sound receiving patterns recorded by the first sound receiving device and the second sound receiving device whose sound receiving sensitivity is ensured in a wider frequency band than the first sound receiving device are cut out for each section of a predetermined time length. It was used in the codebook learning device of the voice recognition device that calculates the feature amount and extracts the feature vector by referring to the codebook, and received the sound by the first sound receiver and the second sound receiver. A step of converting an analog signal into a digital signal at an appropriate sampling frequency and storing it in a sampling data buffer prepared for each, and a step of calculating a data feature amount for each frame of the data stored in each sampling data buffer. And the step of determining the start and end frames of the voice by comparing the power calculated for each frame with the preset threshold, and the start frame of the voice while the appropriate word is spoken. Based on the information of the end frame and the step of storing only the feature amount in the range as the feature vector in the first and second vector buffers, and the feature vector stored in the first vector buffer, a representative feature vector is obtained. The step of generating and storing in the first codebook and the feature vector stored in the first vector buffer are frame-by-frame vectors based on the pre-generated feature vector and stored in the first codebook. A step of quantization, a step of assigning an index of the feature vector having the highest similarity among the feature vectors of the first codebook to the first feature vector, and a step of being stored in the second vector buffer. Corresponds to the step of calculating the difference between the feature vector and the feature vector stored in the first vector buffer for each frame and storing the difference in the feature difference data buffer, and the feature vector stored in the first vector buffer. The step of assigning the same index as the index assigned to the first feature vector to the feature difference data to be performed, and the feature difference data clustered for each index assigned to the feature difference data, and the data included in the cluster. To averageA recording medium in which a step of generating a typical feature correction vector and storing it in the second codebook is recorded. 【請求項9】 第一の受音器と、該第一の受音器より広い周波数帯域で受音感度が確保される第二の受音器で収録された受音パターンを所定時間長の区間毎に切り出して特徴量を算出し、コードブックを参照することによって特徴ベクトルを抽出する音声認識装置のコードブック学習装置に用いられ、前記第一の受音器と、第二の受音器で受音されたアナログ信号を適切なサンプリング周波数でディジタル信号に変換し、それぞれに用意されるサンプリングデータバッファに格納するステップと、それぞれのサンプリングデータバッファに格納されたデータをフレーム毎にデータの特徴量を算出するステップと、フレーム毎に算出されるパワーと事前に設定された閾値とを比較することにより、音声の開始フレーム及び終了フレームを決定するステップと、適切な単語の発声がある間、前記音声の開始フレームと終了フレームの情報に基づき当該範囲の特徴量のみを特徴ベクトルとして第一・第二のベクトルバッファに格納するステップと、前記第一のベクトルバッファに格納された特徴ベクトルから代表的な特徴ベクトルを生成し、第一のコードブックに格納するステップと、前記第一のベクトルバッファに格納された特徴ベクトルを、事前に生成され前記第一のコードブックに格納された特徴ベクトルに基づいてフレーム毎ベクトル量子化するステップと、前記第一のコードブックの特徴ベクトルの中で最も類似度の高い特徴ベクトルのインデックスを前記第一の特徴ベクトルに付与するステップと、前記第二のベクトルバッファに格納された特徴ベクトルと第一のベクトルバッファに格納された特徴ベクトルとの差分をフレーム毎算出し、その差分を特徴差分データバッファに格納するステップと、前記第一のベクトルバッファに格納された特徴ベクトルに対応する特徴差分データに対し、前記第一の特徴ベクトルに付与されたインデックスと同じインデックスを付与するステップと、前記特徴差分データに付与されたインデックス毎に特徴差分データをクラスタリングし、クラスタに含まれるデータを平均化することにより代表的な特徴補正ベクトルを生成して前記第二のコードブックに格納するステップが記録された記録媒体。
- 10The voice input pattern recorded via the sound receiver is cut out for each section of a predetermined time length, the feature amount is extracted, the feature vector is extracted by referring to the code book, and the code vector stored in advance as a dictionary is used. It is used in a voice recognition device that outputs dictionary information with the highest degree of similarity by collating the similarity between the two, and the analog signal received by the sound receiver is measured at an appropriate sampling frequency. The step of converting to a digital signal and storing it in the sampling data buffer, the signal power is calculated each time the signal for one frame is stored in the sampling data buffer, and the power calculated for each frame is spread over an appropriate number of frames. The start frame of the sound is obtained by comparing the step of cumulatively adding and calculating the average value per frame to use as the threshold for detecting the voice section with the power calculated for each frame and the preset threshold. And the step of determining the end frame, the step of extracting the feature amount of the range as a vector based on the information of the start frame and the end frame and storing it in the feature vector buffer, and the step of storing the feature vector in the feature vector buffer in advance. A step of vector quantization for each frame based on the feature vector stored in the first codebook, and an index assigned to the feature vector having the highest similarity among the feature vectors of the first codebook. The step of transferring the feature vector and the correction vector corresponding to the transferred index are extracted from the feature vector stored in the second codebook in advance, and the extracted correction vector is transferred to the transferred feature vector. The step of estimating the feature vector stored in the second codebook by adding the data and the feature vector obtained here are used as input patterns and are registered in the speech dictionary as reference patterns in advance, and are to be recognized. The feature parameters from the voice start frame to the voice end frame of each word are given in sequence, the distance value between the step of collating both patterns for each word and the input pattern for each reference pattern is calculated, and all the distances.A recording medium in which a step of outputting a reference pattern name corresponding to the smallest distance value among the values as a recognition result is recorded. 【請求項10】 受音器を介して収録される音声入力パターンを所定時間長の区間毎に切り出して特徴量を抽出し、コードブックを参照することにより特徴ベクトルを抽出し、予め辞書として記憶されたコードベクトルとの間で類似度を照合することにより、照合された中から最も類似度の高い辞書情報を出力する音声認識装置に用いられ、前記受音器で受音されたアナログ信号を適切なサンプリング周波数によりディジタル信号に変換しサンプリングデータバッファに格納するステップと、サンプリングデータバッファに1フレーム分の信号が格納される毎に信号パワーを算出し、フレーム毎に算出されるパワーを適切なフレーム数に渡って累積加算し、フレームあたりの平均値を計算して音声区間検出のための閾値とするステップと、フレーム毎に算出されるパワーと事前に設定された前記閾値とを比較することにより音声の開始フレーム及び終了フレームを決定するステップと、開始フレームと終了フレームの情報に基づき当該範囲の特徴量をベクトルとして抽出し、特徴ベクトルバッファに格納するステップと、特徴ベクトルバッファに格納された特徴ベクトルを、事前に第一のコードブックに格納された特徴ベクトルに基づいてフレーム毎にベクトル量子化するステップと、前記第一のコードブックの特徴ベクトルの中で最も類似度が高い特徴ベクトルに付与されるインデックスとその特徴ベクトルを転送するステップと、事前に第二のコードブックに格納された特徴ベクトルから前記転送されたインデックスに相当する補正ベクトルを抽出し、この抽出された補正ベクトルを転送された特徴ベクトルに加算することにより前記第二のコードブックに格納される特徴ベクトルの推定を行うステップと、ここで得られる特徴ベクトルを入力パターンとし、参照パターンとして音声辞書に予め登録されてある、認識対象となる各単語の音声開始フレームから音声終了フレームまでの特徴パラメータを順次与え、単語毎両パターンの照合を行なうステップと、各参照パターン毎入力パターンとの距離値を算出し、全ての距離値の中で最小となる距離値に対応する参照パターン名を認識結果として出力するステップが記録された記録媒体。
Independent claims5
61 paragraphs in 1 section, as filed
Description: TECHNICAL FIELD [Detailed description of the invention]
【0001】
[Technical field to which the invention belongs]
The present invention relates to a voice recognition device having high recognition performance even in a noisy environment, a learning method and a learning device used for the device, and a recording medium on which the method is programmed and recorded.
【0002】
[Conventional technology]
Conventionally, a method for realizing a voice recognition device having high recognition performance even in a noisy environment has been proposed. For example, (1) a method of measuring the noise at the time when there is no voice input in advance and reducing the component of the measured noise at the time of voice input, (2) noise using two microphones for voice input and noise input. How to reduce the component of the input signal to the input microphone from the input signal to the voice input microphone, (3) How to learn the dictionary of the voice recognition device with the noise measured in advance, (4) Bone conduction microphone as a sound receiver There was a way to use, etc. However, any of the above-mentioned conventional methods leaves a problem for improving the recognition performance as shown below. Specifically, the method shown in (1) is less effective except when the noise property is always constant. In addition, although the method shown in (2) has some effect regardless of the nature of the noise, if the microphones are installed too close to each other, the sound will be mixed into the noise microphone and the sound will be mixed with the noise component. Even some of the ingredients are reduced. On the other hand, if the microphones are placed too far apart from each other, the nature of the noise input to both microphones will be different, and the noise component cannot be accurately subtracted. Further, there are various problems such as an increase in the scale of the device due to the installation of a plurality of microphones and a limitation on the position of the speaker. In addition, the method shown in (3) is less effective if the noise properties during learning and recognition are different. The method shown in (4) has an advantage that it is not easily affected by noise in principle, but has a problem that voice information is lost because the frequency band of the received voice is narrow.
【0003】
[Problems to be Solved by the Invention]
On the other hand, a method using a filter has been proposed as a method of correcting the difference in microphone characteristics of a voice input microphone and the difference in transmission line characteristics from voice input to voice recognition. Therefore, the characteristics of a bone-conducting microphone with a narrow audio frequency band (compared to an air-conducting microphone that can receive over a wide frequency band, the receivable frequency band is limited, but the influence of noise components propagating in the air is small). There is also a method of using a filter that corrects the characteristics of a microphone with a wide frequency band, but the current situation is that sufficient effects have not been obtained for practical use. The present invention has been made based on the above circumstances, and by using a bone conduction microphone that is not easily affected by noise as a sound receiver and bringing the frequency characteristics closer to the characteristics of the air conduction microphone, speech recognition in a noisy environment The improvement in performance can be easily applied to a conventional voice recognition device without limiting the position of the speaker, without increasing the scale of the device, and the learning method and learning used by the speech recognition device and the device. It is an object of the present invention to provide a device and a recording medium on which the method is programmed and recorded.
【0004】
[Means for solving problems]
The voice recognition device of the present invention is a means for receiving a sound signal by the first sound receiver and extracting a feature vector from the received sound signal for each frame having a predetermined time length, and the extracted features. A means for temporarily storing a vector, a means for storing a finite number of representative feature vectors extracted from an audio signal pre-received by the first sound receiver as a first set, and the first receiving device. A feature vector of a sound signal pre-received by a sound device and a feature vector of a sound signal pre-received by a second sound receiver whose sound receiving sensitivity is ensured in a wider frequency band than that of the first sound receiver. A means for storing a representative finite number of correction vectors calculated by using the difference between the above as a second set, and a means for associating a feature vector belonging to the first set with each correction vector belonging to the second set. And a means for calculating the similarity of the feature vectors extracted from the audio signal received by the first sound receiver for each feature vector belonging to the first set, and the highest degree of similarity. A feature vector is selected from the first set, a means for extracting a correction vector belonging to the second set corresponding to this vector, and a sound signal received by the first sound receiver are extracted. Between the means for calculating the feature vector generated by adding the extracted correction vector to the feature vector for each frame and the feature vector series stored in advance as a dictionary for the series of the feature vectors. It is characterized in that it includes a means for collating the similarity in the above and a means for outputting the information of the dictionary having the highest degree of similarity among the collated.
【0005】
The codebook learning method of the voice recognition device of the present invention is recorded by a first sound receiver and a second sound receiver whose sound reception sensitivity is ensured in a frequency band wider than that of the first sound receiver. The feature amount is calculated by cutting out the received sound pattern for each section of a predetermined length, and the feature amount extracted through one of the sound receivers is compared with the code vector stored in the codebook to obtain the highest degree of similarity. It is characterized by outputting an index of a high feature vector and outputting a correction vector stored in the other codebook corresponding to this index.
【0006】
The codebook learning device of the voice recognition device of the present invention is recorded by a first sound receiver and a second sound receiver whose sound reception sensitivity is ensured in a frequency band wider than that of the first sound receiver. The sound receiving pattern is cut out for each section of a predetermined time length, the feature amount is calculated, the feature vector is extracted by referring to the code book, and the similarity with the code vector stored in advance as a dictionary is collated. As a result, in the voice recognition device that outputs the dictionary information having the highest degree of similarity from the collated as the recognition result, the voice signal is received by the first sound receiver, and the received voice signal is used for a predetermined time. A means for extracting a feature vector for each long frame, a means for temporarily storing the extracted feature vector, and a typical finite number of representative finite numbers extracted from a sound signal pre-received by the first sound receiver. A means for storing the feature vector as a first set, a feature vector of a sound signal pre-received by the first sound receiver, and a feature vector of a sound signal pre-received by the second sound receiver. A means for storing a representative finite number of correction vectors calculated using differences as a second set, and a means for associating each feature vector belonging to the first set with each correction vector belonging to the second set. It is characterized by having.
【0007】
The recording medium of the present invention has a sound receiving pattern recorded by a first sound receiving device and a second sound receiving device whose sound receiving sensitivity is ensured in a wider frequency band than the first sound receiving device for a predetermined time. It is used in the codebook learning device of the voice recognition device that cuts out each long section, calculates the feature amount, and extracts the feature vector by referring to the codebook. A step of converting an analog signal received by a sound device into a digital signal at an appropriate sampling frequency and storing it in a sampling data buffer prepared for each, and data stored in each sampling data buffer for each frame. There is a step of calculating the feature amount of the data, a step of determining the start frame and the end frame of the voice by comparing the power calculated for each frame with the preset threshold, and the utterance of an appropriate word. In the meantime, a step of storing only the feature amount in the range as a feature vector in the first and second vector buffers based on the information of the start frame and the end frame of the voice, and the feature vector stored in the first vector buffer. A step of generating a representative feature vector from the data and storing it in the first codebook, and a feature vector stored in the first vector buffer are generated in advance and stored in the first codebook. The step of performing vector quantization for each frame based on the vector, the step of assigning the index of the feature vector having the highest similarity among the feature vectors of the first codebook to the first feature vector, and the second step. The step of calculating the difference between the feature vector stored in the vector buffer and the feature vector stored in the first vector buffer for each frame and storing the difference in the feature difference data buffer, and in the first vector buffer. A step of assigning the same index as the index assigned to the first feature vector to the feature difference data corresponding to the stored feature vector, and clustering the feature difference data for each index assigned to the feature difference data. And the data included in the clusterIt is characterized in that a step of generating a typical feature correction vector by averaging the data and storing it in the second codebook is recorded.
【0008】
In addition, the voice input pattern recorded via the sound receiver is cut out for each section of a predetermined time length to extract the feature amount, and the feature vector is extracted by referring to the code book, and the code stored in advance as a dictionary. It is used in a voice recognition device that outputs dictionary information with the highest degree of similarity by collating the similarity with a vector, and appropriately samples the analog signal received by the sound receiver. The step of converting to a digital signal according to the frequency and storing it in the sampling data buffer, and calculating the signal power each time the signal for one frame is stored in the sampling data buffer, and adjusting the power calculated for each frame to an appropriate number of frames. By comparing the step of cumulatively adding over and calculating the average value per frame to use as the threshold for detecting the voice section, and the power calculated for each frame with the preset threshold of the voice. The step of determining the start frame and the end frame, the step of extracting the feature amount in the range as a vector based on the information of the start frame and the end frame and storing it in the feature vector buffer, and the feature vector stored in the feature vector buffer. , The step of vector quantization for each frame based on the feature vector stored in the first codebook in advance, and the feature vector having the highest degree of similarity among the feature vectors of the first codebook are given. The step of transferring the index and its feature vector, and the correction vector corresponding to the transferred index is extracted from the feature vector stored in the second codebook in advance, and the extracted correction vector is transferred to the feature. The step of estimating the feature vector stored in the second codebook by adding to the vector, and the feature vector obtained here as an input pattern, which is registered in the speech dictionary as a reference pattern in advance, is a recognition target. The feature parameters from the voice start frame to the voice end frame of each word are sequentially given, and the distance value between the step of collating both patterns for each word and the input pattern for each reference pattern is calculated and all.It is also characterized in that the step of outputting the reference pattern name corresponding to the minimum distance value among the distance values of is output as the recognition result.
【0009】
As a result, by using a bone conduction microphone that is not easily affected by noise as a sound receiver and using a filter that corrects the frequency characteristics, the frequency characteristics can be brought closer to the air conduction voice, and the position of the speaker is limited. Without increasing the scale of the device, the voice recognition performance can be improved in a noisy environment, and it can be easily applied to a conventional voice recognition device.
【0010】
BEST MODE FOR CARRYING OUT THE INVENTION
FIG. 1 is a block diagram showing an embodiment of a codebook learning device for a voice recognition device according to the present invention. In the figure, 101 is a bone conduction microphone and 112 is an air conduction microphone. The air conduction microphone 112 is known to have good sensitivity over a wide frequency band. It can satisfactorily receive acoustic signals in the band of 8 kHz to 12 kHz required for speech recognition. On the other hand, since a noise signal in the same frequency band as the voice can be received without distinguishing it from the voice signal, there is a drawback that it becomes difficult to detect the voice section under high noise. Since the bone conduction microphone 101 uses an acceleration pickup, the frequency band is narrow and the high frequency component of the voice is greatly attenuated. Therefore, when used alone for voice recognition, the performance deteriorates, but it propagates in the air from the outside. It has the advantage that the influence of the noise component is small. In the present invention, the voice input patterns recorded by the bone conduction microphone 101 and the air conduction microphone 112 are cut out at regular time intervals, and the code vectors having the highest degree of similarity as compared with the code books prepared for each are cut out. Vector quantization method (VQ: vector) It will be described below assuming that the voice input pattern is expressed by quantization). A typical example of the method of using the vector stored in the codebook as the centroid in the pattern space is the LBG method (Linde, Y, Buzo, A. and Gray, RM: An Algorithm for vector quantizerdesign. Known as IEEE Trans.Commun., COM-28,1,84-95 (1980)).
【0011】
The audio signals received by the bone conduction microphone 101 and the air conduction microphone 112 are supplied to analog-digital converters (hereinafter, simply referred to as A / D converters) 102 and 113, respectively, and are digitally generated by the A / D converters 102 and 113, respectively. The signals are supplied to the sampling data buffers 103 and 114, respectively. The sampling data buffers 103 and 114 outputs are supplied to the feature extraction unit 104 and the feature extraction unit 115, respectively. Feature extraction unit 104, All the blocks described after 115 are realized by software and are expressed here as functional blocks. The sampling data buffer 103 output is further supplied to the bone conduction voice feature buffer 105 and the air conduction voice feature vector buffer 116 by the paths of the power calculation unit 106 and the voice section detection unit 107. 108 is a changeover switch. The changeover switch 108 connects the output of the bone conduction voice feature vector buffer 105 to the bone conduction voice codebook generation unit 109 or the vector quantization unit 111. Reference numeral 110 denotes a bone conduction voice codebook storage unit that stores the bone conduction voice code generated by the bone conduction voice codebook generation unit 109. On the other hand, the output of the bone conduction voice feature vector buffer 105 is supplied to the feature difference calculation unit 117 in addition to the changeover switch 108. The feature vector differences between the bone conduction voice feature vector buffer 105 and the air conduction voice feature vector buffer 116 calculated by the feature difference calculation unit 117 are supplied to the feature difference data buffer 118. The data supplied to the feature difference data buffer 118 is supplied to the feature correction filter generation unit 120. The difference data supplied to the feature correction filter generation unit 120 is generated as a feature correction vector component by a logic described later, and is supplied to the feature correction filter storage unit 121 and the frame correspondence table storage unit 119. The frame correspondence table storage unit 119 reflects the correction vector generated by the feature correction filter generation unit 120 according to the logic described later in the quantization output of the feature vector in the bone conduction voice output by the vector quantization unit 111.
【0012】
FIG. 3 is a flowchart cited for explaining the operation procedure of the codebook learning device of the voice recognition device shown in FIG. 1, and the procedure is specifically programmed and recorded in the learning device of the present invention. When a CPU (not shown) reads and executes this, the procedure shown below is executed. Hereinafter, the operation of the codebook learning device of the voice recognition device shown in FIG. 1 will be described in detail with reference to the flowchart shown in FIG. The operations are roughly classified into "bone conduction voice codebook generation" and "feature correction vector generation". First, the operation of "generating a bone conduction voice codebook" will be described. The analog signals received by the bone conduction microphone 101 and the air conduction microphone 112 are converted into digital signals at appropriate sampling frequencies by the A / D converters 102 and 113, respectively (steps S31 and S32), and sequentially sent to the sampling data buffers 103 and 114, respectively. Stored (step S33). Here, the appropriate sampling frequency is a frequency that does not impair the characteristics of speech required for speech recognition processing, and is usually set to 8 kHz to 12 kHz. The feature extraction units 104 and 115 calculate the feature amount of the data every time 20 to 30 milliseconds of data is stored in the sampling data buffers 103 and 114. That is, data feature extraction is performed for each frame (step S35). On the other hand, from the sampling data of the bone conduction microphone 101, the signal power is calculated by the power calculation unit 106 for each frame and sent to the voice section detection unit 107. The voice section detection unit 107 determines the start frame and end frame of the voice by comparing the power calculated for each frame with the preset threshold value (step S34). The feature quantity for each frame in the bone conduction microphone 101 and the air conduction microphone 112 is based on the detected start frame and end frame information, and only the feature quantity in the corresponding range is used as a vector for each feature vector buffer 105, Stored in 116 (step S37). This operation is repeated as long as the appropriate word is uttered (step S36).
【0013】
Here, the period during which an appropriate word is uttered is the period during which a group of words with less bias in the utterance frequency of all phonemes is output. In addition, it should be noted here that the feature amount is calculated using the bone conduction voice, and the frame for calculating the power and the frame for calculating the feature amount using the air conduction voice are synchronized. is there. Specifically, synchronization between the frame for calculating the power and the frame for calculating the feature amount using the air conduction voice can be easily realized by synchronizing the sampling clocks of both the A / D converters 102 and 113. .. When generating the bone conduction voice codebook, the changeover switch 108 connects the bone conduction voice feature vector buffer 105 to the bone conduction voice codebook generation unit 109 (step S38). The bone conduction voice codebook generation unit 109 generates a representative feature vector from the feature vector stored in the bone conduction voice feature vector buffer 105 (step S39) and stores it in the bone conduction voice codebook 110 (step S40). .. The typical feature vector described above is a sample in which the feature amount extracted via the voice signal of the bone conduction microphone 101 is accumulated while the appropriate word group is uttered, and the distance between the samples is small. It is obtained by clustering each other and taking the arithmetic mean of the features for each cluster. The representative vector obtained here is recorded and used as a bone conduction voice codebook (bone conduction voice codebook storage unit 110).
【0014】
Next, the operation of "feature correction vector generation" will be described. When the feature correction vector is generated, the changeover switch 108 is connected to the vector quantization unit 111 (step S48). In the vector quantization unit 111, the feature vector stored in the bone conduction voice feature vector buffer 105 is vector-quantized frame by frame based on the feature vector generated in advance and stored in the bone conduction voice codebook storage unit 110 (step). S41), the index (number) of the feature vector having the highest similarity among the feature vectors of the bone conduction voice codebook is assigned to the bone conduction voice feature vector (step S42). On the other hand, the feature difference calculation unit 117 calculates the difference between the feature vector stored in the bone conduction voice feature vector buffer 105 and the feature vector stored in the air conduction voice feature vector buffer 116 for each frame, and calculates the feature difference data buffer 118. Sequentially store in. The frame correspondence table storage unit 117 assigns the same number as the number assigned to the bone conduction voice feature vector to the feature difference data stored in the feature difference data buffer 118 corresponding to the bone conduction voice feature vector. The feature correction filter generation unit 120 clusters the feature difference data for each number assigned to the feature difference data, and generates a typical feature correction vector by averaging the data included in the cluster (step S53). It is stored in the feature correction filter storage unit 121 (step S54). The above-mentioned bone conduction voice codebook and feature correction filter correspond to each number. The above-mentioned feature correction vector is generated from the residual signal (feature difference calculation unit 117) obtained by subtracting the audio signal from the air conduction microphone 112 obtained in synchronization with the audio signal from the bone conduction microphone 101. It is done by extracting the extracted feature amount.
【0015】
FIG. 2 is a block diagram showing an embodiment of the voice recognition device according to the present invention. In the figure, 201 is a bone conduction microphone, 202 is an A / D converter, and 203 is a sampling data buffer. In addition, 204 is a feature extraction unit, 205 is a feature vector buffer, 206 is a power calculation unit, 207 is a voice section detection unit, 208 is a vector quantization unit, 209 is a bone conduction voice codebook storage unit, and 210 is an air conduction voice estimation unit. Part, 211 is a feature correction filter storage unit, 212 is a pattern collation unit, 213 is a speech dictionary storage unit, and 214 is a recognition result output unit, all of which are realized by software and are therefore shown as functional blocks. The details of the operation procedure such as functions will be described later.
【0016】
FIG. 4 is a flowchart cited for explaining the operation procedure of the voice recognition device shown in FIG. 2, in which (a) shows a noise measurement operation and (b) shows a procedure for a voice recognition operation. Specifically, the procedure is programmed and recorded in the speech recognition device of the present invention. When a CPU (not shown) reads and executes this, the procedure shown below is executed. Hereinafter, the operation of the voice recognition device shown in FIG. 2 will be described in detail with reference to the flowcharts shown in FIGS. 4 (a) and 4 (b). The operation of the voice recognition device of the present invention can be roughly classified into "noise measurement" that measures the noise level when no voice is input and determines the threshold value of the voice section, and the uttered voice pattern. Is classified into "speech recognition" that collates with the voice pattern in the registered voice dictionary and outputs the result. First, the "noise measurement" operation will be described. The analog signal received by the bone conduction microphone 201 is converted into a digital signal by the A / D converter 202 and sequentially stored in the sampling data buffer 203 (steps S51 and S52). Every time one frame of signal is stored in the sampling data buffer 203, the power calculation unit 206 calculates the signal power (step S53). Then, the power calculated for each frame is input to the voice section detection unit 207. The voice section detection unit 207 cumulatively adds the power calculated for each frame over an appropriate number of frames, and further calculates the average value per frame. Here, the appropriate number of frames is usually about 4 to 16. By adding an appropriate constant to the calculated average power, it becomes a threshold value for voice interval detection (step S54).
【0017】
Next, the "speech recognition" operation will be described. Voice input becomes possible when the above-mentioned "noise measurement" is completed. First, the signal received by the bone conduction microphone 201 is converted into a digital signal by the A / D converter 202 and sequentially stored in the sampling data buffer 203 (step S61, S62). For each frame, the feature extraction unit 204 calculates the feature amount of the data, and at the same time, the power calculation unit 206 calculates the power of the signal (step S63) and sends it to the voice section detection unit 207. The voice section detection unit 207 determines the start frame and end frame of the voice by comparing the power calculated for each frame with the preset threshold value (step S64). Based on the start frame and end frame information detected here, the feature amount in the range is stored as a vector in the feature vector buffer 205 (step S65). In the vector quantization unit 208, the feature vector stored in the feature vector buffer 205 is vector-quantized frame by frame based on the feature vector stored in the bone conduction voice codebook storage unit 209 in advance (step S66). The number of the feature vector having the highest similarity among the feature vectors of the guide voice codebook and the feature vector stored in the feature vector buffer 205 are transferred to the air guide voice estimation unit 210 (step S67). The air conduction voice estimation unit 210 extracts the correction vector corresponding to the transferred number from the feature vector stored in the feature correction filter storage unit 211 in advance, and adds the extracted correction vector to the transferred feature vector. Estimates to the air-conducted voice feature vector (step S68). The above-mentioned similarity is the distance between the vector consisting of the feature amount of the bone conduction sound and the vector stored in the bone conduction code book. That is, it is a value obtained by adding the squared value of the difference between the two vectors for each element. Here, this similarity is calculated for each vector in the bone conduction codebook, and the smallest vector is selected. The air conduction voice feature vector is estimated by adding the vector selected in this way and obtained from the bone conduction codebook and the correction vector.
【0018】
On the other hand, in the voice dictionary storage unit 213, feature parameters from the voice start frame to the voice end frame of each word to be recognized are registered. Therefore, the air conduction voice feature vector estimated by the air conduction voice estimation unit 210 is given as an input pattern to the pattern matching unit 212, and the voice parameters stored in the voice dictionary storage unit 213 are sequentially given word by word as a reference pattern. As a result, both patterns can be collated (step S69). As a result, the collation result of the input pattern and the reference pattern is output by the distance value. The larger the distance value, the larger the difference between the two patterns. The distance value from the input pattern is calculated for each reference pattern (step S70), and the reference pattern name corresponding to the smallest distance value among all the distance values becomes the recognition result and is displayed on the recognition result display unit 214. (Step S71). The flowcharts shown in FIGS. 3 and 4 are fixedly written in the storage device (not shown) of the learning device and the voice recognition device, respectively, or the magnetic of the semiconductor storage device, floppy disk, hard disk, etc. It is written and distributed as a program on a recording device, CD-ROM, etc., and functions by being taken into a storage device inside the device as needed.
【0019】
The applicant conducted the following speech recognition experiment in order to confirm the effect of the above-described embodiment of the present invention. FIG. 5 is a graph showing the transition of the word speech recognition rate for each number of feature correction filter divisions, which is increased according to the codebook size, in the speech recognition device shown in FIG. The graph shows the relationship between the two, with the number of filter divisions on the X-axis and the speech recognition rate on the Y-axis. For the words, the first 20 cities of the 100 cities of the Electronic Cooperative are selected, and the utterances are uttered twice by two men and women in a predetermined noise environment, and the average at that time is shown as the recognition rate. As a result of the experiment, when a single feature correction filter was used (filter division number 1), the recognition rate was 52.5%, but when the filter division number was increased according to the codebook size, the filter division was performed. The number was 64 and the recognition rate was 80%, 128 was 82.2%, and 256 was 88.5%. This confirmed the effect of the present invention under no noise. Next, the effect of the present invention under noise will be described. In a noisy environment, the word voice recognition performance of the conventional voice recognition device that inputs air-conducted voice and the voice recognition device of the present invention shown in FIG. 2 was compared. The number of divisions of the correction filter used in the voice recognition device of the present invention is 256. In addition, the words were uttered twice in a predetermined environment by selecting the first 20 cities of the 100 electronic cooperative names. As a result, the average value of the voice recognition rate of each of the two men and women was 79% in the device of the present invention, compared with 42.5% in the conventional device when pink noise with a noise of 64 dB was generated from the loudspeaker. It was. In addition, the noise caused environmental noise (unsteady noise, environmental noise 1 along the road, maximum 80 dB, minimum 55 dB, average 66 dB, environmental noise 2 shopping mall, maximum 70 dB, minimum 60 dB, average 64 dB) from the loud speaker. In this case, the conventional devices could not recognize the recognition rate at all as 0%, whereas the device of the present invention had the recognition rates of 82.1% and 80.4%, respectively. This performance comparison confirmed the effect of the present invention in a noisy environment.
【0020】
On the same date, the applicant used a bone conduction microphone that is not easily affected by noise and an air conduction microphone with a wide frequency band, and used the mapping of the feature vector from the bone conduction voice to the air conduction voice to create a noise environment. We have applied for a voice recognition device and a voice learning method in the device and a recording medium in which the device and the method are programmed and recorded in order to improve the voice recognition performance below. On the other hand, the present invention uses a bone conduction microphone that is not easily affected by noise as a sound receiver, and adds a filter that corrects the frequency to bring the frequency characteristics closer to the air conduction voice, so that the frequency characteristics are closer to those of the air conduction voice. This is to improve the voice recognition performance. Therefore, in the learning device of the present invention, the feature difference for calculating the vector difference between the feature vector stored in the bone conduction voice feature vector buffer 105 and the feature vector stored in the air conduction voice feature vector buffer 116 for each frame. A calculation unit 117 and a feature difference data buffer 118 for storing the calculation unit 117 are added, and a number assigned to the feature difference data obtained here (feature data difference data corresponding to the bone conduction voice feature vector in the frame-compatible storage unit). The feature difference data stored in the buffer is given the same number as the number given to the bone conduction voice feature vector), and the feature difference data is clustered, and the data contained in the cluster is averaged to be typical. A feature correction filter generation unit 120 having a logic for generating a feature correction filter is added. Further, in the voice recognition device of the present invention, the feature correction filter storage unit 211 is added, and here, the transfer is performed from the correction filter previously stored in the feature correction filter storage unit 211 by the air conduction voice estimation unit 210. The correction filter corresponding to the number (the number of the feature vector with the highest similarity among the feature vectors of the bone motion speech codebook and the feature vector stored in the feature vector buffer are transferred from the vector quantization unit 208) was extracted and transferred. It has the logic to estimate the air-conducted voice feature vector by adding the extracted correction filter to the feature vector. Due to this, bone conduction
【0021】
[Effect of the invention]
As described above, the present invention outputs a voice input pattern (bone conduction sound) recorded by the first sound receiver by selecting a feature vector using the first codebook, and corresponds to the index. By selecting the correction vector stored in the second codebook, adding the two vectors, and connecting them, the second sound receiving sensitivity is ensured in a wider frequency band than the first sound receiving device. The feature vector of the sound (air conduction sound) recorded by the device is estimated, and the estimated sound is used as the voice recognition target. As a result, the external air that the bone conduction microphone has conventionally had as a feature. Since the acceleration pickup is used while taking advantage of the fact that the influence of the noise component propagating in the inside is small, the frequency characteristics such as the narrow frequency band and the large attenuation of the high frequency component of the sound can be improved. Therefore, this The bone conduction microphone can be used alone as a sound recording microphone for voice recognition. In addition, it has been possible to improve the voice recognition performance in a noisy environment, which can be realized without limiting the position of the speaker and without increasing the scale of the device, and can be easily applied to the conventional voice recognition device. It can be done.
[Simple explanation of drawings]
[Figure 1]
It is a block diagram which shows the embodiment of the learning apparatus of this invention.
[Figure 2]
It is a block diagram which shows the embodiment of the voice recognition apparatus of this invention.
[Fig. 3]
It is a flowchart cited for demonstrating the operation of the Embodiment of this invention shown in FIG.
[Fig. 4]
It is the flowchart cited for demonstrating the operation of the Embodiment of this invention shown in FIG.
[Fig. 5]
It is a graph quoted for demonstrating the effect of embodiment of this invention.
[Explanation of symbols]
101, 201 ... Bone conduction microphone (first receiver), 102, 113, 202 ... Analog-to-digital converter (A / D converter), 103, 114, 203 ... Sampling data buffer, 104, 115, 204 ... feature extraction unit, 105 ... bone conduction audio feature vector buffer, 106, 206 ... power calculation unit, 107, 207 ... audio section detection unit, 108 ... selector switch , 109 ... Bone conduction voice codebook generator, 110, 209 ... Bone conduction voice codebook storage part (first codebook), 111, 208 ... Vector quantization part, 112 ... Qi Guide microphone (second receiver), 116 ... air conduction voice feature vector buffer, 117 ... feature difference calculation unit, 118 ... feature difference data buffer, 119 ... frame correspondence table storage unit, 120 ... Feature correction filter generator, 121, 211 ... Feature correction filter storage (second codebook), 205 ... Feature vector buffer, 210 ... Air conduction voice estimation section, 212 .. .Pattern collation unit, 213 ... Voice dictionary storage unit (voice dictionary), 214 ... Recognition result display unit
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| JP2005157354A | Cited by | Japan | Examiner |
| JP2006276603A | Cited by | Japan | Examiner |
| US6741962B2 | Cited by | United States of America | Applicant |
| JP2011209758A | Cited by | Japan | Examiner |
| KR20030040610A | Cited by | Republic of Korea | Search report |
| EP1536414A2 | Cited by | European Patent Office (EPO) | Applicant |
| EP1536414A3 | Cited by | European Patent Office (EPO) | Search report |
| EP1569422A3 | Cited by | European Patent Office (EPO) | Search report |
2 priority claims, no other members on record
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 4726299 | Japan | A | |
| JP19990047262 | – | – | – |
Numbers
- Publication
- 2000-250577
- Publication, DOCDB
- 2000250577
- Publication, EPODOC
- JP2000250577
- Application
- 11047262
- Application, DOCDB
- 4726299
- Application, EPODOC
- JP19990047262
Titles2
- Japanese
- 【発明の名称】音声認識装置及び同装置に使用される学習方法ならびに学習装置及び同方法がプログラムされ記録された記録媒体
- English
- INDUSTRIAL APPLICABILITY: A voice recognition device, a learning method used in the device, a learning device, and a recording medium in which the learning device is programmed and recorded.
Classification
- IPC, 3
- G10L15 28
- G10L15 02
- G10L15 06