JP2000250577A

Voice recognition device and learning method and learning device to be used in the same device and recording medium on which the same method is programmed and recorded

Abstract

(57) A summary and subject An improvement of the frequency characteristic which a bone conduction microphone has is aimed at, and it uses for speech recognition. It unites and improvement in the recognition performance under noise environment is aimed at. Solution means Using the first code book 110, choose a feature vector and the audio input pattern (bone conduction sound) recorded by first 受音器 101 is outputted, The correction vector memorized by the second code book 121 corresponding to the index is chosen, The feature vector of the sound (air conduction sound) recorded by second 受音器 112 to which the degree of 受音感 is secured by a frequency band larger than the first 受音器 of the above is presumed by adding both the above-mentioned vectors and connecting. Moreover, the speech parameter stored in the voice dictionary 213 as a pattern for reference is given one by one, using the presumed sound as a voice recognition object, and both patterns are compared.

Term

Term ended

Projected expiry passed 24 February 2019, 7.6 years ago.

  1. Priority and filed
  2. Published
  3. Projected expiry
  4. Today

10 claims: 5 independent, 5 dependent

  1. 1
    [Claims] [Claim 1] A means for receiving a sound signal by the first sound receiver and extracting a feature vector from the received sound signal for each frame having a predetermined time length, and a means for temporarily storing the extracted feature vector. A means for storing a representative finite number of characteristic vectors extracted from a sound signal pre-received by the first sound receiver as a first set, and a sound pre-received by the first sound receiver. A representative calculated using the difference between the characteristic vector of the signal and the characteristic vector of the sound signal previously received by the second receiver whose sound receiving sensitivity is ensured in a wider frequency band than the first receiver. A means for storing a finite number of correction vectors as a second set, a means for associating a feature vector belonging to the first set with each correction vector belonging to the second set, and the first set. For each of the feature vectors to which it belongs, a means for calculating the similarity of the feature vectors extracted from the audio signal received by the first sound receiver, and the feature vector having the highest degree of similarity are set in the first set. A means for extracting a correction vector belonging to the second set corresponding to this vector selected from the above, and the extraction for a feature vector extracted from an audio signal received by the first sound receiver. A means for calculating a feature vector generated by adding the corrected correction vectors for each frame, and a means for collating the similarity between the feature vector series stored as a dictionary in advance with respect to the feature vector series. , A sound recognition device including a means for outputting information of a dictionary having the highest degree of similarity from among the collated sounds. 【特許請求の範囲】 【請求項1】 音声信号を第一の受音器により受音し、受音した音声信号から、予め定めた時間長のフレーム毎に特徴ベクトルを抽出する手段と、抽出された特徴ベクトルを一時的に記憶する手段と、第一の受音器で予め受音した音声信号から抽出した代表的な有限個の特徴ベクトルを第一のセットとして記憶する手段と、前記第一の受音器で予め受音した音声信号の特徴ベクトルと前記第一の受音器よりも広い周波数帯域で受音感度が確保される第二の受音器で予め受音した音声信号の特徴ベクトルとの差分を用いて算出した代表的な有限個の補正ベクトルを第二のセットとして記憶する手段と、前記第一のセットに属する特徴ベクトルと前記第二のセットに属する各々の補正ベクトルを対応付ける手段と、前記第一のセットに属する各々の特徴ベクトルに対して、前記第一の受音器で受音した音声信号から抽出された特徴ベクトルの類似度を算出する手段と、類似度の最も高い特徴ベクトルを前記第一のセットの中から選択し、このベクトルに対応する前記第二のセットに属する補正ベクトルを抽出する手段と、前記第一の受音器で受音した音声信号から抽出された特徴ベクトルに対して前記抽出された補正ベクトルを加算して生成される特徴ベクトルをフレーム毎に算出する手段と、この特徴ベクトルの系列に対し、予め辞書として記憶された特徴ベクトル系列との間で類似度を照合する手段と、照合された中から最も類似度の高い辞書の情報を出力する手段とを備えることを特徴とする音声認識装置。
  2. 4
    A sound receiving pattern recorded by a first sound receiving device and a second sound receiving device whose sound receiving sensitivity is ensured in a wider frequency band than the first sound receiving device has a predetermined length. The feature amount is calculated by cutting out for each section, and the feature amount extracted through one of the receivers is compared with the code vector stored in the codebook to output the index of the feature vector with the highest degree of similarity. , A codebook learning method of a speech recognition device, characterized in that a correction vector stored in the other codebook corresponding to this index is output. 【請求項4】 第一の受音器と、該第一の受音器より広い周波数帯域で受音感度が確保される第二の受音器とで収録された受音パターンを所定長の区間毎に切り出して特徴量を算出し、一方の受音器を介して抽出された特徴量をコードブックに記憶されたコードベクトルと比較することにより最も類似度の高い特徴ベクトルのインデックスを出力し、このインデックスに対応する他方のコードブックに記憶された補正ベクトルを出力することを特徴とする音声認識装置のコードブック学習方法。
  3. 7
    The sound receiving patterns recorded by the first sound receiving device and the second sound receiving device whose sound receiving sensitivity is ensured in a wider frequency band than the first sound receiving device are cut out at intervals of a predetermined time length. By calculating the feature quantity, extracting the feature vector by referring to the code book, and collating the similarity with the code vector stored in advance as a dictionary, the highest degree of similarity is obtained from the collated sounds. In a voice recognition device that outputs dictionary information as a recognition result, a means for receiving a voice signal by the first sound receiver and extracting a feature vector from the received voice signal for each frame having a predetermined time length. A means for temporarily storing the extracted feature vectors, and a means for storing as a first set a finite number of representative feature vectors extracted from the voice signal previously received by the first sound receiver. A typical finite number of correction vectors calculated by using the difference between the feature vector of the voice signal previously received by the first receiver and the feature vector of the voice signal previously received by the second receiver. The sound recognition device is provided with a means for storing the sound as a second set and a means for associating each feature vector belonging to the first set with each correction vector belonging to the second set. Codebook learning device. 【請求項7】 第一の受音器と、該第一の受音器より広い周波数帯域で受音感度が確保される第二の受音器で収録される受音パターンを所定の時間長の区間毎に切り出して特徴量を算出し、コードブックを参照することによって特徴ベクトルを抽出し、予め辞書として記憶されたコードベクトルとの間で類似度を照合することにより、照合された中から最も類似度の高い辞書情報を認識結果として出力する音声認識装置において、音声信号を第一の受音器により受音し、受音した音声信号から、所定の時間長のフレーム毎に特徴ベクトルを抽出する手段と、抽出された特徴ベクトルを一時的に記憶する手段と、前記第一の受音器で予め受音した音声信号から抽出した代表的な有限個の特徴ベクトルを第一のセットとして記憶する手段と、前記第一の受音器で予め受音した音声信号の特徴ベクトルと前記第二の受音器で予め受音した音声信号の特徴ベクトルとの差分を用い算出した代表的な有限個の補正ベクトルを第二のセットとして記憶する手段と、前記第一のセットに属する各々の特徴ベクトルと前記第二のセットに属する各々の補正ベクトルを対応付ける手段とを具備することを特徴とする音声認識装置のコードブック学習装置。
  4. 9
    The sound receiving patterns recorded by the first sound receiving device and the second sound receiving device whose sound receiving sensitivity is ensured in a wider frequency band than the first sound receiving device are cut out for each section of a predetermined time length. It was used in the codebook learning device of the voice recognition device that calculates the feature amount and extracts the feature vector by referring to the codebook, and received the sound by the first sound receiver and the second sound receiver. A step of converting an analog signal into a digital signal at an appropriate sampling frequency and storing it in a sampling data buffer prepared for each, and a step of calculating a data feature amount for each frame of the data stored in each sampling data buffer. And the step of determining the start and end frames of the voice by comparing the power calculated for each frame with the preset threshold, and the start frame of the voice while the appropriate word is spoken. Based on the information of the end frame and the step of storing only the feature amount in the range as the feature vector in the first and second vector buffers, and the feature vector stored in the first vector buffer, a representative feature vector is obtained. The step of generating and storing in the first codebook and the feature vector stored in the first vector buffer are frame-by-frame vectors based on the pre-generated feature vector and stored in the first codebook. A step of quantization, a step of assigning an index of the feature vector having the highest similarity among the feature vectors of the first codebook to the first feature vector, and a step of being stored in the second vector buffer. Corresponds to the step of calculating the difference between the feature vector and the feature vector stored in the first vector buffer for each frame and storing the difference in the feature difference data buffer, and the feature vector stored in the first vector buffer. The step of assigning the same index as the index assigned to the first feature vector to the feature difference data to be performed, and the feature difference data clustered for each index assigned to the feature difference data, and the data included in the cluster. To averageA recording medium in which a step of generating a typical feature correction vector and storing it in the second codebook is recorded. 【請求項9】 第一の受音器と、該第一の受音器より広い周波数帯域で受音感度が確保される第二の受音器で収録された受音パターンを所定時間長の区間毎に切り出して特徴量を算出し、コードブックを参照することによって特徴ベクトルを抽出する音声認識装置のコードブック学習装置に用いられ、前記第一の受音器と、第二の受音器で受音されたアナログ信号を適切なサンプリング周波数でディジタル信号に変換し、それぞれに用意されるサンプリングデータバッファに格納するステップと、それぞれのサンプリングデータバッファに格納されたデータをフレーム毎にデータの特徴量を算出するステップと、フレーム毎に算出されるパワーと事前に設定された閾値とを比較することにより、音声の開始フレーム及び終了フレームを決定するステップと、適切な単語の発声がある間、前記音声の開始フレームと終了フレームの情報に基づき当該範囲の特徴量のみを特徴ベクトルとして第一・第二のベクトルバッファに格納するステップと、前記第一のベクトルバッファに格納された特徴ベクトルから代表的な特徴ベクトルを生成し、第一のコードブックに格納するステップと、前記第一のベクトルバッファに格納された特徴ベクトルを、事前に生成され前記第一のコードブックに格納された特徴ベクトルに基づいてフレーム毎ベクトル量子化するステップと、前記第一のコードブックの特徴ベクトルの中で最も類似度の高い特徴ベクトルのインデックスを前記第一の特徴ベクトルに付与するステップと、前記第二のベクトルバッファに格納された特徴ベクトルと第一のベクトルバッファに格納された特徴ベクトルとの差分をフレーム毎算出し、その差分を特徴差分データバッファに格納するステップと、前記第一のベクトルバッファに格納された特徴ベクトルに対応する特徴差分データに対し、前記第一の特徴ベクトルに付与されたインデックスと同じインデックスを付与するステップと、前記特徴差分データに付与されたインデックス毎に特徴差分データをクラスタリングし、クラスタに含まれるデータを平均化することにより代表的な特徴補正ベクトルを生成して前記第二のコードブックに格納するステップが記録された記録媒体。
  5. 10
    The voice input pattern recorded via the sound receiver is cut out for each section of a predetermined time length, the feature amount is extracted, the feature vector is extracted by referring to the code book, and the code vector stored in advance as a dictionary is used. It is used in a voice recognition device that outputs dictionary information with the highest degree of similarity by collating the similarity between the two, and the analog signal received by the sound receiver is measured at an appropriate sampling frequency. The step of converting to a digital signal and storing it in the sampling data buffer, the signal power is calculated each time the signal for one frame is stored in the sampling data buffer, and the power calculated for each frame is spread over an appropriate number of frames. The start frame of the sound is obtained by comparing the step of cumulatively adding and calculating the average value per frame to use as the threshold for detecting the voice section with the power calculated for each frame and the preset threshold. And the step of determining the end frame, the step of extracting the feature amount of the range as a vector based on the information of the start frame and the end frame and storing it in the feature vector buffer, and the step of storing the feature vector in the feature vector buffer in advance. A step of vector quantization for each frame based on the feature vector stored in the first codebook, and an index assigned to the feature vector having the highest similarity among the feature vectors of the first codebook. The step of transferring the feature vector and the correction vector corresponding to the transferred index are extracted from the feature vector stored in the second codebook in advance, and the extracted correction vector is transferred to the transferred feature vector. The step of estimating the feature vector stored in the second codebook by adding the data and the feature vector obtained here are used as input patterns and are registered in the speech dictionary as reference patterns in advance, and are to be recognized. The feature parameters from the voice start frame to the voice end frame of each word are given in sequence, the distance value between the step of collating both patterns for each word and the input pattern for each reference pattern is calculated, and all the distances.A recording medium in which a step of outputting a reference pattern name corresponding to the smallest distance value among the values as a recognition result is recorded. 【請求項10】 受音器を介して収録される音声入力パターンを所定時間長の区間毎に切り出して特徴量を抽出し、コードブックを参照することにより特徴ベクトルを抽出し、予め辞書として記憶されたコードベクトルとの間で類似度を照合することにより、照合された中から最も類似度の高い辞書情報を出力する音声認識装置に用いられ、前記受音器で受音されたアナログ信号を適切なサンプリング周波数によりディジタル信号に変換しサンプリングデータバッファに格納するステップと、サンプリングデータバッファに1フレーム分の信号が格納される毎に信号パワーを算出し、フレーム毎に算出されるパワーを適切なフレーム数に渡って累積加算し、フレームあたりの平均値を計算して音声区間検出のための閾値とするステップと、フレーム毎に算出されるパワーと事前に設定された前記閾値とを比較することにより音声の開始フレーム及び終了フレームを決定するステップと、開始フレームと終了フレームの情報に基づき当該範囲の特徴量をベクトルとして抽出し、特徴ベクトルバッファに格納するステップと、特徴ベクトルバッファに格納された特徴ベクトルを、事前に第一のコードブックに格納された特徴ベクトルに基づいてフレーム毎にベクトル量子化するステップと、前記第一のコードブックの特徴ベクトルの中で最も類似度が高い特徴ベクトルに付与されるインデックスとその特徴ベクトルを転送するステップと、事前に第二のコードブックに格納された特徴ベクトルから前記転送されたインデックスに相当する補正ベクトルを抽出し、この抽出された補正ベクトルを転送された特徴ベクトルに加算することにより前記第二のコードブックに格納される特徴ベクトルの推定を行うステップと、ここで得られる特徴ベクトルを入力パターンとし、参照パターンとして音声辞書に予め登録されてある、認識対象となる各単語の音声開始フレームから音声終了フレームまでの特徴パラメータを順次与え、単語毎両パターンの照合を行なうステップと、各参照パターン毎入力パターンとの距離値を算出し、全ての距離値の中で最小となる距離値に対応する参照パターン名を認識結果として出力するステップが記録された記録媒体。