Language model creation device
Abstract
This record has no abstract on file.
Term
Projected expiry 3 September 2029.
- Priority
- Filed
- Granted
- Today
- Projected expiry
25 claims: 8 independent, 17 dependent
- 1第1の内容を表す単語列において特定の単語が出現する確率を表す第1の内容別言語モデルと、第2の内容を表す単語列において前記特定の単語が出現する確率を表す第2の内容別言語モデルと、を記憶する内容別言語モデル記憶手段と、 音声に対応する単語列を認識する音声認識処理を行うことにより生成された音声認識仮説に含まれる単語列であって入力された単語列である入力単語列の少なくとも一部である処理対象単語列が表す内容が前記第1の内容である確率を表す第1の確率パラメータと、当該処理対象単語列が表す内容が前記第2の内容である確率を表す第2の確率パラメータと、を取得するとともに、当該取得された第1の確率パラメータと、当該取得された第2の確率パラメータと、前記記憶されている第1の内容別言語モデルと、前記記憶されている第2の内容別言語モデルと、に基づいて、前記音声のうちの前記処理対象単語列に対応する部分に対応する単語列において前記特定の単語が出現する確率を表す言語モデルを作成する言語モデル作成処理を行う言語モデル作成手段と、 を備え、前記言語モデル作成手段は、ある処理対象単語列に対して前記取得された第1の確率パラメータが、予め設定された上限閾値よりも大きい場合、その処理対象単語列に隣接する処理対象単語列に対して前記取得された第1の確率パラメータを増加させるように補正する、 言語モデル作成装置。
- 2請求項1に記載の言語モデル作成装置であって、 前記言語モデル作成手段は、 前記取得された第1の確率パラメータ及び前記取得された第2の確率パラメータの両方が、予め設定された下限閾値よりも小さい場合、前記取得された第1の確率パラメータ及び前記取得された第2の確率パラメータの両方を同一の値に設定するように構成された言語モデル作成装置。
- 3第1の内容を表す単語列において特定の単語が出現する確率を表す第1の内容別言語モデルと、第2の内容を表す単語列において前記特定の単語が出現する確率を表す第2の内容別言語モデルと、を記憶する内容別言語モデル記憶手段と、 音声に対応する単語列を認識する音声認識処理を行うことにより生成された音声認識仮説に含まれる単語列であって入力された単語列である入力単語列の少なくとも一部である処理対象単語列が表す内容が前記第1の内容である確率を表す第1の確率パラメータと、当該処理対象単語列が表す内容が前記第2の内容である確率を表す第2の確率パラメータと、を取得するとともに、当該取得された第1の確率パラメータと、当該取得された第2の確率パラメータと、前記記憶されている第1の内容別言語モデルと、前記記憶されている第2の内容別言語モデルと、に基づいて、前記音声のうちの前記処理対象単語列に対応する部分に対応する単語列において前記特定の単語が出現する確率を表す言語モデルを作成する言語モデル作成処理を行う言語モデル作成手段と、 を備え、前記言語モデル作成手段は、前記取得された第1の確率パラメータ及び前記取得された第2の確率パラメータの両方が、予め設定された下限閾値よりも小さい場合、前記取得された第1の確率パラメータ及び前記取得された第2の確率パラメータの両方を同一の値に設定する、 言語モデル作成装置。
- 4請求項1乃至請求項3のいずれか一項に記載の言語モデル作成装置であって、 前記言語モデル作成手段は、前記音声認識処理によって認識された単語列の単語間の境界を境界として前記入力単語列を分割した複数の前記処理対象単語列のそれぞれに対して、前記言語モデル作成処理を行うように構成された言語モデル作成装置。
- 5請求項1乃至請求項3のいずれか一項に記載の言語モデル作成装置であって、 前記言語モデル作成手段は、前記音声認識処理によって認識された単語列の単語間の境界と異なる位置を境界として前記入力単語列を分割した複数の前記処理対象単語列のそれぞれに対して、前記言語モデル作成処理を行うように構成された言語モデル作成装置。
- 6請求項1乃至請求項5のいずれか一項に記載の言語モデル作成装置であって、 前記言語モデル作成手段は、 前記取得された第1の確率パラメータが大きくなるほど大きくなる第1の係数を前記第1の内容別言語モデルが表す確率に乗じた値と、前記取得された第2の確率パラメータが大きくなるほど大きくなる第2の係数を前記第2の内容別言語モデルが表す確率に乗じた値と、の和が大きくなるほど、前記音声のうちの前記処理対象単語列に対応する部分に対応する単語列において前記特定の単語が出現する確率が大きくなる前記言語モデルを作成するように構成された言語モデル作成装置。
- 7請求項1乃至請求項6のいずれか一項に記載の言語モデル作成装置であって、 前記言語モデル作成手段は、条件付確率場の理論に基づいて前記第1の確率パラメータ及び前記第2の確率パラメータを取得するように構成された言語モデル作成装置。
- 8請求項1乃至請求項7のいずれか一項に記載の言語モデル作成装置であって、 前記言語モデル作成手段は、前記処理対象単語列に含まれる単語の属性を表す単語属性情報、及び、前記処理対象単語列を音声から認識する音声認識処理を行う際に取得された音声認識処理情報、の少なくとも1つに基づいて前記第1の確率パラメータ及び前記第2の確率パラメータを取得するように構成された言語モデル作成装置。
- 9請求項8に記載の言語モデル作成装置であって、 前記単語属性情報は、単語表層を表す情報、読みを表す情報、及び、品詞を表す情報の少なくとも1つを含む言語モデル作成装置。
- 10請求項8又は請求項9に記載の言語モデル作成装置であって、 前記音声認識処理情報は、前記音声認識処理による認識結果の信頼度である認識信頼度を表す情報、1つの音が継続する時間である継続時間長を表す情報、及び、先行する無音の有無を表す情報の少なくとも1つを含む言語モデル作成装置。
- 11請求項1乃至請求項10のいずれか一項に記載の言語モデル作成装置であって、 前記言語モデル作成手段は、 前記入力単語列における前記処理対象単語列の位置を表す情報、前記入力単語列が1つの単語を複数含むことを表す情報、前記入力単語列における内容の連接状態を表す情報、及び、前記入力単語列が複数存在する場合における各入力単語列間の関係を表す情報、の少なくとも1つに基づいて前記第1の確率パラメータ及び前記第2の確率パラメータを取得するように構成された言語モデル作成装置。
- 12第1の内容を表す単語列において特定の単語が出現する確率を表す第1の内容別言語モデルと、第2の内容を表す単語列において前記特定の単語が出現する確率を表す第2の内容別言語モデルと、を記憶する内容別言語モデル記憶手段と、 音声に対応する単語列を認識する音声認識処理を行うことにより生成された音声認識仮説に含まれる単語列であって入力された単語列である入力単語列の少なくとも一部である処理対象単語列が表す内容が前記第1の内容である確率を表す第1の確率パラメータと、当該処理対象単語列が表す内容が前記第2の内容である確率を表す第2の確率パラメータと、を取得するとともに、当該取得された第1の確率パラメータと、当該取得された第2の確率パラメータと、前記記憶されている第1の内容別言語モデルと、前記記憶されている第2の内容別言語モデルと、に基づいて、前記音声のうちの前記処理対象単語列に対応する部分に対応する単語列において前記特定の単語が出現する確率を表す言語モデルを作成する言語モデル作成処理を行う言語モデル作成手段と、 前記言語モデル作成手段により作成された言語モデルに基づいて、入力された音声に対応する単語列を認識する音声認識処理を行う音声認識手段と、 を備え、前記言語モデル作成手段は、ある処理対象単語列に対して前記取得された第1の確率パラメータが、予め設定された上限閾値よりも大きい場合、その処理対象単語列に隣接する処理対象単語列に対して前記取得された第1の確率パラメータを増加させるように補正する、 音声認識装置。
- 13第1の内容を表す単語列において特定の単語が出現する確率を表す第1の内容別言語モデルと、第2の内容を表す単語列において前記特定の単語が出現する確率を表す第2の内容別言語モデルと、を記憶する内容別言語モデル記憶手段と、 音声に対応する単語列を認識する音声認識処理を行うことにより生成された音声認識仮説に含まれる単語列であって入力された単語列である入力単語列の少なくとも一部である処理対象単語列が表す内容が前記第1の内容である確率を表す第1の確率パラメータと、当該処理対象単語列が表す内容が前記第2の内容である確率を表す第2の確率パラメータと、を取得するとともに、当該取得された第1の確率パラメータと、当該取得された第2の確率パラメータと、前記記憶されている第1の内容別言語モデルと、前記記憶されている第2の内容別言語モデルと、に基づいて、前記音声のうちの前記処理対象単語列に対応する部分に対応する単語列において前記特定の単語が出現する確率を表す言語モデルを作成する言語モデル作成処理を行う言語モデル作成手段と、 前記言語モデル作成手段により作成された言語モデルに基づいて、入力された音声に対応する単語列を認識する音声認識処理を行う音声認識手段と、 を備え、前記言語モデル作成手段は、前記取得された第1の確率パラメータ及び前記取得された第2の確率パラメータの両方が、予め設定された下限閾値よりも小さい場合、前記取得された第1の確率パラメータ及び前記取得された第2の確率パラメータの両方を同一の値に設定する、 音声認識装置。
- 14請求項12又は請求項13に記載の音声認識装置であって、 前記言語モデル作成手段は、 前記取得された第1の確率パラメータが大きくなるほど大きくなる第1の係数を前記第1の内容別言語モデルが表す確率に乗じた値と、前記取得された第2の確率パラメータが大きくなるほど大きくなる第2の係数を前記第2の内容別言語モデルが表す確率に乗じた値と、の和が大きくなるほど、前記音声のうちの前記処理対象単語列に対応する部分に対応する単語列において前記特定の単語が出現する確率が大きくなる前記言語モデルを作成するように構成された音声認識装置。
- 15請求項12乃至請求項14のいずれか一項に記載の音声認識装置であって、 前記音声認識手段は、前記入力された音声に対応する単語列を認識する前記音声認識処理を行うことにより前記入力単語列を生成するように構成され、 前記言語モデル作成手段は、前記音声認識手段により生成された前記入力単語列に基づいて前記言語モデルを作成するように構成され、 前記音声認識手段は、前記言語モデル作成手段により作成された前記言語モデルに基づいて、前記入力された音声に対応する単語列を認識する前記音声認識処理を再度行うように構成された音声認識装置。
- 16請求項12乃至請求項15のいずれか一項に記載の音声認識装置であって、 前記言語モデル作成手段によって作成された言語モデルに基づいて、前記音声認識手段が前記入力された音声に対応する単語列を認識する前記音声認識処理と、前記音声認識手段によって認識された単語列に基づいて、前記言語モデル作成手段が前記言語モデルを作成する前記言語モデル作成処理と、を交互に繰り返す反復処理を実行するように構成された音声認識装置。
- 17請求項16に記載の音声認識装置であって、 所定の終了条件が成立した場合、前記反復処理を終了するように構成された音声認識装置。
- 18請求項17に記載の音声認識装置であって、 前記終了条件は、前回の前記音声認識処理により認識された単語列と、今回の前記音声認識処理により認識された単語列と、が一致しているという条件である音声認識装置。
- 19請求項17に記載の音声認識装置であって、 前記終了条件は、前記音声認識処理を実行した回数が予め設定された閾値回数よりも大きいという条件である音声認識装置。
- 20第1の内容を表す単語列において特定の単語が出現する確率を表す第1の内容別言語モデルと、第2の内容を表す単語列において前記特定の単語が出現する確率を表す第2の内容別言語モデルと、が記憶装置に記憶されている場合に、 音声に対応する単語列を認識する音声認識処理を行うことにより生成された音声認識仮説に含まれる単語列であって入力された単語列である入力単語列の少なくとも一部である処理対象単語列が表す内容が前記第1の内容である確率を表す第1の確率パラメータと、当該処理対象単語列が表す内容が前記第2の内容である確率を表す第2の確率パラメータと、を取得するとともに、当該取得された第1の確率パラメータと、当該取得された第2の確率パラメータと、前記記憶されている第1の内容別言語モデルと、前記記憶されている第2の内容別言語モデルと、に基づいて、前記音声のうちの前記処理対象単語列に対応する部分に対応する単語列において前記特定の単語が出現する確率を表す言語モデルを作成し、ある処理対象単語列に対して前記取得された第1の確率パラメータが、予め設定された上限閾値よりも大きい場合、その処理対象単語列に隣接する処理対象単語列に対して前記取得された第1の確率パラメータを増加させるように補正する、 言語モデル作成方法。
- 21第1の内容を表す単語列において特定の単語が出現する確率を表す第1の内容別言語モデルと、第2の内容を表す単語列において前記特定の単語が出現する確率を表す第2の内容別言語モデルと、が記憶装置に記憶されている場合に、 音声に対応する単語列を認識する音声認識処理を行うことにより生成された音声認識仮説に含まれる単語列であって入力された単語列である入力単語列の少なくとも一部である処理対象単語列が表す内容が前記第1の内容である確率を表す第1の確率パラメータと、当該処理対象単語列が表す内容が前記第2の内容である確率を表す第2の確率パラメータと、を取得するとともに、当該取得された第1の確率パラメータと、当該取得された第2の確率パラメータと、前記記憶されている第1の内容別言語モデルと、前記記憶されている第2の内容別言語モデルと、に基づいて、前記音声のうちの前記処理対象単語列に対応する部分に対応する単語列において前記特定の単語が出現する確率を表す言語モデルを作成し、前記取得された第1の確率パラメータ及び前記取得された第2の確率パラメータの両方が、予め設定された下限閾値よりも小さい場合、前記取得された第1の確率パラメータ及び前記取得された第2の確率パラメータの両方を同一の値に設定する、 言語モデル作成方法。
- 22請求項20又は請求項21に記載の言語モデル作成方法であって、 前記取得された第1の確率パラメータが大きくなるほど大きくなる第1の係数を前記第1の内容別言語モデルが表す確率に乗じた値と、前記取得された第2の確率パラメータが大きくなるほど大きくなる第2の係数を前記第2の内容別言語モデルが表す確率に乗じた値と、の和が大きくなるほど、前記音声のうちの前記処理対象単語列に対応する部分に対応する単語列において前記特定の単語が出現する確率が大きくなる前記言語モデルを作成するように構成された言語モデル作成方法。
- 23情報処理装置に、 第1の内容を表す単語列において特定の単語が出現する確率を表す第1の内容別言語モデルと、第2の内容を表す単語列において前記特定の単語が出現する確率を表す第2の内容別言語モデルと、を記憶装置に記憶させる内容別言語モデル記憶処理手段と、 音声に対応する単語列を認識する音声認識処理を行うことにより生成された音声認識仮説に含まれる単語列であって入力された単語列である入力単語列の少なくとも一部である処理対象単語列が表す内容が前記第1の内容である確率を表す第1の確率パラメータと、当該処理対象単語列が表す内容が前記第2の内容である確率を表す第2の確率パラメータと、を取得するとともに、当該取得された第1の確率パラメータと、当該取得された第2の確率パラメータと、前記記憶されている第1の内容別言語モデルと、前記記憶されている第2の内容別言語モデルと、に基づいて、前記音声のうちの前記処理対象単語列に対応する部分に対応する単語列において前記特定の単語が出現する確率を表す言語モデルを作成する言語モデル作成手段と、 を実現させるための言語モデル作成プログラムであり、前記言語モデル作成手段は、ある処理対象単語列に対して前記取得された第1の確率パラメータが、予め設定された上限閾値よりも大きい場合、その処理対象単語列に隣接する処理対象単語列に対して前記取得された第1の確率パラメータを増加させるように補正するように構成された、 言語モデル作成プログラム。
- 24情報処理装置に、 第1の内容を表す単語列において特定の単語が出現する確率を表す第1の内容別言語モデルと、第2の内容を表す単語列において前記特定の単語が出現する確率を表す第2の内容別言語モデルと、を記憶装置に記憶させる内容別言語モデル記憶処理手段と、 音声に対応する単語列を認識する音声認識処理を行うことにより生成された音声認識仮説に含まれる単語列であって入力された単語列である入力単語列の少なくとも一部である処理対象単語列が表す内容が前記第1の内容である確率を表す第1の確率パラメータと、当該処理対象単語列が表す内容が前記第2の内容である確率を表す第2の確率パラメータと、を取得するとともに、当該取得された第1の確率パラメータと、当該取得された第2の確率パラメータと、前記記憶されている第1の内容別言語モデルと、前記記憶されている第2の内容別言語モデルと、に基づいて、前記音声のうちの前記処理対象単語列に対応する部分に対応する単語列において前記特定の単語が出現する確率を表す言語モデルを作成する言語モデル作成手段と、 を実現させるための言語モデル作成プログラムであり、前記言語モデル作成手段は、前記取得された第1の確率パラメータ及び前記取得された第2の確率パラメータの両方が、予め設定された下限閾値よりも小さい場合、前記取得された第1の確率パラメータ及び前記取得された第2の確率パラメータの両方を同一の値に設定するように構成された、 言語モデル作成プログラム。
- 25請求項23又は請求項24に記載の言語モデル作成プログラムであって、 前記言語モデル作成手段は、 前記取得された第1の確率パラメータが大きくなるほど大きくなる第1の係数を前記第1の内容別言語モデルが表す確率に乗じた値と、前記取得された第2の確率パラメータが大きくなるほど大きくなる第2の係数を前記第2の内容別言語モデルが表す確率に乗じた値と、の和が大きくなるほど、前記音声のうちの前記処理対象単語列に対応する部分に対応する単語列において前記特定の単語が出現する確率が大きくなる前記言語モデルを作成するように構成された言語モデル作成プログラム。
Independent claims25
140 paragraphs, as filed
The present invention relates to a language model creating device that creates a language model used for performing speech recognition processing that recognizes a word string corresponding to speech.
A voice recognition device that recognizes a word string represented by a voice from a voice (utterance) uttered by a user is known. The voice recognition device described in Patent Document 1 as one of the voice recognition devices of this type performs voice recognition processing for recognizing a word string corresponding to a voice based on a plurality of content-specific language models stored in advance. ..
The content-based language model is a model that represents the probability that a specific word appears in a word string that represents a specific content (topic, keyword, etc.). For example, in a word string containing a TV program, the probability that a program name or a talent name appears is high, and in a word string containing sports, a team name, an exercise equipment name, or a player name is used. The probability of appearing increases.
By the way, the content may change in a series of voices emitted by the user. In this case, if the voice recognition process is performed based on only one content-specific language model, the accuracy of recognizing the word string may be excessively lowered.
Therefore, the voice recognition device is configured to use a different content-specific language model for each predetermined section in one utterance.
<p><patcit num="1"><text>JP-A-2002-229589</text></patcit></p>
<p num="0007"> However, in the voice recognition device, if the content related to the content-based language model used in the above section does not match the content of the actual utterance, the accuracy of recognizing the word string is excessively lowered. was there.</p><p num="0008"> Further, in the voice recognition device, in order to determine which content-specific language model to use, a process of evaluating the recognition result when each content-specific language model is used is performed. Therefore, in the above-mentioned voice recognition device, there is a problem that the processing load for deciding which content-specific language model to use is excessive.</p><p num="0009"> Therefore, an object of the present invention is to solve the above-mentioned problem "the calculation load for creating a language model becomes excessive, and the word string may not be recognized from the voice with high accuracy". The purpose is to provide a language model creation device that can be used.</p>
<p num="0010"> The language model creation device, which is one embodiment of the present invention, is used to achieve such an object. A first content-specific language model that represents the probability that a specific word will appear in a word string that represents the first content, and a second content that represents the probability that the specific word will appear in the word string that represents the second content. Another language model, a content-based language model storage means to memorize, A processing target word string that is at least a part of an input word string that is a word string included in a voice recognition hypothesis generated by performing a voice recognition process that recognizes a word string corresponding to a voice and is an input word string. Acquires a first probability parameter representing the probability that the content represented by is the first content, and a second probability parameter representing the probability that the content represented by the processing target word string is the second content. At the same time, the acquired first probability parameter, the acquired second probability parameter, the stored first content-based language model, and the stored second content-based language model. Based on the above, a language model is created to create a language model representing the probability that the specific word appears in the word string corresponding to the part corresponding to the processing target word string in the voice. Means and To be equipped.</p><p num="0011"> Further, the voice recognition device according to another embodiment of the present invention is A first content-specific language model that represents the probability that a specific word will appear in a word string that represents the first content, and a second content that represents the probability that the specific word will appear in the word string that represents the second content. Another language model, a content-based language model storage means to memorize, A processing target word string that is at least a part of an input word string that is a word string included in a voice recognition hypothesis generated by performing a voice recognition process that recognizes a word string corresponding to a voice and is an input word string. Acquires a first probability parameter representing the probability that the content represented by is the first content, and a second probability parameter representing the probability that the content represented by the processing target word string is the second content. At the same time, the acquired first probability parameter, the acquired second probability parameter, the stored first content-based language model, and the stored second content-based language model. Based on the above, a language model is created to create a language model representing the probability that the specific word appears in the word string corresponding to the part corresponding to the processing target word string in the voice. Means and A voice recognition means that performs voice recognition processing that recognizes a word string corresponding to an input voice based on the language model created by the above language model creation means. To be equipped.</p><p num="0012"> In addition, the language model creation method, which is another form of the present invention, is A first content-based language model that represents the probability that a specific word will appear in a word string that represents the first content, and a second content that represents the probability that the specific word will appear in the word string that represents the second content. When another language model and is stored in the storage device, A word string to be processed that is at least a part of an input word string that is a word string included in a voice recognition hypothesis generated by performing a voice recognition process that recognizes a word string corresponding to a voice and is an input word string. Acquires a first probability parameter representing the probability that the content represented by is the first content, and a second probability parameter representing the probability that the content represented by the processing target word string is the second content. At the same time, the acquired first probabilistic parameter, the acquired second probabilistic parameter, the stored first content-based language model, and the stored second content-based language model. Based on the above, it is a method of creating a language model representing the probability that the specific word appears in the word string corresponding to the part corresponding to the processing target word string in the voice.</p><p num="0013"> In addition, the language model creation program, which is another form of the present invention, For information processing equipment A first content-specific language model that represents the probability that a specific word will appear in a word string that represents the first content, and a second content that represents the probability that the specific word will appear in the word string that represents the second content. Another language model, a content-based language model storage processing means for storing in a storage device, A word string to be processed that is at least a part of an input word string that is a word string included in a voice recognition hypothesis generated by performing a voice recognition process that recognizes a word string corresponding to a voice and is an input word string. Acquires a first probability parameter representing the probability that the content represented by is the first content, and a second probability parameter representing the probability that the content represented by the processing target word string is the second content. At the same time, the acquired first probability parameter, the acquired second probability parameter, the stored first content-based language model, and the stored second content-based language model. Based on the above, a language model creating means for creating a language model representing the probability that the specific word appears in the word string corresponding to the part corresponding to the processing target word string in the voice, and It is a program to realize.</p>
<p num="0014"> The present invention can create a language model capable of recognizing a word string corresponding to a voice with high accuracy while preventing an excessive calculation load by being configured as described above. it can.</p>
<figref num="1">It is a block diagram which shows the outline of the function of the language model making apparatus which concerns on 1st Embodiment of this invention.</figref><figref num="2">It is a flowchart which showed the operation of the language model creation apparatus shown in FIG.</figref><figref num="3">It is explanatory drawing which conceptually showed the example of the word string of the speech recognition hypothesis.</figref><figref num="4">It is explanatory drawing which conceptually showed the example of the content candidate graph.</figref><figref num="5">It is explanatory drawing which showed the example of the feature about the content.</figref><figref num="6">It is explanatory drawing which showed the example of the feature used in CRF which is an example of a content model.</figref><figref num="7">It is explanatory drawing which conceptually showed an example of the score acquired for the processing target word string.</figref><figref num="8">It is explanatory drawing which conceptually showed an example of the score acquired for the processing target word string.</figref><figref num="9">It is a block diagram which shows the outline of the function of the voice recognition apparatus which concerns on 2nd Embodiment of this invention.</figref><figref num="10">It is a flowchart which showed the operation of the voice recognition apparatus shown in FIG.</figref><figref num="11">It is a block diagram which shows the outline of the function of the language model making apparatus which concerns on 3rd Embodiment of this invention.</figref>
Hereinafter, embodiments of a language model creation device, a voice recognition device, a language model creation method, and a language model creation program according to the present invention will be described with reference to FIGS. 1 to 11.
<First Embodiment> (Constitution) The language model creating device 101 according to the first embodiment of the present invention will be described with reference to FIG. The language model creation device 101 is an information processing device. The language model creation device 101 includes a central processing unit (CPU; Central Processing Unit), a storage device (memory and hard disk drive device (HDD; Hard Disk Drive)), an input device, and an output device (not shown).
The output device has a display. The output device displays an image composed of characters, figures, and the like on the display based on the image information output by the CPU.
The input device includes a mouse, a keyboard and a microphone. The language model creating device 101 is configured so that information based on a user's operation is input via a keyboard and a mouse. The language model creating device 101 is configured such that input voice information representing the voice around the microphone (that is, outside the language model creating device 101) is input via the microphone.
In the present embodiment, the language model creating device 101 receives a voice recognition hypothesis (input word string) including a word string as a voice recognition result output by a voice recognition device (not shown), and responds to the accepted voice recognition hypothesis. The language model is configured to be output to the voice recognition device.
(function) Next, the function of the language model creation device 101 configured as described above will be described. As shown in FIG. 1, the functions of the language model creation device 101 include a voice recognition hypothesis input unit (a part of the language model creation means) 11 and a content estimation unit (a part of the language model creation means) 12. Language model creation unit (part of language model creation means) 13, content model storage unit (part of language model creation means) 14, content-specific language model storage unit (content-specific language model storage means, content-specific language model) A storage processing means, a content-specific language model storage processing step) 15, and the like. This function is realized by the CPU of the language model creating device 101 executing a program stored in the storage device. Note that this function may be realized by hardware such as a logic circuit.
The voice recognition hypothesis input unit 11 receives a voice recognition hypothesis (input word string) including a word string as a voice recognition result output by an external voice recognition device (not shown), and estimates the content of the received voice recognition hypothesis. Output to unit 12. The speech recognition hypothesis is information generated by performing a speech recognition process in which a speech recognition device recognizes a word string corresponding to speech. In this example, the speech recognition hypothesis is information that represents a word string consisting of one or more words. Further, the speech recognition hypothesis may be information representing a plurality of word strings (for example, a word graph, N best word strings (best N word strings), or the like).
The content estimation unit 12 divides the voice recognition hypothesis output from the voice recognition hypothesis input unit 11 with the boundary between words of the word string recognized by the voice recognition process as a boundary to divide at least one word string to be processed. Extract (generate) from the speech recognition hypothesis. According to this, when creating a language model, it is possible to use the information acquired when performing the voice recognition process. As a result, the content can be estimated accurately and the language model can be created quickly.
Further, the content estimation unit 12 extracts at least one processing target word string from the voice recognition hypothesis by dividing the voice recognition hypothesis with a position different from the boundary between words of the word string recognized by the voice recognition processing as a boundary. It may be (generated). According to this, even if the boundary between actual words in the utterance is different from the boundary between words of the word string recognized by the speech recognition process, the word string corresponding to the speech is recognized with high accuracy. You can create a language model that allows you to do this.
The content estimation unit 12 has a probability that the content represented by the processing target word string is a specific content (first content, second content, etc.) for each of the generated plurality of processing target word strings. A probability parameter (first probability parameter, second probability parameter, etc.) representing the above is calculated (acquired) based on the content model stored in the content model storage unit 14. For example, the content estimation unit 12 calculates a first probability parameter representing the probability that the content represented by the processing target word string is the first content, and a second probability parameter representing the probability that the content represented by the second content is the second content. Is calculated. Then, the content estimation unit 12 outputs the acquired probability parameter to the language model creation unit 13.
In this example, the probability parameter is the value of the probability that the content represented by the processing target word string is a specific content. The probability parameter may be a value that increases as the probability that the content represented by the processing target word string is a specific content increases. That is, it can be said that the probability parameter represents the plausibility that the content represented by the processing target word string is a specific content. Probability parameters may be referred to as likelihood parameters or weight parameters.
Here, the content is also called a topic. For example, the contents used as search conditions for searching TV programs include personal names (talent names and group names, etc.), program names, program genre names (variety, sports, etc.), broadcasting station names, etc. And time expression (evening and 8 o'clock, etc.), etc. If the contents are different, the probability that a specific word string appears (exists) during the utterance is different.
In this way, the content estimation unit 12 estimates the probability that the content represented by the word string in the section is a specific content for each section (word string to be processed) during utterance. Therefore, even if the content changes during the utterance, the above probability can be estimated with high accuracy for each section.
The content model storage unit 14 stores in advance a content model (information) representing a relationship between a word string and the probability that the content represented by the word string is each of a plurality of contents. In this example, the content model is a random model based on the theory of Conditional Random Fields (CRF). The content model is expressed by the following equation (1).<maths num="1"><img id="000002" he="25" wi="159" file="0005598331.tif" img-format="tif" img-content="drawing" /></maths>
Here, "X" is a word string to be processed, and is "Y" is the content. That is, the right-hand side P (Y | X) of the equation (1) represents the probability that the content represented by the processing target word string X is the content Y. Further, "Φ (X, Y)" is information representing the characteristics (features) of the word string X to be processed. "Λ" is a model parameter (weight value) in the CRF corresponding to each of the features Φ (X, Y). Further, "Z" is a normalization term. Note that "exp ()" indicates a function for finding the power of a numerical value with e as the base.
Therefore, in this example, the content model storage unit 14 stores the feature Φ and the model parameter Λ (weight value) in the storage device.
Now, when the speech recognition hypothesis is a word string and CRF is used as the content model, an example of a method in which the content estimation unit 12 estimates the content represented (belonging to) by each word of the speech recognition hypothesis will be described.
First, the content estimation unit 12 expands the section corresponding to each word included in the word string of the speech recognition hypothesis into candidate content and holds it in a graph format (content candidate graph). FIG. 3 is an example of a word string of the speech recognition hypothesis, and FIG. 4 is an example of a content candidate graph.
For example, suppose that the voice recognition hypothesis of the utterance "I want to see a drama with Inagaki Goro" is "I want to see a drama with a country trip". FIG. 3 is a part of the word string of the speech recognition hypothesis. As shown in FIG. 4, the content estimation unit 12 develops and generates three types of content candidates, "personal name", "program name", and "other" for each section. The arc (arc, edge) A in FIG. 4 indicates that the content represented by the word "country travelogue" in the speech recognition hypothesis is a "personal name" as the content.
Next, the content estimation unit 12 ranks and outputs the content path (content sequence) represented by the content candidate graph based on a predetermined criterion (for example, a score calculated with reference to the content model). Specifically, the content estimation unit 12 obtains a score by referring to the content model in each arc in the graph, and accumulates the score for each pass.
The content estimation unit 12 identifies the path on which the left side P (Y | X) of the above equation (1) is maximized by a search using the Viterbi algorithm. In addition, the content estimation unit 12 identifies the ranked higher-order paths by A * search. Note that the content estimation unit 12 may apply a process or the like to combine the same contents when they are continuous when outputting the information representing the specified path.
The score in each arc in the content candidate graph is the product of the feature (feature) related to each arc and the weight value for each feature which is a model parameter of CRF. Taking the arc A of the content candidate graph of FIG. 4 as an example, an example of a method of obtaining a score in the arc will be described.
FIG. 5 is an example of the features relating to the arc A. FIG. 6 is an example in which the features of FIG. 5 are expressed as the features of the content model. For example, when the speech recognition hypothesis of the interval corresponding to the time interval of a certain arc A has features such as "part of speech = noun" and "co-occurrence = out" when the content is "personal name". Is assumed. In such a case, these features are used as features used in the content model.
Now, it is assumed that the word string corresponding to the arc A has features such as "part of speech = noun" and "co-occurrence = out" as shown in FIG. These features are expressed as CRF features (Φ) as shown in FIG. The score of the arc A is calculated by the product of the values taken by these features and the weight Λ of the "personal name" corresponding to the arc A in the model parameters. The higher this score, the higher the content.
In this example, as the feature (Φ) of the content model, the linguistic features (word surface layer, reading, part of speech, etc.) in the section corresponding to the target arc for which the score is to be obtained are used. In other words, the content estimation unit 12 acquires the probability parameter based on the word attribute information representing the attribute of the word included in the processing target word string. The word attribute information includes at least one of information representing a word surface layer, information representing reading, and information representing part of speech.
As the feature (Φ) of the content model, features related to speech recognition processing (recognition reliability, duration length, presence / absence of preceding silence, etc.) may be used. In other words, the content estimation unit 12 may acquire the probability parameter based on the voice recognition processing information acquired when performing the voice recognition processing for recognizing the processing target word string from the voice. Here, the voice recognition processing information includes information representing the recognition reliability, which is the reliability of the recognition result by the voice recognition processing, information representing the duration time, which is the duration of one sound, and the presence or absence of preceding silence. Contains at least one piece of information representing.
In addition, the above-mentioned features regarding the sections before and after the target arc and the sections overlapping the target arc in the word graph or N best word string can also be used.
In addition, not only local features related to the target section, but also global features related to the entire speech recognition hypothesis (entire utterance), position information within the speech recognition hypothesis (first half, second half, etc.), co-occurrence within the utterance. Word information, information on the structure of the word graph (average number of branches of the arc, etc.), concatenation information of the contents, etc. may be used as the utterance (Φ) of the content model. In other words, the content estimation unit 12 includes information indicating the position of the processing target word string in the input word string, information indicating that the input word string includes a plurality of one word, and information indicating the concatenated state of the contents in the input word string. And, the probability parameter may be acquired based on at least one of the information representing the relationship between each input word string when there are a plurality of input word strings.
The posterior appearance probability (posterior probability) p (Yi = c | X) of each arc in the content candidate graph is calculated by recursive calculation using the Forward algorithm and the Backward algorithm. Here, "Yi = c" indicates that the content represented by the word string in the i-th section is the content c ". The content estimation unit 12 sets this probability p to the appearance probability of each content in the section ( Used as a probabilistic parameter).
The model parameters of the CRF are iteratively calculated according to the criteria for maximizing the log-likelihood of the above equation (1), using the pair of the input (X: word string) and the output (Y: content) associated in advance as learning data. It may be optimized (learned) by a method or the like.
For details on the above-mentioned identification method using CRF, the method of calculating the posterior probability of the identification result, and the method of learning the model parameters, see, for example, the non-patent documents "J. Lafferty, A. McCallum, F. Pereira,". Conditional Random Fields: Probabilistic Models for Segmenting and Labeling Sequence Data ", Proceedings of 18th International Conference on Machine Learning, 28th International Conference on Machine Learning, 28th International Conference on Machine Learning.
The language model creation unit 13 includes content estimation results including probability parameters (for example, a first probability parameter, a second probability parameter, etc.) output from the content estimation unit 12, and a content-specific language model storage unit 15. Represents the probability that a specific word will appear in the word string corresponding to the part of the voice that is the basis of the input word string that corresponds to the processing target word string, based on the content-specific language model stored in. Creating a language model Perform a language model creation process for each of the word strings to be processed. Then, the language model creation unit 13 outputs the created language model. In this example, the content-based language model and the language model are N-gram language models constructed on the assumption that the probability that a certain word appears depends only on the N-1 word immediately before it. Is.
In the N-gram language model, the i-th word w<sub>i</sub>The probability of appearance of is P (w)<sub>i</sub>| W<sub>i-N + 1</sub><sup>i-1</sup>). Here, W of the condition part<sub>i-N + 1</sub><sup>i-1</sup>Represents the (i-N + 1) to (i-1) th word strings. The model with N = 2 is called a bigram model, and the model with N = 3 is called a trigram model. In addition, a model constructed based on the assumption that it is not affected by the immediately preceding word is called a unigram model.
According to the N-gram language model, the word string W<sub>1</sub><sup>n</sup>= (W<sub>1</sub>, W<sub>2</sub>, ..., w<sub>n</sub>) Appears P (W)<sub>1</sub><sup>n</sup>) Is expressed by the following equation (2). Further, such a parameter consisting of various conditional probabilities of various words used in the N-gram language model can be obtained by maximum likelihood estimation or the like for learning text data.<maths num="2"><img id="000003" he="20" wi="159" file="0005598331.tif" img-format="tif" img-content="drawing" /></maths>
The content-based language model storage unit 15 stores a plurality of content-based language models in a storage device in advance. A plurality of content-based language models are models that represent the probability that a specific word appears in a word string that represents different contents. That is, the plurality of content-based language models include the first content-based language model representing the probability that a specific word appears in the word string representing the first content, and the specific word in the word string representing the second content. Includes a second content-based language model that represents the probability that In this example, each content-specific language model is a trigram model.
The language model creation unit 13 stores the score for each content in each section (that is, the probability parameter indicating the probability that the content represented by each processing target word string is each content) and the content-specific language model storage unit 15. Create a language model according to the following equation (3) from a plurality of content-specific language models.<maths num="3"><img id="000004" he="19" wi="159" file="0005598331.tif" img-format="tif" img-content="drawing" /></maths>
In equation (3), P<sub>t</sub>(W<sub>i</sub>) Is the word w<sub>i</sub>Is the probability that appears in the interval "t", α<sub>j</sub>(T) is a probability parameter (score) representing the probability that the content represented by the word string in the interval "t" is the content "j" (probability of appearance of the content), and is P.<sub>j</sub>(W<sub>i</sub>) Is the word w in the content-specific language model for the content "j"<sub>i</sub>Is the probability of appearing. In this example, the language model creation unit 13 sets the probability parameter (probability of appearance of the content in each section (processed word string) in the utterance) acquired by the content estimation unit 12 to α in the equation (3).<sub>j</sub>Used as (t).
In this way, the language model creation unit 13 describes the probability that the first content-based language model represents the first coefficient (for example, the first probability parameter) that increases as the calculated first probability parameter increases (the above). P in equation (3)<sub>j</sub>(W<sub>i</sub>)) And the value obtained by multiplying the calculated second probability parameter by the probability represented by the second content-based language model by the second coefficient (for example, the second probability parameter) that increases as the calculated second probability parameter increases. As the sum of, increases, a language model is created in which the probability that a specific word appears in the word string corresponding to the part corresponding to the processing target word string in the voice that is the basis of the input word string increases.
By the way, "t" in the equation (3) may be a section corresponding to a time frame used in the voice recognition process, or may be a time or the like representing a time point in the utterance.
The content-specific language model storage unit 15 may store a content-specific language model and a list of words (word list) having a high probability of appearing for each content. In this case, the language model creation unit 13 may be configured to increase the probability that a word included in the word list for the content having the highest score appears by a predetermined value in each section in the utterance.
The content estimation unit 12 may use the above-mentioned score (content appearance score) by changing the value estimated for each section without using it as it is. For example, a spoken word string may contain a word string that is not related to any content. In such a case, the content estimation unit 12 estimates the content represented by the word string from, for example, three types of content and a content of "no content", for a total of four types. Then, the content estimation unit 12 sets the scores of the other three types of contents by a predetermined value (for example, a predetermined ratio) in the section in which the content represented by the word string is estimated to be "no content". (For example, a value according to a certain ratio) may be changed.
Further, when all the calculated probability parameters (scores) are smaller than the preset lower limit threshold value, the content estimation unit 12 may set all the calculated probability parameters to the same value.
For example, as shown in FIG. 7A, it is assumed that all of the calculated probability parameters (scores) are smaller than the lower limit threshold value in a certain interval t2. In this case, as shown in FIG. 7B, the content estimation unit 12 sets all the probability parameters for this interval t2 to the same value (lower limit threshold value in this example).
According to this, it is possible to prevent the creation of a language model in which only the influence of any of the content-specific language models is largely reflected in the section in which the content represented by the processing target word string cannot be accurately specified. As a result, it is possible to create a language model that enables recognition of a word string corresponding to speech with high accuracy.
Further, for example, when the content represented by the word string is the content "personal name" related to the search condition of the TV program, the word such as "appearing" or "appearing" follows the word string. It is relatively likely to appear. Therefore, in the section following the section where the content represented by the word string is presumed to be the content "personal name", it is desirable that the score of the "personal name" immediately decreases in order to recognize the following word string with high accuracy. Absent.
Therefore, when the probability parameter (for example, the first probability parameter) acquired for a certain processing target word string is larger than the preset upper limit threshold value, the content estimation unit 12 is adjacent to the processing target word string. The probability parameter (for example, the first probability parameter) acquired for the processing target word string to be processed may be corrected so as to increase.
For example, as shown in FIG. 8A, it is assumed that the calculated probability parameter (score) is larger than the preset upper limit threshold value in a certain interval t2. In this case, as shown in FIG. 8B, the content estimation unit 12 corrects so as to increase the acquired score for the section t1 and the section t3 adjacent to the section t2.
Specifically, the content estimation unit 12 corrects the score so that the closer to the section t2 in the section t1, the closer the score is to the score acquired for the section t2. Similarly, in the section t3, the content estimation unit 12 corrects the score so that the closer to the section t2, the closer the score is to the score acquired for the section t2.
According to this, it is possible to recognize the word string corresponding to the voice with high accuracy even for the section adjacent to the section in which the content represented by the processing target word string is specified with relatively high accuracy. You can create a language model.
When the language model creation unit 13 outputs the created language model, it may output all the information included in the language model, or may output only the information specified from the outside.
(Operation) Next, the operation of the language model creating device 101 according to the first embodiment of the present invention will be described with reference to the flowchart shown in FIG.
As shown in FIG. 2, when the language model creation device 101 is activated, the content model and the content-specific language model are read from the storage device that realizes the content model storage unit 14 and the content-specific language model storage unit 15. , Each of which is initialized for reference from the content estimation unit 12 and the language model creation unit 13 (step S11).
On the other hand, the voice recognition hypothesis input unit 11 receives a voice recognition hypothesis from an external voice recognition device in response to a notification indicating the end of the voice recognition process, and outputs the received voice recognition hypothesis to the content estimation unit 12 (step S12). , Part of the language model creation process). The voice recognition hypothesis input unit 11 may be configured to receive the voice recognition hypothesis input by the user.
When the voice recognition hypothesis is input via the voice recognition hypothesis input unit 11, the content estimation unit 12 uses each processing target word string in the voice recognition hypothesis based on the content model stored in the content model storage unit 14. A probability parameter representing the probability that the content represented by (for example, each word) is a specific content is calculated (step S13, a part of the language model creation step).
Next, the language model creation unit 13 uses the probability parameters output from the content estimation unit 12 and the content-specific language model stored by the content-specific language model storage unit 15 as the basis of the speech recognition hypothesis. A language model representing the probability that a specific word appears in the word string corresponding to the part corresponding to the processing target word string in the sound has been created, and the created language model is output (step S14, language model creation step). Part of).
As described above, according to the first embodiment of the language model creating device according to the present invention, the language model creating device 101 has a probability that the content represented by the processing target word string is the first content and the processing target word. A language model is created based on the probability that the content represented by the column is the second content, the first content-based language model, and the second content-based language model.
As a result, it is possible to avoid creating a language model based only on the content-specific language model relating to the content different from the content represented by the processing target word string. That is, it is possible to create a language model by reliably using the content-specific language model related to the content represented by the processing target word string. As a result, it is possible to create a language model that enables recognition of a word string corresponding to speech with high accuracy.
Further, according to the above configuration, in order to determine which content-specific language model to use, it is not necessary to perform a process of evaluating the recognition result when each content-specific language model is used, so that the language model creation device The processing load of 101 can be reduced.
That is, according to the language model creating device 101, it is possible to create a language model capable of recognizing a word string corresponding to a voice with high accuracy while preventing an excessive calculation load.
Further, according to the first embodiment, the greater the probability that the content represented by the processing target word string is the first content, the greater the probability that the probability represented by the first content-specific language model is reflected in the language model. can do. Similarly, the greater the probability that the content represented by the processing target word string is the second content, the greater the degree to which the probability represented by the second content-specific language model is reflected in the language model. As a result, it is possible to create a language model that enables recognition of a word string corresponding to speech with high accuracy.
<Second Embodiment> Next, the voice recognition device according to the second embodiment of the present invention will be described with reference to FIG. FIG. 9 is a block diagram showing the function of the voice recognition device 201 according to the second embodiment of the present invention.
The voice recognition device 201 is an information processing device having the same configuration as the language model creation device 101 according to the first embodiment. The functions of the voice recognition device 201 include a voice recognition unit (speech recognition means) 21, a voice recognition model storage unit 22, and a language model update unit (language model creation means) 24.
The voice recognition device 201 generates a voice recognition hypothesis as an input word string by performing a voice recognition process that recognizes a word string corresponding to the input voice, and based on the generated voice recognition hypothesis, the first Similar to the language model creating device 101 according to the embodiment, the language model is created, and the voice recognition process is performed again based on the created language model.
The voice recognition unit 21 generates an input word string as a voice recognition hypothesis (for example, a word graph) by performing the voice recognition process for recognizing a word string corresponding to the voice input via the input device. The voice recognition unit 21 may be configured to input voice by receiving voice information representing voice from another information processing device. The voice recognition unit 21 includes a model (a model for performing voice recognition processing, which includes an acoustic model, a language model, a word dictionary, etc.) stored in the voice recognition model storage unit 22 for the entire section of speech. ), The voice recognition process is performed by searching for a word string that matches the voice. In this example, the acoustic model is a hidden Markov model and the language model is a word trigram.
The voice recognition unit 21 refers to the language model output by the language model update unit 24 when performing the voice recognition process. For example, the voice recognition unit 21 refers to the language model of the equation (3) in a certain time frame "f" during the voice recognition process, and refers to the word w.<sub>i</sub>When calculating the probability that will appear, for the section "t" corresponding to that "f", P<sub>t</sub>(W<sub>i</sub>). In this example, the time frame represents a unit for converting the voice to be recognized into a feature amount for recognition.
When the voice recognition process is performed before the language model update unit 24 creates the language model corresponding to the utterance, the voice recognition unit 21 refers to the language model stored in the voice recognition model storage unit 22. Further, the voice recognition unit 21 may be configured to use the sum of the probabilities represented by the plurality of content-specific language models stored in the content-specific language model storage unit 245 as the probability that a word appears.
The voice recognition device 201 is recognized by the voice recognition unit 21 and the voice recognition process in which the voice recognition unit 21 recognizes a word string corresponding to the input voice based on the language model created by the language model update unit 24. Based on the word string, the language model update unit 24 executes a language model creation process for creating a language model and an iterative process for alternately repeating the process.
The higher the accuracy of the input word string (to the extent that it matches the true word string), the higher the accuracy of acquiring the first probability parameter and the second probability parameter. Further, the higher the accuracy of the first probability parameter and the second probability parameter, the more accurately a language model capable of recognizing the word string corresponding to the speech can be created. Therefore, according to the above configuration, it is possible to recognize the word string corresponding to the voice with even higher accuracy.
The voice recognition unit 21 ends the iterative process when a predetermined end condition is satisfied based on the generated voice recognition hypothesis or the language model used in the voice recognition process. When the voice recognition unit 21 finishes the iterative process, the voice recognition unit 21 outputs the latest voice recognition hypothesis acquired at that time as a voice recognition result. The voice recognition unit 21 may select and output the voice recognition result from the voice recognition hypotheses accumulated up to that point.
The end condition is a condition that the word string recognized by the previous voice recognition process and the word string recognized by the current voice recognition process match. The end condition may be a condition that the number of times the voice recognition process is executed is larger than the preset number of times. Further, the end condition may be a parameter of the language model created by the language model creation unit 243, or a condition determined based on the estimation result output by the content estimation unit 242 or its score.
The language model update unit 24 has the same function as the language model creation device 101 according to the first embodiment. The language model update unit 24 includes a voice recognition hypothesis input unit 241 similar to the voice recognition hypothesis input unit 11, a content estimation unit 242 similar to the content estimation unit 12, and a language model creation unit 243 similar to the language model creation unit 13. And the content model storage unit 244 similar to the content model storage unit 14, and the content-specific language model storage unit similar to the content-specific language model storage unit 15 (content-specific language model storage means, content-specific language model storage processing means, content). Another language model storage processing step) 245 and.
When the voice recognition unit 21 determines that the end condition is not satisfied, the language model update unit 24 sets the voice recognition hypothesis output from the voice recognition unit 21, the stored content model, and the stored content. Create a language model based on the language model and output the created language model.
In this example, the content-based language model storage unit 245 stores the word trigram in the storage device as the content-based language model. The language model creation unit 243 uses the score representing the probability that the content represented by the processing target word string is a specific content, the stored language model for each content, and the above equation (3) for each processing target word string. Create a language model based on.
The language model update unit 24 creates a language model based on the received voice recognition hypothesis each time it receives a voice recognition hypothesis from the voice recognition unit 21 until the voice recognition unit 21 determines that the end condition is satisfied. .. In the language model created for the kth time, the word w<sub>i</sub>The probability that<sub>t, k</sub>(W<sub>i</sub>) (See equation (4) below). The voice recognition unit 21 performs the (k + 1) th voice recognition process with reference to this language model, and outputs the voice recognition hypothesis.<maths num="4"><img id="000005" he="22" wi="159" file="0005598331.tif" img-format="tif" img-content="drawing" /></maths>
Then, the content estimation unit 242 inputs this voice recognition hypothesis, and as the (k + 1) th content estimation result, the appearance score α of each content is α.<sub>j, k + 1</sub>(T) is output. The language model creation unit 243 uses this appearance score to (k + 1) the language model P for the third time.<sub>t, k + 1</sub>(W<sub>i</sub>) (See equation (5) below). In this way, by repeatedly updating the speech recognition hypothesis and the content estimation result, the accuracy of each is gradually improved.<maths num="5"><img id="000006" he="21" wi="159" file="0005598331.tif" img-format="tif" img-content="drawing" /></maths>
In the above iterative processing, when the voice recognition unit 21 performs the second and subsequent processing, the voice recognition unit 21 replaces the voice recognition processing in which the voice is input, and instead of the voice recognition processing, the previous voice recognition hypothesis (word graph, etc.) Rescore processing may be performed by inputting.
(Operation) Next, the operation of the voice recognition device according to the second embodiment of the present invention will be described with reference to the flowchart shown in FIG.
As shown in FIG. 10, when the voice recognition device 201 is activated, the voice recognition model and the language model are transferred from the storage device that realizes the voice recognition model storage unit 22 and the content-specific language model storage unit 245. The reading is performed, and the initialization process for referencing each of them from the voice recognition unit 21 and the language model updating unit 24 is performed (step S21).
On the other hand, the voice recognition unit 21 receives the voice input from the outside via the input device in response to the notification indicating the end of the voice input (step S22).
When the voice recognition unit 21 receives the voice, the voice recognition unit 21 is based on the voice recognition model stored in the voice recognition model storage unit 22 and the language model created by the language model update unit 24 for the received voice. Perform voice recognition processing (step S23).
The voice recognition device 201 determines whether or not the end condition is satisfied based on the voice recognition hypothesis output by the voice recognition unit 21 performing the voice recognition process (step S24). When the end condition is satisfied, the voice recognition device 201 determines "Yes" and outputs the latest voice recognition hypothesis acquired at that time as a voice recognition result (step S27).
On the other hand, when the end condition is not satisfied, the voice recognition device 201 creates a language model by determining "No" in step S24 and executing the processes of steps S25 and S26. This process is the same as the process of steps S13 and S14 of FIG.
As described above, according to the second embodiment of the voice recognition device according to the present invention, in the voice recognition device 201, the probability that the content represented by the processing target word string is the first content and the processing target word string are A language model is created based on the probability that the content to be represented is the second content, the first content-specific language model, and the second content-specific language model. Then, the voice recognition device 201 performs a voice recognition process for recognizing a word string corresponding to the voice based on the created language model. As a result, it is possible to recognize the word string corresponding to the voice with high accuracy while preventing the calculation load of the voice recognition device 201 from becoming excessive.
Further, according to the second embodiment, the greater the probability that the content represented by the processing target word string is the first content, the greater the probability that the probability represented by the first content-specific language model is reflected in the language model. can do. Similarly, the greater the probability that the content represented by the processing target word string is the second content, the greater the degree to which the probability represented by the second content-specific language model is reflected in the language model. As a result, the word string corresponding to the voice can be recognized with high accuracy.
In addition, the voice recognition device 201 has a voice recognition process in which the voice recognition unit 21 recognizes a word string corresponding to the input voice based on the language model created by the language model update unit 24, and a voice recognition unit 21. Based on the word string recognized by, the language model update unit 24 executes a language model creation process for creating a language model and an iterative process for alternately repeating.
By the way, as the accuracy of the input word string (the degree of matching with the true word string) becomes higher, the first probability parameter and the second probability parameter can be acquired with higher accuracy. Further, the higher the accuracy of the first probability parameter and the second probability parameter, the more accurately a language model capable of recognizing the word string corresponding to the speech can be created. Therefore, according to the above configuration, it is possible to recognize the word string corresponding to the voice with even higher accuracy.
<Third Embodiment> Next, the language model creating device according to the third embodiment of the present invention will be described with reference to FIG. The function of the language model creating device 301 according to the third embodiment includes a content-based language model storage unit (content-specific language model storage means) 35 and a language model creation unit (language model creation means) 33.
In the content-based language model storage unit 35, the first content-based language model representing the probability that a specific word appears in the word string representing the first content and the specific word in the word string representing the second content are included. A second content-based language model representing the probability of appearance is stored in the storage device.
The language model creation unit 33 is at least one of the input word strings that are the input word strings that are included in the voice recognition hypothesis generated by performing the voice recognition process that recognizes the word strings corresponding to the voice. A first probability parameter representing the probability that the content represented by the processing target word string, which is a part, is the first content, and a second probability parameter representing the probability that the content represented by the processing target word string is the second content. And get.
The language model creation unit 33 includes the acquired first probability parameter, the acquired second probability parameter, the first content-specific language model stored by the content-specific language model storage unit 35, and the content. Based on the second content-specific language model stored by the different language model storage unit 35, a specific word appears in the word string corresponding to the portion of the voice to be processed. Performs language model creation processing to create a language model that represents probability.
According to this, the language model creation device 301 has a probability that the content represented by the processing target word string is the first content, a probability that the content represented by the processing target word string is the second content, and the first content. A language model is created based on another language model and a second content-based language model.
As a result, it is possible to avoid creating a language model based only on the content-specific language model relating to the content different from the content represented by the processing target word string. That is, it is possible to create a language model by reliably using the content-specific language model related to the content represented by the processing target word string. As a result, it is possible to create a language model that enables recognition of a word string corresponding to speech with high accuracy.
Further, according to the above configuration, in order to determine which content-specific language model to use, it is not necessary to perform a process of evaluating the recognition result when each content-specific language model is used, so that the language model creation device The processing load of 301 can be reduced.
That is, according to the language model creating device 301, it is possible to create a language model capable of recognizing a word string corresponding to a voice with high accuracy while preventing an excessive calculation load.
In this case, the language model creation means is The value obtained by multiplying the first coefficient, which increases as the acquired first probability parameter increases, with the probability represented by the first content-based language model, and the value obtained, which increases as the acquired second probability parameter increases. The larger the sum of the second coefficient and the value obtained by multiplying the value represented by the probability represented by the second content-based language model, the more the above-mentioned specification is made in the word string corresponding to the part corresponding to the above-mentioned processing target word string in the above-mentioned voice. It is preferable to be configured to create the above language model in which the probability that the word of is large.
According to this, the greater the probability that the content represented by the processing target word string is the first content, the greater the degree to which the probability represented by the first content-specific language model is reflected in the language model. Similarly, the greater the probability that the content represented by the processing target word string is the second content, the greater the degree to which the probability represented by the second content-specific language model is reflected in the language model. As a result, it is possible to create a language model that enables recognition of a word string corresponding to speech with high accuracy.
In this case, the language model creating means uses the language for each of the plurality of processing target word strings obtained by dividing the input word string with the boundary between words of the word string recognized by the voice recognition process as a boundary. It is preferable that the model creation process is performed.
According to this, when creating a language model, it is possible to use the information acquired when performing the voice recognition process. As a result, the content can be estimated accurately and the language model can be created quickly.
Further, in another aspect of the above language model creation device, The language model creating means has a language for each of a plurality of processing target word strings obtained by dividing the input word string with a position different from the boundary between words of the word string recognized by the voice recognition process as a boundary. It is preferable that the model creation process is performed.
According to this, even if the boundary between actual words in the utterance is different from the boundary between words of the word string recognized by the speech recognition process, the word string corresponding to the speech is recognized with high accuracy. You can create a language model that allows you to do this.
In this case, the language model creation means is When both the acquired first probability parameter and the acquired second probability parameter are smaller than the preset lower limit threshold value, the acquired first probability parameter and the acquired second probability parameter are obtained. It is preferable that both of the probability parameters of are set to the same value.
According to this, it is possible to prevent a language model in which only the influence of one of the content-specific language models is largely reflected in the speech section corresponding to the processing target word string whose content cannot be accurately specified. Can be done. As a result, it is possible to create a language model that enables recognition of a word string corresponding to speech with high accuracy.
In this case, the language model creation means is When the first probability parameter acquired for a certain processing target word string is larger than a preset upper limit threshold value, the acquired first probability parameter for the processing target word string adjacent to the processing target word string is obtained. It is preferably configured to be corrected to increase the probability parameter of 1.
According to this, even for the voice section corresponding to the processing target word string adjacent to the processing target word string whose content is specified with relatively high accuracy, the word string corresponding to the voice is recognized with high accuracy. You can create a language model that allows you to do this.
In this case, it is preferable that the language model creating means is configured to acquire the first probability parameter and the second probability parameter based on the theory of the conditional random field.
In this case, the language model creating means has word attribute information representing the attributes of the words included in the processing target word string, and the voice acquired when performing the voice recognition process for recognizing the processing target word string from the voice. It is preferable that the first probability parameter and the second probability parameter are acquired based on at least one of the recognition processing information.
In this case, it is preferable that the word attribute information includes at least one of information representing a word surface layer, information representing reading, and information representing a part of speech.
In this case, the voice recognition processing information includes information representing recognition reliability, which is the reliability of the recognition result by the voice recognition processing, information representing the duration time, which is the duration of one sound, and preceding silence. It is preferable to include at least one piece of information indicating the presence or absence of.
In this case, the language model creation means is Information indicating the position of the processing target word string in the input word string, information indicating that the input word string contains a plurality of one word, information indicating the concatenated state of the contents in the input word string, and the input word. It is preferable that the first probability parameter and the second probability parameter are acquired based on at least one of the information representing the relationship between each input word string when there are a plurality of columns. is there.
Further, the voice recognition device according to another embodiment of the present invention is A first content-specific language model that represents the probability that a specific word will appear in a word string that represents the first content, and a second content that represents the probability that the specific word will appear in the word string that represents the second content. Another language model, a content-based language model storage means to memorize, A processing target word string that is at least a part of an input word string that is a word string included in a voice recognition hypothesis generated by performing a voice recognition process that recognizes a word string corresponding to a voice and is an input word string. Acquires a first probability parameter representing the probability that the content represented by is the first content, and a second probability parameter representing the probability that the content represented by the processing target word string is the second content. At the same time, the acquired first probability parameter, the acquired second probability parameter, the stored first content-based language model, and the stored second content-based language model. Based on the above, a language model is created to create a language model representing the probability that the specific word appears in the word string corresponding to the part corresponding to the processing target word string in the voice. Means and A voice recognition means that performs voice recognition processing that recognizes a word string corresponding to an input voice based on the language model created by the above language model creation means. To be equipped.
According to this, the voice recognition device has a probability that the content represented by the processing target word string is the first content, a probability that the content represented by the processing target word string is the second content, and a first content-specific language. A language model is created based on the model and the second content-based language model. Then, the voice recognition device performs a voice recognition process for recognizing a word string corresponding to the voice based on the created language model. As a result, it is possible to recognize the word string corresponding to the voice with high accuracy while preventing the calculation load of the voice recognition device from becoming excessive.
In this case, the language model creation means is The value obtained by multiplying the first coefficient, which increases as the acquired first probability parameter increases, with the probability represented by the first content-based language model, and the value obtained, which increases as the acquired second probability parameter increases. The larger the sum of the second coefficient and the value obtained by multiplying the value represented by the probability represented by the second content-based language model, the more the above-mentioned specification is made in the word string corresponding to the part corresponding to the above-mentioned processing target word string in the above-mentioned voice. It is preferable to be configured to create the above language model in which the probability that the word of is large.
According to this, the greater the probability that the content represented by the processing target word string is the first content, the greater the degree to which the probability represented by the first content-specific language model is reflected in the language model. Similarly, the greater the probability that the content represented by the processing target word string is the second content, the greater the degree to which the probability represented by the second content-specific language model is reflected in the language model. As a result, the word string corresponding to the voice can be recognized with high accuracy.
in this case, The voice recognition means is configured to generate the input word string by performing the voice recognition process that recognizes the word string corresponding to the input voice. The language model creating means is configured to create the language model based on the input word string generated by the voice recognition means. The voice recognition means is preferably configured to perform the voice recognition process for recognizing the word string corresponding to the input voice again based on the language model created by the language model creation means. Is.
In this case, the voice recognition device Based on the language model created by the language model creating means, the voice recognition process that recognizes the word string corresponding to the input voice and the word string recognized by the voice recognition means Based on this, it is preferable that the language model creating means is configured to execute the iterative process of alternately repeating the language model creating process for creating the language model.
The higher the accuracy of the input word string (to the extent that it matches the true word string), the higher the accuracy of acquiring the first probability parameter and the second probability parameter. Further, the higher the accuracy of the first probability parameter and the second probability parameter, the more accurately a language model capable of recognizing the word string corresponding to the speech can be created. Therefore, according to the above configuration, it is possible to recognize the word string corresponding to the voice with even higher accuracy.
In this case, it is preferable that the voice recognition device is configured to end the iterative process when a predetermined end condition is satisfied.
In this case, it is preferable that the end condition is a condition that the word string recognized by the previous voice recognition process and the word string recognized by the current voice recognition process match. ..
Further, in another aspect of the voice recognition device, The end condition is preferably a condition that the number of times the voice recognition process is executed is larger than the preset threshold number of times.
In addition, the language model creation method, which is another form of the present invention, is A first content-based language model that represents the probability that a specific word will appear in a word string that represents the first content, and a second content that represents the probability that the specific word will appear in the word string that represents the second content. When another language model and is stored in the storage device, A word string to be processed that is at least a part of an input word string that is a word string included in a voice recognition hypothesis generated by performing a voice recognition process that recognizes a word string corresponding to a voice and is an input word string. Acquires a first probability parameter representing the probability that the content represented by is the first content, and a second probability parameter representing the probability that the content represented by the processing target word string is the second content. At the same time, the acquired first probabilistic parameter, the acquired second probabilistic parameter, the stored first content-based language model, and the stored second content-based language model. Based on the above, it is a method of creating a language model representing the probability that the specific word appears in the word string corresponding to the part corresponding to the processing target word string in the voice.
In this case, the above language model creation method is The value obtained by multiplying the first coefficient, which increases as the acquired first probability parameter increases, with the probability represented by the first content-based language model, and the value obtained, which increases as the acquired second probability parameter increases. The larger the sum of the second coefficient and the value obtained by multiplying the value represented by the probability represented by the second content-based language model, the more the above-mentioned specification is made in the word string corresponding to the part corresponding to the above-mentioned processing target word string in the above-mentioned voice. It is preferable to be configured to create the above language model in which the probability that the word of is large.
In addition, the language model creation program, which is another form of the present invention, For information processing equipment A first content-specific language model that represents the probability that a specific word will appear in a word string that represents the first content, and a second content that represents the probability that the specific word will appear in the word string that represents the second content. Another language model, a content-based language model storage processing means for storing in a storage device, A word string to be processed that is at least a part of an input word string that is a word string included in a voice recognition hypothesis generated by performing a voice recognition process that recognizes a word string corresponding to a voice and is an input word string. Acquires a first probability parameter representing the probability that the content represented by is the first content, and a second probability parameter representing the probability that the content represented by the processing target word string is the second content. At the same time, the acquired first probability parameter, the acquired second probability parameter, the stored first content-based language model, and the stored second content-based language model. Based on the above, a language model creating means for creating a language model representing the probability that the specific word appears in the word string corresponding to the part corresponding to the processing target word string in the voice, and It is a program to realize.
In this case, the language model creation means is The value obtained by multiplying the first coefficient, which increases as the acquired first probability parameter increases, with the probability represented by the first content-based language model, and the value obtained, which increases as the acquired second probability parameter increases. The larger the sum of the second coefficient and the value obtained by multiplying the value represented by the probability represented by the second content-based language model, the more the above-mentioned specification is made in the word string corresponding to the part corresponding to the above-mentioned processing target word string in the above-mentioned voice. It is preferable to be configured to create the above language model in which the probability that the word of is large.
Even the invention of the speech recognition device, the language model creation method, or the language model creation program having the above-described configuration has the same operation as the above-mentioned language model creation device. Can be achieved.
Although the present invention has been described above with reference to each of the above embodiments, the present invention is not limited to the above-described embodiments. Various modifications that can be understood by those skilled in the art can be made within the scope of the present invention in the configuration and details of the present invention.
Further, as another modification of the above embodiment, any combination of the above-described embodiment and modification may be adopted.
Further, although the program is stored in the storage device in each of the above embodiments, it may be stored in a recording medium readable by the CPU. For example, the recording medium is a portable medium such as a flexible disk, an optical disk, a magneto-optical disk, and a semiconductor memory.
The present invention enjoys the benefit of priority claim based on the patent application of Japanese Patent Application No. 2008-304564 filed on November 28, 2008 in Japan, and is disclosed in the patent application. All of the content is included herein.
The present invention can be applied to a voice recognition device or the like that performs voice recognition processing for recognizing a word string represented by the voice from the voice.
11 Speech recognition hypothesis input section 12 Content estimation unit 13 Language Model Creation Department 14 Content model storage unit 15 Content-specific language model memory 21 Speech recognition unit 22 Speech recognition model storage 24 language model update department 33 Language Model Creation Department 35 Content-specific language model storage 101 Language model creation device 201 Speech recognition device 241 Speech recognition hypothesis input section 242 Content estimation unit 243 Language Model Creation Department 244 Content model storage 245 Content-specific language model storage 301 Language model creation device
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| JP2004198597A | Cites | Japan | Search report |
| JP2004198597A | Cites | Japan | Examiner |
| WO2005122143A1 | Cites | World Intellectual Property Organization (WIPO) | Search report |
| WO2005122143A1 | Cites | World Intellectual Property Organization (WIPO) | Examiner |
| JP2005284209A | Cites | Japan | Search report |
| JP2005284209A | Cites | Japan | Examiner |
| WO2008001485A1 | Cites | World Intellectual Property Organization (WIPO) | Examiner |
| WO2008004666A1 | Cites | World Intellectual Property Organization (WIPO) | Examiner |
| JPH11259084A | Cites | Japan | Search report |
| JPH11259084A | Cites | Japan | Examiner |
5 members in 3 offices
Priority claims10
| Document | Office | Kind | Date |
|---|---|---|---|
| 2008304564 | Japan | A | |
| 2008304564 | Japan | A | |
| 2009004341 | Japan | W | |
| 2009004341 | Japan | W | |
| 2010540302 | Japan | A | |
| 20082008304564 | – | – | – |
| 2009004341 | – | – | – |
| JP20080304564 | – | – | – |
| JP20100540302 | – | – | – |
| WO2009JP04341 | – | – | – |
Members5
| Document | Office | Kind | |
|---|---|---|---|
| WO2010061507A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US2011231183A1 | United States of America | A1 | |
| JPWO2010061507A1 | Japan | A1 | |
| JP5598331B2This record | Japan | B2 | |
| US9043209B2 | United States of America | B2 |
9 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 | |
| Notification of extinguishment of power of attorneyJAPANESE INTERMEDIATE CODE: A7427RD07 | RD07 |
Numbers
- Publication, DOCDB
- 5598331
- Publication, EPODOC
- JP5598331B
- Application
- 2010540302
- Application, DOCDB
- 2010540302
- Application, EPODOC
- JP20100540302
Titles
- English
- Language model preparation device
Classification
- CPC, 4
- G10L15/197
- G10L15/1815
- G10L15/183
- G10L15/19
- IPC, 4
- G10L15 065
- G10L15 06
- G10L15 183
- G10L15 197