Natural language interface control system
34 claims: 9 independent, 25 dependent
- 1複数のデバイスを操作する自然言語インターフェースコントロールシステムであって、 三次元マイクロフォンセットと、 上記三次元マイクロフォンセットに接続された特徴抽出モジュール であって、音声検出を行い、発生音の開始部及び終了部における多数の余分サンプルを用いることによって開始部及び終了部の無音を見つけ、開始部及び終了部の無音を除去することにより、発生音の音声部分を取り出す特徴抽出モジュール と、 上記特徴抽出モジュールに接続されている音声認識モジュールであって、隠れマルコフモデルを用いて、異なる音声モデル間、及び、異なる文法間で切り替えを行うことができる音声認識モジュールと、 を有し、少なくとも一つの異なる音声モデル及び少なくとも一つの異なる文法がネットワークを通じてダウンロードされ、 さらに、自然言語インターフェースコントロールシステムは、 上記音声認識モジュールに接続された自然言語インタフェースモジュールと、 上記自然言語インターフェースモジュールに接続されたデバイスインターフェースとを有し、上記自然言語インターフェースモジュールは、ユーザからの非プロンプト式で開放型の自然言語要求に基づいて、上記デバイスインターフェースに接続された一又は二以上のタイプからなる複数のデバイスを操作し、 上記自然言語インターフェースモジュールは、上記複数のデバイスのそれぞれを、上記複数のデバイスのそれぞれに対応する異なる文法のそれぞれと複数の語彙目録のそれぞれに抽象化することを特徴とする自然言語インターフェースコントロールシステム。
- 2上記自然言語インターフェースモジュールに接続された上記複数のデバイスをさらに有することを特徴とする請求の範囲第1項記載の自然言語インターフェースコントロールシステム。
- 3上記音声認識モジュールはNグラムグラマーを用いることを特徴とする請求の範囲第1項記載の自然言語インターフェースコントロールシステム。
- 4上記自然言語インターフェースモジュールは確率的コンテキストフリーグラマーを用いることを特徴とする請求の範囲第1項記載の自然言語インターフェースコントロールシステム。
- 5上記三次元マイクロフォンセットは、平面マイクロフォンセットと、空間的に異なる平面に配置された少なくとも1つのリニアマイクロフォンセットとを備えることを特徴とする請求の範囲第1項記載の自然言語インターフェースコントロールシステム。
- 6上記デバイスインターフェースは無線デバイスインターフェースからなることを特徴とする請求の範囲第1項記載の自然言語インターフェースコントロールシステム。
- 7上記自然言語インターフェースコントロールシステムに接続された外部ネットワークインターフェースをさらに有することを特徴とする請求の範囲第1項記載の自然言語インターフェースコントロールシステム。
- 8上記三次元マイクロフォンセットが有する第1のマイクロフォンセットと、上記特徴抽出モジュールと、上記音声認識モジュールと、上記自然言語インターフェースモジュールとを有するリモートユニットをさらに有することを特徴とする請求の範囲第1項記載の自然言語インターフェースコントロールシステム。
- 9上記リモートユニットに接続されたベースユニットをさらに有することを特徴とする請求の範囲第8項記載の自然言語インターフェースコントロールシステム。
- 10上記ベースユニットは上記三次元マイクロフォンセットが有する第2のマイクロフォンセットを含むことを特徴とする請求の範囲第9項記載の自然言語インターフェースコントロールシステム。
- 11上記第1のマイクロフォンセットと上記第2のマイクロフォンセットとは、上記三次元マイクロフォンセットを実現することを特徴とする請求の範囲第1項記載の自然言語インターフェースコントロールシステム。
- 12複数のデバイスを操作する自然言語インターフェースコントロールシステムであって、 三次元マイクロフォンセットと、 上記三次元マイクロフォンセットに接続された特徴抽出モジュール であって、音声検出を行い、発生音の開始部及び終了部における多数の余分サンプルを用いることによって開始部及び終了部の無音を見つけ、開始部及び終了部の無音を除去することにより、発生音の音声部分を取り出す特徴抽出モジュール と、 上記特徴抽出モジュールに接続されている音声認識モジュールであって、隠れマルコフモデルを用いて、異なる音声モデル間、及び、異なる文法間で切り替えを行うことができる音声認識モジュールと、 上記音声認識モジュールに接続された自然言語インタフェースモジュールと、 上記自然言語インターフェースモジュールに接続されたデバイスインターフェースとを有し、上記自然言語インターフェースモジュールは、ユーザからの非プロンプト式で開放型の自然言語要求に基づいて、上記デバイスインターフェースに接続された一又は二以上のタイプからなる複数のデバイスを操作し、 上記自然言語インターフェースモジュールは、上記複数のデバイスのそれぞれを、上記複数のデバイスのそれぞれに対応する複数の文法のそれぞれと複数の語彙目録のそれぞれに抽象化することを特徴とする自然言語インターフェースコントロールシステム。
- 13複数のデバイスを操作する自然言語インターフェースコントロールシステムであって、 三次元マイクロフォンセットと、 上記三次元マイクロフォンセットに接続された特徴抽出モジュール であって、音声検出を行い、発生音の開始部及び終了部における多数の余分サンプルを用いることによって開始部及び終了部の無音を見つけ、開始部及び終了部の無音を除去することにより、発生音の音声部分を取り出す特徴抽出モジュール と、 上記特徴抽出モジュールに接続されている音声認識モジュールであって、隠れマルコフモデルを用いて、異なる音声モデル間、及び、異なる文法間で切り替えを行うことができる音声認識モジュールと、 上記音声認識モジュールに接続された自然言語インタフェースモジュールと、 上記自然言語インターフェースモジュールに接続されたデバイスインターフェースとを有し、上記自然言語インターフェースモジュールは、ユーザからの非プロンプト式で開放型の自然言語要求に基づいて、上記デバイスインターフェースに接続された一又は二以上のタイプからなる複数のデバイスを操作し、 上記自然言語インターフェースモジュールは、アテンションワードを受け取って認識すると、上記非プロンプト式で開放型のユーザ要求をサーチすることを特徴とする自然言語インターフェースコントロールシステム。
- 14複数のデバイスを操作する自然言語インターフェースコントロールシステムであって、 三次元マイクロフォンセットと、 上記三次元マイクロフォンセットに接続された特徴抽出モジュール であって、音声検出を行い、発生音の開始部及び終了部における多数の余分サンプルを用いることによって開始部及び終了部の無音を見つけ、開始部及び終了部の無音を除去することにより、発生音の音声部分を取り出す特徴抽出モジュール と、 上記特徴抽出モジュールに接続されている音声認識モジュールであって、隠れマルコフモデルを用いて、異なる音声モデル間、及び、異なる文法間で切り替えを行うことができる音声認識モジュールと、 上記音声認識モジュールに接続された自然言語インタフェースモジュールと、 上記自然言語インターフェースモジュールに接続されたデバイスインターフェースとを有し、上記自然言語インターフェースモジュールは、ユーザからの非プロンプト式で開放型の自然言語要求に基づいて、上記デバイスインターフェースに接続された一又は二以上のタイプからなる複数のデバイスを操作し、 上記自然言語インターフェースモジュールは、アテンションワードを受け取って認識すると、文法、音声モデル、語彙目録のコンテキスト切り換えを行うことを特徴とする自然言語インターフェースコントロールシステム。
- 15複数のデバイスを操作する自然言語インターフェースコントロールシステムであって、 三次元マイクロフォンセットと、 上記三次元マイクロフォンセットに接続された特徴抽出モジュール であって、音声検出を行い、発生音の開始部及び終了部における多数の余分サンプルを用いることによって開始部及び終了部の無音を見つけ、開始部及び終了部の無音を除去することにより、発生音の音声部分を取り出す特徴抽出モジュール と、 上記特徴抽出モジュールに接続されている音声認識モジュールであって、隠れマルコフモデルを用いて、異なる音声モデル間、及び、異なる文法間で切り替えを行うことができる音声認識モジュールと、 上記音声認識モジュールに接続された自然言語インタフェースモジュールと、 上記自然言語インターフェースモジュールに接続されたデバイスインターフェースとを有し、上記自然言語インターフェースモジュールは、ユーザからの非プロンプト式で開放型の自然言語要求に基づいて、上記デバイスインターフェースに接続された一又は二以上のタイプからなる複数のデバイスを操作し、 さらに、自然言語インターフェースコントロールシステムは、 上記複数のデバイスそれぞれについての異なる文法を記憶するためのグラマーモジュールを有することを特徴とする自然言語インターフェースコントロールシステム。
- 16複数のデバイスを操作する自然言語インターフェースコントロールシステムであって、 三次元マイクロフォンセットと、 上記三次元マイクロフォンセットに接続された特徴抽出モジュール であって、音声検出を行い、発生音の開始部及び終了部における多数の余分サンプルを用いることによって開始部及び終了部の無音を見つけ、開始部及び終了部の無音を除去することにより、発生音の音声部分を取り出す特徴抽出モジュール と、 上記特徴抽出モジュールに接続されている音声認識モジュールであって、隠れマルコフモデルを用いて、異なる音声モデル間、及び、異なる文法間で切り替えを行うことができる音声認識モジュールと、 上記音声認識モジュールに接続された自然言語インタフェースモジュールと、 上記自然言語インターフェースモジュールに接続されたデバイスインターフェースとを有し、上記自然言語インターフェースモジュールは、ユーザからの非プロンプト式で開放型の自然言語要求に基づいて、上記デバイスインターフェースに接続された一又は二以上のタイプからなる複数のデバイスを操作し、 さらに、自然言語インターフェースコントロールシステムは、 上記複数のデバイスそれぞれについての異なる音声モデルを記憶するための音声モデルモジュールを有することを特徴とする自然言語インターフェースコントロールシステム。
- 17複数のデバイスを操作する自然言語インターフェースコントロールシステムであって、 第1のマイクロフォンと、 上記第1のマイクロフォンに接続された特徴抽出モジュール であって、音声検出を行い、発生音の開始部及び終了部における多数の余分サンプルを用いることによって開始部及び終了部の無音を見つけ、開始部及び終了部の無音を除去することにより、発生音の音声部分を取り出す特徴抽出モジュール と、 上記特徴抽出モジュールに接続されている音声認識モジュールと、 上記音声認識モジュールに接続された自然言語インタフェースモジュールと、 上記自然言語インターフェースモジュールに接続されたデバイスインターフェースとを有し、当該自然言語インターフェースモジュールは、ユーザからの非プロンプト式で開放型の自然言語要求に基づいて、上記デバイスインターフェースに接続された一又は二以上のタイプからなる複数のデバイスを操作し、 さらに、自然言語インターフェースコントロールシステムは、 上記自然言語インターフェースコントロールシステムに接続されている外部ネットワークインタフェースを有し、上記自然言語インターフェースモジュールは、上記複数のデバイスのそれぞれを、上記複数のデバイスのそれぞれに対応する複数の文法のそれぞれと複数の語彙目録のそれぞれに抽象化することを特徴とする自然言語インターフェースコントロールシステム。
- 18上記自然言語インターフェースモジュールに接続されている上記複数のデバイスをさらに有することを特徴とする請求の範囲第17項記載の自然言語インターフェースコントロールシステム。
- 19上記音声認識モジュールはNグラムグラマーを用いることを特徴とする請求の範囲第17項記載の自然言語インターフェースコントロールシステム。
- 20上記自然言語インターフェースモジュールは確率的コンテキストフリーグラマーを用いることを特徴とする請求の範囲第17項記載の自然言語インターフェースコントロールシステム。
- 21上記第1のマイクロフォンは、平面マイクロフォンセットと、空間的に異なる平面に配置された少なくとも1つのリニアマイクロフォンセットとを備える三次元マイクロフォンセットを有することを特徴とする請求の範囲第17項記載の自然言語インターフェースコントロールシステム。
- 22上記自然言語インターフェースモジュールは、アテンションワードを受け取って認識すると、非プロンプト式で開放型のユーザ要求をサーチすることを特徴とする請求の範囲第17項記載の自然言語インターフェースコントロールシステム。
- 23上記自然言語インターフェースモジュールは、アテンションワードを受け取って認識すると、文法、音声モデル、語彙目録のコンテキスト切り換えを行うことを特徴とする請求の範囲第17項記載の自然言語インターフェースコントロールシステム。
- 24上記複数のデバイスのそれぞれについて異なる文法を記憶するグラマーモジュールをさらに有することを特徴とする請求の範囲第17項記載の自然言語インターフェースコントロールシステム。
- 25上記複数のデバイスのそれぞれについて異なる音声モデルを記憶する音声モデルモジュールをさらに有することを特徴とする請求の範囲第17項記載の自然言語インターフェースコントロールシステム。
- 26上記デバイスインターフェースは無線デバイスインターフェースからなることを特徴とする請求の範囲第17項記載の自然言語インターフェースコントロールシステム。
- 27上記第1のマイクロフォンと、上記特徴抽出モジュールと、上記音声認識モジュールと、上記自然言語インターフェースモジュールとを有するリモートユニットをさらに有することを特徴とする請求の範囲第17項記載の自然言語インターフェースコントロールシステム。
- 28上記リモートユニットに接続されたベースユニットをさらに有することを特徴とする請求の範囲第27項記載の自然言語インターフェースコントロールシステム。
- 29上記ベースユニットは第2のマイクロフォンセットを有することを特徴とする請求の範囲第28項記載の自然言語インターフェースコントロールシステム。
- 30上記第1のマイクロフォンは、第1のマイクロフォンセットを有し、当該第1のマイクロフォンセットと上記第2のマイクロフォンセットは、三次元マイクロフォンセットを実現することを特徴とする請求の範囲第29項記載の自然言語インターフェースコントロールシステム。
- 31上記外部ネットワークインタフェースに接続されている中央データベースであって、文法、音声モデル、抽象化されたデバイス、プログラミング情報、語彙目録のうちの少なくとも一つを有する中央データベースをさらに有することを特徴とする請求の範囲第17項記載の自然言語インターフェースコントロールシステム。
- 32上記中央データベースは、外部ネットワークを介して上記外部ネットワークインタフェースに接続されていることを特徴とする請求の範囲第31項記載の自然言語インターフェースコントロールシステム。
- 33上記外部ネットワーク及び上記中央データベースに接続されているリモートサーバをさらに有することを特徴とする請求の範囲第32項記載の自然言語インターフェースコントロールシステム。
- 34他の自然言語インターフェースコントロールシステムと、 上記他の自然言語インターフェースコントロールシステム及び上記外部ネットワークに接続されている他の外部ネットワークインタフェースと、をさらに有することを特徴とする請求の範囲第32項記載の自然言語インターフェースコントロールシステム。
Independent claims34
1 paragraph, as filed
[0001] This application is a US provisional patent filed by Konopka, A NATURAL LANGUAGE INTERFACE FOR PERSONAL ELECTRONIC PRODUCTS on October 19, 1999, under 35 USC § 119 (e), 35 USC § 119 (e). It claims the priority of application No. 60 / 160,281 and is described with reference to this US provisional patent application. [0002]<u style="single">Technical field to which the invention belongs</u>The present invention relates to a speech recognition method, and more particularly to a speech recognition method for recognizing speech in a natural language. Specifically, the present invention relates to a natural language speech recognition system used to control an application. [0003]<u style="single">Background of the invention</u>So far, many have sought a device that could close the gap between the artificial voice produced by the device and the voice produced by humans. In addition, voice recognition technology has made it possible for devices to recognize human voice. For example, speech recognition technology is used in many applications such as document creation processing, device control, and menu-based data input. [0004] Most users prefer to input voice in natural language format. Natural language voice input is written or verbal input in a natural way as if the user were actually talking to the device. On the other hand, speech input in the form of non-natural language has limitations in language syntax and language structure. In order to communicate with a device by voice input in non-natural language, the user needs to remember commands and requirements according to the language syntax and semantic language structure of the speech recognition engine and speak based on them. [0005] The advantage of a natural language interface system is that it is easy to interface with a device or system because the user does not have to remember the proper format to speak a command or request, but simply speaks in a conversational format. is there. On the other hand, the problem with natural language interface systems is that human natural language is difficult to realize because it has complex and variable "rules". [0006] Conventional natural language processing techniques are inefficient and inflexible in their ability to recognize the meaning of utterances in natural language. For this reason, it is necessary to limit the context of the user's natural language input, that is, to facilitate the processing of the input voice, and conventional natural language interface systems have a dialog-based format or prompt-. driven) method. The natural language interface controls the context of the audio being input to the system. For example, the natural language interface is realized by an automatic telephone system such as an automatic airline reservation system in natural language. Such a system prompts or prompts the user to speak within a context. For example, a natural language system asks a user in which city they want to fly. In this case, the system presents the user with the expected audio context. In this way, the natural language interface system looks for a natural language that indicates the city name. The system then prompts the user to tell them how many days they want to fly. Again, the natural language interface system provides the context of the answer. The problem is that users cannot enter open-ended information or requests. If the audio data received by the system is outside the context prompted by the system, the system either ignores the request, notifies the user that it does not understand the answer, or falls within the context of the prompt. May misunderstand the request. [0007] There is a need for an efficient natural language system where the context is not limited by natural language processing but by the user's voice. The present invention can address the above and other requirements. [0008]<u style="single">Outline of the invention</u>The present invention provides an open natural language interface control system that controls multiple devices, where the context is not defined by the natural language interface, but by user instructions and the capabilities of the plurality of devices. , To favor the above and other needs. [0009] In one embodiment, the present invention comprises a first microphone set, a feature extraction module connected to the first microphone set, and a voice recognition module connected to the feature extraction module, the voice recognition module being a hidden Markov model. Can be characterized by a natural language interface control system that operates multiple devices using. In addition, this system is equipped with a natural language interface module connected to the voice recognition module and a device interface connected to the natural language interface module. The natural language interface module is a non-prompt and open natural language from the user. Operate multiple devices connected to the device interface based on the request. [0010] In another embodiment, the present invention has a step of searching for attention words based on a first context having a first set of models, grammar and lexicons, and a second context when the attention word is found. The second context can be characterized by a speech recognition method with a second set of models, grammar and lexicons, with the step of switching to and searching for open user requests. [0011] In yet another embodiment, the invention receives an attention word indicating that an open natural language user request is received, a step of receiving an open natural language user request, and an open natural language user request. To match the most appropriate command corresponding to an open natural language request, and a natural language control method for one or more devices that has a step to send the command to each of the one or more devices, and this method. The means can be characterized. [0012] The above and other aspects, features and advantages of the present invention will be clarified by the following detailed specific description based on the accompanying drawings. [0013] In each drawing, the same reference numerals indicate the same components. [0014]<u style="single">Detailed description of the invention</u>Hereinafter, the best embodiment for carrying out the present invention will be described, but the present invention is not limited to this embodiment and is merely an embodiment of the present invention. The gist of the present invention should be construed with reference to the claims. [0015] FIG. 1 is a system-level block diagram showing a configuration of an embodiment of a natural language interface control system according to the present invention. As shown in FIG. 1, the natural language interface control system 102 (also referred to as NLICS102) includes a remote unit 104 and a base unit 106 (also referred to as a base station unit 106). The remote unit 104 includes a linear microphone set 108 and a speaker 112, and the base unit 106 includes a flat microphone set 110. The remote unit 104 of the natural language interface control system 102 is connected to a plurality of controllable devices 114. Further, the base unit 106 is connected to the external network 116. [0016] The natural language interface control system 102 connects the user to the plurality of devices 114 and controls the plurality of devices 114 during operation. The natural language interface control system 102 provides a natural language interface that allows a user to control one or more of a plurality of devices 114 simply by speaking to the natural language interface control system 102 in a natural conversational manner. NLICS102 can interpret the user's request in the natural language and send the appropriate command to each device to make the user's request. For example, when applying such a natural language interface control system 102 at home, the device 114 may be a television, stereo, videocassette recorder (VCR), digital videodisc (DVD) player, or the like. When the user wants to operate any of the devices 114, he simply says "I wanna watch TV" or speaks a similar natural language. The NLICS102 is a Hidden Markov model known in the art. It is equipped with a speech recognition module that detects speech using Models (HMMs)), interprets natural language using a natural language interface, and appropriately determines the probability of what the user request is. The natural language interface uses stochastic context free grammer (also called PCFG) rules and lexicons stored for each device 114. For this reason, the natural language interface module comprises a device abstraction module with each abstraction device 114 designed to interface with the NLICS102. Therefore, each device 114 is abstracted into a command set understood by each device 114. In addition, each abstraction is associated with an individual grammar and lexicon specific to each device. [0017] Once the user's request is determined at the desired level of trust, the natural language interface module sends a series of commands to the appropriate device to execute the user's request. For example, in response to a user request "I wanna watch TV", the natural language interface module sends a command to the appropriate device, turns on the TV and amplifier, and puts the TV and amplifier in the appropriate mode. Set to and set the volume to the appropriate level. It also updates the state and settings of each abstracted device stored internally. The command can also switch the TV channel to the desired channel known by the NLICS102 or requested by the user with an open natural language request. As yet another example, the user may request specific information such as "Do you have the album'Genesis'?" That the system answers "Yes". it can. The user then "Play that" or "Play the album" You can answer "Genesis (playing the album" Genesis ")". The system responds to this request by turning on the CD jukebox and amplifier, setting the amplifier to the appropriate mode, setting the appropriate volume level, selecting the appropriate album, and playing the album. The system also updates the abstracted device's internal storage state and settings in addition to the user's profile. This command signal is preferably transmitted over a radio frequency (RF) link or an infrared (IR) link, as is known in the art. [0018] Speech recognition technology is well known in the art, and device control based on verbal commands is known. When the user utters a predetermined voice command to the voice recognition control system, for example, the user may say "Turn on" to the controlled television receiver. The TV turns on accordingly. However, such an approach does not utilize natural or conversational languages, but rather a dialog context (dialog). It also does not abstract the device under control to retrieve the context). The system will not issue a command unless the exact prescribed voice command is issued. On the other hand, in this embodiment, a natural language interface module used for probabilistically determining the most appropriate meaning for oral utterance and issuing an appropriate command is realized. Therefore, the instructions from the user are issued in a very conversational form, and the user does not need to remember a specific command signal. For example, the user says "hey, let's watch TV , I wanna watch TV , turn on the TV , whatty a say we watch a little television The system can use the natural language interface module to probabilistically determine that the user wants to watch TV and understand it by the TV and other suitable devices. Issue the appropriate set of commands that you can. [0019] This conveniently removes the physical obstacle between the device 114 and the user. For example, the user does not need to know how to operate the device 114. For example, the user does not need to know how to operate the DVD player. The user simply says "I wanna watch DVD" and a command signal is sent to turn on the DVD player and start playing the DVD in the player. [0020] In addition, the natural language interface module clarifies the user's request when there is no certainty about what the user's request is. For example, suppose the user's request is "I want to watch a movie". However, the natural language interface module does not know which device the user wants to watch the movie on, a DVD player, VCR, or TV. In such cases, the natural language interface module uses a feedback mechanism such as a feedback module (eg, a text-speech module) and a speaker to instruct the user to clarify the request. For example, the natural language interface module responds to such a request by "Do you want to watch a movie on the DVD, VCR or television? (Which DVD, VCR, or TV do you want to watch the movie on?) ". In response, the user replies, for example, "DVD." [0021] [0021] Thus, this system is a true "natural language interface" that can accept "open" demands. The natural language interface control system 102 is not primarily a dialog or prompt "closed-end" system. For example, a known natural language system controls a conversation by prompting the user to provide some information and then the system attempts to identify the information obtained. For example, in an airline reservation system using natural language, the user is guided by a dialog in which the context is limited by a question from the system. For example, the system says "To what city would you like to fly? (Which city do you fly to?) ". The user then answers the destination city in natural language, and the system essentially tries to understand the answer by matching the answer with the city name. The system then prompts the user by asking "What date would you like to leave?" Limits the context of the text string being analyzed, that is, the text string. In NLICS102, on the other hand, the user, not the system, initiates the dialog. The user "I want to hear some" before being prompted by NLICS102 music (I want to hear some music) ". The context searched is not limited by system prompts, but by the capabilities of device 114 controlled by NLICS102. Therefore, the user requests NLICS102 to perform any of the tasks that each device under control can perform. If, for example, the user requests NLICS102 to perform functions that are not available on the device under control. That is, for example, if the user says "Make me some breakfast", such a request is not in the function programmed in the device under control, so NLICS102 makes such a request. Cannot be executed. For example, NLICS102 properly interprets phrases within the capabilities of device 114 and simply ignores other requests. Also, the feedback section of the natural language interface module can warn the user that the request is not available. [0022] In this embodiment, the natural language interface control system 102 is "always on" so that the user can make a request at any time and the system responds. However, to get the attention of NLICS102, the user says the "attention word" and then the request. This helps identify the user, prevent false positives for requests, and distinguish between normal conversation and background noise that is not related to NLICS. Inform NLICS102 in advance that a request will follow the attention word. Therefore, the microphone set used in NLICS only needs to search for attention words in the physical space defined by the microphone set. For example, if the attention word is programmed as "Mona", the user's request is "Mona," I wanna watch TV (Mona, I want to watch TV) ". This significantly reduces the amount of processing and search by the microphone set. [0023] In addition, individual users can have their own user-specific attention words. For example, in the home, the attention word of the first user is "Mona" and the attention word of the second user is "Thor". When NLICS102 hears the attention word "Mona", it considers the first user to be issuing the command. For example, the first user is "Mona, When you say "I wanna watch TV", the system not only turns on the TV (and other suitable devices), but also switches the TV to the channel of your choice of the first user. .. In this case, since the first user can also say the attention word of the second user, the true ID is not given. This mechanism only provides a means to tailor the NLICS102 to the individual user's tastes, pronunciations and habits. [0024] One of the features that allows NLICS102 to function efficiently is that each device 114 connected to NLICS102 is abstracted into an individual abstracted device, so that each individual grammar and lexicon is individually abstracted. It is stored for each device 114. For example, if the natural language interface module determines that the request is for a DVD player, it uses a grammar and lexicon specific to that particular context (ie, the DVD player's context) and inputs within the speech recognition module. Useful for processing voice data. This causes context switching in the speech recognition module. [0025] In some embodiments, the NLICS102 is configured to allow models used in HMMs or speech recognition modules for grammar to be streamed from secondary sources such as hard disks, CD-ROMs, DVDs, etc. at runtime. .. If the data is read, it can be used immediately without preprocessing. Therefore, many models and grammars can be stored separately from the memory of NLICS102, which improves the memory usage of the speech recognition module. [0026] In another embodiment, the NLICS 102 is implemented as two separate units, such as the remote unit 104 and the base unit 106. The base unit 106 is the "docking station" of the remote unit 104. Acting as a "station)", the remote unit 104 is connected to the base unit 106, for example, by a general purpose serial bus (USB) connection. In some embodiments, the remote unit 104 functions as a general purpose remote control for various devices as conventionally practiced by providing a button to be used by the user. In addition, base unit 106 provides the 1NLICS 102 with an external network interface. For example, the external network interface connects NLICS to an external network 116 such as a home local area network (LAN), an intranet, or the Internet. In this case, NLICS102 may newly download the grammar, HMM model, abstracted device, DC, DVD, TV and other programming information and / or lexicon stored in the central database in the external network 116. it can. [0027] Further, the base unit 106 functions as a secondary cache of the remote unit 104. The remote unit 104 includes a feature extraction module, a speech recognition module, and a natural language interface module, in addition to device interfaces for various devices. At this time, the base unit 106 is provided with a memory for storing a new model, grammar, and lexicon used in the remote unit 104. [0028] The remote unit 104 includes two conventional linear microphone sets 108 that receive audio signals. Further, the base unit 106 includes a flat microphone set 110 that takes in sound wave energy from a two-dimensional space. The NLICS102 uses both microphone sets 108 and 110 as appropriate to realize a three-dimensional microphone set that allows audio to be heard in a predetermined three-dimensional physical space by the two sets of microphone sets 108 and 110. In this case, the three-dimensional volume can be defined in a certain space. For example, the NLICS102 can be configured to listen to the volume of the space, including the sofa in the living room, where the user sits when operating each device. Therefore, the phase of the audio data from the source outside the predetermined space is attenuated, and the phases of the audio data from within the predetermined space are summed. [0029] Although this system has been described above, the natural language interface control system will be described in more detail below. [0030] FIG. 2 shows another embodiment of the present invention, and is a block diagram showing a configuration of a remote unit 104 of the natural language interface control system 102 of FIG. As shown in FIG. 2, the remote unit 104 includes a linear microphone set 108, a feature extraction module 202, a voice recognition module 204, a natural language interface control module 206, a system processing controller 208, a device interface 210, and a base. It includes a unit interface 212 (also referred to as a general-purpose serial bus (USB) interface 212) and a speaker 214. Also, each device 114 is shown. The voice recognition module 204 includes a voice decoder 216, an N-gram grammar module 218, and a voice model module 220. The natural language interface control module 206 includes a natural language interface module 222, a probabilistic context-free grammar module 224 (also referred to as PCFG module 224), a device abstraction module 226, and a feedback module 228. [0031] The system is described as two separate elements, the remote unit 104 and the base unit 106, respectively, and some preferred embodiments use the remote unit 104 and the base unit 106 as separate units. The core function of NLICS102 can be realized only in the remote unit 104. Here, first, the remote unit 104 will be described, and then the base unit 106 will be described. [0032] Audio data is input to the remote unit 104 via a linear microphone set 108, which is two narrow cardioid microphones that identify the source or user and distinguish them from interference noise. Such linear microphone sets are well known in the art. The linear microphone set 108 samples the input audio data from each microphone element, then time-matches and sums the data to produce an enhanced signal-to-noise ratio (SNR) of the input audio signal. [0033] Then, the voice data is sent to the feature extraction module 202. The feature extraction module 202 extracts a parameter or a feature vector representing related information of the input voice data. [0034] The feature extraction module 202 performs edge detection, signal conditioning, and feature extraction. According to one embodiment, audio edge detection is the 0th Keptral coefficient (0).<sup>th</sup> It is performed by noise estimation and energy detection based on Cepstral coefficient and zero-crossing statistics. For feature extraction and signal conditioning, Mel-frequency Cepstral coefficient (MFCC), delta information, and acceleration information are extracted. A 38-dimensional feature vector based on a 12.8 ms sample buffer with 50% overlap. Such a feature extraction module 202 and its function are well known in the art, and those skilled in the art can realize the feature extraction module by various methods. As described above, the output of the feature extraction module 202 is a series of feature vectors. [0035] The speech recognition module 204 is then generally referred to as "unmodeled events" such as out-of-vocabulary events, lack of fluency, environmental noise, etc. Acts as a continuous speech recognition device based on the Hidden Markov Model (HMM), which has the ability to remove. The speech recognition module 204 is under the control of the natural language interface module 222 and can switch between different speech models and different grammars based on the speech context determined by the natural language interface control module 206. The speech recognition module 204 has some features that are favorably used in the NLICS 102, but may be entirely conventional. Further, since the memory usage in the voice recognition module 204 is optimized, the required memory mainly reflects the amount of voice model data used. Hereinafter, the voice recognition module 204 and the natural language interface control module 206 will be described in more detail. [0036] The feature vector from the feature extraction module 202 is input to the speech recognition module 204. That is, it is input to the voice decoder 216 of the voice recognition module (SRM) 204. The speech recognition module (SRM) 204 requests the speech feature vector from the feature extraction module (FEM) 202 and uses the speech model set to find the one that best matches the corresponding vocalization and to the Hidden Markov Model (HMM). Its role is to reject non-speech events through a based approach. [0037] The model used by the voice decoder 216 is stored in the voice model module 220. These models are context-dependent or context-independent phonetic models, sub-word models, or whole word. Models), such as monophones, biphones and / or trophones. In one embodiment, the audio decoder 216 can dynamically switch between different models. For example, the audio decoder 216 can switch between a model based on triple sounds and a model based on single sounds. This is different from known systems. In known systems, there are a fixed number of states and a fixed number of Gaussians for each state, that is, the architecture of each phoneme is fixed. In contrast, with selection between models based on single, digraph, and triple notes, the architecture of these phonemes, such as the number of states for each phoneme of each type (single, digraph, triplet) and for each state. The number of Gaussian shapes can be varied to optimize space, speed, and accuracy. As is well known in the art, the input vocal sound is, for example, the Viterbi algorithm. It is analyzed by a model using algorithm) and a score is assigned to indicate how well the vocalized sound matches a given model. In addition, the model used by the speech decoder 216 is under the direct control of the natural language interface control module 206. This will be further described below. [0038] In addition, two garbage-modeling techniques are used. The Garbage filler model is stored in the voice model module 220 to model not only lack of fluency and "silence" but also background noise. These models are used by the voice decoder 216 to eliminate out-of-vocabulary (oov) events (events that are not in the vocabulary). The voice decoder 216 also provides online garbage calculation (online garbage). Use valculation) to eliminate out-of-vocabulary (oov) events. And if the scores are very close, N best candidates will be returned. Elimination of such out-of-vocabulary (oov) events is also well known in the art. [0039] In some embodiments, the rejection technique has been improved compared to techniques known in the art. The basic principle behind an HMM-based speech recognition system is to compare a speech to a number of speech models (from the speech model module 220) and find the model that best matches the speech. This means that the output of the speech recognition module 204 serves as a reference to the optimal match model (eg, a word). However, this is problematic when there is no model for spoken verbal language. In such cases, HMM-based systems typically try to find the closest match between the vocalization and the model and report the result. In many cases this is not desirable. That's because any sound picked up by an open mic is a reference to the model it produces. To prevent this, it may be preferable to determine if the vocalization is included in an in-vocabulary word. For example, the Viterbi score If score) exceeds the threshold, the vocalization is considered an in-vocabulary reward. If the Viterbi score of the vocalization does not exceed the threshold, the vocalization is considered out of vocabulary. Such a Viterbi score is obtained using the Viterbi algorithm. This algorithm considers a series of observations to calculate a single optimal state sequence by the HMM and its corresponding probabilities. However, experiments have shown that this is not a very accurate exclusion method. Instead, in many systems, by reprocessing vocal sounds by another HMM whose task is to represent all out-of-vocabulary events or filler sounds, i.e. by using a garbage model. It relies on comparing the original Viterbi score with the other Viterbi score obtained. Garbage score) can be defined as the difference between the logarithms of each of the two Viterbi scores divided by the number of frames in the vocalization by Equation 1 below. The garbage score indicates whether the vocalizations show a closer match to the word model or a closer match to the out-of-vocabulary model. Many variants have been proposed for how to eliminate out-of-vocabulary events. It is generally known that the silence time during vocalization also produces a high Viterbi score for the model in which the high energy speech portion should be modeled. This can be prevented to some extent by providing a new feature representing the energy of the audio signal in the feature extraction module 202. However, it still results in an inaccurate garbage score measurement. It has been found that if there is silence at the beginning or end of the vocalization and this silence at the beginning or end is not modeled, it will have a significant effect on the garbage score. The feature extraction module 202 performs voice detection so that the silence at the start and end parts is not included in the sample sent to the voice decoder 216 of the voice recognition module 204. However, finding the beginning and end of a vocal sound is a complex task for vocal sounds that begin or end with low energy. Fricatives are an example of a voice group where this is a problem. Fricatives are, for example, white noise. Characterized as wideband, low energy noise such as noise). As is known in the art, fricatives are sounds represented by phonemes such as "th" and "sh". Feature extraction module 202 attempts to solve this problem by making the best effort to find samples at the beginning and end. To ensure that low-energy sounds are included in the speech sample, the feature extraction module 202 includes a large number of extra samples at the beginning and end of the vocalization. If there is no low-energy sound at the beginning or end of the speech, the speech is considered to be spoken in isolation and silence is created and added to the speech sample, so that the speech decoder 216 The quarantine score will be distorted. In order to solve this problem, in one embodiment, a single state silence model that consumes the silence frame sent from the feature extraction module 202 is arranged before and after each model. The speech decoder 216 finds the closest matching series of models, and in addition to the word model, the silent model optimally matches the vocalization. In this way, the start and end indexes of the start and end silence parts of the vocal sound can be obtained and removed. In addition, the optimal matching word model is retained and reprocessed using only the pure speech portion of the vocalization, without the preceding and following silence models. Next, the out-of-vocabulary HMM processes the same part of the vocalization sound, and the garbage score can be calculated by the following equation (1). Retention and reprocessing are performed using only the pure voice portion of the vocalized sound, without a sound model. Next, the out-of-vocabulary HMM processes the same part of the vocalization sound, and the garbage score can be calculated by the following equation (1). Retention and reprocessing are performed using only the pure voice portion of the vocalized sound, without a sound model. Next, the out-of-vocabulary HMM processes the same part of the vocalization sound, and the garbage score can be calculated by the following equation (1). [0040] [Number 1]<img file="JP5118280B2_D0001.tif" />[0041] Here, w is the logarithm of the Viterbi score for the in-vocabulary reward voice model without the front and back silence models, and the vocalization does not include silence. Similarly, g is the logarithm of the corresponding score for the Out of Vocabulary HMM model. Also, n is the total number of frames in the vocal sound, and m is the number of frames consumed by the front and rear silence models. That is, using this exclusion technique, the system will be able to accurately extract the voice portion of the vocalized sound. As a result, the extraction of in-vocabulary rewards is improved, and the elimination of out-of-vocabulary events starting or ending with low-energy sounds such as fricatives is improved as compared with the conventional exclusion method. [0042] The N-gram grammar module 218 has the grammar used by the voice decoder 216. These grammars are the rules for building a lexicon, which is a dictionary of words and their pronunciation inputs. Also, the specific grammar used by the speech decoder 216 is controlled by the natural language interface module 222. In this example, the N-gram grammar is configured to use multiple grammar types or combinations of grammar types. When using complex languages (eg, controlled devices with a large number of controls and functions), it is advantageous to use the trigram grammer option. For smaller systems (eg devices with very simple controls and features), the bigram grammer option (bigram grammer) With option), memory and accuracy can be balanced. To obtain memory-efficient representations of bigrams and trigramgrammers, possible combinations of lexicon inputs may be represented by specific lexicon input labels or word groups. If you want to allow any lexicon entry to follow the lexicon entry, you can use the ergodic grammar option. [0043] In general, it is not intuitively possible to use the N-gram glamor in a device having a small signal reception range. The small signal receivable range means that the system only needs to recognize the audio for the controlled device 114 connected to the remote unit 104, and the rest of the audio may be classified as out of vocabulary. That is. However, the N-gram grammar module 218 allows the use of multiple grammars and types even in the case of the speech recognition module 204, which has a small signal receivable range. [0044] Another grammar used primarily in the speech decoder 216 as an exclusion method is word list grammer. The word list grammar is used to recalculate the Viterbi score for a fixed set of words and a subset of vocal sounds. [0045] The system incorporates various grammars in a way that allows immediate switching between grammar types and grammar rule sets under the control of a "context switch" or natural language interface module. It is important to be able to do this, as what humans say is heavily influenced by the context. For example, it is assumed that only one phrase (eg, the attention word mentioned above) initiates a dialog, and other phrases only occur following a question (eg, a natural language interface that clarifies an ambiguous request). This is especially true when the speaker is dealing with a different audience, or in the case of consumer electronics, i.e. different products such as televisions, DVD players, stereos, VCRs, etc. In an attempt to improve speech recognition accuracy and reduce the amount of processing required, the system provides a way to define a context to which only certain grammatical rules apply. If the context is known, the natural language interface module 222 can command the speech recognition module 204 to listen only to the expected phrases. For example, when the natural language interface module 222 determines that the user intends to operate the DVD player, the speech recognition module 204 is instructed to use the grammar type and grammar corresponding to the DVD player. Therefore, the voice decoder 216 extracts the appropriate grammar from the N-gram grammar module 218. It is also possible to switch contexts with a high precision level that uses a flag for each grammar rule or lexicon input to indicate which rule or word is available or unavailable for each rule or word. Further, depending on the system settings and grammar mode, it may be preferable to use only the lexicon input set when searching for the optimal guess word. You only need to define several lexicons and refer only to the relevant lexicons. [0046] The speech recognition module 204 can dynamically change the grammar to be used in consideration of the context of the input speech, and the lexicon depends on the selected grammar, so that the lexicon is dynamically changed. Be changed. [0047] Depending on the size of the system, that is, how much search is required in the audio decoder 216, the processing time can be reduced. For medium to large natural language interface control systems 102 (which have a large number of devices 114 under control), efficient execution of the Beam Search algorithm significantly reduces processing time. Will be done. This beam search algorithm keeps the number of guess words to a minimum in the Viterbi Search algorithm. In this case, all active guess words are compared at each discrete time step and the Viterbi score for the optimal guess word is calculated. Then, pruning, that is, pruning, can be performed by discarding all the guess words whose score is less than the value obtained by subtracting the predetermined exclusion threshold function from the maximum guess word score. This limits searches based on inferred words that are pruned and are not considered in the following time steps until the score of the corresponding model state exceeds the threshold. [0048] Another problem associated with large speech recognition systems is the amount of memory required to store a speech model. Fortunately, the number of sub word units (eg phonemes) used in NLICS102 is generally fixed, so more speech models refer to the same subword model for larger lexicon inputs. It will be. The required memory can be kept to a minimum by referencing the same model elements, such as subword models, model states and / or Gaussian forms, in the lexicon input. In exchange, the computational resources required will increase slightly. This indirect model reference can be used to represent speech at any level of abstraction (eg, phrases, words, subwords). You can combine these abstractions to form more abstraction units according to the lexicon, which you can refer to in your grammar definition. [0049] Token Passing is a well-known approach for searching for optimal word guessing by HMM. As is known in the art, in a connected word recognition system, once the processing of the pre-frame of the vocalization is complete, it is easy to find the last model state for the state sequence with the highest Viterbi score. be able to. However, this does not always give the optimal state (or word) sequence. To find the optimal state sequence, you need to do "back tracing". The traditional way to do this is for each state to have a pointer back to the previous optimal state for each frame. Backtracing can be done by moving the pointer back from the last model state of the state sequence with the highest Viterbi score. This means that if the system uses N states over T discrete time steps, the number of back pointers required is usually NT. As soon as you do this, the numbers will grow and the memory required will increase. Various methods have been proposed to minimize the memory required for the storage of such back pointers, but some of them do not allocate memory for each state, but "token (token)" to various states. It is based on the idea of patrolling "tokens)". [0050] According to an embodiment of the present invention, instead of storing one token pointer in each state, the voice decoder 216 holds the token pointer in each state using two columns S1 and S2. Column S1 holds a token pointer for each state and the previous frame, and column S2 holds a token pointer for each state and the current frame. When each state i goes back to find the previous optimal state j, there are two possible consequences. If the previous optimal state j is an element of the same speech model as i, the token pointer of state j in S1 is copied to position i in S2. Otherwise, a new token is created and stored at position i in S2. The new token has the same contents as the token i of S1, and a reference to the models m and iεm is added in the token history. When all states have been processed for the current frame, the pointers to structures S1 and S2 are swapped or swapped, and processing is repeated for the next frame. Therefore, token passing technology can solve the well-known problems of HMM-based speech recognition systems with very high memory efficiency. That is, it stores a back pointer from which the optimal word sequence guess can be found after all speech data has been processed. [0051] In some embodiments, a caching scheme is used for the lexicon stored in the memory of the remote unit, for example by the Ngram grammar module 218. As mentioned above, the lexicon is a dictionary consisting of words and their pronunciation inputs. These pronunciations refer to either the phonetic model or the full-word model phonetic spelling. It can be realized as spellings). A given word input may have multiple alternative pronunciation inputs, but most of these pronunciation inputs are rarely used by any speaker. This redundancy is repeated in the abstraction of each audio part, increasing the input not used by a given speaker. In other words, if the lexicon input is classified according to the frequency of use, there is a high possibility that the word in the vocal sound can be found from the top n lexicon inputs. In this case, the cache is divided into different levels according to the frequency of use. For example, frequently used lexicon entries are stored at a higher level in the cache. The cache method may be devised so that the top 10% of the cache is used, for example, at 90% of the time. Therefore, according to one embodiment, a multi-pass search is performed and the most appropriate input is examined in the first pass. If the garbage score from this path is high enough to assume that the actual spoken language was included in the most appropriate set of spellings, the voice decoder 216 reports the result to the calling function. If this score is low, the system returns to a wider range of spelling considerations. If the score from the first pass is high, but not high enough to determine if the most appropriate set of spellings contained the correct spelling for the vocal element, this is also a call feature. Reported and the calling function prompts the user for clarification. If lexicon spelling for a given audio part is not used and some of its alternative spellings are frequently used, the spelling is "trash". Can) and will not be considered further for that user. In this case, spellings that are rarely used are not considered, and the possibility of confusing the uttered sounds of similar sounds with any of those spellings is reduced, thus improving recognition accuracy. In addition, the cache method allows the system to consider a small amount of data, which greatly improves processing speed. [0052] Next, the natural language interface control module 206 will be described in detail. The natural language interface control module 206 includes a natural language interface module 222, a probabilistic context free grammar (PCFG) module 224, a device abstraction module 226, and a feedback module 228. In general, the Natural Language Interface Module (NLIM) 222 is a range of user usage history defined by the context of each device 114 under control and a set of probabilistic context free grammar (PCFG) rules and abstracted devices. Its role is to interpret the user's request within. In this case, the natural language interface module 222 controls the search of the speech recognition module 204 and the microphone set 108. This is done by controlling the grammar of the speech recognition module 204 and the lexicon under consideration. The natural language interface module 222 also controls system parameters in addition to the current state and current language reference of the abstracted device. [0053] As mentioned above, the user initiates a dialog with NLICS by saying an attention word. A preferred method of identifying attention words will be described with reference to FIG. The user issues an open request after the attention word, which is limited only by the capabilities of each device connected to the remote unit 104. The attention word alerts the natural language interface module 222 to the user's identity so that the speech decoder is instructed to use the appropriate grammar and model based on the attention word. Therefore, the system can be set in advance according to the user's speaking pattern (for example, pronunciation, structure, habit, etc.). [0054] Speech recognition module 204 documents user requests made in conversational natural language. Vocal sounds are documented as a set of alternative guess words ordered by probability. For example, the speech decoder 216 sends N optimal text strings to the natural language interface module 222 for analysis to determine the plausible meaning of the vocalized sound. [0055] Then, the natural language interface module 222 parses the input string by applying the stochastic context free grammar (PCFG) rule from the PCFG module 224, and the probability of the string, the user history, and the current system. Find the most suitable string considering the context. These PCFG rules reflect the context of the user (based on the attention word) and the context of the device being manipulated (if already determined). PCFGs are initially ordered by frequency of use as well as frequency of use. Over time, it will track individual user habits and improve the rule's probability estimation to reflect this data. This data can be shared or combined with data from other systems and redistributed via collaborative corpus (linguistic material). [0056] In addition, NLICS has two grammar sets. One is the N-gram grammar of the speech recognition module 204, and the other is the stochastic context-free glamor module 224 of the natural language interface control module 206. Traditional systems use only one grammar set and do not use the combination of N-gram grammar and PCFG rules inferred using data collected from man-machine dialogs in the field of personal electronics. [0057] By using PCFG rules for incoming text strings, the natural language interface module 222 reaches one of three conclusions: That is, (1) it is possible to clearly understand and respond to the user's request, (2) it is possible to clearly understand the user's request but cannot respond, and in this case, the user is notified of the conclusion, (3). The ambiguity of the request cannot be resolved, in which case the user is required to clarify. [0058] [0058] For example, in the case of (1), the natural language interface module 222 interprets the incoming string as a "Turn on the television" request with a sufficiently high level of confidence. In this case, the appropriate command in the device abstraction module 226 is retrieved and sent to the device 114 under control (ie the television). The device abstraction module 226 has all the commands to execute the appropriate user's request in a format that the television itself can understand. Commands are typically sent to the television via device interface 210, such as an IR transmitter. The TV turns on accordingly. The second case is when the user requests NLICS a task that NLICS cannot perform. For example, when a user requires a television to explode. [0059] The feedback module (eg, text-voice) 228 is instructed to play an audible message through the speaker to inform the user that the request cannot be made. Note that the feedback module 228 may simply display a notification on the screen instead of reproducing the audio signal via the speaker 214. [0060] In the third case, the ambiguity is resolved according to the type of ambiguity that has arisen. In this way, the natural language interface module 222 clarifies unclear requirements. If the ambiguity is caused by unreliability, the natural language interface module 222 requires the user to affirm the conclusion. For example, speaker 214 says "Did you mean play the CD?". Alternatively, the natural language interface module 222 requires the user to repeat the request. If the ambiguity is caused by a set of choices, the natural language interface module 222 wants these choices, for example, "Did you want to watch a movie on the VCR or the DVD?" Or is it a DVD?) , Presenting the user with the option. If the ambiguity is caused by the current context, the user is informed. For example, the user may request playback of a DVD player when it is already playing. [0061] In the first two obscure cases, the system adjusts the user profile to reflect reliability when decisions are made, in addition to preferences that take into account a set of choices. In some embodiments, over time, these statistics are used to reorder the inputs in the PCFG rules and the appropriate lexicon. As a result, the most suitable inputs are always checked early, and these suitable inputs provide high reliability, resulting in faster and more accurate systems. [0062] If the natural language interface module 222 commands the feedback module 228 to clarify the request, for example, if the speaker 214 tells "Did you mean play the CD?" The natural language interface module 222 switches context and grammatical rules based on what it expects to receive on the microphone set 108. For example, the system switches to a context that predicts that it will receive a "yes" or "no" or a variant thereof. If the user answers "yes", the natural language interface module 222 switches the context back to its original state. [0063] In this case, when the context changes, the natural language interface module 222 commands the speech recognition module 204 to switch grammars. Since the grammar controls the lexicon used, this indirectly modifies the lexicon. [0064] The natural language interface control module 206 also includes a device abstraction module 226. The device abstraction module 226 stores an abstraction for each device 114. In this case, the command for each device 114 and the objects that can be operated by each device 114 are stored here. Device abstraction module 226 also associates these controls with what the device can do and what it can do. The content of the device abstraction module 226 depends on the various devices connected to the remote unit 104. The device abstraction module 226 also has commands for other devices to operate another device. For example, when a user requests DVD playback, a command is issued to turn on the DVD player and play the DVD. In addition, if the TV is not turned on yet, a command signal is sent to turn it on. [0065] The commands stored in the device abstraction module 226 are sent to each device 114 under control via the device interface 210. In some embodiments, the device interface 210 is an IR or RF interface. [0066] NLICS can be configured to control any device that can be controlled via such an IR link. As long as the device abstraction remembers the commands for manipulating a particular device, it does not understand that the device is controlled by a natural language interface. The device simply thinks that the remote control of the device or the general purpose remote control sent the signal. [0067] The system processing controller 208 operates as a controller and processor for various modules in NLICS. Its function is well known in the art. Further, an interface 212 is connected to the system processing controller 208. This allows connection to the base unit 106 or a computer. Interface 212 may be any kind of wired or wireless link, as is known in the art. [0068] The various components of the system, such as the feature extraction module 202, the speech recognition module 204, and the natural language interface control module 206, are software or firmware using, for example, an application specific integrated circuit (ASIC) or digital signal processor (DSP). It may be realized by. [0069] Next, with reference to FIG. 3, a functional block diagram of the base unit or base station of the natural language interface control system of FIG. 1 according to still another embodiment of the present invention is shown. FIG. 3 illustrates a base unit 106 (also referred to as a base station 106) and a remote unit 104 with a linear microphone set 108. The base unit 106 includes a planar microphone set 110, a frequency localization module 302, a time search module 304, a remote interface 306 (also referred to as interface 306), an external network interface 308, and a secondary cache 310. .. The linear microphone set 108 and the planar microphone set 110 are combined to form a three-dimensional microphone set 312 (also referred to as a 3D microphone set 312). FIG. 3 also shows an external network 116 connected to the external network interface 308. [0070] During operation, the base unit 106 serves as a docking station for the remote unit 104 (similar to a general purpose remote control). Base unit 106 includes an external network interface 308 that allows NLICS to interface to an external network 116 such as a home LAN or the Internet, either directly or through a hosted Internet portal. .. In this case, a new grammar, voice model, programming information, IR code, abstracted device, etc. can be downloaded to the base unit 106 and stored in, for example, the secondary cache 310. [0071] In addition, the NLICS102 sends and stores its grammar, model, and lexicon to a remote server on the external network. This remote storage serves as a storage for information that can be retrieved by other similar devices. In this case, the lexicon is constantly updated with the latest pronunciation and usage so that the system does not age. Therefore, since the plurality of natural language interface control systems individually contribute to the external database in the remote server, it becomes possible to construct a joint lexicon and / or a joint corpus. [0072] In addition, NLICS102 can download command signals to the device abstraction module of remote unit 104. For example, suppose a user wants to operate an older VCR with an IR remote control manufactured by a manufacturer different from NLICS. Base unit 106 only downloads the commands stored for any number of devices. Then, these commands are stored in the device abstraction module. NLICS can also send feature vector data and labels related to reliable vocalizations to the joint corpus. This data can be combined with other data and used to tailor an improved model that is subsequently redistributed. This approach can be used to incorporate new words into the joint corpus by sending feature vector data and its labels. The sent feature vector data and the label are combined with other data and documented phonetically using the forward-backward algorithm. This input is then added to the lexicon and can be redistributed. [0073] The base unit 106 includes a flat microphone set 110. The planar microphone set 110 and the linear microphone set 108 of the remote unit 104 are combined to form the three-dimensional microphone set 312. Both sets consist of conventional point source specific microphones. As is known in the art, a 3D set is constructed by first constructing a plane set (eg, plane microphone set 110) and then adding one or two microphone elements to the plane of the plane set. Will be done. In this case, the linear microphone set 108 becomes an additional one or two elements. This allows the NLICS 102 to determine the 3D search volume. The device only searches for audio energy within its volume. Therefore, microphone sets 108, 110 localize points within the range of the search volume. Voice energy outside the search volume range, background noise, etc. are attenuated, and voice energy within the search volume range is totaled. In practice, the user needs to be within a specific volume range to control various devices. For example, the search volume is configured to be the volume near the sofa in the user's living room. [0074] Both the linear microphone set 108 and the planar microphone set 110 are controlled by the natural language interface module 222. The frequency localization module 302 and the time search module 304 are connected to the 3D microphone set 110. The time search module 304 receives the control signal from the natural language interface module 222 in the remote unit 104 via the remote interface 306. The time search module 304 puts together a timed buffer given by the microphone. This helps the time search module 304 locate the estimated hit and point the 3D microphone set 110 in the direction of the hit. The function of the time search module 304 is well known in the art. [0075] The frequency localization module 302 is also under the control of the natural language interface module 222. The frequency localization module 302 executes a localization algorithm as is known in the art. Localization algorithms are used to localize voice energy within a predetermined volume range. In this case, the audio energy emitted from other than the localization point within the volume range is attenuated (out of phase), and the audio energy from within the localization point is summed (in phase). phase) Yes). Therefore, localization utilizes constitutive and destructive interference in the frequency domain. In operation, the search module is used to perform a coarse search for attention words. If the voice energy exceeds the threshold, the localization module will perform a precise search. If it passes the precise search, the word is sent to the recognition and NLI module. From this coarse search to a precise search, it is very helpful in reducing the processing associated with localization. For example, such localization requires a great deal of computation because the energy must be converted back and forth in the frequency domain. Therefore, processing is reduced by eliminating a large number of estimated hits in the coarse search. When the SR module identifies an estimated hit as an attention word, the estimated hit is sent to the natural language interface module 222 for analysis to determine which attention word was said. The context of the natural language interface module is initially the context of the attention word. That is, the system is searching for attention words to activate the system. When the attention word is found, the NLICS context is changed to the request context, searching for requests restricted by the device connected to NLICS. [0076] The secondary cache of base unit 106 is used to store the secondary model, grammar and / or lexicon used by remote unit 104. This complements speech recognition modules designed to read (stream) speech models and grammars from secondary storage devices or secondary caches (eg, hard disks, CD-ROMs, DVDs) at run time. Once the data is read, it can be used immediately without any preprocessing. This is effectively linked to context switching. Due to the grammar context switching feature, the required processing is reduced and the speech recognition accuracy is improved. In addition, for grammars that are not frequently used, the memory in the remote unit 104 is not occupied and is secondary. Since it can be stored in the cache 310 and read when needed, the required memory is greatly reduced. In addition, more speech data can be used to improve speech recognition accuracy, and the secondary storage device can hold a large number of basic models for various dialects and accents, so that it can be adapted to the speaker. Various approaches can be implemented efficiently. Further, the secondary cache may be a storage device for models, grammars, etc. downloaded from the external network 116. [0077] Next, with reference to FIG. 4, a flowchart of each step performed by the natural language interface algorithm of the natural language interface control system of FIGS. 1 to 3 is shown. First, the speech recognition module 204 and the natural language interface module 222 are initialized in the context of searching for attention words (step 402). This allows NLICS to accept non-prompt user requests, but first must notify the system that the user request is coming in. The attention word does this work. In this case, the existence of the attention word is specifically identified using the grammar and model of the hidden Markov model. Next, the remote unit receives the voice data with the microphone set (step 404). Audio data is separated into 12.8 msec frames with 50% overlap. Extract the 38-dimensional feature vector from the audio data. These features consist of a Mel frequency Keptral coefficient 1-12 and an MFC coefficient 0-12 first-order and second-order derivatives. In this way, a feature vector is created from the voice data (step 406). This is done in the feature extraction module 202. [0078] The speech recognition module 204 then applies the Hidden Markov Model (HMM) and N-gram grammar to the incoming feature vector (specified by the natural language interface) and in-vocabulary (IV) Viterbi. (Likelihood) Extract the score (step 408). Then, using a model of the OOV event, for example, a single-note model ergodic bank, the feature data is reprocessed and the out-of-vocabulary (OOV) Viterbi score is extracted (step 410). Calculate the garbage score from the IV and OOV scores. For example, the garbage score is equal to [Ln (IV score) -Ln (OOV score)] / number of frames (step 411). If the score is low, it means that it is a Garbage vocalization. N optimally documented text strings and their corresponding garbage scores are sent to the natural language interface module 222 (step 412). Natural language interface module 222 parses the input string using stochastic context free grammar (PCFG) ruleset in addition to device context information about attention vocalizations (step 414). As mentioned above, the natural language interface module 222 requires an attention method. For example, you need to receive your own attention word (Mona), or the speaker's ID associated with acceptable grammar rules. [0079] When the user draws the attention of the system, that is, when the natural language interface module 222 detects the attention word (step 416), the natural language interface module knows the user's identity. This is done by configuring the system according to the user. This is done by changing the appropriate system parameters and instructing the speech recognition module 204 to switch to a grammar that is appropriate for accepting commands and requests and that is appropriate for the user. The speech recognition module 204 modifies the lexicon according to the grammatical rules and individual users. Thus, the speech recognition module 204 and the natural language interface module 222 change the context to search for user requests (step 418). The natural language interface module also directs the microphone set to narrow its focus in order to successfully eliminate environmental noise. In addition, if there is a device under NLICS control (TV, CD, etc.) playing at high volume, the natural language interface module will instruct the amplifier to turn down the volume. Then, the natural language interface module 222 starts the timer and waits for the user's request until the timeout time expires. When the system times out, the natural language interface module 222 reconfigures the system by resetting the speech recognition module rules and lexicon appropriate for searching for attention words. Also, when they are adjusted, the microphone set and amplifier volume are reset. These reset steps are similar to those performed in step 402. [0080] [0080] After switching to the context of user request search (step 418), steps 404 to 414 are repeated in this path, except that the voice is a request to operate one or more of the devices under control. [0081] When the natural language interface module 222 detects a user request (step 416), that is, when a user request (determined by the PCFG grammar system and device context) is received, it draws one of three conclusions (step 420, step 420, 422, 424). In the case of step 420, the user request is clearly understood and the natural language interface module can respond to the user request. Therefore, the natural language interface module 222 executes a command by sending an appropriate signal through the device interface 210, as indicated by the abstracted device. Then, after switching the context of the speech recognition module 204 and the natural language interface control module 206 to the context of the attention word search (step 426), the process proceeds to step 404. [0082] In the case of step 422, the user request is clearly understood, but the natural language interface module cannot meet the user request. In this case, the user is notified and prompted for further instructions. The system waits for further user requests or times out and proceeds to step 426. [0083] In the case of step 424, the ambiguity cannot be resolved for the request, in which case the natural language interface module 222 requests clarification from the user, for example using the feedback module 228 and the speaker 214. The ambiguity is resolved according to the type of ambiguity that arises. If the ambiguity is caused by unreliability, the natural language interface module 222 affirms the conclusion to the user (eg, "Did you mean play the CD?"). When the user confirms the conclusion, the command is executed and the system is reset (step 426). The system adjusts the user profile to reflect reliability when decisions are made, in addition to preferences that take into account a set of choices. In some embodiments, over time, these statistics are used to reorder the inputs in the PCFG rules and the appropriate lexicon. As a result, the most suitable inputs are always checked early, and these suitable inputs provide high reliability, resulting in a faster and more accurate system. [0084] If the ambiguity is caused by a set of choices, the natural language interface module 222 presents these choices to the user (eg, "Did you want to watch a movie on the DVD player or the VCR?" Did you want to watch a movie or a VCR) "). When the user chooses from the options given, the natural language interface module 222 executes the command, otherwise the system is reset (step 426). In either case, the user profile is updated as described above. [0085] If the ambiguity is caused by the current context (eg, when the user requests the TV to stop, but the TV is off), the user is informed. [0086] Although the present invention has been described above with reference to specific examples and examples, various modifications have been made by ordinary engineers in the art to the extent that the gist of the present invention described in the claims is not deviated. It can be carried out. [Simple explanation of drawings] FIG. 1 is a system-level block diagram of a natural language interface control system (NLICS) according to an embodiment of the present invention. FIG. 2 is a functional block diagram of a remote unit of the natural language interface control system (NLICS) of FIG. 1 according to another embodiment of the present invention. FIG. 3 is a functional block diagram of a base station unit of the natural language interface control system (NLICS) of FIG. 1 according to still another embodiment of the present invention. 4 is a flowchart showing each step performed in the natural language interface algorithm of the natural language interface control system of FIGS. 1 to 3. FIG.
1 sheet
Sheet 1
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| JP11237896A | Cites | Japan |
| JP327698A | Cites | Japan |
| JP8185196A | Cites | Japan |
| JP10254477A | Cites | Japan |
| JP1185183A | Cites | Japan |
| JP8123476A | Cites | Japan |
| JP8223309A | Cites | Japan |
| JP6274190A | Cites | Japan |
| JP61285495A | Cites | Japan |
13 members in 7 offices
Priority claims9
| Document | Office | Kind | Date |
|---|---|---|---|
| 16028199 | United States of America | P | |
| 16028199 | United States of America | P | |
| 60160281 | United States of America | – | |
| 0029036 | United States of America | W | |
| 0029036 | United States of America | W | |
| 1999160281 | – | – | – |
| 2000029036 | – | – | – |
| US19990160281P | – | – | – |
| WO2000US29036 | – | – | – |
Members13
| Document | Office | Kind | |
|---|---|---|---|
| CA2387079A1 | Canada | A1 | |
| CA2748396A1 | Canada | A1 | |
| WO0129823A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU8030300A | Australia | A | |
| EP1222655A1 | European Patent Office (EPO) | A1 | |
| KR20020071856A | Republic of Korea | A | |
| JP2003515177A | Japan | A | |
| US2008059188A1 | United States of America | A1 | |
| KR100812109B1 | Republic of Korea | B1 | |
| US7447635B1 | United States of America | B1 | |
| CA2387079C | Canada | C | |
| JP2011237811A | Japan | A | |
| JP5118280B2This record | Japan | B2 |
22 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Re-examination (zenchi) completed and case transferred to appeal boardAppealJAPANESE INTERMEDIATE CODE: A912A912 | A912 | |
| Transfer to examiner for re-examination before appeal (zenchi)AppealJAPANESE INTERMEDIATE CODE: A911A911 | A911 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Decision of refusalJAPANESE INTERMEDIATE CODE: A02A02 | A02 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written permission of extension of timeJAPANESE INTERMEDIATE CODE: A602A602 | A602 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Notification of resignation of power of attorneyJAPANESE INTERMEDIATE CODE: A7424RD04 | RD04 | |
| Notification of acceptance of power of attorneyJAPANESE INTERMEDIATE CODE: A7422RD02 | RD02 | |
| Notification of appointment of power of attorneyJAPANESE INTERMEDIATE CODE: A7423RD03 | RD03 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 5118280
- Publication, DOCDB
- 5118280
- Publication, EPODOC
- JP5118280B
- Application
- 2001532534
- Application, DOCDB
- 2001532534
- Application, EPODOC
- JP20010532534
Titles2
- Japanese
- 自然言語インターフェースコントロールシステム
- English
- Natural language interface control system
Classification
- CPC, 5
- G10L15/22
- G10L15/18
- G10L15/1815
- G10L25/78
- G10L15/14
- IPC, 11
- G06F17 28
- G10L15 20
- G10L15 14
- G10L15 18
- G10L15 187
- G10L15 197
- G10L15 28
- G10L15 22
- G10L11 02
- G10L15 00
- G10L15 08
