Distributed real time speech recognition system
23 claims: 1 independent, 22 dependent
- 1クライアントコンピューティング装置に一体化されて、一乃至複数の発声単語で構成された文を含む連続 発声 音声を受信し、関連する発声音声信号を発生するサウンド処理回路と、 クライアントコンピューティング装置の一部を構成し、前記 発声 音声信号からベクトル係数の組である第一の組の音声信号を発生する第一の信号処理回路と、 サーバーコンピューティング装置の一部を構成する第二の信号処理回路と、 前記クライアントコンピューティング装置に結合されたものであって、前記第一の組の音声信号をフォーマット化し、通信チャンネルを通じて、音声発生時にリアルタイムに、前記第二の信号処理回路へ送信する送信回路とを有し、 前記送信回路では、前記第一の組の音声信号が断続することなく連続して送信され、前記第二の信号処理回路では、連続して受信した前記第一の組の音声信号に基づいてベクトル係数の組である第二の組の音声信号が生成され、 前記第一の組の音声信号と前記第二の組の音声信号を生成するのに必要とされる信号処理機能が、前記第一の信号処理回路と前記第二の信号処理回路のそれぞれにおいて利用可能とされた計算リソースに基づいて、前記送信回路で送受信可能な交信チャンネルを通じ前記第一の信号処理回路と前記第二の信号処理回路に割り当てられ、 前記信号処理機能に基づいて、前記第一の信号処理回路での前記第一の 組の 音声信号の生成と、前記第二の信号処理回路における単語の認識動作及び文の認識動作とが協働して行われることを特徴とする音声認識装置。
- 2前記第一の信号処理回路で生成される前記第一の組の音声信号が、分散認識システムの処理に基づくデータコンテントを含み、 前記第二の信号処理回路では、前記第一の組の音声信号と前記第二の組の音声信号とから複合音声 信号 を発生し、 前記第二の信号処理回路に、前記複合音声 信号 を使用して 発声 音声の単語の認識を行う単語認識回路が設けられている請求項1記載の音声認識装置。
- 3前記第二の信号処理回路に、連続 発声 音声に含まれる文を認識する文認識回路を有する請求項2に記載の音声認識装置。
- 4前記文認識回路では、予め定義された多数の文から可能性のある文の候補の組を識別し、可能性のある文の候補の組の各エントリを前記連続 発声 音声に含まれた文と比較して、各エントリの一致を判定する請求項3に記載の音声認識装置。
- 5前記文認識回路が、認識された単語上で動作する自然語エンジンを含む請求項4に記載の音声認識装置。
- 6連続 発声 音声に含まれる文は、名詞句を調べることによって可能性のある文の候補の組と比較される請求項5に記載の音声認識装置。
- 7可能性のある文の候補の組は、ユーザによって与えられた動作環境に応じて前記文認識回路によってロードされるコンテックスト辞書によって決定される請求項4に記載の音声認識装置。
- 8前記第一の信号処理回路が、前記連続 発声 音声を格納する記憶手段と、連続 発声 音声に対応する音声信号を捕捉する手段と、連続 発声 音声に含まれる前記一乃至複数の単語に相関させるには不十分な音声からある程度認識された音声 信号 を発生する第一の処理手段と、を有し、 前記第二の信号処理回路が、ある程度認識された音声 信号 から一乃至複数の単語に相関された認識可能な音声 信号 を発生する第二の処理 手段、を 有し、 前記第一及び第二の信号処理回路の処理動作に基づいて、一乃至複数の単語に相関された音声情報が得られる請求項1ないし7のいずれかに記載の音声認識装置。
- 9前記送信回路は、前記第一と第二の信号処理回路を結合する非永久データ送信接続手段である請求項1ないし8のいずれかに記載の音声認識装置。
- 10前記非永久データ送信接続手段は、交換された回路又はパケット交換接続手段である請求項9に記載の音声認識装置。
- 11前記非永久データ送信接続手段は、インターネットのネットワークを含んでいる請求項9に記載の音声認識装置。
- 12前記非永久データ送信接続手段は、ワイヤレス通信チャンネルである請求項9に記載の音声認識装置。
- 13前記第一の信号処理回路の処理動作が、連続 発声 音声からスペクトルパラメータベクトル を 抽出する動作を含む請求項1ないし12のいずれかに記載の音声認識装置。
- 14前記第一の信号処理回路の処理動作が、前記スペクトルパラメータベクトルに関するMFCC係数を得るため に、 Mel周波数転送処理によりスペクトルパラメータベクトルを分解する動作を含んでいる請求項13に記載の音声認識装置。
- 15ある程度認識された音声 信号 は、スペクトルパラメータベクトルより得られるMFCC係数とデルタ及び加速係数を含む観察ベクトルO t で構成される請求項13に記載の音声認識装置。
- 16ある程度認識された音声 信号 は、観察ベクトルO t を含み、前記第二の信号処理動作は、音声 信号 シンボルにより一連の観察ベクトルO t をマッピングするためのビタビ復号動作を含んでいる請求項8に記載の音声認識装置。
- 17前記第二の信号処理回路の処理動作が、前記音声 信号 シンボルを一乃至複数の単語テキストに変換する変換動作を含んでいる請求項1ないし16のいずれかに記載の音声認識装置。
- 18前記第二の信号処理回路の処理動作が、一乃至複数の単語のいずれ が、 前記一乃至複数の単語テキストに対応しているかを判定するクエリーを実行する請求項17記載の音声認識装置。
- 19前記第二の信号処理回路の処理動作が、一乃至複数の単語テキストのいずれが音声情報に対応しているかを判定するための環境変数を用いる請求項8に記載の音声認識装置。
- 20前記第一の信号 処理 回路では、前記 発声 音声信号から前記 発声 音声の単語の認識を可能 にする ためにそれ自体では不十分な第一の音声 信号 の組を発生し、 第一の組の音声信号を 前記送信回路の通信チャンネルの通信プロトコルと互換 性を有する フォーマット となるように 生成するとともに、サーバーコンピューティング装置に送信し、 前記第二の信号処理回路で、前記 発声 音声信号の単語を抽出する単語認識エンジンによって使用するのに十分な情報を含む第二の組の音声信号を生成し、 前記第一の組の音声信号と前記第二の組の音声信号を用いて分散音声認識が行なわれる請求項1記載の音声認識装置。
- 21前記第二の音声信号の音声 信号 が、前記第一の組の音声信号から 取り出された 音声 信号 に基づいて計算される請求項20に記載の音声認識装置。
- 22前記第一と第二の信号処理回路の信号処理の分担として、前記第一の信号処理回路は、前記 発声 音声信号から単語認識エンジンによって使用可能な形式に変換するために必要な信号処理動作の約1/2未満の処理を行うように割り当てられる請求項1に記載の音声認識装置。
- 23前記第一の信号処理回路は、第二の組の音声信号を発生するために必要な信号処理演算によって前記第二の信号処理回路を補助するように構成される請求項1に記載の音声認識装置。
Independent claims23
1 paragraph, as filed
[Technical field to which the invention belongs] [0001]<u style="single">this</u>The invention relates to a system and an interactive method that responds to a user's voice input and inquiry given on a distributed network such as the Internet or a local intranet. When implemented on the Internet's Worldwide Web (WWW) service, this dialogue system allows clients or users to ask questions in natural languages such as English, French, German, Spanish, Japanese, etc. It functions so that the appropriate answer in her native language can be received on his or her computer or its peripherals. This system is particularly useful for distance learning, electronic commerce, technical e-support services, Internet searches, and the like. [Conventional technology] [0002] The Internet, especially the World Wide Web (WWW), continues to grow in terms of both popularity and commercial and recreational use, and this trend is expected to continue. Increasing and widespread use of personal computer systems and low-cost internet connection<u style="single">But</u>Being able to do so contributes to the acceleration of this phenomenon. The advent of inexpensive Internet connectivity and high-speed access technologies such as ADSL, cable modems, and satellite modems is expected to further accelerate the mass use of WWW. [0003] Therefore, it is expected that the number of businesses providing services, products, etc. on WWW will increase dramatically in the last few years. However, to date, users' Internet "experiences" are mostly limited to non-voice inputs and outputs such as keyboards, intelligent electronic pads, mice, trackballs, printers, and monitors. This is, in many ways, a bottleneck for dialogue on the WWW. [0004] First of all, there is the problem of skill level. Many types of applications are trying to adapt more naturally and smoothly to voice-based environments. For example, for most people who want to buy audio recordings, it is very pleasant to ask a clerk at a record store or the like for information on the title of a particular author that can be found at that store. It is often possible to browse and search web pages on your own media to find items on the Internet, but it is usually easier, with few exceptions, to get human help first. Yes and efficient. This request for assistance is made in the form of a voice inquiry. Moreover, due to physical or psychological barriers, many people cannot or do not use any of the conventional I / O devices mentioned above. For example, many seniors cannot easily read the text that exists on a WWW page, or do not understand the layout or hierarchy of menus, or move well to display selected items. I can't operate the mouse. Many other people have computer systems, WWW pages, etc.<u style="single">State</u>And complicated<u style="single">Sa</u>I'm fooled and refuse to use online services for this reason as well. [0005] Therefore, applications that can imitate normal human dialogue are online.<u style="single">I</u>It tends to be preferred by users who want to shop and find information on WWW. In addition, the use of voice systems is expected to increase the number of people who want to participate in electronic commerce and e-learning. However, to date, there is no or, if any, system that enables this type of dialogue.<u style="single">hand</u>The performance is very limited, even if there are few and dialogues take place. For example, various commercial programs sold by IBM (Via Voice) and Kurzweil (Dragon) allow some control (opening and closing files) and searching (using pre-learned URLs) by the user of the interface. However, it does not provide a flexible solution that can be used by many users of different cultures without the need for time-consuming voice training. Conventional general attempts to perform voice functions on the Internet can be found in US Pat. No. 5,819,220, incorporated herein by reference. [0006] Another problem with the lack of voice-based systems is efficiency.<u style="single">。</u>Many companies today provide technical support over the Internet, some of which provide human operator assistance to respond to these inquiries. This is very beneficial (for the reasons mentioned above), but it is very costly and inefficient because it requires hiring humans to handle these inquiries. This has practical limitations, requires long wait times to respond, and increases employer overhead. An example of this method is set forth in US Pat. No. 5,802,526, incorporated as part of the disclosure herein. In general, the services offered on the WWW are "extensible" or, in other words, to the expected users.<u style="single">Therefore,</u>Even if there is a perceived delay or confusion, this is negligible<u style="single">In a certain state</u>A user<u style="single">Amount of information</u>It is possible to handle the increase<u style="single">That</u>, Highly desirable. [0007] In a similar sense, distance learning has become a rapidly growing and popular option for many students, making it virtually impossible for instructors to answer questions from more than one person at a time. Even in that case, due to the time constraints of other instructors, such dialogue will only take place for a very limited time. [0008] However, to date, there is no practical way for students to continue a human question-and-answer dialogue after a learning session or without the need for an instructor to personally respond to these questions. [0009] On the other hand, another aspect of emulating human questions and answers is<u style="single">、</u>Includes dictated feedback. In other words, for many people it is preferable to receive answers and information in a voice form. This form of functionality has been adopted by several websites to exchange information with website visitors, but it is not a real-time, interactive question-and-answer session and its effectiveness. And its usefulness is limited. [0010] Another area where voice interaction is useful is used by Internet users to search for information of interest on websites, such as those available on YAHOO.com, nzetacraaler.core, Excite.com, etc. It is a so-called "search" engine. These tools use a combination of keywords or metacategories to form a search query through a database of web pages that has an index of text related to one or more different web pages. Allows you to search. After processing a user's request, search engines are generally search engines.<u style="single">To</u>Therefore, the search processing logic used<u style="single">To</u>User specific question based on<u style="single">To</u>On the other hand, the search engine returns the URL pointer that detected the closest match and the number of hits corresponding to the excerpt from the web page. The structure and operation of such conventional search engines, including a mechanism for constructing a database of web pages and interpreting search questions, has been well known. To date, as far as the applicant knows, there is no search engine that can easily and reliably search and search for information based on the search input from the user. [0011] Despite the many benefits that come from these capabilities, there are many reasons why voice-based interfaces are not used in the above environments (e-commerce, electronic support, distance learning, internet search, etc.). .. First of all, it is due to the clear requirement that the output of the speech recognition device must be as accurate as possible. One of the more reliable methods for speech recognition currently in use is the Hidden Markov Model. (HMM)-A model in which all time sequences are used for arithmetic representation. Conventional uses of this technique are disclosed, for example, in US Pat. No. 4,587,670, which is incorporated herein by reference. Since the spoken language is considered to have a potential sequence of one or more symbols, the HMM model corresponding to each symbol is learned from the vector from the spoken language waveform. The Hidden Markov model is a finite set of states, each set associated with a (usually multidimensional) probability distribution. In a particular state, the result or observation occurs according to the associated probability distribution. This finite state device changes the state for each time unit, and at each time point t entering the state j, the variable vector O of the spectrum<u style="single">t</u>Is generated at the probability density Bj (Ot). This is the only result, not the state visible to the outside observer, and therefore the state is "hidden" to the outside.<u style="single">)</u>Therefore, it is called the Hidden Markov model. The basic theory of HMMs was published in a series of classical literature by Baum and his colleagues in the late 1960s and early 1970s. The HMM was used for dictation by Baker at Carnegie Mellon, used by Jelenik and his colleagues at IBM in the late 1970s, and by Steve Young and his colleagues at Cambridge University in the United Kingdom in the 1990s. Some representative publications and texts are as follows. [0012] 1.LE Baum, T. Petrie, "Statistical inference for probabilistic fuinctions for finite state Markov chains" Ann. Math. Stat: 1554-1563, 1966 2. LE Baum, "An inequyality and associated maximation technique in statistical estimation fpr probabilistic functions of Markov processes" Inequalities 3:18, 1972 3. JH Baker, "The dragon system-An Overview", IEEE Trans. On ASSP Proc., ASSP-23 (1), 24-29, Feby, 1975 4. F. Jeninek et al., "Continuous Speech Recognition: Statistical method", Handbook of Statistics, II , PR Kristnaiad, Ed. Amsterdam, The Netherlands. North-Holland 1982 6. JD Ferguson, "Hidden Markov Analysis: An Introduction" Hidden Markov Model for Speech, Institute of Defense Analyses, Princeton, NJ., 1980 7. HR Rabiner and BH Juang, "Fundamentals of Speech Recognition" Prentice Hall, 1993 8. HR Rabiner, Digital Processing of Speech Signals, Prentice Hall, 1978 Various research institutes are conducting more recent research on extending HMMs with neural networks and combining them with HMMs for speech recognition applications. The following are typical publications. 9. Nelson Morgan, Herve Bourlard, Steve Renals, Michael Cohen and Horacio Franco (1993), Hybrid Neural Network / Hidden Markov Model System for Continuous Spoeech Recognition, Journa; of Pattern Recognition and Artificial Intelligence, Vol. 7, No. 4, pp 899-916 and I. Guyon and P. Wang, Advance in Pattern Recognition System using Neutral Networks, Vol. 7, of Series in Machine Perception and Artificial Intelligence, World Scientific, Feb. 1994 [0013] All of the above publications are incorporated herein by reference. Although good results can be obtained with HMM-based speech recognition, recent changes in this technology are 100% similar to those required for WWW applications in all anticipated user and environmental conditions. It does not guarantee the exact requirements of an accurate and consistent language. [0014] Therefore, although speech recognition technology has been available and has been significantly improved over the last few years, the accuracy of speech recognition required for the combined use of speech recognition and natural language processing to be fully functional. There are strict technical requirements in the specifications. [0015] Contrary to language recognition, natural language processing (NLP) concerns the interpretation, understanding and heading of transcribed statements and large linguistic units. Natural speech contains surface phenomena that cannot be handled by common speech recognition devices, such as stuttering-hesitation, corrections and restarts, and discourse signs such as "well", which is a problem. It is the cause of the big gap that separates speech recognition and natural language processing technology. Another problem is, except for silence during vocalization.<u style="single">word</u>To segment the audio into significant units such as<u style="single">use</u>There are no possible punctuation marks. For optimal NLP performance, these types of phenomena must be annotated in their input. However, the most continuous speech recognition systems speak the raw language sequence. Examples of conventional systems using NLP are given in US Pat. Nos. 4,991,094, 5,068,789, 5,146,405 and 5,680,628. All of these are incorporated as part of the disclosure of the specification. [0016] Second, most highly reliable speech recognition systems require a speaker-dependent, time-consuming "trained" interface by the user, which allows the user to do so.<u style="single">C</u>It is highly undesirable from the standpoint of a WWW environment where the website is accessed only a few times. In addition, speaker-dependent systems usually require a large user dictionary (one for each user) that reduces speech recognition. It makes it very difficult to implement a real-time interactive interface with sufficient response performance (ie, -3 to 5 seconds, which reflects a natural conversation, is probably ideal). Today, typical commercial speech recognition application software includes those provided by IBM (Via Voice) and Dragon (Dragon). Most of these applications are sufficient for dictation and other transcription applications, but are extremely inadequate for applications such as NLQS where a recognition error rate close to 0% is required. In addition, these provided systems require long training times and are generally non-client-server structures. Another type of training-requiring system is US Pat. No. 5,231,670 assigned to Kurzweil, which is also incorporated herein by reference. [0017] Another critical problem faced by distributed speech-based systems is the lack of universality / control in speech recognition processing. General speech recognition system<u style="single">Alone</u>In practice, the entire search engine runs on a single client. A well-known system of this type is disclosed in US Pat. No. 4,991,217, which is incorporated herein by reference. These clients can take various forms (desktop PCs, laptop PCs, PDAs, etc.) with various audio signal processing and communication performances. Therefore, from a server-side perspective, these users have very different speech recognition and error rate performance, and it is easy to standardize the processing of all users who access web pages that can be operated by voice. is not it. The prior art of Gould et al.-US Pat. No. 5,915,236-generally describes the concept of preparing cognitive processing for a set of available computational resources.<u style="single">Is</u>How to optimize resources in a distributed environment like a client-server model<u style="single">Ka</u>Has not been solved or attempted to solve the problem<u style="single">。</u>To reiterate, in order to enable such voice-based technology on a widely distributed scale, it is possible to support even the lowest performing clients and remote to carry out e-commerce, electronic support and / or distance learning. It is preferable to have a system that harmonizes and considers the differences between individual systems so that they can interact with the server in a satisfactory manner. [0018] Two references regarding decentralized methods for speech recognition include US Pat. Nos. 5,956,683 and 5,956,683, which are incorporated as part of the disclosure herein. The first of these, US Pat. No. 5,956,683-Distributed Speech Recognition System (transferred to Qualsomm), discloses the implementation of a distributed speech recognition system between a telephone-based handset and a remote station. .. In this practice, all word recognition operations appear to be performed in the handset. This patent provides the advantage of placing the system in a portable or mobile phone for the extraction of acoustic properties in a portable or mobile phone to limit the degradation of acoustic properties due to quantization distortion resulting from narrow bandwidth telephone channels. Since it is explained, it is presumed as above. This document is not about the question of how to get sufficient performance for very low performance client platforms. Moreover, it is difficult to determine how the system performs real-time word recognition, and there is no compelling explanation for how to combine the system with a natural language processing system. [0019] Second of these references-US Pat. No. 5,956,683-Kula<u style="single">I</u>Ant / Server Speech Processing / Recognition (transferred to GTE) describes the implementation of an HMM-based distributed speech recognition system. This reference is not useful in many respects, but how to optimize the extraction of acoustic properties for different client platforms, such as by performing partial word recognition processing where appropriate. Includes. More importantly, only primitive syllable-based recognizers that only recognize the user's voice and return certain keywords such as the user's name and travel destination to fill out a dedicated form on the user's machine. It is disclosed. Further, since the streaming of the acoustic parameters is performed only after the silence state is detected, it is not considered that the sound parameters are streamed in real time. Finally, the bibliography mentions the availability of natural language processing (column 9), but no explanation is given as to how it is done in real time to give the user a sense of dialogue. Not. [Problems to be Solved by the Invention] [0020] Therefore, an object of the present invention is to provide an improved system and method capable of overcoming the above-mentioned limitations in the prior art. [0021] [0021] A primary object of the present invention is a flexible and optimally distributed word and phrase recognition system across client / platform arithmetic architectures that can achieve improved accuracy, speed and universality for a wide range of user groups. Is to provide. [0022] Another object of the present invention is to provide a speech recognition system that efficiently combines distributed word recognition with a natural language processing system so that individual words and entire remarks can be recognized quickly and accurately in any number of languages. There is. [0023] A related object of the present invention is to respond to voice inquiries with a very accurate and real-time set of appropriate answers.<u style="single">To</u>The purpose is to provide an efficient inquiry response system that can be given. [0024] Yet another object of the present invention is to provide an interactive real-time teaching / learning system distributed on a client / server architecture. [0025] A related object of the present invention is to provide interactivity with the ability to respond clearly so that the user can experience human interaction. [0026] Yet another object of the present invention is to provide voice recognition performance to enable voice data and commands to interact with such sites and to enable easily expandable voice electronic commerce and electronic support services. The purpose is to provide a complete internet website. [0027] Another of the present invention<u style="single">of</u>The purpose is to implement a distributed speech recognition system that uses environment variables as part of the recognition process to improve accuracy and speed. [0028] Yet another object of the present invention is to provide an extensible query / answer database system that supports any number of query subjects and users required for a particular application and momentary request. [0029] Yet another object of the present invention is to identify the first step of narrowing the list of expected responses in a relatively short time to a smaller set of candidates, and the best choice to respond to and return a question from candidates and the like. The purpose is to provide an inquiry recognition system that employs a two-step method that includes a more intensive second step as the calculation to be performed. [0030] Yet another object of the present invention is to query by extracting the word component of a statement, which can be used to quickly identify a suitable set of expected response candidates for such statements. The purpose is to provide a natural language processing system that facilitates recognition. [0031] A related object of the present invention is a natural language that facilitates the recognition of a query by comparing the word component of the statement with a set of expected response candidates to provide a very accurate and best response to such a query. It is to provide a processing system. [Means to solve the problem] [0032] Therefore, one general feature of the present invention is the Natural Language Questioning System (NLQS), which provides a fully interactive way to answer user's questions on a distributed network such as the Internet or a local intranet. is there. This dialogue system, when implemented on the Internet's Worldwide Web (WWW) service, allows clients or users to ask questions in natural languages such as English, French, German, and Spanish. Receive an appropriate answer on his or her personal computer in the natural language of his or her native language. [0033] The system is distributed and consists of a set of integrated software modules in the client's equipment and another set of integrated software programs residing on the server or servers. The client-side software program includes a voice recognition program, an agent and its control program, and a communication program. The server-side program is a communication program, natural language engine (NLE), database processing device (DB processing).<u style="single">、</u>It consists of an interface program for interacting DB processing with NLE and a SQL database. In addition, the client's device will be equipped with a microphone and speakers. The processing of remarks is divided between the client and server side in order to optimize processing and transmission latency and also provide support for very poor client platforms. [0034] In the context of interactive learning applications, the system is specifically used to give a single best answer to a user's question. Questions asked on the client's device are clearly pronounced by the speaker and, in the case of notebook computers, etc., are captured by a microphone built-in or usually provided as a peripheral accessory. The question is caught<u style="single">When</u>, Questions are partially handled by the NLQS client-side software residing on the client's equipment. The output of this partial processing is a set of voice vectors that are delivered to the server over the Internet to complete the recognition of the user question. This recognized voice is then converted to text on the server. [0035] After the user's question is decoded by a speech recognition engine located on the server, the question is translated into a structured query language (SQL) query. This query is then given simultaneously to software processing in the server called DB processing for pre-processing, and to the Natural Language Engine (NLE) module to extract the noun phrase (NP) of the user's question. .. During the noun phrase extraction process within the NLE, the user's question token is tagged. The tagged tokens are then grouped so that the NP list can be determined. This information is stored and sent to the DB processing process. [0036] In DB processing, SQL queries are completely customized using NP extracted from the user's question and other environment variables suitable for the purpose. For example, in training applications, user choices for courses, chapters and / or sections constitute environment variables. SQL queries are constructed using the extended SQL full-statement predicates-CONTAINS, FREETEXT, NEAR, AND. The SQL query is then sent to a full-text search engine in the SQL database, where the full-text search procedure begins. The result of this search procedure is a recordset of answers. This recordset contains memory questions that are linguistically similar to the user's question. Each of these memorized questions has a pair of answers stored in different text files, and this path is stored in a table in the database. [0037] The entire recordset of returned memory answers is then<u style="single">Array</u>Returned to the NLE engine in the form of (array). And each memory question of that is processed linguistically one by one in succession. This linguistic process constitutes the second step of a two-step algorithm for determining a single best answer to a user's question. This second step proceeds as follows: For each memory question returned to the recordset, the memory question NP is compared to the user question NP. After all of that memory question is compared to the user question, the memory question that best matches the user question is selected as the best memory question that matches the user question. The criterion used to determine the best memory question is the number of noun phrases. [0038] The memory answer paired with the best memory question is chosen to answer the user question. Then, the ID tag of the question is passed to the DB process. DB processing returns the question stored in the file. [0039] The communication link is reestablished and the response is sent back to the client in compressed form. Once received by the client, the answer is decompressed and clearly communicated to the user by the text-speech engine. As described above, the present invention can be used for many various uses including an interactive learning system, an Internet-related commercial transaction site, an Internet search engine, and the like. [0040] Computer-based educational environments often require the help of a counselor or a living teacher to answer questions from students. This help is often done by setting up a pre-determined special forum, or a meeting time as a chat session or a live telephone participation session so that questions can be answered at a designated time. Due to the immediacy and on-demand or asynchronous nature of online training that allows students to log on and be educated anytime, anywhere, questions are answered in a timely and cost-effective manner, thus providing users or students. It is important to be able to get the maximum benefit from what is given. [0041] The present invention solves this problem. It provides the user or student with answers to questions normally given to a living teacher or counselor. The present invention provides a single best answer to a question asked by a student. Students ask questions in his or her own voice in the language of choice. Speech is recognized and answers to questions are found using many technologies, including distributed speech recognition, full-text search database processing, natural language processing, and text-speech technology. Answers are given to the user in clear pronunciation by the counselor or agent imitating the teacher, as in the case of a living teacher, and in the selected language-English, French, German, Japanese, or other natural colloquialism. Be done. The user can choose the gender of the agent as well as some voice parameters such as pitch, loudness, and speed of the character's voice. [0042] Another use that benefits from NLQS is e-commerce use. In this application, user questions about the price of books and compact discs, or the availability of purchased items can be searched without having to select through various lists on successive web pages. Instead, the answer is delivered directly to the user, without any extra user input. [0043] Similarly, the system could be used to provide answers to frequently asked questions (FAQs) and as a diagnostic service tool for electronic support. These questions are common on certain websites and are provided to help users find information related to payment procedures or product / service specifications or problems they have encountered. The NLQS architecture is applicable to all of these applications. [0044] Numerous creative methods associated with these architectures are also usefully used in a variety of Internet-related applications. [0045] The present invention will be described below in a group of preferred embodiments, but many need to perform rapid and accurate speech recognition and / or provide the intellectual system with human dialogue capabilities. It will be apparent to those skilled in the art that the present invention may be usefully used in the environment of. BEST MODE FOR CARRYING OUT THE INVENTION [0046] Overview As described above, the present invention allows the user to use a client computer system (simple personal digital assistant, mobile phone, sophisticated high-end desktop PC, etc.) to naturally speak English, French, German, Spanish, Japanese, etc. It allows you to ask questions in words and receive appropriate answers from remote servers in the natural language of his or her native language. Therefore, the embodiment of the invention shown in FIG. 1 is a natural language questioning system configured to interact in real time to provide human dialogue ability / experience for e-commerce, electronic support and electronic learning applications. It is used advantageously to give an overview of NLQS) 100. [0047] The outline of the processing of the NLQS100 will be described over the client-side system 150, the data link 160, and the server-side system 180. These components are conventionally known and, in preferred embodiments, include a personal computer system 150 and an internet connection 160A, 160B and a large computer system 180. It will be appreciated by those skilled in the art that these are exemplary components and that the present invention is by no means limited by the particular implementation or combination of such systems. For example, the client-side system 150 can be implemented as a peripheral device of a computer, a PDA, a part of a mobile phone, a part of an Internet-compatible device, a public telephone linked to the Internet, or the like. Similarly, the internet connection is shown as data link 160A, but a suitable route for data transfer between client system 150 and server system, such as wireless link, RF link, IR link, LAN, is sufficient. Finally, the server system 180 can be a single large computer system or a concatenated smaller set of systems that support a large number of expected network users. [0048] First, the voice input is given as a utterance in the form of a question or inquiry clearly expressed by the speaker in the client's device or personal accessory. This statement is captured and partially processed by NLQS's client-side software 155, which resides on the client machine. To facilitate and improve the human aspect of the dialogue, questions are raised in the presence of the video character 157, which can be viewed by the user assisting the user as a personal information retrieval agent. The agent can also interact with the user in audible form using a visible text output and / or text-speech engine 159 on a monitor / display (not shown). The output of the partial processing performed by the SRE155 is a set of audio vectors that links the user's device or personal auxiliary device to a server or multiple servers via the Internet or a wireless gateway linked to the Internet as described above. Communication channel 160<u style="single">A</u>Sent through. In the server 180, the partially processed voice signal data is handled by the server-side SRE182, and the recognized voice text corresponding to the user's question is output. Based on the user's questions related to the text, the text-question transducer 184 forms the appropriate question to be used as input to the database processor 186. Based on the question, database processor 186 then searches and searches database 188 for the appropriate answer using customized SQL questions. The natural language engine 190 facilitates the construction of questions to database 188. After the machine answers the user's question, the former sends text over datalink 160B, which is converted to voice by the text-speech engine 159 and expressed as colloquial feedback by the video character agent 157. .. [0049] Since speech processing is decomposed in this way, it is possible to compose interactive human questions and answers in real time that compose large, controllable question / answer pairs. With the help of Video Agent 157, the experience will be improved and made natural and comfortable for novice users. To make speech recognition processing more reliable, context-specific grammars and dictionaries are used in NLE190 to analyze the vocabulary of user questions, along with natural language processing routines. Processes that identify the context of audio data have been known in the past (see US Pat. Nos. 5,960,394, 5,867,817, 5,758,322, and 5,384,892 as part of the disclosure herein. ), The inventor and the like do not know the embodiment as carried out by the present invention. The text of the user's question is D<u style="single">B</u>Processing device / engine (DBE<u style="single">)</u>Compared to the text of other questions to identify the question posed by the user by 186. By optimizing the dialogue and relationships of SR engines 155 and 182, NPL routine 190, dictionaries and grammars, very fast and accurate matches can be obtained, providing users with unique and responsive answers. Can be done. [0050] At server-side 180, alternating processing further accelerates speech recognition processing. In simplified terms, the question is given to NLE190 and to DBE186 after the question is formed. NLE190 and SRE182 perform complementary functions in all recognition processes. In general, SRE182 primarily determines the uniqueness of words expressed by the user, while NLE190 performs a linguistic morphological analysis of both on the search results returned after the user's question and database query. [0051] After the user's question is analyzed by NLE190, some parameters are extracted and sent to the DB process. Additional statistics are stored in an array for the second step of processing. During the second step of the two-step algorithm, the results of the primary search for the recordset are NLE1<u style="single">9</u>Sent for processing 0. At the end of this second step, a single question that matches the user's question goes through the process of getting a companion answer to the single stored question.<u style="single">B</u>Sent to processing. [0052] Therefore, the present invention is used to form natural language processing (NLP) and achieves optimum performance in speech-based web application systems. Although NLP has been known in the past, traditional attempts at natural language processing (NLP) work have not been well combined with speech recognition (SR) technology to achieve reasonable results in web-based application environments. .. In speech recognition, the result is generally a grid of recognized words that are expected to have some probability of being compatible with each speech recognition device. As mentioned above, the input to a typical NLP system is primarily a large unit of language. The NLP system is responsible for interpreting, understanding, and indexing large word units or pairs of transcribed statements. The result of this NLP processing is word recognition<u style="single">In contrast</u>, Understand the entire unit of words linguistically or morphologically. In other ways, sentence output of linguistic or concatenated words by SRE is<u style="single">simply</u>What was "recognized"<u style="single">In contrast</u>Understood linguistically. [0053] As mentioned earlier, speech recognition technology has been available for several years, but the technical requirements for the invention of NLQS are the speech required in applications that combine speech recognition and natural language processing to function well. There are very strict restrictions on the specifications for obtaining recognition accuracy. In this realization, it is not possible to achieve the required completely 100% speech recognition accuracy even in the best conditions, and the present invention is unable to achieve complete speech recognition for each word in question. In any case, an algorithm that balances the predicted risk of speech recognition processing with the requirements of natural language processing is adopted so that the recognition of the entire question itself can be sufficiently accurate. [0054] This recognition accuracy is 100 to 25 for responding to voice questions and expected.<u style="single">0</u>It can even meet the stringent user constraints of a short latency of 3-5 seconds (ideally ignoring fluctuating transmit latency) for this question. This short response time gives the overall impression a more natural and favorable real-time interaction experience from the user's point of view. Of course, even in non-real-time applications such as translation services, HMMs, grammars, dictionaries, etc. are aggregated and held, so this technique is advantageous. Overview of speech recognition used in the present invention [0055] General background technical information on speech recognition can be found in the references incorporated as part of the disclosure above. Nevertheless, specific examples of speech recognition structures and techniques conforming to NLQS100 will be described to better demonstrate some of the properties, qualities and features of the present invention. [0056] Speech recognition techniques are generally of two types-speaker-independent and speaker-dependent. In a speaker-dependent type of speech recognition technology, each user has a voice file that stores a sample of words that may be recognized. Speaker-dependent speech recognition systems generally have large vocabularies and dictionaries to make them suitable for dictation and text transcription applications. In addition, memory and processor resources are usually large, potentially intensive, and common for speaker-dependent systems. [0057] vice versa,<u style="single">Non-speaker</u>Dependent speech recognition techniques allow a single vocabulary file to be used by many groups of users. The accuracy achieved depends on the size and complexity of the grammar and dictionaries that can be supported for a given language. When used for NLQS, in NLQS due to the use of small grammars and dictionaries<u style="single">Non-speaker</u>Dependent type speech recognition technology can be implemented. [0058] [0058]<u style="single">Non-speaker</u>An important issue or requirement for both dependent and speaker dependent types is accuracy and speed. As the size of the user dictionary increases, the speech recognition accuracy increases, and the false recognition rate (WER) and speed decrease. This is because the search time increases and the pronunciation matching becomes complicated as the size of the dictionary increases. [0059] The basis of the NLQS speech recognition system is a series of Hidden Markov models (HMMs), which are mathematical models that characterize signals that change over time, as described above. Since part of the speech is based on a potential sequence of one or more symbols, the HMM model corresponding to each symbol is trained as a vector from the speech waveform. The Hidden Markov model is a finite set of states, each associated with a (usually multidimensional) probability distribution. Transitions between states are governed by a set of probabilities called transition probabilities. In certain conditions, results or findings occur according to the associated probability distribution. This finite set device changes the state for each time unit, and at the time t when entering the state j, the spectral parameter vector O<sub>t</sub><u style="single">To</u>, Probability density B<sub>j</sub>(O<sub>t</sub>). The condition was named the Hidden Markov model because it is "hidden" in the result, as this is the only result and the condition is invisible to outside observers. [0060] In isolated speech recognition, a series of spoken speech vectors corresponding to each word is represented by the following Markov model. [0061]<u style="single">O</u>= o<sub>1</sub>, o<sub>2</sub>, ......... o<sub><u style="single">t</u></sub> (1-1) [0062] Where o<sub>t</sub>Is the speech vector spoken at time point t. The isolated word recognition is calculated as follows. [0063] arg max [P (w)<sub><u style="single">i</u></sub>|<u style="single">O</u>)] (1-2) [0064] By using Bayes' theorem [P (w)<sub><u style="single">i</u></sub>|<u style="single">O</u>)] = [P (P ()<u style="single">O</u>| w<sub><u style="single">i</u></sub>) P (w)<sub><u style="single">i</u></sub>)] / P (<u style="single">O</u>) (1-3) [0065] In the general case, when applied to speech, the Markov model also assumes a finite state device, changes the state for each time zone, and at each point in the state j, the speech vector o<sub>t</sub>But the probability density b<sub>j</sub>(o<sub>t</sub>). Furthermore, the transition from state i to state j is also stochastic, and the discrete probability a<sub>ij</sub>Is dominated by. [0066] For the state sequence X, the composite probability that O is generated by the model M moving through the state sequence X is the product of the transition probability and the output probability. If only the speech sequence is known, the state sequence is hidden as described above. [0067] X is known<u style="single">Absent</u>Then, the required likelihood is calculated by adding all the predicted state sequences X = x (1), x (2), x (3), .... x (T). [0068]<img file="JP4987203B2_D0001.tif" />[0069] Word w<sub><u style="single">i</u></sub>The set of models corresponding to is M<sub>i</sub>If so, the formula<u style="single">(</u>1-2<u style="single">)</u>Is an expression<u style="single">(</u>1-3<u style="single">)</u>Also using P (<u style="single">O</u>| w<sub>i</sub>) = P (<u style="single">O</u>| M<sub>i</sub>) It is solved by assuming that. [0070] All of these are parameters<u style="single">{</u>a<sub>ij</sub><u style="single">}</u>And {b<sub>j</sub>(o<sub>t</sub>)} Is each model<u style="single">Ru</u>M<sub>i</sub>It is assumed that you are known about. This is done using a learning example that corresponds to a particular model, as described above. The model parameters are then automatically determined by a robust and efficient reprediction procedure. Therefore, once a sufficient number of representative examples of each word have been collected, an HMM can be constructed, simply modeling all of the many causes of unavoidable changes in real speech. Since this training has been well known in the past, detailed description is omitted, but since the HMM is derived and configured on the server side rather than the client side, the distributed architecture of the present invention improves the quality of the HMM. In this way, appropriate samples from users in different map regions can be easily compiled and analyzed to optimize the changes that may occur in a particular language to be recognized. Since each predicted user uses the same set of HMMs during the recognition process, the uniformity of the speech recognition process is well maintained and simplifies error diagnosis. [0071] To determine the parameters of the HMM from the set of training samples, the first step is to make a rough guess of what they are. The refinement is then performed using the Baum-Welch estimation formula. With these equations, the maximum likelihood μ<sub>j</sub>(Μ<sub>j</sub>Is the mean vector<u style="single">And</u>Σ<sub>j</sub>Is a variance matrix) μ μ<sub>j</sub>= Σ<sup>T</sup><sub>t = 1</sub>L<sub>j</sub>(t) o<sub>t</sub>/ [Σ<sup>T</sup><sub>t = 1</sub>L<sub>j</sub>(t) o<sub>t</sub>] [0072] Next<u style="single">To</u>The forward-backward algorithm is used and the state occupancy probability L<sub>j</sub>(t) is calculated. Forward probability α for model M with state N<sub>j</sub>(t) is α<sub>j</sub>(t) = P (o)<sub>1</sub>, ......, o<sub>t</sub>x (t) = j | M) It is represented by. This probability can be calculated using induction. [0073] α<sub>j</sub>(t) = [Σ<sup>N-1</sup><sub>j = 2</sub>α (t-1) a<sub>ij</sub>] b<sub>j</sub>(o<sub>t</sub>) [0074] Similarly, the backward probabilities can be calculated using induction. [0075] β<sub>j</sub>(t) =<u style="single">Σ</u><sup>N-1</sup><sub>j = 2</sub> a<sub>ij</sub>b<sub>j</sub>(o<sub>t + 1</sub>) (T + 1) [0076] By realizing that the forward probability is a compound probability and the backward probability is a conditional probability, the probability of state occupancy is the product of two probabilities. [0077] α<sub>j</sub>(t) β<sub>j</sub>(t) = P (<u style="single">O,</u> x (t) = j | M) [0078] Therefore, the probability of state j at time point t is L<u style="single">j</u>(t) = 1 / P [α<sub>j</sub>(t) β<sub>j</sub>(t)] [0079] Where P = P (<u style="single">O</u>| M) [0080] [0080] To generalize the above for continuous speech recognition, we assume a maximum likelihood state sequence in which the sum is replaced by the maximum motion. Therefore, for a given model M, φ<sub>j</sub>Speaking speech vector o used in state j where (t) is at time point t<sub>1</sub>~ O<sub>t</sub>Assuming that it shows the maximum likelihood of φ<sub>j</sub>(t) = max [φ<sub>j</sub>(t) (t-1) α<sub>ij</sub>] β<sub>j</sub>(o<sub>t</sub>) [0081] If you use the log representation to avoid underflow, the maximum likelihood is ψ<sub>j</sub>(t) = max [ψ<sub>i</sub>(t-1) + log (α)<sub>ij</sub>)] + Log (b)<sub>j</sub>(o<sub>t</sub>)) [0082] This is also known as the Viterbi algorithm. The vertical dimension indicates the state of the HMM, and the horizontal dimension is visualized by searching for the optimal path through the audio frame, that is, the matrix indicating time. To complete the extension to concatenated speech recognition, it is assumed that each HMM showing a further potential sequence is concatenated. Therefore, training data for continuous speech recognition consists of concatenated utterances, but does not need to know the boundaries between words. [0083] To improve computational speed / efficiency, the Viterbi algorithm sometimes achieves convergence using a model known as a Token Passing Model. The token passing model is a remark sequence o<sub>1</sub>~ O<sub>t</sub>And a partial match between specific models, and is constrained to be the state j model at time t. This token passing model can be easily extended to connected speech recognition if it can be defined as a sequence finite state network of HMMs. A complex network containing phoneme-based HMMs and complete words is configured to recognize a single best word to form a concatenated speech with N best extracted words from the probability grid. This composite HMM-based articulated speech recognizer is the basis of the NLQS speech recognizer module. In any case, the present invention is not limited to a special form of speech recognition device, but is compatible with the architecture of the present invention, and the required ability of accuracy and speed to provide the user with a real-time dialogue experience. Other techniques for speech recognition can also be adopted as long as the criteria are met. [0084] The representation of speech for a speech recognition system based on HMM according to the present invention is that speech is essentially either a quasi-periodic pulse sequence (of spoken speech) or a random noise source (of unspoken speech). These are modeled as two audio sources, one as an impulse train generator with a pitch period of P and a random noise generator that can be controlled by a vocal / non-vocal switch. The output of the switch is supplied to the gain function predicted from the voice signal and expanded to supply the digital filter H (z) controlled by the vocal tract parameter characteristics of the generated voice. All parameters for this model, such as vocal / non-vocal switching, vocal pitch period, gain parameters for audio signals and digital filter coefficients-change slowly over time. In extracting acoustic parameters from the user's voice input to allow evaluation in terms of the HMM set, cepstrum analysis is commonly used to separate vocal tract information from vibration information. The cepstrum of a signal is calculated by adopting the Fourier (or similar) transform of the log spectrum. The main advantage of extracting cepstrum coefficients is that they are uncorrelated, allowing diagonal covariances to be used with the HMM. Since the human ear eliminates frequency non-linearity across the speech spectrum, it has been shown that a similarly non-linear front end improves speech recognition performance. [0085] Therefore, regardless of analysis based on general linear prediction, the front end of the NLQS speech recognition engine is approximately equal on the Mel-scale.<u style="single">Resolution</u>Performs a simple Fast Fourier Transform based on a filter bank designed to give degrees. To implement this filter bank, the window of audio data (for a particular time frame) is transformed using a software Fourier transform and a possible size. The magnitude of each FFT is multiplied by the gain of the corresponding filter and the result is accumulated. The cepstrum coefficient obtained from this filter bank analysis by the front end is calculated in the first partial processing step of the audio signal by using the discrete cosine transform of the amplitude of the log filter bank. These cepstrum coefficients are the Mel-Frequency Cepstral. Called Coeficient (MFCC), it shows the user some of the audio parameters sent by the client to characterize the acoustic characteristics of the audio signal. These parameters are selected and operational for a variety of reasons, including the fact that they are quickly and universally determined in systems with different performance (ie, for everything from low-performance PDAs to high-performance desktops). It fits for a number of related useful perceptions and is relatively small and compact so that it can be quickly delivered over a relatively narrow band link. Therefore, these parameters indicate the minimum amount of information that can be used so that the recognition process can be completed sufficiently and quickly on the subsequent server side. [0086] Energy conditions in the form of signal energy loggerism are added to enhance the voice parameters. Therefore, RMS energy is added to 12MFCC to produce 13 coefficients. These coefficients generate partially processed audio data that is promoted in a compressed state from the user's client system to a remote server. [0087] The performance of this speech recognition system is greatly enhanced by calculating the time derivative on the server side and adding it to the basic static MFCC parameters. These two other sets of functions-deltas and acceleration factors that indicate changes in 13 values from frame to frame (actually measured over several frames) are used to complete the initial processing of the audio signal. It is calculated during the second partial audio signal processing stage and added to the original set of coefficients after the latter is received. These MFCCs, along with the delta and acceleration factors,<u style="single">voice</u>data<u style="single">To</u>Used to determine the appropriate HMM for<u style="single">Remark vector O mentioned above</u><sub><u style="single">t</u></sub><u style="single">To compose</u>.. [0088] The delta and acceleration factors are as follows<u style="single">Regression</u>Calculated using the formula. [0089] d<sub>t</sub> = Σθθ<sub>=1</sub>[c<sub>t +</sub>θ --c<sub>t-</sub>θ] / 2Σθθ<sub>=1</sub>θ<sup>2</sup>[0090] Where d<sub>t</sub>Is the delta coefficient of time t calculated for the corresponding static coefficient [0091] d<sub>t</sub> = [c<sub>t +</sub>θ --c<sub>t-</sub>θ] / 2θ [0092] In the general independent implementation of speech recognition systems, the entire SR engine is run by a single client. In other words, the first and second partial processing steps described above are performed by a DSP (or microprocessor) running in a ROM or software code routine on the client's computer. [0093] Conversely, for several reasons, especially cost, technical performance and client hardware uniformity, the NLQS system employs a split or distributed approach. Some processing is done on the client side, but the main speech recognition engine runs on an intensively deployed server or a large number of servers. More specifically, as described above, the capture of the audio signal, the extraction and compression of the MFCC vector is performed by the client machine in the first partial processing step. Routines are thus streamlined and simplified enough to run browser programs (eg, plugin modules, even as downloadable applets), maximizing ease of use and use. Therefore, it is possible to support even a client platform with very low performance, and it is possible to use this system at a large number of expected sites. The primary MFCC is then sent to the server via dial-up channels such as Internet connection, LAN connection, and wireless connection. After decompression, the delta and acceleration factors are calculated in the server to complete the initial speech processing step and the resulting speech vector O<sub>t</sub>Is determined. [0094] Overview of speech recognition engine The speech recognition engine is also located on the server and is based on an HTK-based recognition network compiled from a set of word-level networks, dictionaries and HMMs. The recognition network is composed of a set of nodes connected by an arc. Each node is an HMM model or word end. The nodes of each model are themselves configured to be connected by an arc. Thus, when compilation is complete, the speech recognition network consists of HMM states linked by transitions. For unknown speech inputs in T-frames, all paths from the network's entry node to exit node go through the THMM state. Each of these paths is calculated by adding the log probabilities of the individual transitions within each path and the log probabilities of the emission states that generate the corresponding statements. The function of the Viterbi decoder is to find the path with the highest log probability on the network. A token passing algorithm is used for this detection. Allow the propagation of these potential winners in a network with many nodes<u style="single">thing</u>Only reduces the calculation time. This process is called batch erasure. [0095] Natural language processing device<u style="single">Day</u>Tabass<u style="single">Typical to</u>In the natural language interface, the user inputs a question in his / her natural language, such as English. The system interprets this and translates it into a query language expression. The system then processes the question using the query language representation, and if the search is successful, the recordset showing the results is formatted in either raw text or image format and displayed in English. A well-working natural language interface has many technical requirements. [0096] For example, this must be sturdy, "what's the department<u style="single">s</u> In the "turnover)" sentence, determine the word whats = what's = what is. You also have to determine departments = department's. In addition to being robust, the natural language interface needs to identify some form of ambiguity in natural language, such as linguistic, structural, reference and abbreviation ambiguities. All of these requirements are carried out by the present invention in addition to the general ability to perform basic linguistic morphological actions of tokenization, tagging and grouping. [0097] Tokenization is done by a text analyzer that treats the text as a series of tokens that are larger than individual characters but smaller than a phrase or sentence, or as a significant unit available. These include words, word separable parts and punctuation. Each token is associated with an offset or length. The first step of tokenization is a segmentation process that extracts individual tokens from the input text and records the offset of each token derived within the input text. The output of the tokenizing device lists the offset and category of each token. In the next stage of text analysis, the tagging device uses a built-in morphological analyzer to describe each word / token in the sentence.<u style="single">Search</u>And list all parts of the audio internally. Output is voice recording<u style="single">of</u>It is an input string by each token with a tag attached to the part. Finally, as a phrase extractor or phrase analyzer<u style="single">function</u>The grouping device that forms the phrase determines the group of words that form the phrase. These three behaviors, which are the basis of all modern language processing systems, implement all optimized algorithms to determine the single best answer to a user's question. [0098] SQL database and full text query Another important component of the system is the SQL database. This database is used to store text, especially answer-question pairs, which are stored in the database's full-text table. In addition, the full-text search capability of the database makes it possible to perform full-text search. [0099] All digitally stored information is in the form of unstructured data, primary text, but it is possible to store this source data in a publicly known database with text-based columns such as varchar and text. Become. In order to effectively retrieve source data from the database, the technology needs to perform the search for answers in a meaningful way, as in the case of questioning and answering the source data, as in the case of the NLSQ system. [0100] There are two main types of source search. Property-This search technique first filters documents to extract properties such as author, subject, type, number of words, number of pages printed, and last written date, and then for these properties. Do a search. Full Text-This search technique first indexes non-noise words in a document, which in turn supports language and proximity searches. [0101] Two additional techniques are implemented in this particular RDBM. SQL server is integrated. Search services-, index engines and parsing tools that accept full-text indexing and search services and full-text SQL extensions and maps, called search, make it possible for search engines to process. [0102] Four main aspects of performing a full-text search of plain text from a full-text-capable database<u style="single">To</u>Includes. Managing the definitions of registered tables and fields for full-text search; indexing the data in the registered fields-the indexing process scans strings and breaks words. (Called) is determined, all noisy words are eliminated (this is called stop words), and the remaining words are given a full-text index; for the registered fields for the full-text index to occupy. Ask a question; ensure that subsequent changes to the data in the fields created to keep the full-text index synchronized are propagated to the index engine. [0103] A potential design principle for indexing, querying and synchronization processing is the presence of a full-text unique key field (or single field basic key) in every table created for full-text search. The full-text index has an entry for a non-noise word on each row, along with the value of the key field on each row. [0104] When processing a full-text search, the search engine returns the key values of the rows that match the search criteria to the database. [0105] The full-text management process is started by specifying the relevant table and its fields for full-text search. The customized NLQS stored procedure is first used to create tables and fields that are convenient for a full search. After that, another request is made according to the stored procedure, and the full-text index is stored. As a result, the potential index engine is called and the asynchronous index storage is started. Full-text indexing searches for which significant words were used and where they are located. For example, in the full-text index, the word "NLQS" can be found in word number 423 and word number 982 in the Abstract column of the DevTools table in the row associated with ProductID6. This index structure supports efficient searching of all items, including indexed words, as well as advanced searches such as phrase search and proximity search. (An example of a phrase search is "white" An example of a proximity search, where "elephant" is searched and "white" is followed by "elephant", is to search for "big" and "house". "Large" occurs in close proximity to the "house". ) To prevent the full-text index from becoming huge, "a", "and", "the", etc. are ignored. [0106] Transact-The extension to the SQL word is used to construct a full-text query. The two key predicates used in NLQS are CONTAINS and FREETEXT. [0107] CONTAIN<u style="single">S</u>Whether the predicate contains several words and phrases in the full text registration field<u style="single">Ka</u>Used to determine. In particular, this predicate Word or phrase · Word or phrase prefix · Adjacent words or phrases -Words that are inflectional forms of other words (for example, "drive" is the uninflected form of "drives", "drove", "driving" and "driven") -A set of words or phrases that are assigned different weights to each Used for searching for. [0108] SQL server relational engine is CONTAIN<u style="single">S</u>And FREETEXT predicates are recognized, and the minimum syntax and meaning are checked, such as confirmation that the fields referenced in the predicate are registered for full-text search. During the execution of the query, the full-text predicate and other related information are sent to the full-text search section. After checking the syntax and meaning, the search engine is started and returns a unique key value that identifies the row in the table that satisfies the full-text search condition. In addition, FREETEXT and CONTAINS are combined with other predicates such as AND, LIKE, and NEAR to generate a customized NLQS SQL structure. [0109] SQL database full-text query architecture The full-text query architecture consists of several components: the full-text query part, the SQL server relational engine, the full-text provider and the search engine. [0110] The full-text query part of the SQL database is a full-text predicate or rowset-valued function from SQL Server. function) is accepted, a part of the predicate is converted to the internal format, and this is sent to the search service that returns a low set match. The low set is then sent back to the SQL server. SQL Server uses this information to generate a result set and returns it to the sender of the query. [0111] SQL Server Relational Engine includes CONTAINSTABLE () and FREETEXTT along with CONTAINS and FREETEXT predicates.<u style="single">A</u>BLE () Accepts low set value functions. During the analysis period, this code checks the status such as making an inquiry to a column that is not registered for full-text search. If valid, forward search conditions (ft_search_condition) and contextual information are sent to the full-text provider at runtime. After all, the full-text provider returns the low set to SQL Server and is used for the join (identified or implied) to the original query. The full-text provider interprets and confirms the forward search condition, constructs an appropriate internal representation of the full-text search condition, and sends it to the search engine. The result is returned to the relational engine by a low set of rows that satisfy the forward search criteria. [0112] Client side system 150 The architecture of the client-side system 150 of the natural language query system 100 is shown in detail by FIG. With respect to FIG. 2, the three main processes achieved by the client-side system 150 are shown below. Initialization process 200A consisting of SRE201, communication 202 and MS agent 203 routines; a) User voice reception 208 and b consisting of SRE204 and communication 205; b) Receive response from server 207 and uninitialization (un- Iterative processing 200B consisting of two subroutines of initialization) processing. Finally, the uninitialization process 200C is composed of three subroutines, SRE212, communication 213, and MS agent 214. Each of the above three processes will be described in detail in the following paragraphs. Execution of such processing and routines varies from client platform to client platform, and in some environments such processing can be performed by hard-coded routines executed by a dedicated DSP, but by a shared host processor. Those skilled in the art will appreciate that it can be implemented as software to be executed, or a combination of the two can be used. [0113] Initialization on client system 150 The initialization of the client-side system 150 is shown in Figure 2-2, which generally consists of three separate initialization processes: the client-side speech recognition engine 220A, the MS agent 220B, and the communication process 220C. [0114] Initialization of voice recognition engine 220A Speech recognition engine 155 initially uses the routine shown in 220A<u style="single">Being</u>, Consists of. First, the SRECOM library is initialized. Memory 220 is then allocated to hold the source and a coder object is created by routine 221. The configuration file 221A from the configuration data file 221B is also loaded at the same time as the SRE library is initialized. In configuration file 221B, the type of input to the coder and the type of output to the coder are declared. Such routines and actions are well known and they are performed using a number of fairly simple methods. Therefore, they will not be described in detail here. The voice and silence components of the speech are then scaled using routine 222 by a well-known procedure. In order to scale the voice and silence components, the user preferably speaks a sentence and displays it in a text box on the display screen. The SRE library predicts the noise and other parameters needed to detect the silence and audio components of future user speech. [0115] Initialization of MS Agent 220B The software code used to initialize and set up the MS Agent 220B is shown in Figure 2-2. The MS Agent 220B routine is responsible for coordinating and handling the operation of Video Agent 157 (Fig. 1). This initialization consists of the following steps. [0116] 1. Initialize COM library 223. This part of the code initializes the COM library needed to use the well-known control ActiveX control. 2. Instantiate Agent Server 224-This part of the code creates an instance of Agent Active X Control. 3. Load MS Agent 225-Load MS Agent characteristics from the identified file 225A, which contains general parameter data about agent characteristics such as full picture, shape, size, etc. 4. Acquisition of Character Interface 226-This part of the code acquires the appropriate interface for the identified character: the character has different control / interaction capabilities that can be provided to the user. 5. Add a command to agent character option 227-This part of the code appends the command to the agent's property sheet accessible by clicking the icon in the system tray when the agent character is loaded. For example, a character can speak, how he / she moves, TTS properties, etc. 6. Show Agent Character 228-This part of the code shows the agent character so that it can be seen by the user. 7.AgentNofifySink-to handle events<u style="single">Things</u>.. The code in this part creates an AgentNotifySink object 229, registers it with 230, and acquires the agent property interface 231. The property sheet for the agent character is assigned to routine 232 to be used. 8. Animate the character 233-This part of the code causes a specific character to animate to reach the user's NLQS100. [0117] The above configures the entire sequence required to initialize the MS agent. Similar to SRE routines, MS agent routines can be performed by one of ordinary skill in the art based on the present art by any suitable known technique. The specific structure and behavior of these routines are not important and will not be discussed in detail. [0118] In a preferred embodiment, the MS agent has a suitable appearance and appearance for a particular application.<u style="single">ability</u>Is configured to include. For example, when used for distance learning, the agent may look like a university professor or<u style="single">habit,</u>Attitude, gesture<u style="single">To</u>Have. Other visual props (blackboard, textbooks, etc.)<u style="single">To</u>Agent<u style="single">Is</u>It can be used and reminds the user of the experience of being in a real educational environment. Agent features consist of 150 client-sides and / or parts of code executed by a browser program (not shown) in response to configuration data and commands from a particular web page. For example, certain websites that provide medical care<u style="single">Then</u>, It is preferable to use a visual image of a doctor. These and many other changes will be adopted by those skilled in the art to improve the user's experience of human and real-time interaction. [0119] Initialization of communication link 160A The initialization of the communication link 160A will be explained with reference to the process 220C in Fig. 2-2. For Figure 2-2, this initialization consists of the following code components: Internet Connection 234-This part of the code makes an internet connection and sets the parameters for the connection. The set callback state routine 235 then sets the callback state to inform the user of the connection state. Finally, the new HTTP Internet session 236 starts a new Internet session. The details of communication link 160 and the setup process 220C are not important and will vary from platform to platform. In some cases, the user uses a slow dial-up connection, a dedicated high speed exchange connection (eg T1), an always-on xDSL, a wireless connection, and so on. [0120] Query / answer iteration When initialization is complete, iterative query / answer processing begins when the user presses the start button to start the query, as shown in Figure 3. As for FIG. 3, the iterative query / response process consists of two main sub-processes: reception of user voice 240 and reception of user response 243, which execute the routines of the client-side system 150. The receive user voice 240 routine receives voice from the user, while the receive response to the user 243 routine receives the user in the form of text from the server for conversion to voice to the user by the text-voice engine 159. Receive answers to your questions. As used here, the term "query" is used in the broadest sense to mean some form of input used as a control variable by a question, command or system. For example, a query consists of questions directed to a particular topic, such as "What is a network?" In a distance learning application. In an e-commerce application, a query is, for example, a command such as "List all books by Mark Twain." Similarly, answers in distance learning applications consist of text that is audible by the text-speech engine 159, including graphic images, sound files, video files, etc., as required by the particular application. It can be multimedia information in the format of. Given the technique of the present invention with respect to the structure, function, performance, etc. required for the client-side user voice reception 240 and user response reception 243 routines, those skilled in the art can implement this in a variety of ways. [0121] Receiving User Voice-As shown in FIG. 3, the user voice receiving routine 240 consists of SRE241 and communication 242 processing, both of which receive the user's speech and partially process the client-side system 150. It is configured as a routine of. SRE routine 241 is a coder object<u style="single">Is so</u>Use a coder 248 prepared to receive audio data from the voice object. Then the start source 249 routine is started. The code for this part is<u style="single">Da</u>Start a data search using the source object given to the object. The MFCC vector 250 is then continuously extracted from the generated voice until silence is searched. As mentioned above, this indicates the first step of processing the input audio signal, and in the preferred embodiment, it is limited to the calculation of the MFCC vector for the reasons already described above. These vectors contain 12 cepstrum coefficients and RMS energy conditions for 13 different numbers of the partially processed audio signal. [0122] Computational resources available in an environment,<u style="single">server</u>Transmission band in data link 160A available in side system 180, in data<u style="single">To</u>The MFCC delta parameter and MFCC acceleration parameter are calculated by the client-side system 150 according to the speed of the transmitter used for transmitting data. These parameters are automatically determined by the client-side system or direct control by the user during SRE155 initialization (using some type of grading routine to measure resources) and share the signal processing. It can be optimized on a case-by-case basis. Server for some applications<u style="single">~ side</u>System 180 may lack the resources and routines to complete the processing of the input audio signal. Therefore, in some applications, the signal processing load allocation can be varied and the audio signal processing<u style="single">Both</u>The steps can be performed on the client-side system 150, and the audio signal can be processed completely rather than partially and sent on the server-side system 180 for conversion into a query. [0123] Sufficient accuracy from a question / answer standpoint in preferred embodiments<u style="single">and</u>In order to ensure real-time performance, the client-side system will be able to use sufficient resources to partially process 100 frames per second of audio data so that it can be transmitted over link 160A. Minimal information required to complete the speech recognition process (13)<u style="single">Person in charge</u>Other latency (ie, client-side computational latency, packet formation latency, transmission latency) is minimal as only numbers are transmitted)<u style="single">Ri</u>, The system achieves highly optimized real-time performance. The principle of the present invention is an input audio signal by SRE (ie, non-MFCC based).<u style="single">To disassemble</u>It is clear that it can be extended to other SR applications where other methods are used. Only<u style="single">Standards of</u>SR processing is similarly in multiple stages<u style="single">Split</u>Possible<u style="single">Is</u>By that<u style="single">The load at different stages can be handled on both sides of the link 160A, depending on the goals and requirements of the entire system.</u>To. Functions of the present invention<u style="single">sex</u>Is for each system<u style="single">Based on</u>Achieved and predicted required for each concrete implementation<u style="single">Typical</u>Optimization is required. [0124] Accordingly, the present invention achieves response rate performance that is calculated, encoded, and adjusted according to the amount of information transmitted by the client-side system 150. In applications where real-time performance is paramount, the minimum amount of extracted audio data is transmitted to reduce latency, and in other applications it is processed.<u style="single">、</u>The amount of coded and extracted audio that is transmitted can be varied. [0125] The communication-transmission module 242 is used in preferred embodiments to transmit data from the client to the server over the data link 160A posted on the Internet. As described above, the data composed of the encoded MFCC vectors is used by the server-side speech recognition engine to complete the speech recognition decoding. The communication sequence is as follows. [0126] OpenHTTPRequest251-This part of the code first converts the MFCC vector to a sequence of bytes and then processes the bytes for compatibility with a protocol known as HTTP. This protocol is well known and other data links and other suitable protocols can be used. [0127] 1. MFCC byte string coding 251-This part of the code encodes the MFCC vector so that it can be sent to the server over HTTP. 2. Data transmission 252-This part of the code sends the MFCC vector to the server using internet connection and HTTP protocol. [0128] Waiting for server response 253-This part of the code is the data link and server-side system 1<u style="single">8</u>Monitor the response coming from 0. In general, the MFCC parameters are from the input audio signal.<u style="single">on</u>The fly<u style="single">so</u>Extraction<u style="single">Be done</u>、<u style="single">Or</u>Observation<u style="single">Be done</u>To. They are then encoded into HTTP bytes and sent to the server in a streaming manner before silence is detected, i.e. to the server-side system before the utterance is complete. This aspect of the invention facilitates real-time operation, as data can be transmitted and processed while the user is speaking. [0129] Receive the answer from the server<u style="single">To do</u>243 consists of the following modules shown in Figure 3: MS Agent 244, Text-Voice Engine 245 and Receive Module 246. These three modules work together to receive a response from the server-side system 180. As shown in Figure 3, the reception process<u style="single">、</u>It consists of three separate processes that act as receive routines on the client-side system 150. Receive the best answer<u style="single">To do</u>The 258 receives the best answer over the data link 160B (HTTP communication channel). The answer is decompressed at 259, then sent by code 260 to MS agent 244 and received by its code part 254. Routine 255 is then uttered with an answer using the text-speech engine 257. Of course, the text can be displayed for feedback on the monitor used by the client-side system. Text-speech engines are suitable for specific language applications (English, French, German, Japanese, etc.)<u style="single">Painful</u>Use natural language audio data file 256. As described above, when the answer is more than text, the answer is given to the user by a graphic image, sound, a video chip, or the like. [0130] Figure 4 shows the procedure and process of deinitialization. These functional modules are used to uninitialize the basic components of the client-side system 150, which include the SRE270, Communication 271 and MS Agent 272 uninitialization procedures. .. To uninitialize the SRE220A, the memory allocated during the initialization phase is freed by code 273, and the objects created during this initialization phase are deleted by code 274. Similarly, as shown in FIG. 4, in order to uninitialize the communication module 220C, the Internet connection already established with the server is closed by the code part 275 of the communication uninitialization procedure 271. The internet session created during the initialization stage is then also closed by step 271. For MS Agent 220B uninitialization, MS Agent Uninitialization Step 272 first releases command interface 227 using Step 277, as shown in FIG. This allows the property during loading of the agent character by step 225.<u style="single">C</u>A command to be added to the command is issued. The character interface initialized by step 226 is then released by step 278 and the agent is unloaded in step 279. In addition, the sink object interface is released in step 280, and then the property sheet interface is released in step 281. The agent notification sink then unregisters the agent in step 282, and finally the agent interface is unregistered in step 283, which allows all resources allocated in the initialization step shown in Figure 2-2. It will be released. [0131] It is well-versed in the art that the specific implementation of the deinitialization process and procedure, as shown in Figure 4, will vary depending on the client and server platform, as well as the other procedures described above. It is obvious to the person. The structure, operation, and the like of such a procedure are known in the prior art examples, and these can be implemented by using many direct methods. Therefore, these changes will not be mentioned in detail. [0132] Server-side system<u style="single">Te</u>Description of Mu 180 preface 11A to 11C show high-level flowcharts of an embodiment of the processing group implemented in the server-side system 180 of the natural language query system 100. In this embodiment, this process consists of a two-step algorithm for processing a voice input signal, recognizing the meaning of a user's inquiry, and obtaining an appropriate answer or response to each inquiry. [0133] The first step shown in FIG. 11A is as a high-speed first pruning mechanism and includes the following processes. After processing the voice input signal, the user's inquiry is recognized in step 1101, the string of the inquiry text is sent to the natural language engine 190 (see Figure 1) in step 1107, and at the same time DB engine 186 (as well) in step 1102. (See Figure 1). Here, "recognized" means that the user's inquiry is converted into a character string of a unique natural language sentence by the HMM method described above. [0134] In NLE190, the lexical string undergoes morpheme language analysis processing in step 1108, the lexical string is tokenized, tagged, and the tagged token is classified. Next, in step 1109, the noun phrase (NP) of the string is stored and copied and transferred in step 1110 so that the DB engine 186 can use it in the DB process. In step 1102, as shown in Figure 11A<u style="single">D</u>The string corresponding to the user query transferred to B engine 186 is used with the NP received from NLE190 to query the SQL in step 1103.<u style="single">-</u>Is configured. Then in step 1104, SQL query<u style="single">-</u>Is executed, and a record set of plausible questions in response to user inquiries is provided in step 1105.<u style="single">all</u>Obtained as a result of a statement search and further returned to NLE190 in the form of an array in step 1106. [0135] As mentioned above, the first step of server-side processing acts as an efficient and fast pruning mechanism, and in a very short time, accurate search results corresponding to the user's actual inquiry can be a plausible sentence candidate. It is intended to narrow down to. [0136] The second step shown in FIG. 11B can be regarded as a more accurate selection processing unit in the recognition processing when compared with the first step described above. This step begins with linguistic processing of each of the questions stored in an array, obtained by full-text search as potential candidates for a user's inquiry. The processing of these stored question sentences proceeds in NLE190 as follows. In step 1111, morpheme language analysis is performed for each question sentence in the array of question sentences corresponding to the record set obtained by the SQL full-text search. In this process, the character string corresponding to the searched question sentence candidate is tokenized, tagged, and the tagged token is classified. Next, in step 1112, the noun phrase in the character string is extracted and stored. This process is repeatedly determined by the determination point 1113, and steps 1118, 1111, 1112, and 1113 are repeated so as to obtain and store the NP for the obtained inquiry sentence candidates. When the NP is extracted for each of the query statement candidates in the array, step 1114 compares each of the query statement candidates in the array with the user's query based on the magnitude of the NP value. If step 1117 determines that there are no more query statements in the array to process, step 1117A identifies the stored query with the highest NP value for the user query and best fits the user query. Judge as a stored inquiry. [0137] In particular, it can be seen that the second step of the recognition process requires a larger amount of calculation than the previous first step because a plurality of character strings are tokenized and it is necessary to compare a plurality of NPs. However, this is not realistic unless the first step is to quickly and efficiently narrow down the candidates to be evaluated to a considerable extent. Therefore, since this large amount of calculation in the present invention contributes to bring about a more accurate recognition result in the entire recognition process of the inquiry sentence, its feature is rather valuable. Therefore, in this regard, the second step of inquiry statement recognition aims to ensure accuracy throughout the system, whereas the first step provides the user with sufficient speed to provide a real-time response sensation. The purpose is to do. [0138] Figure 11C shows the final part of the inquiry and response process, which is the process of providing the user with an appropriate and matched answer and response. First, in step 1120, identify the conforming one in the stored query text. Then, in step 1121, the file path corresponding to the answer to the identified conformance question is retrieved. Further processing is continued, with step 1122 extracting the answer based on this file path, and finally step 1123 compressing the answer and sending it to the client-side system 150. [0139] The above content is intended to give an overview of the basic elements, behavior, functions and characteristics of each part of the NLQS system of the server-side system 180. Below, for each subsystem<u style="single">Details</u>Will be explained. [0140] Software used in server-side system 180<u style="single">D</u>A module Major software used in the server-side system 180 of the NLQS system<u style="single">D</u>A module is shown in Figure 5. These generally include the following elements: The communication module 500 includes the communication server ISAPI500A (executed by the SRE server side 182 in FIG. 1 and detailed below) and the database process DB process module 501 (executed by the DB engine 186 in FIG. 1). Includes a natural language engine module 500C (executed by NLE190 in Figure 1) and an interface 500B between the NLE process module 500C and the DB process module 500B. As shown here, the communication server ISAPI500A includes a server-side speech recognition engine and a suitable communication interface located between the client-side system 150 and the server-side system 180. Furthermore, according to FIG. 5, it can be said that the server-side logic of the natural language query system 100 is characterized by including two dynamic link library components, the communication server ISAPI500 and the DB process 501. The communication server ISAPI500 is composed of three submodules: a server-side speech recognition engine module 500A, an interface module 500B between the natural language engine module 500C and the DB process 501, and a natural language engine module 500C. [0141] DB process 501 connects to the SQL database and configures SQL queries in response to user queries.<u style="single">-</u>It is a module whose basic function is to execute. In addition, this module has an interface that connects to the logic to retrieve this exact answer, based on the file path once the answer was obtained from the natural language engine module 500C. [0142] Speech recognition subsystem on server-side system 180 182 The server-side speech recognition engine module 500A is a group of distributed components that perform the necessary functions and operations of the speech recognition engine 182 (Fig. 1) on the server-side 180. As is known, these components can be implemented as software processing executed on the server side 180. Figure 4A shows a more detailed development of the behavior of the server-side speech recognition component, which will be described below. [0143] Inside part 601 of server-side SRE module 500A, the client-side system 150 extracts and communicates channel 160.<u style="single">Assistance</u>Receives a binary MFCC vector byte stream corresponding to the acoustic characteristics of the transmitted audio signal. The MFCC acoustic vector is decoded from the encoded HTTP byte stream as follows. Since the MFCC vector contains a built-in null character, these vectors cannot be sent in this format to the server side that uses the HTTP protocol. Therefore, first of all, prior to the transfer, the client side 150 encodes the MFCC vector and converts all the audio data into a byte stream so that the data does not contain null characters. A single null character is inserted at the end of the byte stream to indicate the end of the byte stream that is transferred to the server over the Internet 160A using the HTTP protocol. [0144] As mentioned above, a smaller number of bytes of data (13 MFCC coefficients) are sent from the client-side system 150 to the server-side system 180 in order to maintain and store the latency between the client and server. This can be done automatically or adjusted for a particular application environment to ensure uniformity for each platform. For example, the server may calculate the delta and acceleration factors (26 more calculations), but the client may want to encode them, transfer them, and determine if they can be done in less time than decoding from an HTTP stream. That's right. This is preferred because the server-side system 180 is typically well equipped to calculate the MFCC delta and acceleration parameters. In addition, server resources are easier to manage than client resources, which means that future upgrades, optimizations, etc. will be enjoyed throughout the system to provide more reliable and predictable overall system performance. There is a situation that it is easy. Therefore, the present invention can be implemented even in the worst-case scenario where the client machine is extremely poor and has only enough resources to input voice input data and perform only minimal processing. [0145] Dictionary preparation and grammar files In Figure 4A, code block 605 receives various options selected by the user (or sought out from the user's state within a particular application). For example, in the example of a distance learning system, course, chapter and / or section data are communicated. For other applications (eg e-commerce), other data options that the user uses to browse in his / her browser, such as product type, product category, product brand, etc., are communicated. To. These selected options are based on the background that the user has experienced in the interactive process, thus limiting and defining the scope of the search. That is, the background is a grammar and dictionary for Viterbi decoding when dynamically loading the speech recognition engine 182 (Fig. 1) and analyzing the user's speech utterance. In order to optimize speech recognition, both grammar and dictionary files are used in this example. The grammar file provides a world of user queries available. That is, it provides all possible terms to be recognized. The dictionary file is based on the phonemes of each word contained in the grammar file (information on how to pronounce the words, which is based on the specific natural language file installed, for example UKEnglish or USEnglish (US English). (US English) is like). If all the sentences in a specific recognizable environment are contained in a single grammar file, the accuracy of the recognition deteriorates, and the loading time of this grammar and dictionary file alone is the speech recognition process. It is clear that it will be shorter than the time of. [0146] To avoid this problem, dynamically load a specific grammar and act as the current grammar depending on the user's background of use, for example, the selection of courses, chapters and sections in a distance learning system. It is conceivable to configure the system. Grammar and dictionary files can be dynamically loaded according to a given course, chapter and section that the user is reading and writing, or automatically loaded by an application program run by the user. [0147] The second code block 602 is a part that implements the initialization of the speech recognition engine 182 (Fig. 1). The MFCC vector received from the client-side system 150 along with the grammar file name and dictionary file name is input to this block to initialize the voice decoder. [0148] As shown in FIG. 4A, the initialization process 602 uses the following subroutines. First, routine 602<u style="single">A</u>Now load the SRE library. Then code 602<u style="single">B</u>Then, the received MFCC vector is used to generate an object that is identified as an external source. Code 602<u style="single">C</u>Allocates memory to store recognized objects. Then routine 602<u style="single">D</u>Also creates and initializes the objects needed for recognition. These objects are the source, coder, and code 602<u style="single">E</u>Dictionary recognition and result loading generated by, code 602<u style="single">F</u>Hidden Markov Model (HMM) generated by, and Routine 602<u style="single">G</u>Is the loading of the grammar file generated by. [0149] As shown in FIG. 4A, the voice recognition 603 is the process to be called next, and is usually executed in response to the completion of the process of the user's voice input signal on the client side 150, but as described above. , When the MFCC vector is transmitted over the link 160, it is preferred that it be partially processed (ie, only the MFCC vector is calculated in the first phase). Subroutine 602<u style="single">B</u>Using the functions generated in the external source by, this code is from the external source 603<u style="single">A</u>Read MFCC vectors one at a time from and block these 603<u style="single">B</u>The audio pattern represented by the MFCC vector processed by<u style="single">-</u>Recognize the words in During this second phase, 13 additional delta coefficients and 13 acceleration coefficients were added to the calculation as part of the recognition process, for a total of 39 aforementioned observation vectors O.<u style="single">t</u>Is obtained. In addition, a previously defined set of hidden Markov models (HMMs) is used to determine the words that correspond to the user's utterances, as described above. As a result, the processing of "recognizing" the word in the inquiry processing is performed, and the processing result is used in the inquiry processing below. [0150] The distributed configuration and high-speed performance of word recognition processing is particularly beneficial in itself, and incorporating other query processing or performing it in combination with other environments that do not require it is for those familiar with the technical field. , Obvious. For example, in some applications, individual recognized words can simply be used to fill in data items in computer-generated formats, and the systems and processing methods described above do this. It is possible to provide a high-speed and highly reliable mechanism. [0151] When the user's voice is recognized, the processing flow of SRE182 proceeds to SRE uninitialization routine 604, where the voice engine is uninitialized as shown. In this block, routine 604<u style="single">A</u>Deletes all the objects previously created in the initialization block, and the memory allocated by the initialization block in the initialization phase is routine 604.<u style="single">B</u>Will be deleted by. [0152] Further, the above-mentioned contents are merely an example of an embodiment in which a specific routine used in the server-side speech recognition system of the present invention is implemented. According to the disclosure of the present invention, it is clear that other modifications are possible in order to achieve the predetermined functionality and purpose of the present invention as well. [0153] Database processor 186 behavior-DB process User inquiry processing<u style="single">of</u>SQL query used as part<u style="single">-</u>The structure of is shown in FIG. 4B, where the SELECT SQL statement is preferably constructed using known CONTAINS predicates. Module 950 is a SQL query based on this SELECT SQL statement<u style="single">-</u>Configure this query<u style="single">-</u>Is used to respond to inquiries uttered by the user (here, referred to as a question) and to search for the optimum inquiry text stored in the database. Routine 951 then adds a table name to the constructed SELECT statement. Routine 952 then calculates the number of noun phrases in the question asked by the user. In addition, routine 953 allocates the memory needed to accommodate all the words given by the NP. Routine 954 then finds a list of words (identifying all the individual words given by the NP). Then, in Routine 955, a SQL query that separates this set of individual words with the NEAR () keyword.<u style="single">-</u>In addition to. Then, in routine 956, SQL query<u style="single">-</u>Add the AND keyword after each NP in. Finally, code 957 frees memory resources and allocates memory to store words received from the NP for the next iteration. As described above, as a result of this processing, a complete SQL query corresponding to the question uttered by the user<u style="single">-</u>Is generated. [0154] Connect to SQL server<u style="single">―</u>SQL query in routine 710, as illustrated in Figure 4C<u style="single">-</u>After configuring, routine 711 connects to query database 717 and continues to process user queries. This connection procedure and the accompanying set of search records are implemented by the following procedure. [0155] 1. In routine 711A, assign server and database names to DB process member variables. 2. Routine 711B generates a connection string. 3. Connect the SQL server database under the control of code 711C. 4. SQL query in routine 712A<u style="single">-</u>To receive. 5. SQL query with code 712B<u style="single">-</u>To execute. 6. Query in routine 713<u style="single">-</u>Retrieves all records searched by. 7. In routine 713, allocate memory to store the entire set of questions. 8. In routine 713, store the entire number of pairs of questions in the form of an array. [0156] When routine 716 in Figure 4C receives the best answer ID from NLE14 (Figure 5), the code corresponding to 716C receives it and then sends it to code 716B, where it uses the record number to pass to the answer file. Is determined. The 716C then opens this file with the path to this file and reads the contents of the file corresponding to the answer. In addition, this answer is compressed by the 716D code and processed for transmission over communication channel 160B (Figure 1). [0157] NLQS database 188-table structure FIG. 7 shows an example of the logical structure of a table used in a typical NLQS database 188 (FIG. 1). When the NLQS database 188 is used as part of the NLQS query system 100 implemented as a distance learning / training environment, this database typically contains course 701 consisting of several chapters 702, 703, 704. , Will include an organized multi-level hierarchy. Each of these chapters has one or more sections 705, 706, 707, shown as Chapter 1. Chapter 2, Chapter 3 ... Chapter N has the same structure. Each section has one or more question / answer pairs 708, 709, 710 stored in a table detailed below. This configuration is appropriate and optimal for educational / training applications, but other implementation methods are possible and are whole for other applications such as e-commerce, electronic support, internet browsing, etc. It is clear that more suitable configurations can be made depending on the parameters of the system. [0158] The structure of the NLQS database 188 is intricately linked to the switchable grammatical structure described above. In other words, the background (or environment) that the user embodies is always determined based on the choices made at that section level, for example, only a subset of a limited set of questions / answers 708 is section 705. It is considered to be suitable for handling in. That is, while the user is faced with such a background, only a specific appropriate grammar is switched and selected for such a question / answer pair in order to handle the user's inquiry. Similarly, e-commerce applications for commerce over the Internet include, first-level Home page 701, which identifies options that users can select (product type, service, contact information, etc.). The second level consists of a hierarchical structure that includes specific "product types" 702, 703, 704, etc., and the third level contains specific product models 705, 706, 707, etc. Cass to handle inquiries for pairs 708, 709 and such product models<u style="single">Ta</u>It may consist of a mized grammar. Note that depending on the application, the specific implementation method will be changed according to the needs and desires of the business, and the processing routine is appropriate for such individual applications.<u style="single">To</u>It needs to be optimized. [0159] Table structure In this embodiment, an independent table is used for each course. That is, each database contains the following three types of tables. A master table as shown in FIG. 7A, at least one chapter table as shown in FIG. 7B, and at least one section table as shown in FIG. 7C. [0160] As shown in FIG. 7A, one embodiment of the master table has six columns. That is, field name 701A, data type 702A, size 703A, Null704A, base key 705A, and index 706A. These parameters are known in the field of database design and configuration. The master table has only two fields, chapter name 707A and section name 708A. Chapter names and section names are usually indexed. [0161] An example of a chapter table is shown in FIG. 7B. Like the master table, the chapter table has six columns. That is, field name 720, data type 721, size 722, Null723, base key 724, and index 725. But the data is 9<u style="single">Tsu</u>Contains the line. That is, in this case, chapter ID 726, answer ID 727, section name 728, answer title 729, paired question 730, answer path 731, creator 732, data generation date and time 733, and data modification date and time 734. [0162] A description of the fields in the chapter table is shown in Figure 7C. Each of the eight fields 730 has a description field 731 and stores the data corresponding to: [0163] Answer ID 732-An integer that is automatically updated for each answer at your convenience. Section Name 733-The name of the section to which the particular record belongs. This field and answer ID are used together as a basic key. [0164] Answer Title 734-A brief description of the title of the answer to the user's inquiry. Paired Questions 735-A set of one or more questions corresponding to an answer with the path stored in the next column, the path to the answer. [0165] Path to Answer 736-Contains the path to the file containing the answer to the related question stored in the previous column. For simple question / answer applications, this file is a text file, but it can also be a multimedia file that can be transcribed in some way via the data link 160 as described above. [0166] Creator 737-The name of the creator that generated the contents of the data. Data generation date and time 738-The date and time when the contents of the data were generated. Data modification date 739-Data<u style="single">contents</u>The date and time when was changed or modified. [0167] Of the section table<u style="single">one</u>An example is shown in FIG. 7D. The section table has 6 columns. That is, field name 740, data type 741, size 742, Null743, base key 744 and index 745. The data has seven rows: answer ID 746, answer title 747, paired questions 748, path to answer 749, creator 750, data generation date 751 and data modification date 752. These names correspond to fields and rows similar to the master and chapter tables described above. [0168] It should be noted that this example is an embodiment for a specific education / training application described above. There are a wide variety of applications for which the present invention may be available, depending on the application.<u style="single">Ta</u>Because of the ability to mize, other applications (including other teaching / training applications) may require or have a better implementation of other table, row, and field configurations and hierarchies. In some cases. [0169] Search Service and Search Engine-The query text search service is executed by the SQL search system 1000 shown in FIG. The system supports queries and handles full-text search. Here is the full text index. [0170] SQL search systems typically determine which entry in the database index matches the selection criteria specified by a particular text query configured in response to a clear user statement. The index engine 1011B collects indexes corresponding to indexable units of text for stored questions and corresponding answers in a full-text index table. It scans the string, determines word boundaries, removes all noise words, and collects the remaining words in the full-text index. Unique key field values and ranking values are returned as well to facilitate entry into the full-text database that matches the selection criteria. Catalog set 1013 is a file-system directory, accessible only by the administrator and search service 1010. The full-text index 1014 is organized into full-text catalogs, which are referenced by easy-to-use names. In general, full-text index data for all databases is placed in a single full-text catalog. [0171] The diagram of the full-text database described above (Figs. 7, 7A, 7B, 7C, 7D) is stored in Table 1006 shown in Fig. 10. These tables, for example, need to describe the structure of the memorized question / answer pairs required for a particular course. For each table-course table, chapter table, and section table there is field-column information, which defines each parameter that makes up the logical structure of the table. This information is stored in user and system table 1006. The key values corresponding to these tables are stored in the full-text catalog 1013. Therefore, when processing a full-text search, the search engine returns the key value of the row that matches the search criteria to SQL Server. The relational engine then uses this information to answer the question. [0172] As shown in FIG. 10, the full-text query processing is performed as follows. 1. Submit query 1001 using the SQL full-text structure formed by DB processor 186 to SQL relational engine 1002. 2. A query containing a CONTAINS or FREETEXT predicate is rewritten by routine 1003, and the responsive low set later returned by the full-text provider 1007 is automatically joined to the table on which the predicate operates. This rewrite is the mechanism used to ensure that these predicates are seamless extensions to traditional SQL servers. 3. After this, full-text provider 1007 is called and sends the following information to the query. Forward search condition parameter (this is a logical flag indicating a full-text search condition) b. The name of the full-text catalog where the full-text index of the table resides c. Local ID used for the language (eg word division) d. Database, table, and column identities used for this query e. If the query consists of two or more full-text structures; in this case, full-text provider 1007 is called individually for each structure. 4. SQL relational engine 1002 does not check the contents of forward search conditions. Instead, this information is passed to the full-text provider 1007, which validates the query and then constitutes the appropriate internal representation of the full-text search criteria. 5. The query request / command is then passed to query support 1011. 6. Query support 1012 returns a low set 1009 from the full-text catalog 1013, which contains its own key field values for every row that matches the full-text search criteria. The rank value is also returned for each row. 7. A low set of key field values 1009 is passed to the SQL relational engine 1002. If the processing of the query is related to the CONTAINSTABLE () or FREETEXTTABLE () function, the rank value is returned, otherwise the rank value is filtered out. 8. The low set value 1009 is added to the first query along with the value obtained from the relational database 1006, and the result set 1015 is returned for further processing to produce a response to the user. [0173] At this stage of the query recognition process, user statements are quickly transformed into already carefully crafted text queries, and this text query is suitable.<u style="single">To</u>The initial match group of results is processed first so that it can be further graded for the final decision of the matching question / answer pair. The fundamental principle that makes this possible is the existence of a full-text unique key field in each table registered for full-text search. In this way, when processing full-text search, SQL searcher<u style="single">Screw</u>1010 returns the key value of the row that matches the database to SQL Server 1002. In maintaining the full-text database 1013 and the full-text index 1014, the present invention does not update the full-text index 1014 immediately after the full-text registration field is updated.<u style="single">Say</u>It has unique characteristics. This operation is excluded in order to shorten the identification waiting time and increase the reaction speed again. Thus, this update of the full-text index table, which would otherwise take a considerable amount of time, is instead done synchronously at a better time than other database architectures. [0174] Interface between NLE190 and DB processing device 188 The question candidate result set 1015 corresponding to the user query statement is sent to the NLE190 for further processing as shown in Figure 4D to determine the "best" matching question / answer pair. The NLE / DB processor interface module handles user queries, analyzes noun phrases (NPs) in a set of search questions from SQL queries based on user queries, and compares the NPs of search questions with the NPs of user queries. Is adjusted between NLE190 and DB processing device 188. Therefore, this part of the server-side code contains a function that mediates the processing that resides in both the NLE block 190 and the DB processor block 188. This function is shown in Figure 4D; as you can see, Code Routine 800 serves to extract a noun phrase (NP) list from a user's question. This part of the code interacts with NLE190 to get a list of noun phrases in sentences that are clearly pronounced by the user. Similarly, routine 813 searches the NP list from the list of corresponding candidate / pair questions 1015 and ranks these questions (ranked by NP value).<u style="single">) Column</u>Remember to. Thus, at this point, NP data is being generated for the user query, as well as for question candidate 1015. A sentence such as "What issues have guided the President in considering the impact of foreign trade policy on American businesses?" As an example, NLE190 would return the following as nomenclature: President, issues, impact of foreign trade policy, American businesses, impact: ), Impact of foreign trade, foreign trade, foreign trade policy, trade, trade policy, policy, businesses. The methodology used by NLE190 is therefore apparent to those skilled in the art from this noun phrase set and noun sub-phrases generated in response to the exemplary query. [0175] The function identified as Get Best Answer ID 815 is then executed. This part of the code gets the best answer ID for the user query. To do this, routines 813A and 813B first find the number of noun phrases for each entry in search group 1015 that matches the noun phrase in the user query. Then routine 815<u style="single">A</u>Selects the final result record from the search candidate group 1015 that contains the maximum number of matching noun phrases. [0176] Traditionally, nouns are generally considered to be "naming" words, specifically the names of "people, places, or things." Nouns such as John, London, and computer certainly fit this description, but the types of words classified as nouns by the present invention are much broader. The nomenclature is also birth, happiness, evolution, technology, management, imagination, revenge, politics, hope, It can also show abstract and intangible concepts such as coffee, sport, and literacy. Applicants have found that it is far more appropriate to consider noun phrases as a key to linguistic standards, as nouns are so diverse compared to other parts of the conversation. And, a great many items classified as nouns by the present invention are useful for selecting and identifying individual remarks more easily and quickly than the prior art known in the art. [0177] Following this same idea, the present invention adopts and implements another linguistic element-word phrase, facilitating speech query recognition. Whether it's a noun phrase, a verb phrase, or an adjective phrase, the basic structure of a phrase has three parts: "pre-Head string," "Head," and "after the main word." Consists of a "post-Head string". For example, in the minimal noun phrase "the children," "children" is classified as the main word of the noun phrase. In summary, because of the variety and frequency of noun phrases, the adoption of noun phrases as a linguistically selected criterion for memorized answers is, like any other natural language, the application of this technology to English natural languages. Has a solid basis in. And that is, all noun phrases in the overall speech work very well as a unique type of speech query fingerprint. [0178] Next, the ID corresponding to the best answer to the selected final result record question is generated by routine 815, which is then returned to the DB processing shown in Figure 4C. As you can see here, the best answer ID<u style="single">Ha le</u>Received by routine 716A only taken to find answers file path used by the routine 716B. Routine 716C then opens and reads the answer file and tells Routine 716D the same content. The latter then compresses the answer file data and sends it to the client-side system 150 over the data link 160 for the aforementioned processing (ie for audible feedback, conversion to visible text / images, etc.). .. Repeatedly for learning / educational use<u style="single">To</u>The answer file may consist of a single text phrase, but for other uses the content and format will be formatted into a specific question in a suitable format. For example, an "answer" can consist of a list of entries that correspond to a list of responding category elements (ie, a list of books by a particular author), and so on. Other changes will be apparent depending on your own circumstances. [0179] Natural language engine 190 The schematic structure of the NL engine 190 will be described again with reference to FIG. 4D. This engine performs word analysis, that is, word morphological analysis that composes a user query, along with phrase analysis of phrases extracted from the query. [0180] As shown in FIG. 8A, the functions used for morphological analysis include tokenizers 802A, stemmers 804A, and morphological analyzers 806A. The functions that make up the phrase analysis include tokenizers, taggers, and groupers, the relationships of which are shown in Figure 8. [0181] The tokenizer 802A is a software module that functions to decompose the text of the input sentence 801A into a list of tokens 803A. In performing this function, the tokenizer 802A scans the input text 801A and treats it as a group of tokens, that is, useful meaningful units that are generally larger than individual characters but smaller than phrases and sentences. These tokens 803A include words, word separable parts, and punctuation. Each token 803A is given an offset and a length. The first step in tokenization is segmentation, which extracts individual tokens from the input text and maintains a track of offsets for each token derived within the input text. The category is then associated with each token based on its shape. The tokenization process is technically known and can therefore be carried out by any conventional application suitable for the present invention. [0182] Following tokenization, stemmer processing 804A is performed, which includes two separate forms for parsing the token and determining the respective stems 805A-the lexical form and the derivative form. The refraction stemmer recognizes the affix and returns the underlying word. Derived stemmers, on the other hand, recognize derived affixes and return one or more base words. The Stemmer 804A associates an input word with its base word, but has no audio information part. Analyzer 806B picks up a word regardless of context and returns a set of possible parts of statement 806A. [0183] As shown in FIG. 8, lexical analysis 800 is the next step, which takes place after tokenization. Token 803 is assigned to a part of the voice tag by tagger routine 804, and grouper routine 806 recognizes a group of words as some syntactic type phrase. These syntactic types include, for example, the noun phrases described above, but can optionally include other types such as verb phrases, adjective phrases, and the like. Specifically, the Tagger 804 removes ambiguity from parts of the conversation (parts-of-speech). disambiguator), which parses words in context. It has a unique morpheme analyzer (not shown) that allows each token to identify all possible parts of the conversation. The output of the tagger 804 is a string, and each token is tagged with a parts-of-speech label 805. The final step in language processing 800 is word grouping, forming phrase 807. This function is performed by the grouper 806 and, of course, is highly dependent on the performance and output of the tagger component 804. [0184] Therefore, at the end of language processing 800, a list of noun phrases (NP) 807 is generated in response to the user's query remarks. This group of NPs generated by the NLE190 significantly helps refine the search for the best answer, thus allowing a single best answer to the user's question later. .. [0185] The unique components of NLE190 are shown in Figure 4D and include several components. Each of these components fulfills some of the different functions required of the NLE190 as just described. [0186] Initialization of Grouper Resource Objects and Libraries 900-This routine initializes the structure variables needed to create Grouper Resource Objects and Libraries. Specifically, it initializes a specific natural language used by NLE190 to create a noun phrase, for example, a system offered on the English market initializes an English natural language. It then also uses the routines 900A, 900B, 900C and 900D, respectively, to the objects (routines) needed for the tokenizer, tagger and grouper (as described above).<u style="single">)</u>Create and initialize these objects with suitable values. It also allocates memory and stores all recognized noun phrases for the searched question pair. [0187] Word tokenization from a given text (from a query or pair of questions) is performed using routine 909B-where all words are tokenized with the help of a local dictionary used by the NLE190 resource. To. The resulting tokenized word is passed to the tagger routine 909C. In routine 909C, all tokens are tagged and the output is passed to grouper routine 909D. [0188] Grouping of all tagged tokens to form an NP list is performed by routine 909D so that the grouper groups all tagged token words and outputs a noun phrase. [0189] Deinitialization of grouper resource objects and release of resources are performed by routines 909EA, 909EB, and 909EC. These include token resources, tagger resources, and grouper resources, respectively. After initialization, resources are released. The memory used to store all noun phrases is also reassigned. [0190] Further Examples In the e-commerce embodiment of the invention shown in FIG. 13, the web page 130 has standard visible links such as books 131, music 132, etc., and the customer can be directed to those pages by clicking the appropriate links. Will be done. Web pages can be executed using HTML, java applets, or similar coding techniques that interact with the user's browser. For example, if a customer wants to buy Album C by an artist named Albert, he can take a closer look at some of the web pages as follows: He first clicks on music (Figure 13, 136). Then click on the record (Figures 14, 145). As shown in FIG. 15, this launches another web page 150 with a link for record 155 and has subcategories-artist 156, song 157, song title 158, genre 159. The customer then has to click on artist 156 to select an artist from the selection. This displays another web page 160 as shown in FIG. This page lists various artists 165 as shown-categories Albert 165, Brooks 166, Charlie 167 and White 169 are listed under Artist 165. The customer must now click on Albert 166 to browse the albums available on Albert. When this is done, another web page will be displayed as shown in Figure 17. This web page 170 again represents a similar look and feel, but under the heading title 175, the available albums 176, 177, and 178 are listed. Customers can also read additional information 179 for each album. This album information is similar to the commentary on shrink-wrapped albums purchased at retail stores. Once an album A is identified, the customer must click on album A176. This usually launches another text box with information about its availability, prices, shipping and handling, etc. [0191] When the web page 130 has the above-mentioned type of NLQS function, the web page interacts with the above-mentioned client-side and server-side speech recognition modules. In this case, the user may be a button called Contact Me for Help 148 (eg, this may be a link button on the screen or a key on the keyboard. You can start a query by simply clicking), and the character 144 will teach you how to get the information you need. If the user wants Album A of an artist named Albert, the user pronounces "Do you have Album A of Brooks?" In much the same way as when asking a human clerk at a brick or mortar facility. Due to the rapid recognition performance of the present invention, the user's question is answered in real time from the character 144 who speaks the answer in the user's own language. Easy-to-read words so that you can see the character's answer and run save / print options if needed<u style="single">Speech balloon</u>149 may also be displayed. For each page of the website, similarly appropriate question / answer pairs can be constructed based on this teaching, so that the customer can have a normal conversational, human question-and-answer dialogue in all aspects of the website. Is provided with an environment that emulates. Character 144 is modified by the user's own preference to take a specific voice style (male, female, young, elderly, etc.) according to a specific commercial use or to improve the customer's experience. -Adjustable. [0192] In a similar way, clearly pronounced user queries are received as part of traditional search engine queries, looking for interesting information on the Internet in a similar way to traditional text queries. If a reasonably close question / answer pair is not available on the server side (for example, if one of the appropriate pairs to the user's question does not reach a certain level of trust), the user has the option of expanding the scope. Assigned, and queries are simultaneously served to one or more different NLEs across multiple servers, improving the likelihood of finding a well-matched question / answer pair. Moreover, if desired, it is possible to find one or more "fits" in the same way that traditional search engines return a large number of potential "hits" in response to a user's question. For such questions, of course, real-time operation does not seem possible (due to scattered and decentralized processing), but the benefits provided by a wide range of ancillary question / answer database systems are desirable for some users. Will. [0193] Equally obvious, the NLQS of the present invention is very natural and saves a lot of time for the user, and for the e-commerce operator as well. In the electronic support embodiment, the customer can quickly and efficiently retrieve information without the need for a live customer agent. For example, on consumer computer system vendor related support sites, a simple diagnostic page is provided to the user with a visible support character to assist him / her. The user then simply pronounces the item (ie, "monitor" problem, "keyboard" problem, "printer" problem, etc.) from the Symptoms page, clearly pronouncing such symptoms at the recommendation of the supporting character. You can choose. The system then guides the user in real time to more detailed submenus, possible solutions, etc. for the particular disease recognized. The use of programmable characters can thus make a website a huge number of hits, or customer-acceptable scale, which requires and accompanies a corresponding increase in the number of human resources. There are no training issues. [0194] As a further embodiment, information retrieval on a particular website can be accelerated using the NLQS of the present invention. Even more very useful, the information is provided in a user-friendly manner through the natural interface of conversation. The vast majority of websites now employ a list of frequently asked questions that users typically work on one by one to get answers to their questions or questions. For example, as shown in Figure 13, the customer clicks Help 133 to start the interface with the lists. As shown in Figure 18, a website plan for a typical web page is displayed. This shows how many pages you have to go through to get to the list of frequently asked questions. Once on this page, the user has to scroll and manually identify the question that matches his / her inquiry. This process is usually a daunting task and may or may not provide information to answer user inquiries. Current technology for displaying this information is shown in Figure 18. This figure identifies how information is organized on a typical website: The help links (FIGS. 13, 133) commonly shown on the home page of a web page are illustrated in FIG. 18 as 180. Seeing Figure 18 again, each subcategory of information is listed on a separate page. For example, 181 lists sub-topics such as "first-time users", "search information", "orders", "shipments", and "your account". Other pages deal with "Account Information" 186, "Rates and Policies" 185, etc. At another level, "For First-time Visitors" 196, "Frequently Asked Questions" 195, "Safety Purchase Guarantee" 194, etc. Some pages on a particular page deal only with lower-lower topics. So, if a customer has a query that gives the best answer by going to a frequently asked question link, he or she has three levels to reach the frequently asked question page 195. You have to go past the crowded and cluttered screen pages. There is usually a large list of questions 198 that must be manually scrolled through. During the visual scroll, the customer must visually and intellectually match his or her questions with each of the listed questions. If a possible match is found, the question is clicked, and the answer appears in text format and is read. [0195] In contrast, the process of answering a question using the web pages available to this NLQS can be performed efficiently with much less effort. The user clearly pronounces the word "help" (Figs. 13, 133). This instantly spawns a character (Figures 13, 134) and makes a friendly response to "Would you like to help? Please tell me your question?" When the customer asks a question, the character behaves vigorously or returns, "Thank you, I'll be back soon with the answer." After a short period of time (preferably no more than 5-7 seconds), the character speaks the answer to the user's question. As shown in FIG. 18, the answer is answer 199, which is returned to the user in the form of a conversation, which is a paired answer to question 195. For example, Answer 199: "I accept Visa, MasterCard, and Discover credit cards" is a response to Inquiry 200 "What form of payment do you accept?" [0196] Another embodiment of the present invention is shown in FIG. This web page shows a typical website that employs NLQS in a web-based learning environment. As shown in FIG. 12, the web page in the browser 120 is divided into two or more frames. Character 121, pretending to be an instructor, is available on the screen and speaks the word "help" to the microphone (Figures 13, 134) or "Click to Speak". Appears when a student enters query mode by either clicking a link (Figures 12, 128). Character 121 then prompts the student to select course 122 from the drop-down list 123. If the user selects the course "C Plus Plus", the character verbally confirms that the course "C Plus Plus" has been selected. The character then urges the student to make the next choice from the drop-down list 125, which contains options for chapter 124 from which the question is available. After the student makes a selection, character 121 confirms the selection again in conversation. The character 121 then prompts the student to select a chapter "section" 126 from which questions are available from the drop-down list 127. After the student makes the selection again, the character 121 confirms the selection by clearly pronouncing the selected "section" 126. As a reminder to students, a list of possible questions appears in list box 130. In addition, information 129 for using the system is displayed. Once all the selections have been made, the student will be prompted by the character to ask the following questions: "Ask your question here." The student then speaks his question, and after a short period of time, the character responds with a pre-question answer: "The answer to your question ... is: . This approach allows students to quickly find answers to questions anywhere in the course, replacing the boredom, citations or indexes of instruction books. That is, it is useful for many purposes because it is a virtual teacher that answers ongoing questions, a flashcard alternative. [0197] Estimating from the interim data available to the inventor, the system can easily accommodate 100-250 question / answer pairs, while at the same time using the structures and methods described above to give the user a real-time feel and appearance (ie, latency). Less than 10 seconds, transmissions are not counted) can be provided. Of course, it is expected that these numbers will improve as more processing speeds become available and routine optimizations are made to the various components mentioned for each particular environment. [0198] Again, the above is merely an example of the many possible uses of the invention, and like other consumer uses (such as intelligent interactive toys), more web-based companies will provide this teaching. Expected to be used. Although the present invention has been described so far based on preferred embodiments, it will be apparent to those skilled in the art that many modifications and modifications can be made to the embodiments without departing from the teachings of the present invention. It will also be apparent to those skilled in the art that many aspects of the present disclosure have been simplified in order to reasonably emphasize and focus on the more appropriate aspects of the invention. The microcode and software routines performed to accomplish the methods of the invention can be implemented in various forms in permanent magnetic media, non-volatile ROMs, CD-ROMs, or other suitable machine-readable formats. .. Accordingly, all such changes / amendments shall be included within the scope and spirit of the invention, as provided by the accompanying claims. [Simple explanation of drawings] [0199] FIG. 1 is a block diagram of a preferred embodiment of the Natural Language Query System (NLQS) of the present invention, which is distributed across client / server computer architectures, interactive learning systems, electronic commerce systems, and more. It can be used as an electronic support system. FIG. 2 is a block diagram of a preferred embodiment of a client-side system, including a voice acquisition module, a partial voice processing module, a coding module, a transmission module, an agent control module, and an answer / voice feedback module. , This can be used for the above NLQS. FIG. 2-2 is a block diagram of a preferred embodiment of a set of initialization routines and procedures used in the client-side system of FIG. FIG. 3 is a block diagram of a preferred embodiment of a set of routines and procedures used to handle a set of repetitive speeches in the client-side system of FIG. Send voice data of remarks and receive a favorable response from such a server. FIG. 4 is a block diagram of a preferred embodiment of a set of initialization routines and procedures used to deinitialize the client-side system of FIG. FIG. 4A is a block diagram of a preferred embodiment of a set of routines and procedures used to execute the distributed components of the speech recognition module for the server-side system of FIG. FIG. 4B is a block diagram of a preferred set of routines and procedures used to execute the SQL Query Builder for the server-side system of FIG. FIG. 4C is a block diagram of a preferred embodiment of a set of routines and procedures used to execute the database control processing module for the server-side system of FIG. Figure 4D shows the preferred implementation of a set of routines and procedures used to run a natural language engine that provides interfaces to query building support, query response modules, and database control processing modules for the server-side system of Figure 5. It is a block diagram of a form. FIG. 5 is a block diagram of a preferred embodiment of a server-side system, a speech recognition module, an environment and grammar control module, a query construction module, a natural language engine, and a database control module for completing speech processing. , And the query response module that can be used with the above NLQS. FIG. 7 shows the configuration of a full-text database used as part of the server-side system shown in FIG. FIG. 7A shows the configuration of a full-text database course table used as part of the server-side system shown in FIG. 5 for the interactive learning embodiment of the present invention. FIG. 7B shows the configuration of the full-text database chapter table used as part of the server-side system shown in FIG. 5 for the interactive learning embodiment of the present invention. FIG. 7C describes the fields used in the chapter table used as part of the server-side system shown in FIG. 5 for the interactive learning embodiment of the present invention. FIG. 7D describes the fields used in the section table used as part of the server-side system shown in FIG. 5 for the interactive learning embodiment of the present invention. FIG. 8 is a flow diagram of a first-stage operation performed on a lexical according to a preferred embodiment of a natural language engine, including tokenization, tagging, and grouping. FIG. 8A is a flow diagram of operations performed on remarks according to a preferred embodiment of a natural language engine, including stemming and lexical analysis. FIG. 10 is a block diagram of a preferred embodiment of the SQL database search and support system of the present invention. 11A-C are flow diagrams showing steps performed in a preferred two-step process performed for question recognition by the NLQS of FIG. FIG. 12 is an explanatory diagram of a further embodiment of the present invention implemented as part of a web-based conversation-based learning / training system. FIG. 13 is an explanatory diagram of a further embodiment of the present invention implemented as part of a web-based electronic commerce system. FIG. 14 is an explanatory diagram of a further embodiment of the present invention implemented as part of a web-based electronic commerce system. FIG. 15 is an explanatory diagram of a further embodiment of the present invention implemented as part of a web-based electronic commerce system. FIG. 16 is an explanatory diagram of a further embodiment of the present invention implemented as part of a web-based electronic commerce system. FIG. 17 is an explanatory diagram of a further embodiment of the present invention implemented as part of a web-based electronic commerce system. FIG. 18 is an explanatory diagram of a further embodiment of the present invention implemented as part of a voice-based help page for an e-commerce website.
32 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US2022392447A1 | Cited by | United States of America | Search report |
| US2022301561A1 | Cited by | United States of America | Search report |
| US12205585B2 | Cited by | United States of America | Search report |
| JP2002540479A | Cites | Japan | Examiner |
| JPH0573094A | Cites | Japan | Examiner |
| JPH10214258A | Cites | Japan | Examiner |
| JPH1115574A | Cites | Japan | Examiner |
| JPH11249867A | Cites | Japan | Examiner |
| JP9507105A | Cites | Japan | – |
| JP10214258A | Cites | Japan | – |
| JP11249867A | Cites | Japan | – |
| JP2002540479A | Cites | Japan | – |
| JP573094A | Cites | Japan | – |
| JP1115574A | Cites | Japan | – |
71 members in 6 offices
Priority claims9
| Document | Office | Kind | Date |
|---|---|---|---|
| 09439145 | United States of America | – | |
| 43914599 | United States of America | A | |
| 43914599 | United States of America | A | |
| 0030918 | United States of America | W | |
| 0030918 | United States of America | W | |
| 1999439145 | – | – | – |
| 2000030918 | – | – | – |
| US19990439145 | – | – | – |
| WO2000US30918 | – | – | – |
Members71
| Document | Office | Kind | |
|---|---|---|---|
| WO0135391A1 | World Intellectual Property Organization (WIPO) | A1 | |
| EP1245023A1 | European Patent Office (EPO) | A1 | |
| JP2003517158A | Japan | A | |
| US6615172B1 | United States of America | B1 | |
| US6633846B1 | United States of America | B1 | |
| US6665640B1 | United States of America | B1 | |
| US2004030556A1 | United States of America | A1 | |
| US2004117189A1 | United States of America | A1 | |
| US2004236580A1 | United States of America | A1 | |
| US2004249635A1 | United States of America | A1 | |
| US2005080614A1 | United States of America | A1 | |
| US2005080625A1 | United States of America | A1 | |
| US2005086046A1 | United States of America | A1 | |
| US2005086049A1 | United States of America | A1 | |
| US2005086059A1 | United States of America | A1 | |
| US2005119896A1 | United States of America | A1 | |
| US2005119897A1 | United States of America | A1 | |
| US2005144001A1 | United States of America | A1 | |
| US2005144004A1 | United States of America | A1 | |
| EP1245023A4 | European Patent Office (EPO) | A4 | |
| US7050977B1 | United States of America | B1 | |
| US2006200353A1 | United States of America | A1 | |
| US2006235696A1 | United States of America | A1 | |
| US7139714B2 | United States of America | B2 | |
| US7203646B2 | United States of America | B2 | |
| US2007094032A1 | United States of America | A1 | |
| US7225125B2 | United States of America | B2 | |
| US2007179789A1 | United States of America | A1 | |
| US2007185716A1 | United States of America | A1 | |
| US2007185717A1 | United States of America | A1 | |
| US7277854B2 | United States of America | B2 | |
| US2008021708A1 | United States of America | A1 | |
| US2008052063A1 | United States of America | A1 | |
| US2008052077A1 | United States of America | A1 | |
| US2008052078A1 | United States of America | A1 | |
| US2008059153A1 | United States of America | A1 | |
| US7376556B2 | United States of America | B2 | |
| US7392185B2 | United States of America | B2 | |
| US2008215327A1 | United States of America | A1 | |
| US2008255845A1 | United States of America | A1 | |
| US2008300878A1 | United States of America | A1 | |
| US2009157401A1 | United States of America | A1 | |
| US7555431B2 | United States of America | B2 | |
| US7624007B2 | United States of America | B2 | |
| US2010005081A1 | United States of America | A1 | |
| US7647225B2 | United States of America | B2 | |
| US7657424B2 | United States of America | B2 | |
| US7672841B2 | United States of America | B2 | |
| US7698131B2 | United States of America | B2 | |
| US7702508B2 | United States of America | B2 | |
| US7725307B2 | United States of America | B2 | |
| US7725320B2 | United States of America | B2 | |
| US7725321B2 | United States of America | B2 | |
| US7729904B2 | United States of America | B2 | |
| US2010228540A1 | United States of America | A1 | |
| US2010235341A1 | United States of America | A1 | |
| US7831426B2 | United States of America | B2 | |
| US7873519B2 | United States of America | B2 | |
| EP1245023B1 | European Patent Office (EPO) | B1 | |
| AT500587T | Austria | T | |
| ATE500587T1 | Austria | T1 | |
| US7912702B2 | United States of America | B2 | |
| DE60045690D1 | Germany | D1 | |
| US8229734B2 | United States of America | B2 | |
| JP4987203B2This record | Japan | B2 | |
| US2012265531A1 | United States of America | A1 | |
| US8352277B2 | United States of America | B2 | |
| US8762152B2 | United States of America | B2 | |
| US2014316785A1 | United States of America | A1 | |
| US9076448B2 | United States of America | B2 | |
| US9190063B2 | United States of America | B2 |
20 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Transfer to examiner for re-examination before appeal (zenchi)AppealJAPANESE INTERMEDIATE CODE: A911A911 | A911 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Decision of refusalJAPANESE INTERMEDIATE CODE: A02A02 | A02 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written permission of extension of timeJAPANESE INTERMEDIATE CODE: A602A602 | A602 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Written permission of extension of timeJAPANESE INTERMEDIATE CODE: A602A602 | A602 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 4987203
- Publication, DOCDB
- 4987203
- Publication, EPODOC
- JP4987203B
- Application
- 2001537046
- Application, DOCDB
- 2001537046
- Application, EPODOC
- JP20010537046
Titles2
- Japanese
- 分散型リアルタイム音声認識装置
- English
- Distributed real-time speech recognition device
Classification
- CPC, 7
- G10L15/30
- G10L15/142
- G10L15/1815
- G10L15/183
- G10L15/285
- H04M2250/74
- G06F16/24522
- IPC, 13
- G06F3 16
- G10L15 28
- G06F17 28
- G06F17 30
- G09B7 02
- G10L13 00
- G10L15 00
- G10L15 02
- G10L15 14
- G10L15 18
- G10L15 20
- G10L15 22
- G10L21 02
