Speech recognition with parallel recognition tasks
9 claims: 1 independent, 8 dependent
- 1コンピュータで実施される方法であって、 コンピュータシステムにおいて、音声信号を受け取るステップと、 前記コンピュータシステムにより、前記音声信号に対する複数の音声認識タスクを開始するステップとを備え、前記音声認識タスクは、複数の言語モデルのうち異なる1つをそれぞれ使用し、 前記方法は、 前記複数の音声認識タスクの完了した一部分を検出するステップを備え、前記複数の音声認識タスクの残りの部分は、依然として完了しておらず、 前記方法は、 前記複数の音声認識タスクの前記一部分に関する認識結果および信頼値を取得するステップを備え、前記認識結果は、前記音声信号の1つまたは複数の候補表現を特定するとともに、前記信頼値は、前記認識結果が正しいことの1つまたは複数の可能性を特定し、 前記方法は、 前記コンピュータシステムにより、前記1つまたは複数の信頼値のうち少なくとも1つが信頼閾値に対してより大きいまたは等しいかどうかを判定するステップと、 前記1つまたは複数の信頼値のうち少なくとも1つが前記信頼閾値に対してより大きいまたは等しいとの判定に応答して、完了した前記複数の音声認識タスクに対する前記残りの部分が完了する前に、前記認識結果と前記1つまたは複数の信頼値とに基づいて、前記音声信号に対する最終的な認識結果を提供するステップとを備え 、 前記言語モデルが、複数の言語のうち異なる1つにそれぞれ関連づけられる 、コンピュータで実施される方法。
- 2前記言語モデルが、複数のレベルの細粒度のうち異なる1つをそれぞれ有する、請求項1に記載のコンピュータで実施される方法。
- 3前記言語モデルが、複数の地理的位置のうち異なる1つにそれぞれ関連付けられる、請求項1に記載のコンピュータで実施される方法。
- 4前記言語モデルが、複数のアーキテクチャのうち異なる1つをそれぞれ有する、請求項1に記載のコンピュータで実施される方法。
- 5前記言語モデルが、複数のトレーニング手順のうち異なる1つに基づいてそれぞれ生成された、請求項1に記載のコンピュータで実施される方法。
- 6前記最終的な認識結果が、前記音声認識タスクの一部分から特定の音声認識タスクによって生成された前記認識結果から特定の認識結果を含み、前記特定の音声認識タスクが前記複数の言語モデルから特定の言語モデルを使用するとともに、 前記特定の音声認識タスクまたは前記特定の言語モデルを識別する情報が、前記最終的な認識結果とともに提供される、請求項1に記載のコンピュータで実施される方法。
- 7前記複数の音声認識タスクが、複数の音声認識システムによって開始されるとともに、前記複数の音声認識システム上で実行される、請求項1に記載のコンピュータで実施される方法。
- 8前記1つまたは複数の信頼値のうち前記少なくとも1つが前記信頼閾値に対してより大きいまたは等しいという判定に応答して、前記複数の音声認識タスクの前記残りの部分が完了する前に、完了した前記複数の音声認識タスクに対する前記残りの部分を取り消すステップをさらに含む、請求項1に記載のコンピュータで実施される方法。
- 9前記1つまたは複数の信頼値のうち前記少なくとも1つが、前記信頼閾値に対してより大きいまたは等しいという判定に応答して、完了した前記複数の音声認識タスクに対する前記残りの部分が完了する前に、前記複数の音声認識タスクの前記残りの部分を休止するステップをさらに含む、請求項1に記載のコンピュータで実施される方法。
Independent claims9
136 paragraphs, as filed
0001The present specification relates to speech recognition.
0002Many applications benefit from receiving input in the form of voice commands or voice queries. This is especially true on mobile devices such as mobile phones, where entering input through a small keypad or other device operated by the user's finger can be difficult due to the reduced size of the device. This is true for applications running on. Similarly, when a mobile device is used to access an application running on another device (eg, an email server, map / direction server, or phonebook server), via a small keypad, etc. Sending voice commands to an application instead of typing commands can be much easier for the user.
0003In order for the application to operate according to the voice input, the voice recognizer can convert the voice input into the symbolic representation used in the application. Some current speech recognizers may use a single recognition system that attempts to identify the expected speech in the speech input. The use of a single speech recognition system may limit the accuracy of speech identification to the accuracy associated with a single speech recognition system. Other current speech recognizers may use continuous speech recognition, in which continuous speech recognition involves two or more passes to the speech input, and which speech may be represented by the speech input. The highest is determined. The use of multiple passes may increase the time required to produce the final speech recognition result.
0004In yet another current speech recognizer, each of the speech recognition systems can fully process the speech input and then output the result. Even with this use of multiple speech recognition systems, the length of processing time is determined by the slowest speech recognition system (and / or the slowest computer running the speech recognition algorithm) to produce the final result. The time required may increase.
<p num="0005"> In general, this document uses multiple speech recognition systems (SRS) in parallel to recognize speech, but explains that if the generated recognition results meet the desired threshold, some will be stopped before completion. To do. For example, each SRS may have different latency and accuracy in performing speech recognition tasks. When an SRS with a short waiting time outputs a speech recognition result and a reliability value representing a high reliability of the result, the speech recognition task executed by the remaining SRS can be stopped. If the confidence value is too low relative to the confidence threshold, you can allow another SRS to produce the result. If these results meet the confidence threshold, the SRS that has not yet completed the speech recognition task can be stopped, and so on.</p><p num="0006"> The first general aspect describes a method performed on a computer. This method involves receiving a speech signal and initiating a speech recognition task with multiple speech recognition systems (SRS). Each SRS is configured to generate a recognition result that specifies the voice that is expected to be included in the voice signal and a confidence value that indicates the reliability of the accuracy of the voice result. The method also includes one or more recognition results and one or a plurality for one or more recognition results to generate a confidence value for the number, whether one or more confidence values meets a confidence threshold Determining, stopping the rest of the speech recognition task for SRS that has not generated recognition results, and making the final recognition result based on at least one of the generated speech results. It also includes completing some of the speech recognition tasks, including outputting.</p><p num="0007"> A second general aspect describes a system comprising a plurality of speech recognition systems that initiate a speech recognition task to identify expected speech encoded in a received speech signal. Each speech recognition system (SRS) is configured to generate a recognition result and a confidence value indicating the reliability of the accuracy of the recognition result. The system also includes a recognition management module that receives the recognition result when the recognition result is generated by SRS and receives the trust value associated with the generated recognition result. If one or more of the received confidence values meet the confidence threshold, the recognition management module stops the incomplete speech recognition task by the SRS that is not producing the recognition result. The system includes an interface that sends the final recognition result selected based on the confidence value of the generated recognition result.</p><p num="0008"> A third general aspect describes a system that includes multiple speech recognition systems that initiate a speech recognition task for the received speech signal, with each speech recognition system (SRS) identifying the expected speech in the speech signal. It is configured to generate a recognition result and a confidence value indicating the reliability of the accuracy of the recognition result. The system receives one or more recognition results and one or more corresponding trust values from each SRS when one or more recognition results are generated by SRS, and of the received trust values. A means of stopping an incomplete speech recognition task by SRS that has not generated a recognition result and selecting the final recognition result based on the confidence value of the generated recognition result if one or more meet the confidence threshold. including. The system also includes an interface that sends the final recognition result representing the expected speech in the speech signal.</p><p num="0009"> The systems and techniques described herein can realize one or more of the following advantages: First, a system that uses multiple speech recognition systems to decode speech in parallel can stop unfinished recognition tasks if it receives a satisfactory result, so wait time and Improvements in joint optimization can be achieved. In addition, systems that use multiple recognition systems can improve the rejection rate (ie, reduce the rejection rate). The system can also improve accuracy by comparing the recognition results output by multiple recognition systems. A framework for scaling (eg, improving) the amount of computational resources used to achieve improved recognition performance can also be provided.</p><p num="0010"> Details of one or more embodiments will be described in the accompanying drawings and in the description below. Other features and advantages will become apparent from this description and drawings as well as the claims.</p>
0011<figref num="1">It is the figure of the exemplary system which recognizes a voice.</figref><figref num="2">FIG. 6 is a more detailed diagram of an exemplary system for decoding voice embedded in voice transmission.</figref><figref num="3">It is a flowchart of an exemplary method of recognizing speech using parallel decoding.</figref><figref num="4A">It is a figure which shows the execution of an exemplary speech recognition task.</figref><figref num="4B">It is a figure which shows the execution of an exemplary speech recognition task.</figref><figref num="4C">It is a figure which shows the execution of an exemplary speech recognition task.</figref><figref num="5A">It is a figure of the method of selecting the exemplary recognition result and confidence value generated by SRS, and the final recognition result.</figref><figref num="5B">It is a figure of the method of selecting the exemplary recognition result and confidence value generated by SRS, and the final recognition result.</figref><figref num="5C">It is a figure of the method of selecting the exemplary recognition result and confidence value generated by SRS, and the final recognition result.</figref><figref num="6">FIG. 6 is an exemplary graph of the distribution of confidence values used to weight the values used in the final recognition result selection.</figref><figref num="7A">It is a Venn diagram showing an exemplary recognition result set output by SRS and the correlation between the sets, which can be used to weight the recognition results.</figref><figref num="7B">It is a Venn diagram showing an exemplary recognition result set output by SRS and the correlation between the sets, which can be used to weight the recognition results.</figref><figref num="7C">It is a Venn diagram showing an exemplary recognition result set output by SRS and the correlation between the sets, which can be used to weight the recognition results.</figref><figref num="7D">It is a Venn diagram showing an exemplary recognition result set output by SRS and the correlation between the sets, which can be used to weight the recognition results.</figref><figref num="7E">It is a Venn diagram showing an exemplary recognition result set output by SRS and the correlation between the sets, which can be used to weight the recognition results.</figref><figref num="8A">It is a Venn diagram showing how the intersection between SRS can be adapted or changed during the run-time operation of a speech decoding system.</figref><figref num="8B">It is a Venn diagram showing how the intersection between SRS can be adapted or changed during the run-time operation of a speech decoding system.</figref><figref num="9">It is a graph which shows the exemplary correlation between the SRS error rate and the weight related to the recognition result.</figref><figref num="10">FIG. 6 is a block diagram of a computing device that can be used to implement the systems and methods described in this document.</figref>
0012Similar reference symbols in various drawings indicate similar elements.
0013This document describes systems and techniques for decoding utterances using multiple speech recognition systems (SRS). In some implementations, each SRS has different characteristics such as accuracy, latency, dictionary, etc., so that some of the multiple SRSs output the recognition result before the other SRSs. If the output recognition result meets certain requirements (for example, one or more of the generated results is associated with a specified confidence value that meets or exceeds the threshold confidence), the speech decoding system , The remaining SRS can be stopped before the remaining SRS complete the speech recognition task.
0014FIG. 1 is a diagram of an exemplary system 100 that recognizes speech. In general, the system 100 includes a plurality of SRSs that process, for example, an audio signal received from a mobile phone. In this example, the user calls a voice-enabled phonebook service, and the voice-enabled phonebook service transfers a voice signal including the user's voice to a voice recognizer having a plurality of SRSs.
0015Multiple SRSs can process audio signals in parallel, but some SRSs can generate recognition results before other SRSs. If the SRS that produces the recognition results represents a sufficiently high degree of confidence in these results, then the remaining unfinished speech recognition tasks can be stopped and all of the SRSs wait for the speech recognition task to complete. Instead, the final recognition result can be determined based on the currently generated recognition result.
0016An exemplary system 100 includes a mobile phone 102 that sends a voice input in the form of a voice signal 104 to a voice-enabled phonebook information server 106, in which the voice-enabled phonebook information server 106 verbally requests phonebook information from a mobile phone user. Allows you to respond with the requested information.
0017In the example of FIG. 1, the information server 106 sends the voice signal 104 to the voice recognizer 108 in order to decode the voice embedded in the voice signal 104. In one application, speech recognizer 108 operates in parallel to decode the speech in the speech signal 104 with multiple SRSs.<sub>AE</sub>including.
0018The Speech Recognition System (SRS) Management Module 110 monitors whether any of the SRSs have produced recognition results and collects the trust values associated with those results. Monitoring is shown in Diagram 112, which shows the parallel execution of SRS. Diagram 112 shows the SRS<sub>A</sub>First produces a recognition result with a confidence value of 0.3. Next, SRS<sub>E</sub>However, a recognition result is generated with a confidence value of 0.6. Shortly after that, SRS<sub>B</sub>However, a recognition result is generated with a confidence value of 0.8. The recognition result of the SRS management module 110 is SRS.<sub>B</sub>After being generated in SRS<sub>C, D</sub>You can stop the remaining speech recognition tasks performed in. In this implementation, SRS<sub>B</sub>This is because the recognition result generated in (1) has a reliability value satisfying a predetermined reliability threshold value.
0019In one implementation, the final result selection module 113 in the SRS management module 110 can select the final recognition result 114 and output it to the voice-enabled phonebook information server 106. For example, the SRS management module 110 has completed a speech recognition task.<sub>A, B, E</sub>The final recognition result can be selected based on the generated set of recognition results and the associated confidence value 116 output by. In one implementation, the final recognition result 114 is a symbolic representation of the estimated speech decoded from the speech signal 104. For example, the phonebook information server 106 may have prompted the user to say the city and state names. The spoken city name and state name can be encoded as a voice signal 104, and the voice signal 104 is received from the user of the mobile phone 102 and decoded by the voice recognizer 108. In one implementation, the end result is the voice that the decoder has determined most likely to be represented by the voice signal 104.
0020The confidence value 116 output by SRS can be generated according to several methods. In one implementation, the first SRS can generate some hypotheses, or guesses, regarding the identification of utterances. The generated best hypothesis can be compared to the competing hypothesis generated by SRS, and the confidence value varies based on the difference in recognition score between the generated best hypothesis and the competing hypothesis. can do.
0021In yet another example, the confidence value for the first recognition result can be based on the signal or characteristic used in the generation of the recognition result or in the calculation of the front-end signal analysis. For example, the signal can contain some competing hypotheses used in the search, or the density of the search graph to be investigated, and front-end examples are the estimated signal for noise characteristics, or the estimated channel type (eg, for example). These channel types can be included based on a match for existing models (hands-free and cellular and land). These signal combinations can be conditionally optimized based on the data provided.
0022Confidence can also be estimated as a non-linear coupling of scores from acoustic and linguistic models. For example, given the best hypothesis, the system will have separate language model scores (eg, pre-estimation of recognition results before speech is processed), and acoustic model scores (eg, this utterance is the best result). How well it matches the acoustic unit associated with) can be extracted. The system can then estimate the total confidence result as a non-linear combination of these two scores conditionally optimized over the data provided.
0023In addition to the score, another signal that can be used to optimize confidence is based on the type of path transmitted through the language model. For example, in the n-gram language model, when the system does not encounter a particular 3-word sequence, the system can "back off", that is, the prior to the 2-word to 3-word sequence encountered by the system. Can be estimated. Counting the number of times a search must go through a backoff estimate for a given hypothesis gives another signal that can be used to conditionally estimate confidence in a given utterance.
0024In yet another implementation, the confidence value may be a posterior probability that the recognition result is correct. In some implementations, posterior probabilities can be calculated based on grid density calculations. In another example, posterior probabilities use less specific acoustic models such as monophone loops, all-speech gmm trained with fewer parameters than the main acoustic model, for all audio. It can be found by comparing the more general model with the best hypothesis. Both of these post-estimation methods for reliability are well known in the art, as are combinations of such estimates.
0025In some implementations, confidence values are calculated using multiple techniques. For example, confidence values are based on posterior probabilities as well as the similarity of results compared to other recognition results generated during speech recognition tasks.
0026The selection module 113 can send the final recognition result 114 to the interface 118, and the interface 118 can send the final recognition result 114 to the voice-enabled phonebook information server 106. In one implementation, interface 118 uses a set of APIs that interface with software running on information server 106. For example, the information server 106 can run software that has a common way to enter textual representations of cities, states, and trade names. In another implementation, interface 118 can include a networking protocol (eg TCP / IP) that sends information to information server 106 over the network.
0027Figure 1 shows the phone book information server 106 and voice recognizer on separate computing devices, but this is not necessary. In one implementation, both systems can be implemented on a single computing device. Similarly, each system can be implemented using several computing devices. For example, each SRS can be implemented using one or more computers as shown in Figure 2.
0028FIG. 2 is a diagram of an exemplary system 200 that decodes voice embedded in voice transmission. For illustration purposes, the system 200 is divided into two segments, a voice transmission segment 202 and a voice recognizer segment 204. The voice transmission segment 202 illustrates an exemplary architecture for transmitting a voice signal from a user to an application server. Speech recognizer segment 204 illustrates an exemplary architecture for interpreting or decoding speech represented by voice transmission. In this implementation, decryption is performed on behalf of the application server and the decrypted voice is sent back to the application server for use in processing the user's request.
0029In one implementation, system 200, voice transmission segment 202, includes a telephone device such as a mobile phone 206, which transmits a voice signal 208 to the telephone server 210 over a network (eg, POTS, cellular, internet, etc.). To do. The telephone server can send the voice signal to another computing device, such as the software application server 212, or directly to the voice recognition system described below.
0030The software application server 212 may include software applications with which the user is verbally interacting. For example, the software application server may be a calendar application. The user can call the calendar application and request that the calendar application create an event from 1:30 pm to 2:00 pm on May 16, 2012. The software application server 212 can transfer the received voice input requesting event creation to the voice recognizer segment 204 for decryption.
0031In one implementation, the speech recognizer segment 204 is a speech recognition system, SRS.<sub>AD</sub>, And a voice recognition system (SRS) management module, which can adjust the SRS for use in determining which utterance is most likely represented by the voice signal 208. it can.
0032Each SRS can differ in one or more schemes. In some implementations, the SRS may vary depending on the underlying acoustic model. For example, different acoustic models can target specific conditions, such as the user's gender, accent, age range, or specific background and foreground noise conditions, or specific transmission channels. .. Acoustic models can vary depending on their architecture and size, for example, smaller models with fewer parameters can produce faster recognition, and larger models with more parameters can produce more accurate results. Can be done. In other examples, the acoustic model may vary depending on its training procedure (eg, different probabilistic training sets can be used to train the model, or different training algorithms can be used).
0033In some implementations, SRS may vary depending on its language model. For example, the model can target different types of data, such as language models specific to different regions, different granularity, or different geographic locations. In another example, the model may vary depending on its architecture, size, training procedure, and so on.
0034In yet another implementation, the SRS can vary by other components such as end pointers, front ends, dictionaries, reliability estimation algorithms, and search configurations.
0035For illustration purposes, SRS<sub>D</sub>The language model 252, the acoustic model 254, and the speech recognition algorithm 256 for are shown in Figure 2.
0036In one implementation, when the SRS management module 250 receives an audio signal 208, the SRS management module 250 starts a process, which recognizes incoming utterances in parallel using two or more of the SRSs. .. For example, four speech recognition tasks attempt to recognize the same utterance represented by the speech signal 208, and four SRSs (SRS).<sub>AD</sub>) To be executed in parallel.
0037In some implementations, each SRS may have a certain latency. Latency may depend on the SRS architecture (eg, acoustic model, language model, or other component), but may vary based on the specific embodiment of the speech recognition task. For example, if the SRS has information that indicates that the utterance is contained within a group of words (eg, yes, no, nope, yeah, positive, negative, no way, yipper, etc.), then for a particular model. The wait time can be much shorter than when the SRS does not have information indicating the restricted context in which the utterance was made (eg, the utterance was not in the context of a general interrogative sentence).
0038In one implementation, each SRS has a recognition result (for example, what the SRS has determined what the incoming utterance said) and how confident the SRS is in the accuracy of the result at the completion of its speech recognition task. Outputs the scale of.
0039In one implementation, the SRS management module 250 has a recognition result monitor 258 that tracks the recognition result 262 generated by SRS. The result monitor 258 can also monitor the confidence value 264, or score associated with the recognition result 262.
0040In one implementation, the result monitor 258 can calculate a combined confidence score for each recognition result. For example, SRS<sub>A</sub>And SRS<sub>B</sub>May generate a recognition result "May 16" for incoming utterances. SRS<sub>A</sub>May associate the result with a confidence value of 0.8, SRS<sub>B</sub>May associate the result with a confidence value of 0.7. The result monitor 258 can calculate that the current running average for the result "May 16" is 0.75 (ie, (0.8 + 0.7) / 2). The combination trust value can be updated each time the recognition result (and the corresponding trust value) is generated by SRS.
0041The SRS management module 250 can also include a latency monitor 260 that tracks the latency for each SRS (eg, the actual time or estimated time to complete the speech recognition task). For example, the latency monitor 260 can track how long a particular speech recognition task takes for the SRS to generate a recognition result.
0042The latency monitor 260 can also monitor whether the SRS has completed a speech recognition task (eg, by monitoring whether the SRS has output a recognition result). In addition, the latency monitor 260 can estimate the estimated latency for the SRS to complete the speech recognition task. For example, the latency monitor 260 allows the SRS to decode utterances spoken in a similar context, such as how long it took for the SRS to complete a similar speech recognition task (for example, an answer to a particular question prompted). You can access the empirical information gathered about SRS that shows how long it took before to become.
0043The latency monitor 260 can also access information about the characteristics of the underlying model for the SRS to determine the estimated time to complete the speech recognition task (eg, the latency monitor 260 recognizes). Due to the large dictionary of words that must be searched to identify the results, it is possible to determine if SRS can take longer to complete speech recognition).
0044The SRS management module communicates with the latency monitor 260 and / or the recognition result monitor 258 to determine whether to send a stop command 266 for an SRS that has not yet completed decoding of the received audio signal 208. aborter) 270 can be included. For example, the SRS Avoter 270 can determine if the tracked confidence value and / or latency meets an operating point or motion curve. If so, all remaining speech recognition can be stopped.
0045In a simple example, the SRS management module 250 can determine that a confidence threshold of 0.75 for recognition results should be reached before stopping any incomplete speech recognition task. In some implementations, the confidence threshold can vary for different speech recognition tasks. For example, contextual information related to a particular speech recognition task may be limited to utterances with relatively few recognition results (eg, the recognition task is the context of an answer to a general interrogative presented to the user). When indicated, the SRS Avoter 270 can determine that the confidence value should be relatively high (eg 0.8, 0.9, 0.95).
0046When the context indicates that the recognition result may contain one of many expected utterances (for example, the user asks a free-answer question such as "what do you want to do today?" The SRS Avoter 270 may have a relatively low confidence threshold for recognition results (eg 0.49, 0.55, 0.61) and can still be determined to be acceptable to the SRS Management Module 250.
0047In some implementations, the Avoter 270 can send a stop command 266 to any unfinished SRS if the threshold confidence point (eg 0.75) is met by one of the recognition results. For example, SRS<sub>A, B</sub>If the combinatorial confidence value of is greater than or equal to 0.75, the Avoter 270 can send a stop command 266 to an SRS that has not yet generated a recognition result.
0048In another example, if one of the recognition results has a confidence value of 0.75 or higher, the Avoter 270 can send a stop command 266. In this case, the confidence value may not be a combination confidence value, but instead may be a single confidence value associated with the recognition result generated by a single SRS.
0049In another implementation, the SRS Avoter 270 can send a stop command based on the actual or estimated latency for the SRS. For example, SRS<sub>A</sub>And SRS<sub>B</sub>Can generate recognition results and the SRS Avoter 270 can stop the remaining unfinished speech recognition tasks if the recognition results are associated with a very low confidence value. In this case, the rest of the speech recognition tasks are performed under the assumption that SRSs that have not completed the recognition task will not produce such recognition results because no other SRS has produced a recognition result with a high confidence value. It can be canceled.
0050Instead of waiting for the rest of the SRS to finish, in one implementation the SRS Avoter 270 can send a stop command 266 to start the process where the user is requested to repeat the utterance 274. The SRS can then attempt to decrypt the new utterance.
0051In another implementation, if the recognition result is not satisfactory based on confidence values or other parameters, the SRS management module 250 can initiate a process in which a request is made to switch to a human operator. For example, a human operator can intercept the audio signal in response to the user, as indicated by the arrow 276 indicating that the audio signal is sent to the call center 278. The human operator can handle the requests or information transmitted by the user.
0052In one implementation, the SRS Avoter 270 can first query the latency monitor 260 to decide whether to send a stop command 266 to an unfinished SRS. For example, if the SRS Avoter 270 queries the latency monitor 260 and determines that one of the unfinished SRS is likely to complete in a relatively short amount of time, the SRS Avoter 270 will postpone it. Results can be obtained from an almost complete SRS. After the results are generated, the SRS Avoter 270 can send a stop command 266 to stop the remaining unfinished SRS from completing speech recognition.
0053In some implementations, additional recognition results and related information may be worth enough to delay sending the stop command until the nearly finished SRS is complete. For example, if the characteristics associated with a nearly finished SRS indicate that the recognition result is often more accurate than the result of a previously completed SRS, the Avoter 270 performs the remaining unfinished speech recognition tasks. Before stopping, you can wait until the almost finished SRS produces a recognition result.
0054In some implementations, a function with one or more variables is used to determine the confidence threshold. For example, a confidence function can have variables that include a confidence value and a latency. If the generated confidence value and the observed latency satisfy the confidence function, the Avoter 270 can cancel any unfinished speech recognition task. For example, within a short latency, the confidence function may indicate that the confidence value should be very high in order for the confidence function to be satisfied. This is partly based on the assumption that if the Avoter 270 issues a stop command quickly, no other potential recognition results will be generated, so the results produced should be very likely to be correct. it can. Speech recognition tasks that take a long time to process can be more difficult, and therefore the required confidence value decreases as latency increases, assuming that the results are likely to be less reliable. There is.
0055The SRS management module 250 can also include a final result selection module 280, and in some implementations the final result selection module 280 selects the final recognition result from the generated recognition results. For example, the selection module 280 can select the recognition result with the highest combination confidence value or the recognition result with the highest single confidence value.
0056In some implementations, the final recognition result selection can be affected based on which SRS produced the recognition result. For example, if the recognition result is generated by an SRS that has relatively different components (eg, a language model, an acoustic model, a speech recognition algorithm, etc.) and / or an SRS that normally produces different recognition results, the selection module 280 The selection of recognition results can be weighted or supported.
0057In one implementation, the SRS Correlation Monitor 282 can track the correlation between output recognition results for SRS. The output recognition results show that the two SRSs are not highly correlated, but if the two SRSs produce the same recognition result for a particular recognition task, the final recognition result selection will make the result stronger. Can be weighted or supported. Alternatively, if the SRS that produces the recognition result is highly correlated, the recognition result may not be discounted or weighted so that the final recognition result selection does not necessarily support the result.
0058The correlation monitor 282 can dynamically update the correlation value that specifies the correlation between the two SRSs based on the recognition result generated by the SRS. For example, two SRSs associated with a low correlation value can initiate the generation of similar recognition results. The correlation monitor 282 can update (eg, increase) the correlation value to reflect the increased duplication of recognition results between SRSs.
0059After the final result selection module 280 identifies the final result, the SRS management model can send the result back to the software application server that requested the decoding of the audio signal. The software application server can use the decrypted audio signal to process the user's request.
0060FIG. 3 is a flowchart of an exemplary method 300 that recognizes speech using parallel decoding. Method 300 can be implemented in systems such as, for example, systems 100 and 200, and for clarity of presentation, the following description uses systems 100 and 200 as the basis for an example illustrating this method. However, method 300 can also be implemented using different systems or combinations of systems.
0061In step 302, the audio signal is received. For example, the voice recognizer 108 can receive the voice signal 104. In one implementation, the audio signal 104 has already been sampled and segmented into digital frames for processing before being sent to the speech recognizer 108. In another implementation, speech recognizer 108 can also perform these functions.
0062In some implementations, the audio signal can be preprocessed to identify which portion of the signal contains audio and which portion is determined to be noise. The received voice signal 104 can include only a portion determined to have voice, which can then be decoded by the voice recognizer 108 in the following steps.
0063In steps 304A-N, a speech recognition task (SRT) is started. In one implementation, the SRT is started at about the same time and the decoding of the voice represented by the voice signal 104 is started. SRS in Figure 2<sub>AD</sub>SRS, such as, may have different latency in processing the audio signal, and as a result, the SRT may require different amounts of time to complete.
0064In step 306, the progress of the SRT is monitored. For example, the latency monitor 260 can track the latency associated with each SRS (both actual latency and estimated latency).
0065In step 308, SRT<sub>1-N</sub>It is determined whether any one of them has generated the recognition result. For example, SRS can output a recognition result (or an instruction that a result exists) to the recognition result monitor 258 after it is generated. If none of the SRS has produced a recognition result, Method 300 can return to step 306 and continue to monitor the progress of the SRT. If the SRS produces one or more recognition results, the method can proceed to step 310.
0066In step 310, it is determined whether any confidence value associated with the generated recognition result satisfies the confidence threshold. For example, the SRS Avoter 270 can compare a confidence value (single confidence value or combination confidence value) with respect to a recognition result to a confidence point or reliability function, as described above. If the current confidence value does not meet the confidence threshold, method 300 can return to step 306 and the progress of the SRT is monitored. If the confidence threshold is met, method 300 can proceed to step 312.
0067In step 312, the unfinished SRT is stopped. For example, if you have 10 SRTs running in parallel and 4 are complete, you can undo or stop the remaining 6 SRTs. In one implementation, the SRS Avoter 270 can send a stop command 266 to the SRS so that the appropriate SRS aborts the speech recognition task.
0068In some implementations, one or more of the speech recognition tasks are not stopped, but simply "paused" (for example, the state of a processing task can be saved and restarted later). For example, if the recognition result is found to be inaccurate (for example, when the software application server prompts the user to confirm that the audio was decoded correctly, the user responds with a negative expression). , You can restart the "paused" speech recognition task.
0069In some implementations, the SRT can be selectively paused, for example, based on the accuracy of the SRS running the SRT. For example, if the recognition result is associated with a confidence value that barely meets the confidence threshold, the Avoter 270 can selectively pause the SRT for a more accurate SRS and stop the rest of the SRT. If the recognition result is found to be inaccurate, the paused SRT of the more accurate SRS can be restarted.
0070In some implementations, previously completed SRTs and previously stopped SRTs can start at the same time as the "unpaused" SRT. This gives a more accurate SRT more time to complete than if the SRT were completely restarted. In yet another implementation, the information inferred or determined based on the confirmation of the user's inaccurate perception can be integrated with the unpaused SRT as well as the restarted task. For example, in a new round of speech decoding, erroneous utterances can be taken out of consideration. In addition, in the second round of recognition processing, some voices, words, etc. used to obtain erroneous results can be discounted or excluded from consideration.
0071In step 314, the final recognition result is selected based on the generated result. For example, the final result selection module 280 can identify the recognition result associated with the highest average confidence score. In some implementations, selection can also be weighted based on the accuracy of the SRS that produces the result, and results from the usually accurate SRS are favored over the less accurate SRS. In yet another implementation, the choice can also be based on the correlation between the machines that generate the result or the frequency of occurrence associated with the result. The selected result can be output to the application that requested the decoding of the audio signal. Then the method can be terminated.
00724A-C show diagrams showing the execution of an exemplary speech recognition task. Figure 4A shows the execution of four SRTs with four SRSs. In the implementation shown, SRT is started in parallel and SRS<sub>A</sub>Generates the recognition result first. SRS<sub>A</sub>Finds a confidence value of 0.7 for the recognition result. In some implementations, the SRS management module 110 can compare the confidence value with the confidence threshold. If the confidence value does not meet the threshold, the remaining tasks are allowed to perform. For example, if the confidence threshold is a fixed constant of 0.9, the initial recognition result 0.7 does not meet the threshold, so the SRS management module allows the rest of the SRS to continue.
0073Next, SRS<sub>B</sub>Produces a recognition result and an associated value of 0.85. This confidence value also does not meet the confidence threshold of 0.9, so the rest of the tasks are allowed to continue.
0074In addition, the SRS management system can also track the latency associated with each SRS and compare these latency with the allowed latency thresholds. In some implementations, SRS (eg, SRS), as shown in Figure 4A.<sub>C</sub>And SRS<sub>D</sub>) Does not generate a recognition result before the latency threshold, the SRS management module 110 can send a stop command to the SRS.
0075In one implementation, if the SRT is stopped before a recognition result that meets the confidence threshold is generated, the SRS management module 110 selects the result with the highest confidence value, even if the confidence threshold is not met. can do. In some implementations, the next highest confidence value may have to be within the determined confidence threshold range (eg 10%) in order to be selected. In yet another implementation, the SRS management module 110 can send a request to repeat the voice input if no recognition result is selected.
0076FIG. 4B is a diagram showing that the SRS stops the unfinished SRT after generating a recognition result having a confidence value that satisfies the confidence threshold. In this example, the confidence threshold is 0.9. SRS<sub>A</sub>First produces the recognition result, but SRS<sub>A</sub>Assigns a confidence value of 0.7, which is lower than the confidence threshold, to the result. Therefore, the SRS management module 110 is an SRS.<sub>BD</sub>Allows the execution to be performed.
0077SRS<sub>B</sub>Then generates a recognition result and assigns it a confidence value of 0.9. The SRS management module 110 compares this confidence value with the confidence threshold and determines that the threshold is satisfied. Then the SRS management module is SRS<sub>C</sub>And SRS<sub>D</sub>A stop command can be sent to SRS<sub>C</sub>And SRS<sub>D</sub>Stops each SRT without producing a recognition result.
0078FIG. 4C is a diagram showing that the unfinished SRT is stopped based on the low confidence value of the generated recognition result. In this example, the confidence threshold can be set to fixed point 0.9. SRS<sub>A</sub>And SRS<sub>B</sub>Produces recognition results, both of which are associated with relatively low confidence values of 0.3 and 0.25, respectively. Given that both confidence values are relatively low, the SRS management module 110 generated a recognition result in which the previous SRS had a confidence value significantly lower than the confidence threshold, so the SRS<sub>C</sub>And SRS<sub>D</sub>Can be stopped and commanded to these SRSs under the assumption that is unlikely to produce a recognition result with a confidence value that meets the confidence threshold.
0079In one implementation shown in Figure 4C, the SRS management module 110 may wait for the requested amount of time before sending a stop command based on the low confidence of previously generated recognition results. it can. In one implementation, the SRS management module 110 starts a time frame based on when the final recognition result is generated. The sought time frame can allow another SRS to complete that SRT, but if no results are produced during the allowed time frame, a command to stop any unfinished SRT Can be sent.
0080In some implementations, the determination of the waiting time frame can be based on the estimated latency of one or more of the SRSs that are not producing recognition results. For example, the SRS management module 110 is an SRS.<sub>C</sub>Can be determined to have the shortest estimated latency of the remaining SRS. For example, SRS<sub>C</sub>May have a typical wait time of 0.5 seconds. SRS<sub>B</sub>If the voice recognition management module 100 produces a recognition result after 0.4 seconds, the speech recognition management module 100 is delayed by 0.1 seconds and SRS before sending the stop command.<sub>C</sub>Can determine whether to produce a recognition result.
0081In another implementation, a stop command can be sent immediately. For example, the SRS management module 110 can send a stop command after the desired number of SRSs, also associated with low confidence values, have generated recognition results. In the case shown in Figure 4C, a stop command is sent as soon as half of the SRS returns a recognition result associated with a low confidence value.
0082In some implementations, if the confidence value is low, the system will continue to receive more recognition results until the system verifies that the composite confidence value (eg, total / cumulative confidence value) is above a certain threshold. For some recognition tasks, confirmation is never done and the system can terminate the recognition process by rejecting the utterance. Therefore, in one implementation, there are three types of confidence, the first is the original confidence from each recognition process, and the second is the cumulative total confidence obtained from the original confidence from each recognition process. Third is the expectation that the total confidence will change (eg increase) as the system waits for more recognition events.
0083In some cases, the system receives a sufficient number of consistently unreliable results across the uncorrelated recognizers and prompts them to stop all recognition tasks and reject the utterance. If a denial occurs, the system can prompt the user to repeat the utterance. The case of rejection can occur, for example, when the individual original confidence values are consistently low, the cumulative total confidence is low, and the expectation that the total confidence may change with more perceptions is low. is there.
0084In one implementation, training on estimated expected confidence changes given a particular set of confidence values counts the distribution of final recognition confidence given a partial recognition confidence training example. (For example, after checking 20 confidence values less than 0.1 from the first 20 recognizers, the system uses a combination confidence value with more than 20 recognizers to make the total confidence value greater than 0.5. Never face the increasing example above, so the system is trained to refuse to speak when this situation arises).
0085In some implementations, the combinatorial confidence associated with the final recognition result may be a function of the individual confidence values from the individual SRS. Results with high confidence values from many recognizers that match each other are given high combinatorial confidence values. The weighting of individual contributions for each recognizer can be based on experimental optimization of recognition of test data during the training process.
00865A-C are diagrams of various methods for selecting exemplary recognition results and confidence values generated by SRS, as well as final recognition results. Specifically, Figures 5A to 5C show SRS.<sub>A</sub>SRS from<sub>A</sub>Output 502, SRS<sub>B</sub>SRS from<sub>B</sub>Output 504, and SRS<sub>C</sub>SRS from<sub>C</sub>Shows output 506. In this example, the output is generated in response to each SRS attempting to decode the audio signal representing the word "carry". Since each SRS can be different, the recognition results produced by each SRS can be different, as shown in Figures 5A-C.
0087In some implementations, the SRS output contains the top N recognition results (where N can represent any positive integer or 0), and of the N recognition results, which recognition result has the highest confidence value. Selected based on whether it is related to. For example, SRS<sub>A</sub>Output 502 is SRS<sub>A</sub>Includes the top four recognition results for and associated confidence values Result = carry, Confidence = 0.75; Result = Cory, Confidence = 0.72; Result = query, Confidence = 0.6; and Result = hoary, Confidence = 0.25.
0088SRS<sub>B</sub>Output 504 includes Result = quarter, Confidence = 0.64; Result = Cory, Confidence = 0.59; Result = hoary, Confidence = 0.4; and Result = Terry, Confidence = 0.39.
0089SRS<sub>C</sub>Output 506 includes Result = tary, Confidence = 0.58; Result = Terry, Confidence = 0.57; Result = Cory, Confidence = 0.55; and Result = carry, Confidence = 0.2.
0090Figure 5A shows an exemplary selection algorithm that selects the recognition results associated with the highest confidence value. For example, the final result selection module 113 can compare all of the recognition results and select the recognition result associated with the highest confidence value. In this example, the result "carry" is associated with the highest confidence value of 0.75 of all confidence values, which is selected as the final recognition result. The selection module can then output a recognition result "carry" for further processing by the application requesting audio decoding.
0091FIG. 5B shows an exemplary selection algorithm that selects recognition results based on which result has the highest combined confidence value. For example, multiple SRSs may produce the same recognition result, but may assign different confidence values to the result. In one implementation, multiple confidence scores for the same result can be averaged (or combined) to generate a combined confidence score. For example, the recognition result "carry" is SRS<sub>A</sub>And SRS<sub>C</sub>Generated by both, but SRS<sub>A</sub>Assigns a 0.75 confidence value to the result, SRS<sub>C</sub>Assigns a 0.2 confidence value to the result. The average of these confidence values is 0.475.
0092Similarly, the average combination reliability score for the recognition result "Cory" is 0.61 and the combination reliability score for "quarry" is 0.62. In this example, the selection module 113 can select "quarry" as the final recognition result because the combinatorial confidence value of "quarry" is higher than the combinatorial confidence value of the other results. Note that this selection algorithm produces a different final result than the algorithm shown in Figure 5A, even though the selection was made from the same pool of recognition results.
0093Figure 5C shows an exemplary selection algorithm that takes weight factors into account in the selection of recognition results. In some implementations, the weights can be based on the frequency of occurrence of recognition results. For example, Table 550 lists three weights that can be multiplied by the combined confidence score discussed above to generate a new weighted confidence score.
0094In this example, if the recognition result is generated by a single SRS (eg, if the result occurs with a frequency of "1"), the combination confidence score can be multiplied by the weight "1". Therefore, if the recognition result occurs only once, the recognition result does not benefit from the weighting. If the recognition result occurs twice, the recognition result can be weighted with a factor 1.02 that supports this recognition result with some priority over another recognition result that occurs only once. If the recognition result occurs three times, the recognition result can be weighted by factor 1.04.
0095In the example of FIG. 5C, the combinatorial confidence value for the recognition result Cory is weighted against the factor 1.04, resulting in a weighted value of 0.6344. The combination confidence value for the recognition result "quarry" is weighted against factor 1.02, resulting in a weighting value of 0.6324. In this case, even if the unweighted combination reliability score of the result "Cory" is lower than the unweighted combination reliability score of the result "quarry", the weighted combination reliability score of the result "Cory" is higher than the result "quarry". , The selection module 113 can select the result "Cory" in preference to the result "quarry".
0096The values used to select the final recognition result are, but are not limited to, the distribution of the confidence score generated by the SRS, the characteristics of the SRS that generated the recognition result (eg, overall accuracy, specific context). It can be weighted based on several criteria, including accuracy in, accuracy over a defined time frame, etc.), as well as similarities between SRSs that produce the same recognition result.
0097In another implementation, the correlation between the recognition confidence value and the recognition error for the recognizer and the final composite recognizer can be used to weight the final recognition result. For example, during training, the system can count the number of times a particular recognizer responds with a confidence value of 0.3, and how often these "0.3 confidence recognition results" are incorrect for that recognizer. And how often the final combination recognition is also a recognition error can be counted. The system can use the same normalization counting when combining similar recognition results. The combinatorial reliability can be estimated from the number of times the recognizer obtained the same result (having a given confidence value) and the number of times the common result was correct.
0098FIG. 6 is an exemplary graph 600 of the confidence value distribution used to weight the values used in the final recognition result selection. The y-axis of the graph shows where a particular confidence value is located along the normalization scale (0.0 to 1.0). The x-axis of the graph represents which particular SRS produces which recognition result. In this example, SRS<sub>A</sub>Produces five recognition results, four of which concentrate relatively close to each other towards the expected confidence values in the medium to low range. A single recognition result quarry is located far away from other recognition results and has a relatively high confidence value. This means that the result "quarry" is significantly better than the other results, which are more substitutable for each other.<sub>A</sub>May indicate that it has a higher degree of trust.
0099In some implementations, outliers, that is, high confidence values that are far apart, can be weighted to support the selection of relevant recognition results. For example, selection module 113 can weight a confidence value of 0.9 for the result "quarry" with a constant of 1.05. In that case, the confidence value obtained for "quarry" increases to 0.945.
0100Alternatively, confidence values placed at more even intervals may not receive additional weighting (or may receive lower weighting). For example, SRS<sub>B</sub>The confidence values for the recognition results generated in are arranged at more uniform intervals with no significant outliers. In this case, the top-level recognition result "quarry" is unlikely to be correct (for example, "quarry" does not stand out from the top-level result in the resulting cluster with lower confidence values), so the selection module. 113 may not add weight to the confidence value for "quarry".
01017A-E are Venn diagrams showing the exemplary recognition result sets output by SRS and the correlations between the sets, which can be used to weight the recognition results. Figure 7A shows SRS<sub>A</sub>Recognition result generated by<sub>A</sub>, SRS<sub>B</sub>Recognition result generated by<sub>B</sub>, And SRS<sub>C</sub>Recognition result generated by<sub>C</sub>It is a Venn diagram 700 including the three recognition result sets.
0102Results as shown in Venn diagram 700<sub>A</sub>,result<sub>B</sub>, And results<sub>C</sub>Partially overlap. In this example, the result<sub>A</sub>And results<sub>B</sub>Is the result<sub>A</sub>And the result<sub>C</sub>Overlap or result<sub>B</sub>And the result<sub>C</sub>Has more overlapping results than overlapping. This is SRS<sub>A</sub>And SRS<sub>B</sub>Often produce the same recognition results, whereas SRS<sub>C</sub>The result is SRS less often<sub>A</sub>Or SRS<sub>B</sub>May indicate that it does not correspond to the result of.
0103In some implementations, the intersection of results is based on which SRS produces the same recognition result in response to a particular speech recognition task. For example, if two SRSs generate a top-level recognition result for a particular task, this result can be added to the intersection.
0104In another example, the first SRS produces the recognition result "Cory" as its best result, and the second SRS produces the recognition result "Cory" as the result of its fourth rank (out of the five generated results). If so, the result "Cory" is added to the intersection. In some implementations, results that are neither associated with the top level can be added to the intersection results, but they can also be associated with discount factors that indicate that they differ in terms of ranking. For example, the difference between two rankings can be used to discount the weighting factors associated with the intersection (eg, each difference in rankings can be associated with a discount factor). For example, if the rankings are 1 and 4, the absolute value of the difference is 3, which can be associated with the discount factor 0.3, which is multiplied by the weight associated with the intersection. For example, if the weight is 1.03 and the discount factor is 0.3, the total weight can be multiplied by the "boost" factor of weight 1.03, i.e. 0.03. This results in a smaller new boost factor of 0.01, so the new total weight value is 1.01.
0105In some implementations, recognition result overlap between SRS can be used to weight recognition results so that the recognition results are supported or unsupported in the final recognition result selection. For example, if the recognition results are often produced by two matching SRSs, the recognition results can be weighted less (or disapproved) than the recognition results produced by two less matching SRSs. .. Figures 7B-E show this in more detail.
0106Figure 7B shows the results from the Venn diagram 700 in Figure 7A.<sub>A</sub>And results<sub>B</sub>Venn diagram 710 containing only. As mentioned above, SRS<sub>A</sub>And SRS<sub>B</sub>Can be classified in the same way to some extent based on the similarity of the recognition results. In some implementations, weight factors can be assigned to recognition results that are in the overlap between two (or three or more) SRSs. For example, the weight factor 0.01 can be associated with the recognition results in this set.
0107In some implementations, this weight factor is small when the overlap is large, and the weight factor is large when the overlap is small. This can reflect the assumption that these overlapping results should be supported, as the results produced by the less consistent SRS are more likely to be correct. For example, SRSs that produce different results may have different underlying architectures and may be susceptible to different types of misrecognition.
0108Figure 7C shows the results from the Venn diagram 700 in Figure 7A.<sub>A</sub>And results<sub>C</sub>A Venn diagram 720 containing only is shown. In this example, the overlap between the results is less than the overlap shown in Figure 7B. Therefore, in this implementation, the weight factor 0.6 is larger for the results in the overlap than in the intersection shown in FIG. 7B.
0109Similarly, Figure 7D shows the results.<sub>B</sub>And results<sub>C</sub>A Venn diagram 730 including is shown. The intersection of these results is the size between the intersections of Figures 7B and 7C. Therefore, in this implementation, the weight factor is also the size between the weight factors associated with the intersection of Figures 7B and 7C (eg 0.03).
0110Figure 7E shows the Venn diagram 700, which is also shown in Figure 7A, but all SRS<sub>AC</sub>The intersection between the results of is highlighted. The intersection reflects the set of recognition results generated by each SRS. Given that the match between the three SRSs is relatively rare (in this example), the recognition results within this set can be associated with a higher weight than the other weights, namely 0.1.
0111Figures 8A and 8B show Venn diagrams 800 and 810 showing how the intersection between SRS can be adapted or changed during run-time operation of the system. In some implementations, when the intersection of recognition results changes, so does the weight associated with the intersection.
0112Figure 8A shows SRS<sub>A</sub>And SRS<sub>B</sub>The first intersection of the recognition results generated by is shown. The first intersection is associated with a weight of 0.01. In one implementation, speech recognizer 108 performs additional speech decoding to produce additional recognition results. The SRS Correlation Monitor 282 can monitor the results and identify the intersection of the results between the various SRSs.
0113Correlation monitor 282 can dynamically update intersection calculations as more results are generated. This is shown in Figure 8B, which shows the same SRS as Figure 8A, except that the intersection has changed.<sub>A</sub>And SRS<sub>B</sub>Is shown. In this example, the number of times the SRS matched for a particular speech recognition task increased compared to the number of tasks performed by the SRS, so the intersection increased.
0114Weights can also be reduced in response to an increase in the intersection. For example, a smaller weight of 0.001 can be associated with the intersection result set of Venn diagram 810. In some implementations, changes in weight values can be linearly associated with changes in the size of the intersection result set. For example, the system can reduce the weighting or support of results from a recognizer when the recognizer is similar to another recognizer. In Figures 8A and 8B, the similarity of the recognition results for the two recognizers is represented as the intersection between the two recognizers, and if the intersection is large, the system when both recognizers produce the same result. Reduces the weight that can be associated with the recognition result. On the other hand, when the two recognizers are very different (for example, different recognition algorithms are generally generated because of different speech recognition algorithms), the intersection of the results may be small. Then, when these two different recognizers match in terms of utterance, the match indicates that the result is more likely to be correct, so the system weights the result so that it is more important in the system. Can be done.
0115FIG. 9 is Graph 900 showing an exemplary correlation between the SRS error rate and the weights associated with the recognition result. In some implementations, the final recognition result selection can be more weighted to the recognition results produced by SRS with a low error rate. For example, if the SRS has a high error rate, the recognition result can be discounted (or not heavily weighted) compared to the recognition result produced by the very accurate SRS.
0116Graph 900 shows an exemplary function or algorithm for assigning weights to a particular SRS. The y-axis of Graph 900 shows the error rate associated with SRS, and the x-axis shows the weights associated with SRS. In this example, an SRS with an error rate higher than the calculated threshold (eg, SRS)<sub>A</sub>, SRS<sub>E</sub>, SRS<sub>C</sub>) Is weighted, discount weights (eg 0.9, 0.95, 0.8) are used. SRS with an error rate lower than the threshold (eg SRS)<sub>B</sub>) Is weighted by boost weights (eg 1.01, 1.04, 1.1). In this example, the SRS above the error threshold (eg SRS)<sub>D</sub>) Is weighted by a neutral weight (eg 1).
0117In some implementations, the error rate associated with each SRS can be updated based on the confirmation that the recognition result is inaccurate (eg, the result is selected as the final recognition result, rejected by the user, and so on. 1 The result is selected as the final recognition result and is determined to be correct based on the user's acceptance, so the unselected result is recorded as an erroneous result, etc.). Selection module 113 can dynamically change the weights based on the updated error rate associated with each SRS.
0118FIG. 10 is a block diagram of computing devices 1000, 1050 that can be used as clients, or servers, or multiple servers, to implement the systems and methods described in this document. The computing device 1000 shall represent various forms of digital computers such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. Computing device 1050 shall represent various forms of mobile devices such as personal digital assistants, mobile phones, smartphones, and other similar computing. In addition, the computing device 1000 or 1050 can include a universal serial bus (USB) flash drive. USB flash drives can store operating systems and other applications. A USB flash drive can include input / output components such as a wireless transmitter and a USB connector that can be plugged into the USB port of another computing device. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the invention described and / or claimed in this document.
0119The computing device 1000 includes a high-speed interface connected to a processor 1010, a memory 1020, a storage device 1030, a memory 1020 and a high-speed expansion port, and a low-speed interface connected to a low-speed bus and a storage device 1030. The components 1010, 1020, and 1030 are interconnected using various bus 1050s and can be mounted on a common motherboard, or otherwise optionally mounted. Processor 1010 is a computing device that contains instructions stored in memory 1020 or on storage device 1030 to display graphical information for a GUI on an external input / output device 1040, such as a display coupled to a high-speed interface. Can process instructions to execute within 1000. In another implementation, multiple processors and / or multiple buses can be appropriately used with multiple memories and multiple types of memory. In addition, multiple computing devices 1000 can be connected, each device achieving each part of the required operation (eg, as a server bank, a group of blade servers, or a multiprocessor system).
0120The memory 1020 stores information in the computing device 1000. In one implementation, memory 1020 is a volatile memory unit. In another implementation, memory 1020 is a non-volatile memory unit. The memory 1020 may also be another form of computer-readable medium, such as a magnetic disk or optical disc.
0121The storage device 1030 can provide a large amount of storage for the computing device 1000. In one implementation, the storage device 1030 includes a floppy (registered trademark) disk device, a hard disk device, an optical disk device, a tape device, a flash memory or other similar solid-state memory device, a device in a storage area network, or other configuration. It may be a computer-readable medium, such as an array of, or may include a computer-readable medium. Computer program products can be tangibly implemented as information carriers. Computer program products can also include instructions that, when executed, perform one or more of the methods described above. The information carrier is a computer-readable or machine-readable medium such as memory 1020, storage device 1030, memory on processor 1010.
0122The high-speed controller manages the bandwidth-intensive operation of the computing device 1000, and the low-speed controller manages the bandwidth-heavy operation. The allocation of such functions is only exemplary. In one implementation, a high-speed controller is coupled to a memory 1020, a display (eg, via a graphics processor or accelerator), and a high-speed expansion port that can accept various expansion cards (not shown). In this implementation, the slow controller is coupled to the storage device 1030 and the slow expansion port. One or more slow expansion ports that can include various communication ports (eg USB, Bluetooth®, Ethernet®, Wireless Ethernet®), keyboards, pointing devices, scanners, etc. It can be coupled to an input / output device 1040 or a networking device such as a switch or router, for example via a network adapter.
0123The computing device 1000 can be implemented in several different formats, as shown in the figure. For example, the computing device 1000 can be implemented as a standard server, or can be implemented multiple times as a group of such servers. The computing device 1000 can be implemented as part of a rack server system. Further, the computing device 1000 can be implemented as a personal computer such as a laptop computer. Alternatively, the components of the computing device 1000 can be combined with other components within the mobile device (not shown). Each such device can include one or more computing devices 1000, and the entire system can consist of multiple computing devices 1000 communicating with each other.
0124In another embodiment, device 1000 includes processor 1010, memory 1020, and input / output devices 1040 such as displays, communication interfaces, transceivers, among other components. The device 1000 can also include a storage device 1030, such as a microdrive or other device, to provide additional storage. Each component 1000, 1010, 1020, 1030, and 1040 are interconnected using various bus 1050s, some of which can be mounted on a common motherboard, or in other ways. It can be attached as appropriate.
0125Processor 1010 can execute instructions in computing device 1000, including instructions stored in memory 1020. Processor 1010 can be implemented as a chipset of chips containing multiple separate analog and digital processors. In addition, processor 1010 can be implemented using any of several architectures. For example, the processor 1010 may be a CISC (Complex Instruction Set Computers) processor, a RISC (Reduced Instruction Set Computer) processor, or a MISC (Minimal Instruction Set Computer) processor. Processor 1010 can implement applications running on device 1000 and wireless communication by device 1000 to coordinate with other components of device 1000, such as controlling the user interface.
0126Processor 1010 can communicate with the user via a control interface and a display interface coupled to the display. The display may be, for example, a TFT (thin film transistor liquid crystal display) display or an OLED (organic light emitting diode) display, or other suitable display technology. The display interface may include suitable circuits for driving the display to present graphical or other information to the user. The control interface can receive commands from the user and translate them to submit to processor 1010. In addition, the external interface can implement communication with processor 1010 so as to allow near-range communication with other devices in device 1000. External interfaces can, for example, implement wired communication in one implementation, or wireless communication in another implementation, and can use multiple interfaces.
0127The memory 1020 stores information in the computing device 1000. The memory 1020 can be implemented as one or more of a computer-readable medium, a volatile memory unit, or a non-volatile memory unit. An expansion memory can be provided and connected to the device 1000 via an expansion interface. The expansion memory is, for example, SIMM (Single In Line Memory). Module) Can include a card interface. Such extended memory can provide additional storage space for device 1000, or can also store applications or other information for device 1000. Specifically, the extended memory can include instructions that execute or supplement the above-mentioned processes, and can also include secure information. Thus, for example, extended memory can be provided as a security module for device 1000 and can be programmed with instructions permitting the safe use of device 1000. In addition, secure applications can be provided with additional information via the SIMM card, such as placing identification information on the SIMM card in a non-hackable manner.
0128Memory 1020 can include, for example, flash memory and / or NVRAM memory, as discussed above. In one implementation, the computer program product can be tangibly implemented as an information carrier. Computer program products include instructions that, when executed, perform one or more of the methods described above. The information carrier is a computer-readable or machine-readable medium such as memory 1020, extended memory, memory on processor 1010.
0129The device 1000 can communicate wirelessly via a communication interface, which may include a digital signal processing circuit if desired. Communication interfaces, among others, communicate under various modes or protocols such as GSM® voice calling, SMS, EMS, or MMS messaging, CDMA, TDMA, PDC, WCDMA®, CDMA2000, or GPRS. Can be realized. Such communication can be done, for example, via a radio frequency transceiver. In addition, short-range communications such as using Bluetooth®, WiFi, or other transceivers such as those (not shown) can be performed. In addition, the GPS (Global Positioning System) receiver module can provide additional navigation-related wireless data and location-related wireless data to device 1000, which data is used as appropriate by the application running on device 1000. can do.
0130The device 1000 can also use an audio codec to audibly communicate, and the audio codec 1060 can receive utterance information from the user and convert it into usable digital information. The audio codec can also generate audible sound to the user through a speaker, such as a speaker in the handset of device 1000. Such sounds can include sounds from voice phone calls, include recorded sounds (eg, voice messages, music files, etc.), and also include sounds produced by applications running on device 1000. be able to.
0131The computing device 1000 can be implemented in several different forms, as shown in the figure. For example, the computing device 1000 can be implemented as a mobile phone. The computing device 1000 can also be implemented as part of a smartphone, personal digital assistant, or other similar mobile device.
0132Various implementations of the systems and techniques described here are implemented as digital electronic circuits, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. can do. These various implementations are combined to receive data and instructions from a storage system, at least one input device, and at least one output device, and send the data and instructions to them, at least one program that may be dedicated or general purpose. It can include implementations as one or more computer programs that are executable and / or interpretable on a programmable system that includes capable processors.
0133These computer programs (also called programs, software, software applications, or code) include machine language instructions for programmable processors and are implemented as high-level procedural and / or object-oriented programming languages and / or assembly / machine language. can do. As used herein, the terms "machine readable medium" and "computer readable medium" provide machine language instructions and / or data to a programmable processor, including machine readable media that receive machine language instructions as machine readable signals. Refers to any computer program product, device, and / or device used in (eg, magnetic disk, optical disk, memory, programmable logic device (PLD)). The term "machine-readable signal" refers to any signal used to provide machine-language instructions and / or data to a programmable processor.
0134To enable user interaction, the systems and techniques described here include a display device that displays information to the user (eg, a CRT (cathode tube) monitor or LCD (liquid crystal display) monitor) and the user to a computer. It can be implemented on a computer with a keyboard and pointing device (eg mouse or trackball) that can give input. Other types of devices can also be used to achieve user interaction, for example, the feedback provided to the user can be any form of sensory feedback (eg, visual feedback, auditory feedback, or tactile feedback). The input from the user can be received in any form including acoustic input, voice input, or tactile input.
0135The systems and techniques described herein include a back-end component (eg, as a data server), or a middleware component (eg, an application server), or a front-end component (eg, a system described herein by a user). And a computing that includes a client computer with a graphical user interface or web browser capable of interacting with one implementation of the technique), or any combination of such back-end, middleware, or front-end components. It can be implemented as a client system. The components of the system can be interconnected by digital data communication (eg, a communication network) of any form or medium. Examples of telecommunications networks include local area networks (LAN), wide area networks (WAN), peer-to-peer networks (with ad hoc or static members), grid computing infrastructure, and the Internet.
0136The computing system can include clients and servers. Clients and servers are generally separated from each other and usually interact over a communication network. The client-server relationship is created by computer programs that are running on their respective computers and have a client-server relationship with each other.
0137Some embodiments of the present invention have been described. Nevertheless, it will be appreciated that various modifications can be made without departing from the spirit and scope of the invention. For example, a combination score, a binding score, and a confidence score of multiple SRSs can include features such as hypothesis consistency and guessing about utterance identification. For example, three SRSs that output a first result with a reliability of 0.8 may be more reliable than one SRS that outputs a second result with a reliability of 0.9.
0138In some implementations, a given set of SRS can be selected for use based on latency or other factors. For example, if audio is received in response to prompting the user for an answer to a general interrogative, instead of allowing all available SRS to process the answer, the two fastest SRS You can choose to process the answer.
0139In addition, in some implementations, the overall confidence in the final recognition result may decrease when the individual recognition results generated by SRS do not match. One exemplary algorithm for selecting the "best" current result when the recognition results do not overlap at all is to select the recognition result with the highest reliability of each individual. In this example, the combinatorial confidence is the expected number of correct recognition results counted during training when the system has similar conditions and similar given confidence values without duplication. Similar counts and statistics can be estimated for partial duplication of recognition results for a given amount. Therefore, if the degree of duplication correlates / correlates with the reduction of all recognition errors during training, the entire system can assign higher confidence values to the combination of partially overlapping recognition results.
0140For example, the various forms of flow shown above can be used by reordering, adding, or removing steps. In addition, although some applications and methods of using multiple speech recognition systems in speech decoding have been described, it should be understood that many other applications are also conceivable. Therefore, other embodiments are within the scope of the following claims.
0141100, 200 Illustrative system 102, 206 mobile phones 104, 208 audio signal 106 Voice-enabled phonebook information server 108 Speech recognizer 110, 250 SRS management module 113, 280 Final result selection module 114 Final recognition result 116,264 Confidence value 202 Voice transmission segment 204 Speech recognizer segment 210 phone server 212 Software application server 252 language model 254 Acoustic model 256 Speech recognition algorithm 258 Recognition result monitor 260 Wait time monitor 262 Recognition result 266 Stop command 270 SRS Avota 282 SRS Correlation Monitor
19 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US10049672B2 | Cited by | United States of America | Applicant |
| JP2017076139A | Cited by | Japan | Search report |
| JP2005031758A | Cites | Japan | – |
| JP2005524859A | Cites | Japan | – |
| US20020055845A1 | Cites | United States of America | – |
| JP2002150039A | Cites | Japan | – |
| JP2005266192A | Cites | Japan | – |
| JP5451933B2 | Cites | Japan | – |
35 members in 6 offices
Members35
| Document | Office | Kind | |
|---|---|---|---|
| US2010004930A1 | United States of America | A1 | |
| WO2010003109A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2010003109A3 | World Intellectual Property Organization (WIPO) | A3 | |
| EP2301012A2 | European Patent Office (EPO) | A2 | |
| KR20110043644A | Republic of Korea | A | |
| CN102138175A | China | A | |
| JP2011527030A | Japan | A | |
| EP2301012A4 | European Patent Office (EPO) | A4 | |
| US8364481B2 | United States of America | B2 | |
| US2013138440A1 | United States of America | A1 | |
| US8571860B2 | United States of America | B2 | |
| JP5336590B2 | Japan | B2 | |
| JP2013238885A | Japan | A | |
| CN102138175B | China | B | |
| US2014058728A1 | United States of America | A1 | |
| JP5451933B2 | Japan | B2 | |
| JP2014056278A | Japan | A | |
| CN103761968A | China | A | |
| EP2301012B1 | European Patent Office (EPO) | B1 | |
| KR20150103326A | Republic of Korea | A | |
| KR101605148B1 | Republic of Korea | B1 | |
| US9373329B2 | United States of America | B2 | |
| KR101635990B1 | Republic of Korea | B1 | |
| KR20160079929A | Republic of Korea | A | |
| US2016275951A1 | United States of America | A1 | |
| JP2017076139A | Japan | A | |
| JP6138675B2This record | Japan | B2 | |
| KR101741734B1 | Republic of Korea | B1 | |
| CN103761968B | China | B | |
| US10049672B2 | United States of America | B2 | |
| US2018330735A1 | United States of America | A1 | |
| JP6435312B2 | Japan | B2 | |
| US10699714B2 | United States of America | B2 | |
| US2020357413A1 | United States of America | A1 | |
| US11527248B2 | United States of America | B2 |
21 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Written notification of registration of transferJAPANESE INTERMEDIATE CODE: R350R350 | R350 | |
| Written request for registration of change of nameJAPANESE INTERMEDIATE CODE: R313533S533 | S533 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Transfer to examiner for re-examination before appeal (zenchi)AppealJAPANESE INTERMEDIATE CODE: A911A911 | A911 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Decision of refusalJAPANESE INTERMEDIATE CODE: A02A02 | A02 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A132A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 6138675
- Application
- 268860
Titles2
- Japanese
- 並列認識タスクを用いた音声認識
- English
- Speech recognition using parallel recognition tasks
Classification
- CPC, 6
- G10L15/32
- G10L15/30
- G10L15/34
- G10L15/00
- G10L15/26
- G10L15/01
- IPC, 1
- G10L15 34
