Distributed speech recognition using one way communication
Abstract
This record has no abstract on file.
Term
2.9 yearsto projected expiry
Projected expiry 31 August 2029, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
1 claim: 1 independent, 0 dependent
- 1Zastrzeżenia patentowe 1. Sposób wdrożony na komputerze, obejmujący:(A) po stronie klienta (106), przesyłanie strumienia mowy i strumienia kontrolnego do serwera (118) rozpoznającego mowę przy użyciu protokołu przesyłania dokumentów hipertekstowych, HTTP, mającego pierwszy okres przeterminowania;(B) na serwerze rozpoznawania mowy, przy użyciu automatycznego silnika rozpoznawania mowy, inicjowanie rozpoznawania strumienia mowy;(C) po stronie klienta, przesyłanie pierwszego żądania wyniku rozpoznawania mowy do serwera rozpoznawania mowy za pomocą HTTP;i (D) na serwerze rozpoznawania mowy, przesyłanie powiadomienia do klienta wskazującego, że nie ma dostępnych wyników rozpoznawania mowy w trakcie drugiego czasu przeterminowania, który różni się od pierwszego czasu przeterminowania;i (E) po stronie klienta, w odpowiedzi na otrzymanie powiadomienia, przesyłanie drugiego żądania wyniku rozpoznawania mowy do serwera rozpoznawania mowy za pomocą HTTP. 2. Sposób według zastrz. 1, obejmujący ponadto: (F) na serwerze (118) rozpoznawania mowy, rozpoznanie pierwszego fragmentu strumienia mowy dla wygenerowania pierwszego wyniku rozpoznawania mowy;i (G) przesłanie pierwszego wyniku rozpoznawania mowy do klienta (106) przy użyciu HTTP, w odpowiedzi na drugie żądanie. 3. Sposób według zastrz. 2, w którym (G) obejmuje: (G)(1) określanie czy dostępny jest jakikolwiek wynik rozpoznawania mowy;(G)(2) jeśli żaden wynik rozpoznawania mowy nie jest dostępny, powrót do (G) (1);(G)(3) w innym wypadku, przesłanie pierwszego wyniku rozpoznawania mowy do klienta (106). 4. Sposób według zastrz. 3, w którym serwer (118) rozpoznawania mowy przeprowadza równolegle (F) i (G). 5. Sposób według zastrz. 1, w którym (A) obejmuje przesłanie strumienia mowy i strumienia kontrolnego za pomocą szyfrowanego protokołu przesyłania dokumentów hipertekstowych, HTTPS, i w którym (C) obejmuje przesłanie pierwszego żądania za pośrednictwem HTTPS. 6. System zawierający urządzenie klienckie (106) i serwer (118) rozpoznawania mowy: przy czym urządzenie klienckie zawiera: środki (110, 112) do przesyłania strumienia mowy i strumienia kontrolnego do serwera rozpoznawania mowy przy użyciu protokołu HTTP, mającego pierwszy okres przeterminowania;środki do przesyłania pierwszego żądania wyniku rozpoznawania mowy do serwera rozpoznawania mowy przy użyciu HTTP;i przy czym, serwer rozpoznawania mowy zawiera: środki stosujące automatyczny silnik (120, 218) rozpoznawania mowy dla zainicjowania rozpoznawania strumienia mowy;środki do przesyłania powiadomienia do urządzenia klienckiego wskazującego, że nie ma dostępnych wyników rozpoznawania mowy w trakcie drugiego czasu przeterminowania, który różni się od pierwszego czasu przeterminowania;i przy czym urządzenie klienckie ponadto zawiera środki do przesłania, w odpowiedzi na otrzymanie powiadomienia, drugiego żądania wyniku rozpoznawania mowy do serwera rozpoznawania mowy przy użyciu HTTP. 7. Sposób wdrożony na komputerze, realizowany przez serwer (118), przy czym sposób obejmuje: (A) odbieranie strumienia mowy i strumienia kontrolnego od klienta (106) przy użyciu protokołu HTTP, mającego pierwszy okres przeterminowania;(B) zastosowanie automatycznego silnika rozpoznawania mowy dla zainicjowania rozpoznawania strumienia mowy;(C) odbieranie pierwszego żądania wyniku rozpoznawania mowy od klienta przy użyciu HTTP;i (D) przesyłanie powiadomienia do klienta, wskazującego, że nie ma dostępnych wyników rozpoznawania mowy w trakcie drugiego czasu przeterminowania, który różni się od pierwszego czasu przeterminowania. 8. Sposób według zastrz. 7, obejmujący ponadto: (E) odbieranie drugiego żądania wyniku rozpoznawania mowy od klienta (106) przy użyciu HTTP;(F) rozpoznawanie pierwszego fragmentu strumienia mowy dla wygenerowania pierwszego wyniku rozpoznawania mowy;i (G) przesyłanie pierwszego wyniku rozpoznawania mowy do klienta za pomocą HTTP, w odpowiedzi na drugie żądanie. 9. Urządzenie zawiera: środki do odbierania strumienia (110) mowy i strumienia kontrolnego (112) od klienta (106) przy użyciu protokołu HTTP, mającego pierwszy okres przeterminowania;środki do stosowania automatycznego silnika rozpoznawania mowy dla zainicjowania rozpoznawania strumienia mowy;środki do odbierania pierwszego żądania wyniku rozpoznawania mowy od klienta przy użyciu HTTP;i środki do przesyłania powiadomienia do klienta wskazującego, że nie ma dostępnych wyników rozpoznawania mowy w takcie drugiego czasu przeterminowania, który różni się od pierwszego czasu przeterminowania. Sporządziła i zweryfikowała Anna Stenzel Rzecznik patentowy 'r·* FIG. 4
65 paragraphs, as filed
[0001] Many automatic speech recognition (ASR) modules are available that convert speech into text and control the computer in response to verbal commands. Some applications of automatic speech recognition applications require shorter processing times (the amount of time between saying a phrase and generating text by the speech recognition module) than others so that the end user will see their response. For example, the speech recognition module used for "live" speech recognition, for example controlling the movement of the cursor on the screen, may require a shorter processing cycle time (also called "response time") than the recognition module used to perform the transcript of the medical report.
[0002] The desired processing time may depend, for example, on the content of the speech which is processed by the speech recognition module. For example, in the case of short control and checking texts, for example, "close window", the processing time of about 500 ms may seem too long for the user. Unlike the long dictated sentences that the user wants to translate into text, a response time of 1000 ms may be acceptable to the user. In fact, in the second case, users may require a longer time because they may feel that their speech is being disturbed by the immediate appearance of the text in response to speech. For longer statements, for example entire paragraphs, the end user may allow even longer response times of several seconds.
[0003] In typical speech recognition systems according to the state of the art, increasing the response time while maintaining recognition accuracy requires increased computational capabilities (processing cycles and / or memory) dedicated to performing speech recognition. As a result, many applications that require short response times require the speech recognition system to run on the same computer as the applications themselves. Despite the fact that such collocations can eliminate the delay that could occur if speech recognition results had to be sent over the network to the requesting application, this type of collocation also has many disadvantages.
[0004] For example, collocation requires that a speech recognition system be installed on every user device, for example, on every desktop computer, laptop, cell phone and personal digital assistant (PDA) that requires speech recognition functionality. Installation and maintenance of such speech recognition systems on such a large number of devices can be persistent and time consuming for the user and system administrators. For example, maintenance requires binary systems to be updated when new versions of speech recognition are introduced. Over time, user data is created on individual devices, such as speech models, which take up valuable space and require synchronization with multiple devices used by the same user. Maintenance can be extremely difficult because users strive to make speech recognition systems work on a larger number and on other devices.
[0005] In addition, the mounting of the speech recognition system on the user's device means that the system occupies valuable computing resources, for example CPU processing cycles, main memory and disk space. These types of resources are usually insufficient on mobile devices such as mobile phones. Generating speech recognition results with fast processing cycles, using this type of device, is carried out at the expense of accuracy and reduction of available resources for other applications working on the same device.
[0006] One known technique that overcomes resource constraints in the context of embedded devices is the delegation of some or all of the speech processing activities to a speech recognition server that is positioned differently in relation to the device embedded, and which has much more computing resources than this type of device. In this case, when the user speaks to the embedded device, the device does not attempt to recognize speech using its resources. Instead, the embedded device sends speech (or its processed form) via a network connection to the speech recognition server, which recognizes speech using larger computing resources, and thus generates faster recognition results than the embedded device, maintaining the same accuracy. The speech recognition server then sends the results back over the network to the embedded device. In the perfect case, this technique provides very good speech recognition results that are obtained faster than the built-in device.
[0007] In practice, the "server-side speech recognition" technique has many disadvantages. In particular, because server-side speech recognition requires a fast and reliable connection, it fails in the absence of such a connection. For example, the potential reduction in time due to speech recognition by the server can be negated if it is implemented using a connection with insufficient bandwidth. For example, a typical network delay when calling HTTP to a remote server may be in the range of 100 to 500 ms. When the spoken data arrives at the speech recognition server after 500 ms after it has been spoken, the server will not be able to generate results quickly enough to meet the minimum processing cycle requirement (500 ms) required for control and control applications. As a result, even a fast speech recognition server will generate results that may seem delayed if connected in a slow network.
[0008] Furthermore, the usual server-side speech recognition techniques assume that the network connection established between the client (e.g. embedded device) and the speech recognition server is constant during the entire recognition process. Although this requirement may be likely to be met on a local area network (LAN) or when both the client and server are managed by the same entity, this requirement may not be met or insufficiently met when the client and server are working in WAN or the Internet, in which case network interference may be widespread and unavoidable.
[0009] In addition, organizations often limit the communication opportunities that their users make to public networks such as the Internet. For example, organizations may allow clients on their own network to make outbound communications. This means that the client can contact the external server on the specified port, but the server cannot initiate contact with the client. This is an example of one-way communication.
[0010] Another common limitation imposed on clients is that they can only use a limited range of outgoing ports to communicate with external servers. In addition, outgoing communication on these ports may require encryption. For example, clients can often use the standard HTTP port (port 80) or the secure HTTPS encrypted standard port (port 443).
[0011] What is needed are improved techniques for generating speech recognition results with a short response time without overloading the limited computing resources of client devices.
The device US 7,330,815 describes a speech recognition system, mainly in terms of language learning, including many clients who can communicate via the Internet with a server having speech recognition services. The client user selects the URL of the file on the server that supports speech recognition, and the client search engine establishes a TCP / IP connection to the Internet and provides a URL using that connection. The user submits a request for an HTML file containing a web page to use for speech processing and the page returns to the client. The client search engine displays the content of the file on the client screen, and from the displayed file the user can choose a speech processing exercise. The Java script associated with the selected exercise activates a search engine element that deals with window-level control to capture user speech. The speech is sent to the server and the server sends the response to the client accordingly. Java script sets the clock associated with querying a search engine item to see if it received a response from the server. The response is sent from the server to Java script, which then displays it to the user. One specific speech recognition engine can be a Microsoft ™ engine called SAPI. It processes the speech packets from the client, but may also exceed a time limit. If SAPI exceeds the time limit, information indicating that the time limit exceeds is saved on the output buffer and returned as text to the client search engine element.
SUMMARY [0012] The speech recognition client sends the speech stream and the control stream in parallel to the server-side speech recognition module via the network. The network can be unreliable and slow. The server-side speech recognition module continuously recognizes the speech stream. The speech recognition client receives the recognition results sent by the server-side recognition module in response to a request from the client. During recognition, the client can remotely configure the status of the server-side speech recognition module.
[0013] Other features and benefits of various objects and embodiments of the invention will become apparent from the following description and the appended claims. The invention is described in independent claims 1, 6, 7 and 9.
BRIEF DESCRIPTION OF THE DRAWINGS [0014] FIG. 1 is a data flow diagram of a system for performing speech recognition in a network with a low delay time according to one embodiment of the invention;
[0015] FIG. 2A is a schematic of the method implemented by the system shown in FIG. 1 according to one embodiment of the invention;
[0016] FIG. 2B is a schematic of a method implemented by an automatic server-side speech recognition module that recognizes a speech segment according to one embodiment of the invention;
[0017] FIG. 2C is a schematic of the method implemented by the server-side automatic speech recognition module as speech recognition fragments on speech segments according to one embodiment of the invention;
[0018] FIG. 2D is a schematic of the method implemented by the server-side recognition module to ensure that the recognition module will be reconfigured after obtaining specific recognition results and before further recognition, according to one embodiment of the invention;
[0019] FIG. 3 is a diagram of a speech stream according to one embodiment of the invention; and [0020] FIG. 4 is a diagram of a command and speech stream according to one embodiment of the invention.
DETAILED DESCRIPTION [0021] Referring to FIG. 1, a data flow diagram of a speech recognition system 100 according to one embodiment of the invention is shown. Referring to FIG. 2A, a flow chart of the method 200 implemented by the system 100 shown in FIG. 1 according to one embodiment of the invention.
[0022] The user 102 of the client device 106 speaks and thus provides speech 104 to the client device 106 (step 202). The client device 106 may be a device such as, for example, a desktop computer or laptop, mobile phone, personal digital assistant (PDA) or telephone. Examples of the invention are particularly useful in conjunction with resource-limited clients, for example, by computers or mobile computing devices with slow processors or a small amount of memory, or computers running software that consumes large amounts of resources. The device 106 can freely receive speech 104 from the user 102, e.g. via a microphone connected to the sound card. Speech 104 may be included in an audio signal that is actually stored on a computer-readable medium and / or transmitted via a network connection or other channel. Speech 104 may, for example, include multiple audio streams, as in the case of push-totalk applications, in which each press initiates a new audio stream.
[0023] The client device 106 includes an application 108, for example a transcription or other application that requires speech recognition 104. Although the application 108 may be any type of application that uses speech recognition results, it should be assumed for the purposes of the following discussion that the application 108 is a live recognition application that transfers speech. Fragments of speech 104 provided by user 102, in this context, may fall into one of the following categories: dictated speech that is subject to transfer (e.g., "patient is a 35-year-old man") or command (for example, "delete this" or " sign and submit ').
[0024] The client device 106 also includes a speech recognition client 140. Despite the fact that the speech recognition client 140 is shown in FIG. 1 as a separate module to the application 108, alternatively the speech recognition client 140 may be part of the application 108. The application 108 provides speech 104 to the speech recognition client 140. Alternatively, the application 108 may somehow process speech and forward the processed version of speech 104 or other speech-derived data to the speech recognition client 140. The speech recognition client 140 can independently process speech 104 (in addition or in exchange for speech-based processing by application 108) by preparing the speech 140 to be sent for recognition.
[0025] Speech recognition client 140 sends speech 104 via network 116 to server-side speech recognition engine 120 located on server 118 (step 204). Although the client 140 can send all speech 104 to server 118 using one server configuration, this may result in less than expected results. To improve recognition accuracy or change the context of the speech recognition engine 120, the client 140 may configure the speech recognition engine 120 at different stages of speech transmission 104, and thus, at different stages of speech recognition 104 by the engine. Basically, the configuration commands sent by the client 140 to the speech recognition engine 120 determine the expectations of the recognition module 120 in terms of context and / or speech content that will be sent next. Various types of systems according to the state of the art implement this configuration function by configuring the server-side recognition engine based on the initial configuration, then send a fragment of speech to the server, then reconfigure the server-side recognition engine, send more speech, and so on. This allows the server-side recognition engine to recognize different speech fragments in configurations and on the basis of contexts that are to ensure the best result for subsequent speech fragments than the one that would be generated at the initial configuration.
[0026] It is undesirable for the speech recognition client 140 to wait for confirmation from the server 118 that the previous configuration command has been processed by the server 118 before sending the next speech fragment 104 to the server 118, since such a requirement would introduce a significant delay for speech recognition 104, in particular when the network connection is slow and / or unreliable. It is also undesirable to stop server-side speech processing until it receives instructions from the client-side application 108 on how to process subsequent speech fragments. In prior art systems, the server requires that speech processing be stopped until client instructions such as reconfiguration commands are received.
[0027] Embodiments of the invention relate to these and other problems. Speech recognition client 140 sends speech 104 to server 118 in speech stream 110 via network 116 (FIG. 2, step 204). As shown in FIG. 3, speech stream 110 may be divided into segments 302a-e, each of which may represent a fragment of speech 104 (e.g., 150-250 ms of speech 104). The segmented speech transfer 104 allows the speech recognition client 140 to send speech fragments 104 to the server 118 relatively quickly after these fragments become available to the speech recognition client 140, thereby enabling the recognition module 120 to start recognizing these fragments with minimal delay. Application 108 may, for example, send the first segment 302a immediately after it becomes available, even when the second segment 302b is generated. In addition, the client 140 may send individual portions of speech stream 110 to the server 118 without using a permanent connection (e.g., socket). As a result, speech recognition client 140 for transmitting speech stream 110 to server 118 may use a connectionless or stateless protocol, such as HTTP.
[0028] Despite the fact that in FIG. 2A shows, for simplicity, only five representative segments 302a-e, in practice speech stream 110 may contain any number of segments that can grow as the user continues speech 102. Application 108 can use any procedure to divide speech 104 into segments or to send speech 104 to server 118 via, for example, an HTTP connection.
[0029] Each of the speech segments 302a-e includes data 304a defining a respective user speech fragment 104. The speech data 304a may be in a suitable format. Each of the speech segments 302a-e may contain different information, e.g., start time 304b and end time 304c of respective speech data 304a and tag 304d, which will be described in more detail below. Specific fields 304a-d shown in FIG. 3 are merely examples and do not constitute a limitation of the invention.
[0030] Basically, server side recognition module 120 queues at server 118 (FIG. 2, step 216) segments from speech stream 110 to form queue 124 first in - first out. With some exceptions, which will be described in detail below, server-side recognition module 120 extracts segments from processing queue 124 as soon as they become available and performs speech recognition on these segments to generate speech recognition results (step 218) that the server 120 queues to queue 134 first in entry - first in exit (step 220).
[0031] The application 108, thanks to the speech recognition client 140, can also send the control stream 112 to the server side recognition module 120 via the network 116 as part of step 204. As shown in FIG. 4, control stream 112 may include control messages 402a-c, transmitted sequentially to recognition module 120. Despite the fact that in FIG. 4 only three representative control messages 402a-c are shown, in practice control stream 112 may contain any number of control messages. As detailed below, each of the control messages 402a may contain a plurality of fields, e.g., command field 404a, specifying the command executed by the server-side recognition module 120, configuration object field 404b, specifying the configuration object, and field 404c, specifying the past due value . Specific fields 304a-d shown in FIG. 3 are merely examples and are not limiting of the invention.
[0032] As shown in FIG. 1, speech recognition client 140 may treat speech stream 110 and control stream 112 as two different data streams (steps 206 and 208) transmitted in parallel from the speech recognition client 140 to the engine 120. However, assuming that only one output port is available for the speech recognition client 140 to establish communication with the server 118, the client 106 may multiplex speech stream 110 and control stream 112 to a single data stream 114 sent to server 118 (step 210). Server 118 demultiplexes signal 114 into components of speech stream 110 and server-side control stream 112 (step 214).
[0033] Any multiplexing scheme can be used. For example, if HTTP is used as the transport mechanism, the HTTP client 130 and HTTP server 132 may transparently perform multiplexing and demultiplexing functions, respectively, for client 106 and server 118. In other words, speech recognition client 140 may treat speech stream 110 and control stream 112 as separate streams even if they are sent as a single multiplexed stream 114 because the HTTP client 130 multiplexes the two streams with each other in an automatic and transparent manner as speech recognition client 140. Similarly, server side recognition module 120 may treat speech stream 110 and control stream 112 as separate streams even if they are received by server 118 as a single multiplexed stream 114 because the HTTP server 132 demultiplexes the combined stream 114 into two streams automatically and transparently for the module 120 server-side recognition.
[0034] Accordingly, by default, server-side recognition module 120 extracts the speech segments from the processing queue 124 in the correct order, performs speech recognition on them, and queues the speech recognition results to the output queue 134. The speech recognition client 108 receives the speech recognition results. Speech recognition client 140 sends, in control stream 112, a control message whose command field 404a calls a method called "DecodeNext" here. This method adopts configuration update object 404b (which determines how the configuration state 126 of the server-side recognition module 120 is updated) and the real-time deadline value 404c. Despite the fact that the speech recognition client 140 can send other commands in control stream 112, hereinafter only the DecodeNext command will be described for the purpose of explanation.
[0035] Server-side recognition module 120 extracts control messages sequentially from control stream 112 as soon as they are received and in parallel with processing speech segments in speech stream 110 (step 222). Server-side recognition module 120 executes the command in each control message in turn (step 224).
[0036] Looking at FIG. 2B, a schematic diagram of the method implemented by the server-side recognition module 120 executing a DecodeNext command contained in control stream 112 is shown. If there is at least one speech recognition result in output queue 134 (step 240), the recognition module 120 sends subsequent results 122) queued 134 to a speech recognition client 140 via network 116 (step 242). If more than one result is available in queue 134 at the time step 242 is performed, all available results in queue 134 are sent in the result stream 122 to the speech recognition client 140. (Although the results 122 are shown in FIG. 1 as transmitted directly from the recognition module 120 to the speech recognition client 140, for clarification, the results 122 can be sent via HTTP server 132 via network 116 to HTTP client 130 on client device 106). The DecodeNext method then returns control to application 108 (step 246) and exits.
[0037] It should be remembered that the recognition module 120 constantly performs speech recognition on the speech segments located in the processing queue 124. Therefore, if the output queue 134 is empty when the recognition module 120 starts the implementation of the DecodeNext method, the DecodeNext method it blocks until at least one result is available in output queue 134 (e.g. one word) or until the amount of time determined by the expiration value 404c is reached (step 248). If the result appears in output queue 134 before the expiration value 404c is reached, the DecodeNext method sends this result to the speech recognition client 140 (step 242), returns control to the speech recognition client 140 (step 246) and exits. If the result does not appear in output queue 134 before the 404c overdue value is reached, then the DecodeNext method informs the speech recognition client 140 that there are no results available (step 244), returns control to the speech recognition client 140 (step 246) and exits without returning any recognition results to the client 140 speech recognition.
[0038] When the control returns to the speech recognition client 140 (after the DecodeNext method either returns the recognition result to the speech recognition client 140 or informs the speech recognition client 140 that such results are not available), the speech recognition client 140 can immediately send another DecodeNext message to server 120 trying to get another recognition result. Server 120 may process a DecodeNext message as described above with reference to FIG. 2B. This process may be repeated in the case of subsequent recognition results. As a result, control stream 112 can generally always block on the server side (in the loop described in steps 240 and 248, in FIG. 2B), pending recognition results, and return them to client application 108 as soon as they become available.
[0039] The expiration value 404c may be selected to be shorter than the expiration value of the underlying communication protocol used between the client 140 and the server 120, e.g. As a result, if the client 140 receives a notification from the server that no recognition results were generated before the overdue value 404c was reached, the client 140 may conclude that the overdue was due to the server 120's inability to generate any recognition results before the 404c overdue value was reached , not a problem related to the communication network. Regardless of the reason for expiration, the client 140, after expiration, can send another DecodeNext message to server 120.
[0040] The above described examples include two completely unsynced data streams 110 and 112. However, it may be desirable to perform certain types of synchronization on two data streams 110 and 112. For example, it may be useful for the speech recognition client 140 to ensure that the recognition module 120 is in a specific configuration state prior to recognizing the speech stream 110. For example, recognition module 120 may use the text context of the current cursor position in the text editing window to guide recognition of the text that is entered at that cursor position. Because the cursor position can often change due to mouse or keyboard movements, it may be beneficial for the application 108 to delay the transmission of the text context to the server 120 until the user 102 presses the "start recording" button. In this case, the server-side recognition module 120 must be protected from starting speech recognition transmitted to the server 120 until the server 120 receives the correct text context and the server 120 updates its configuration state 126.
[0041] In another example, some recognition results may trigger the need to change configuration state 126 of recognition module 120. As a result, when the server-side recognition module 120 generates this type of result, it must wait until reconfigured before generating another result. For example, if the recognition module 120 generates a "delete all" result, the application 108 may then attempt to verify the user's intent by querying user 102: "Do you really want to delete everything? "Answer YES or NO". In this case, application 108 (through the speech recognition client 140) should reconfigure the recognition module 120 relative to the "YES / NO" response before the recognition module 120 attempts to recognize the next segment in the speech stream 110.
[0042] These results can be obtained as shown in the diagram of FIG. 2C, which shows a method that can be implemented by server-side recognition module 120 as part of the implementation of speech recognition on audio segments in the processing queue (FIG. 2A, step 218). Each recognition state of the recognition module is assigned a unique configuration state identifier (ID). The speech recognition client 140 assigns integer values to configuration state IDs, e.g., ID1> ID2, then the configuration status associated with ID1 is more current than the configuration status associated with ID2. Accordingly, with reference to FIG. 3, speech recognition client 140 also provides tags 304d in each of the speech stream segments 302a-e, which show the minimum configuration state ID number required that is required before recognition of such segment.
[0043] When the server side recognition module 120 searches for the next audio segment from the processing queue 124 (step 262), the recognition module 120 compares the configuration state ID 136 of the current configuration state 126 of the recognition module with the minimum required configuration ID determined by the searched segment audio tag 304d . If the current configuration ID 136 is at least as large as the minimum required configuration ID (step 264), then the server 120 starts recognizing the searched audio segment (step 266). Otherwise, the server 120 waits until the configuration ID 136 reaches the minimum required ID before it begins recognizing the current speech segment. Because the method shown in FIG. 2C may be implemented in parallel with the method 200 shown in FIG. 2A, ID 136 configuration of the server side recognition module 120 can be updated by performing control messages 224, even when the method shown in FIG. 2C will lock in the loop of step 264. Furthermore, it should be remembered that although server 120 waits for speech processing from processing queue 124, server 120 continues to receive additional segments from speech stream 110 and queues these segments in processing queue 124 (FIG. 2A, step 214-216).
[0044] In another example, in which speech stream 110 can be synchronized with control stream 112, the application 108, through speech recognition client 140, may first instruct the recognition module 120 to interrupt recognition of speech stream 110 or take other actions after generating any result recognition or after generating a recognition result that will meet certain criteria. Such criteria can effectively serve as breakpoints that the application 108, through the speech recognition client 140, can use to proactively control how far ahead the recognition module 120 will generate recognition results.
[0045] For example, consider the context in which the user 102 may issue any of the following voice commands: "delete", "next", "select" and "open file search engine" In this context, a possible configuration that can be determined by configuration update object 404b, would be: <delete, continue>, <next, continue>, <select all, continue>, <open file search engine, stop>. This type of configuration instructs the server-side recognition module 120 to continue recognizing the speech stream 110 after obtaining the "delete", "next" or "select all" recognition results and stop recognizing the speech stream 110 after obtaining the "open file search" recognition result. The reason for this configuration of recognition module 120 is that generating a "delete", subsequent "or" select all "result does not require the recognition module 120 to be reconfigured before generating another result. Therefore, recognition module 120 may continue to recognize speech stream 110 after generating any of the "delete", subsequent "or" select all results, allowing recognition module 120 to continue speech recognition 104 at full speed (see FIG. 2D, step 272 ). In contrast, generating the "open file search" result requires that the recognition module 120 be reconfigured (e.g. to expect results such as "OK", select 1.xml file "or" New Folder ") before recognizing subsequent segments in speech stream 110 (see FIG. 2C, step 274). Therefore, if the application 108, through the speech recognition client 140, is informed by the recognition module 120 that a "open file viewer" result has been generated, the application 108, through the speech recognition client 140, can configure the recognition module 120 to a configuration state that is appropriate for controlling the file search engine. Allowing the application 108 to pre-configure the recognition module 120 restores the balance between maximizing the recognition module response time and ensuring that the recognition module 120 uses the appropriate configuration state to recognize the different speech fragments 104.
[0046] Note that even if the recognition module 120 interrupts speech recognition from the processing queue 124 as a result of the "stop" configuration command (step 274), the recognition module 120 may continue to receive speech segments from speech stream 110 and queue these segments in queue 124 processing (FIG. 2A, step 214, 216). As a result, the additional segments of the speech stream 110 are ready to be processed as soon as the recognition module 120 resumes performing speech recognition.
[0047] Accordingly, the technique presented herein can be used in conjunction with unidirectional communication protocols, for example HTTPS. These types of communication protocols are easy to set up in wide area networks, but offer poor protection against malfunctions. The faults may occur during a request between client 130 and server 132 that may leave the application 108 in an ambiguous state. For example, the problem may arise when a fault occurs halfway through a call on either side (client application 108 or server side recognition module 120). Other problems may occur, for example, due to lost messages sent to or from server 118, messages arriving at client 106 or server 118 out of order, or messages incorrectly sent as duplicates. In general, in prior art systems, the speech recognition client 140 is responsible for ensuring the stability of the entire system 100, because the underlying communication protocol does not guarantee such stability.
[0048] Embodiments of the invention are resistant to such problems due to the fact that all messages and events exchanged between the speech recognition client 140 and the speech recognition module 120 are idempotent. An event is idempotent if several instances of the same event give the same result as a single occurrence of such an event. For this reason, if the speech recognition client 140 detects a fault, e.g., a command forwarding fault to the server side recognition module 120, the speech recognition client 140 may resend the command either immediately or after a suitable downtime. The speech recognition client 140 and recognition module 120 may use an application program communication interface (API) that ensures that retrying will allow the system 100 to remain in a consistent state.
[0049] In particular, the API for speech stream 110 forces the speech recognition client 140 to pass the speech stream 110 in segments. Each segment can have a unique ID 304e except the index of the initial byte 304b (initially 0 for the first segment) and either the end byte 304c or the segment size. Server-side recognition module 120 can confirm that it has received a segment by returning the segment's end byte index, which should be equal to the start byte plus the segment size. The end byte index sent by the server may be a lower value if the server could not read the entire audio segment.
[0050] The speech recognition client 140 then sends another segment starting at the point where it terminated server-side recognition module 120, so that the new start byte index corresponds to the end byte index returned by the recognition module 120. This process is repeated for the entire speech stream 110. If the message is lost (on the way to or from server 118), the speech recognition client 140 repeats the transfer. If the server side recognition module 120 does not receive this speech segment before, the server side recognition module 120 will simply process the new data. If the recognition module 120 has previously processed such a segment (which may have occurred in the event of loss of results back to the client 106), then the recognition module 120 may, for example, acknowledge receipt of the segment and abandon it without reprocessing.
[0051] For control stream 112, all control messages 402a-c may be present on server 118, because each message may contain the ID of the current session. For the DecodeNext method, the speech recognition client 140 may pass, as part of the DecodeNext method, the current unique identifier to identify the current method call. Server 118 tracks this type of identifier to determine if the current message received in control stream 112 is new or has already been received and processed. If the current message is new, the recognition module 120 normally processes the message as described above. If the current message has already been processed, recognition module 120 may re-deliver previously returned results in exchange for generating them again.
[0052] If one of the control messages 402a-c has been sent to server 118 and the server 118 has not acknowledged receipt of the control message, the client 140 may save the control message. When the client 140 has a second message to send to server 118, the client 140 can send both the first (unconfirmed) control message and the second control message to server 118. The client 140 may alternatively obtain the same result by combining the status changes of the first and second control messages into a single control message, which the client 140 may then send to server 140. In this way, the client 140 may combine any number of control messages into the form of a single control message. , until the messages are acknowledged by server 118. Similarly, server 118 may combine speech recognition results that have not been confirmed by the client 140 into individual results in the results stream 122 until such results are confirmed by the client.
[0053] The advantages of the invention include the following features. Embodiments of the invention allow the spread of speech recognition anywhere on the Internet, without the need for any special network. In particular, the techniques disclosed herein may operate using unidirectional communication protocols, e.g., HTTP, thereby enabling operation even in restrictive environments in which clients are limited to establishing an outgoing (unidirectional) connection only. As a result, embodiments of the invention are useful for various types of networks without sacrificing security. In addition, the techniques disclosed herein may use existing network security mechanisms (e.g., SSL and in the development of HTTPS) to ensure secure communication between client 106 and server 118.
[0054] Accordingly, a common limitation imposed on clients is that they can only use a limited range of outgoing ports to communicate with external servers. Examples of the invention can be implemented in such systems by multiplexing speech stream 110 and control stream 112 to a single stream 114 that can be transmitted through a single port.
[0055] In addition, outgoing communication may require encryption. For example, clients can often use the secure encrypted HTTPS standard port (port 443). Embodiments of the invention can work in either (unsecured) HTTP or secured HTTPS standard for all communication needs - both audio transfer 110 and control flow 112. As a result, the techniques disclosed herein can be used in conjunction with systems that allow clients to communicate using an unsecured HTTP protocol and systems that require or enable clients to communicate using the secure HTTPS protocol.
[0056] The techniques disclosed herein are also flexible with respect to temporary network faults because they use a communication protocol in which messages are idempotent. This is particularly useful when embodiments of the invention are used for networks such as WAN, in which dips and network pulses are common. Although such cases may lead to defects in conventional server-side speech recognition systems, they do not affect the results generated by embodiments of the invention (except for the possibility of increasing processing time).
[0057] Embodiments of the invention allow speech 104 from client 106 to be sent to server 118 as fast as network 116 allows, even if server 118 cannot process speech continuously. In addition, server-side recognition module 120 can process speech from processing queue 124 as quickly as possible, even when network 116 cannot send results and / or application 108 is not ready to receive results. These and other features of the embodiment of the invention allow speech and speech recognition results to be transmitted and processed as soon as the individual components of the system 100 allow, so that problems occurring in individual components of the system 100 only have a minimal effect on the performance of other components of the system 100 .
[0058] In addition, embodiments of the invention allow server-side recognition module 120 to process speech as quickly as possible, but without going too far ahead of client application 108. As described above, application 108 may use control messages in control stream 112 to generate reconfiguration commands to recognition module 120, which cause recognition module 120 to reconfigure itself to recognize speech in the appropriate configuration state and temporarily suspend recognition upon the occurrence of a particular condition. so that application 108 can reconfigure the status of recognition module 120 accordingly. These types of techniques allow speech recognition to be carried out as quickly as possible, without using defective configuration states.
[0059] It should be remembered that although the invention has been described above with reference to specific embodiments, the above examples are provided for illustrative purposes only and do not limit or define the scope of the invention. Various other embodiments, including but not limited to the features below, are also within the scope of the claims. For example, the elements and components described here can be further divided into further elements or combined to create fewer elements to perform the same functions.
[0060] As described above, the various methods carried out by the embodiments of the invention can be carried out in parallel with others, in whole or in part. Those skilled in the art will recognize the possibility of using the fragments of the methods disclosed herein in various combinations to obtain the stated benefits.
[0061] The techniques described above can be implemented, for example, in hardware, software, firmware or combinations thereof. The above techniques can be implemented in one or more computer programs working on a programmable computer equipped with a processor, on a processor-readable medium (including for example volatile and non-volatile memory and / or memory elements), at least one input device and at least one output device. Program parameters can be applied to the parameters entered using the device to perform the described functions and generate a result. The result can be sent to one or more output devices.
[0062] Any computer program within the scope of the following claims may be implemented in any programming language, for example, a symbolic language, an internal language, a high-level programming language or an object-oriented programming language. The programming language can be, for example, a compiled or interpreted language.
[0063] Any such computer program may be implemented into a computer program product actually installed on a device having computer readable memory for execution by a computer processor. The steps of the method of the invention can be performed by a computer processor implementing the program actually installed on a computer readable medium to perform the functions of the invention by working on the input data and generating a result. Suitable processors include, for example, both general and special microprocessors. Basically, the processor receives instructions and data from read memory and / or random access memory. The memories specific to the actual recording of computer program instructions include all forms of volatile memory, for example semiconductor memories, including EPROM, EEPROM and flash devices; magnetic disks, for example internal hard disks and removable disks; magneto-optical discs and CD-ROMs. All the above-mentioned issues can be supplemented with or implemented in special ASIC systems application-specific integrated circuits, or integrated circuits for specific applications) or FPGA (Field-Programmable Gate Arrays). The computer can also receive programs and data from storage media such as an internal disk (not shown) or a removable disk. These elements can also be found in an ordinary desktop computer as well as in other computers suitable for operating computer programs using the methods described herein that can be used in conjunction with any digital printing or marking engine, monitor or other raster output device capable of generating colored pixels or pixels in grayscale on paper, film, screen or other output media.
Prepared and verified
Anna Stenzel Patent Attorney
24 members in 8 offices
Priority claims8
| Document | Office | Kind | Date |
|---|---|---|---|
| 9322108 | United States of America | P | |
| 55038109 | United States of America | A | |
| 09810710 | European Patent Office (EPO) | A | |
| 2009055480 | United States of America | W | |
| EP20090810710 | – | – | – |
| US20080093221P | – | – | – |
| US20090550381 | – | – | – |
| WO2009US55480 | – | – | – |
Members24
| Document | Office | Kind | |
|---|---|---|---|
| CA2732256A1 | Canada | A1 | |
| US2010057451A1 | United States of America | A1 | |
| WO2010025441A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2010025441A3 | World Intellectual Property Organization (WIPO) | A3 | |
| WO2010025441A3 | World Intellectual Property Organization (WIPO) | A3 | |
| EP2321821A2 | European Patent Office (EPO) | A2 | |
| US8019608B2 | United States of America | B2 | |
| US2011288857A1 | United States of America | A1 | |
| JP2012501481A | Japan | A | |
| US8249878B2 | United States of America | B2 | |
| US2012296645A1 | United States of America | A1 | |
| EP2321821A4 | European Patent Office (EPO) | A4 | |
| US8504372B2 | United States of America | B2 | |
| EP2321821B1 | European Patent Office (EPO) | B1 | |
| DK2321821T3 | Denmark | T3 | |
| ES2446667T3 | Spain | T3 | |
| JP2014056258A | Japan | A | |
| PL2321821T3This record | Poland | T3 | |
| US2014163974A1 | United States of America | A1 | |
| JP5588986B2 | Japan | B2 | |
| US2015170647A1 | United States of America | A1 | |
| JP5883841B2 | Japan | B2 | |
| US9502033B2 | United States of America | B2 | |
| CA2732256C | Canada | C |
Numbers
- Publication, DOCDB
- 2321821
- Publication, EPODOC
- PL2321821T
- Application
- 810710
- Application, DOCDB
- 09810710
- Application, EPODOC
- PL20090810710T
Titles2
- English
- DISTRIBUTED SPEECH RECOGNITION USING ONE WAY COMMUNICATION
- Polish
- Rozproszone rozpoznawanie mowy z użyciem komunikacji jednokierunkowej
Classification
- CPC, 3
- G10L15/30
- G10L15/22
- G10L15/32
- IPC, 1
- G10L15 30