Hotword detection on multiple devices
Abstract
Disclosed are methods, systems, and devices that include computer programs encoded on computer recording media for hotword detection on multiple devices. In one aspect, the method comprises the operation of receiving audio data corresponding to an utterance by a first computing device. The method further includes the action of determining a first value corresponding to the possibility that the utterance contains a hot word. The method further comprises the action of receiving a second value corresponding to the possibility that the utterance contains a hotword, the second value being determined by the second computing device. The method further includes the operation of comparing the first value with the second value. The method further includes an operation of initiating a speech recognition process on the audio data based on the result of comparison between the first value and the second value.

Term
9 yearsto projected expiry
Projected expiry 29 September 2035, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
20 claims: 5 independent, 15 dependent
- 1コンピュータによって実施される方法であって、 第1コンピューティングデバイスにより、発話に対応するオーディオデータを受信するステップと、 前記発話がホットワードを含む可能性に対応する第1の値を決定するステップと、 前記発話が前記ホットワードを含む可能性に対応する第2の値を受信するステップであって、前記第2の値は第2コンピューティングデバイスによって決定される、ステップと、 前記第1の値と前記第2の値とを比較するステップと、 前記第1の値と前記第2の値との比較結果に基づいて、前記オーディオデータに対する音声認識処理を開始するステップと を有する方法。
- 2前記第1の値がホットワードスコアしきい値に達していると判定するステップ をさらに有する、請求項1に記載の方法。
- 3前記第1の値を前記第2コンピューティングデバイスに送信するステップ をさらに有する、請求項1に記載の方法。
- 4前記第1の値と前記第2の値との比較結果に基づいて、前記第1コンピューティングデバイスの起動状態を決定するステップ をさらに有する、請求項1に記載の方法。
- 5前記第1の値と前記第2の値との比較結果に基づいて、前記第1コンピューティングデバイスの起動状態を決定する前記ステップが、 前記起動状態がアクティブ状態であると判定するステップ を含む、請求項4に記載の方法。
- 6前記第1コンピューティングデバイスにより、追加的な発話に対応する追加的なオーディオデータを受信するステップと、 前記追加的な発話が前記ホットワードを含む可能性に対応する第3の値を決定するステップと、 前記追加的な発話が前記ホットワードを含む可能性に対応する第4の値を受信するステップであって、前記第4の値は第3コンピューティングデバイスによって決定される、ステップと、 前記第3の値と前記第4の値とを比較するステップと、 前記第3の値と前記第4の値との比較結果に基づいて、前記第1コンピューティングデバイスの起動状態が非アクティブ状態であると判定するステップと をさらに有する、請求項1に記載の方法。
- 7前記第1の値を前記第2コンピューティングデバイスに送信する前記ステップが、 前記第1の値を、ローカルネットワークを介して又は短距離無線通信を介して、サーバに送信するステップ を含み、 前記発話が前記ホットワードを含む可能性に対応する第2の値を受信するステップであって、前記第2の値は第2コンピューティングデバイスによって決定される、前記ステップが、 前記第2コンピューティングデバイスによって決定された前記第2の値を、前記ローカルネットワークを介して又は前記短距離無線通信を介して、前記サーバから受信するステップ を含む、請求項3に記載の方法。
- 8前記第2コンピューティングデバイスを識別するステップと、 前記ホットワードを含む発話に応答するように前記第2コンピューティングデバイスが構成されていると判定するステップと をさらに有する、請求項1に記載の方法。
- 9前記第1の値を前記第2コンピューティングデバイスに送信する前記ステップが、 前記第1コンピューティングデバイスに対する第1識別子を送信するステップ を含み、 前記発話が前記ホットワードを含む可能性に対応する第2の値を受信するステップであって、前記第2の値は第2コンピューティングデバイスによって決定される、前記ステップが、 前記第2コンピューティングデバイスに対する第2識別子を受信するステップ を含む、請求項3に記載の方法。
- 10前記起動状態がアクティブ状態であると判定する前記ステップが、 前記発話に対応する前記オーディオデータを受信してから所定の時間が経過したと判定するステップ を含む、請求項5に記載の方法。
- 11前記起動状態がアクティブ状態であると判定したことに基づいて、所定の時間、前記第1の値を送信し続けるステップ をさらに有する、請求項5に記載の方法。
- 12コンピューティングデバイスであって、 命令を格納した1つ又は複数のストレージデバイスを具備し、 前記命令は、前記コンピューティングデバイスによって実行されたとき、前記コンピューティングデバイスに、 第1コンピューティングデバイスにより、発話に対応するオーディオデータを受信する手順と、 前記発話がホットワードを含む可能性に対応する第1の値を決定する手順と、 前記発話が前記ホットワードを含む可能性に対応する第2の値を受信する手順であって、前記第2の値は第2コンピューティングデバイスによって決定される、手順と、 前記第1の値と前記第2の値とを比較する手順と、 前記第1の値と前記第2の値との比較結果に基づいて、前記オーディオデータに対する音声認識処理を開始する手順と を含む動作を実行させる、デバイス。
- 13前記動作が、 前記第1の値がホットワードスコアしきい値に達していると判定する手順 をさらに含む、請求項12に記載のデバイス。
- 14前記動作が、 前記第1の値を前記第2コンピューティングデバイスに送信する手順 をさらに含む、請求項12に記載のデバイス。
- 15前記動作が、 前記第1の値と前記第2の値との比較結果に基づいて、前記第1コンピューティングデバイスの起動状態を決定する手順 をさらに含む、請求項12に記載のデバイス。
- 16前記第1の値と前記第2の値との比較結果に基づいて、前記第1コンピューティングデバイスの起動状態を決定する前記手順が、 前記起動状態がアクティブ状態であると判定する手順 を含む、請求項15に記載のデバイス。
- 17前記動作が、 前記第1コンピューティングデバイスにより、追加的な発話に対応する追加的なオーディオデータを受信する手順と、 前記追加的な発話が前記ホットワードを含む可能性に対応する第3の値を決定する手順と、 前記追加的な発話が前記ホットワードを含む可能性に対応する第4の値を受信する手順であって、前記第4の値は第3コンピューティングデバイスによって決定される、手順と、 前記第3の値と前記第4の値とを比較する手順と、 前記第3の値と前記第4の値との比較結果に基づいて、前記第1コンピューティングデバイスの起動状態が非アクティブ状態であると判定する手順と をさらに含む、請求項12に記載のデバイス。
- 18前記第1の値を前記第2コンピューティングデバイスに送信する前記手順が、 前記第1の値を、ローカルネットワークを介して又は短距離無線通信を介して、サーバに送信する手順 を含み、 前記発話が前記ホットワードを含む可能性に対応する第2の値を受信する手順であって、前記第2の値は第2コンピューティングデバイスによって決定される、前記手順が、 前記第2コンピューティングデバイスによって決定された前記第2の値を、前記ローカルネットワークを介して又は前記短距離無線通信を介して、前記サーバから受信する手順 を含む、請求項14に記載のデバイス。
- 19前記動作が、 前記第2コンピューティングデバイスを識別する手順と、 前記ホットワードを含む発話に応答するように前記第2コンピューティングデバイスが構成されていると判定する手順と をさらに含む、請求項12に記載のデバイス。
- 201つ又は複数のコンピュータによって実行可能な命令を含むソフトウェアを格納した、非一時的なコンピュータ読み取り可能な記録媒体であって、 前記命令は、その実行時に、前記1つ又は複数のコンピュータに、 第1コンピューティングデバイスにより、発話に対応するオーディオデータを受信する手順と、 前記発話がホットワードを含む可能性に対応する第1の値を決定する手順と、 前記発話が前記ホットワードを含む可能性に対応する第2の値を受信する手順であって、前記第2の値は第2コンピューティングデバイスによって決定される、手順と、 前記第1の値と前記第2の値とを比較する手順と、 前記第1の値と前記第2の値との比較結果に基づいて、前記オーディオデータに対する音声認識処理を開始する手順と を含む動作を実行させる、非一時的なコンピュータ読み取り可能な記録媒体。
Independent claims20
59 paragraphs, as filed
0001Generally, this specification relates to a system and a technique for recognizing a spoken word, which is also referred to as speech recognition.
0002A real-life voice-enabled home or other environment, i.e., the user only needs to issue a query or command, and a computer-based system receives and answers the query and / or triggers the execution of the command. The environment exists. A voice-operated environment (eg, home, work, school, etc.) can be implemented by using a network of linked microphone devices distributed across various spaces or areas of the environment. Through such a microphone network, users can verbally contact the system from almost anywhere in the environment without having to have a computer or other device in front of or near them. For example, even while cooking in the kitchen, the user asks the system "how many milliliters in three cups?" And in response the system gives an answer, eg synthetic speech output. Can be received in the form of. Alternatively, the user tells the system, "when does my nearest gas station. You can ask the question "close)" or when you are preparing to go out, you can ask the question "should I wear a coat today?".
0003In addition, users can query the system and / or issue commands regarding their personal information. For example, a user may ask the system "when is my meeting with John?" Or "remind me to call John when I get home." You can tell the system to call John when I get back home).
<p num="0004"> The method by which the user interacts with the system for a system capable of voice operation is not limited to this, and is designed mainly by voice input. As a result, a system that potentially picks up all utterances that occur in the surrounding environment, including those that are not directed to the system, is one in which any given utterance is directed, for example, to someone present in the environment. There must be some means of identifying that it is directed at the system. One way to achieve this is to use hotwords. Hotwords are reserved as predetermined words issued to draw the attention of the system by arrangement with users in the environment. In one example of the environment, the hotword used to draw the attention of the system is "OK Computer (OK). computer) ". Therefore, every time the word "OK computer" is spoken, it is picked up by the microphone and transmitted to the system. The system runs a speech recognition technique to determine if a hotword has been issued and, if so, waits for a subsequent command or query. Therefore, utterances directed at the system take the general form of [HOTWORD] [QUERY]. Here, "HOTWORD" in this example is "OK computer", and "QUERY" is any question that can be voice-recognized, analyzed, and performed by the system alone or in collaboration with the server over the network. It may be a command, notification, or other request.</p><p num="0005"> According to an innovative aspect of the subject matter described herein, the user device receives utterances uttered by the user. The user device determines whether the utterance contains a hot word and calculates a hot word reliability score indicating that the utterance may contain a hot word. The user device sends this score to other nearby user devices. Perhaps other user devices are receiving the same utterance. Other user devices calculate hotword reliability scores and send those scores to the user device. The user device compares the hotword reliability scores. If the user device has the highest hotword reliability score, it will continue to operate and be prepared to process additional audio. The user device does not process additional audio if it does not have the highest hotword reliability score.</p><p num="0006"> In general, another innovative aspect of the subject matter described herein may be implemented as a method, the method being the step of receiving audio data corresponding to an utterance by a first computing device and the utterance. Is the step of determining the first value corresponding to the possibility that the utterance contains the hot word, and the step of receiving the second value corresponding to the possibility that the utterance contains the hot word, and the second value is the second value. Speech recognition processing for audio data based on the steps determined by the computing device, the step of comparing the first and second values, and the result of the comparison of the first and second values. Has a step to start with.</p><p num="0007"> These and other embodiments may selectively include one or more of the following features, respectively. The method further comprises a step of determining that the first value has reached the hotword score threshold. The method further comprises the step of transmitting the first value to the second computing device. The method further comprises a step of determining the boot state of the first computing device based on the result of comparison between the first value and the second value. The step of determining the boot state of the first computing device based on the result of comparison between the first value and the second value further includes a step of determining that the boot state is the active state. The method is the step of receiving additional audio data corresponding to the additional utterance by the first computing device and the step of determining the third value corresponding to the possibility that the additional utterance contains a hot word. And the step of receiving a fourth value corresponding to the possibility that the additional utterance may contain a hotword, the fourth value being determined by the third computing device, the step and the third value. And the step of comparing the fourth value with, and the step of determining that the boot state of the first computing device is inactive based on the comparison result of the third value and the fourth value. Have.</p><p num="0008"> The step of transmitting the first value to the second computing device further includes the step of transmitting the first value to the server via a local network or short-range wireless communication. The step of receiving a second value corresponding to the possibility that the utterance contains a hotword, the second value being determined by the second computing device, the step being determined by the second computing device. It further includes the step of receiving the second value from the server via the local network or via short-range wireless communication. The method further comprises a step of identifying the second computing device and a step of determining that the second computing device is configured to respond to an utterance containing a hotword. The step of transmitting the first value to the second computing device further includes the step of transmitting the first identifier to the first computing device. The step of receiving a second value corresponding to the possibility that the utterance contains a hotword, the second value is determined by the second computing device, the step is the second identifier for the second computing device. Includes further steps to receive. The step of determining that the activated state is the active state further includes a step of determining that a predetermined time has elapsed since receiving the audio data corresponding to the utterance. The method further comprises a step of continuing to transmit the first value for a predetermined time based on determining that the activation state is the active state.</p><p num="0009"> Another embodiment of this aspect includes a computer program recorded on a corresponding system, device, and computer storage device, each configured to perform the operation of the method described above.</p>
<p num="0010"> Certain embodiments of the subject matter described herein may be implemented to achieve one or more of the following advantages: Multiple devices can detect the hotword and only one device will respond to the hotword.</p><p num="0011"> Details of one or more embodiments of the subject matter described herein will be described by the accompanying drawings and the detailed description of the invention below. Other features, aspects, and advantages of the subject matter become apparent from the detailed description of the invention, the drawings, and the description of the claims.</p>
0012<figref num="1">It is a figure which shows an example of the system for hot word detection.</figref><figref num="2">It is a figure which shows an example of the process for hot word detection.</figref><figref num="3">It is a figure which shows an example of a computing device and a mobile computing device.</figref>
0013In the figure, similar symbols refer to similar elements.
0014In the not too distant future, many devices will always listen to hotwords. If a user has multiple devices trained to respond to his or her voice (eg, phones, tablet computers, TVs, etc.), the user is hot with a device that he or she did not intend to speak to. You would want to suppress the response to the ward. For example, if a user issues a hotword to one device and another device of the user is nearby, they will probably also trigger a voice search. In many cases, such behavior is not intended by the user. Therefore, it can be an advantage that only one device, especially the device spoken by the user, is started. This specification deals with the problem of selecting an appropriate device for responding to a hot word and suppressing the reaction of other devices to the hot word.
0015FIG. 1 is a diagram showing an example of a system 100 for hot word detection. In general, system 100 illustrates a situation in which user 102 makes an utterance 104 detected by the microphones of computing devices 106,108,110. The computing devices 106,108,110 process the utterance 104 to determine the possibility that the utterance 104 contains a hotword. The computing devices 106, 108, 110, respectively, transmit data to each other indicating that the utterance 104 may contain a hot word. The computing devices 106, 108, 110 compare their data, respectively, and the computing device that calculates the highest possibility that the utterance 104 contains a hot word starts voice recognition for the 104 in the utterance. The computing device for which utterance 104 has not calculated the highest probability of including a hot word does not initiate speech recognition for the speech following utterance 104.
0016Computing devices located close to each other identify each other before sending data indicating that utterance 104 may correspond to a hotword to other computing devices. In some embodiments, computing devices identify each other by searching the local network for other devices that are configured to respond to hotwords. For example, the computing device 106 may search the local area network for other devices configured to respond to hotwords to identify the computing device 108 and the computing device 110.
0017In some embodiments, the computing device identifies another nearby computing device that is configured to respond to a hotword by identifying the user who has logged in to each device. For example, user 102 is logged in to three computing devices 106,108,110. User 102 holds the computing device 106 in his hand. The computing device 108 is placed on a table and the computing device 110 is located near the wall. The computing device 106 detects the computing devices 108,110, and each computing device shares information about the user logged in to the computing device, such as a user identifier. In some embodiments, the computing device becomes a hotword by identifying the computing device that is configured to respond when the hotword is issued by the same user through speaker identification. It may identify other nearby computing devices that are configured to respond. For example, user 102 responds to his or her own voice when he or she utters a hotword, respectively, in computing devices 106,108, Configure 110. Computing devices share speaker identification information by providing each other with user identifiers for user 102. In some embodiments, the computing device may identify other computing devices that are configured to respond to hotwords via short-range wireless communication. For example, the computing device 106 may transmit a signal over short-range wireless communication to search for other computing devices configured to respond to hotwords. Computing devices may utilize one or a combination of these technologies to identify other computing devices that are configured to respond to hotwords.
0018When computing devices 106,108,110 identify other computing devices that are configured to respond to hotwords, they share and store device identifiers for the identified computing devices. The identifier may be based on the type of device, the IP address of the device, the MAC address, the name given to the device by the user, or any similar uniquely defined identifier. For example, the device identifier 112 for the computing device 106 may be "phone". The device identifier 114 for the computing device 108 may be "tablet". The device identifier 116 for the computing device 110 may be "thermostat". Computing devices 106,108,110 store device identifiers for other computing devices that are configured to respond to hotwords. Each computing device has a device group that stores the device identifier. For example, the computing device 106 has a device group 118 listing "tablet" and "thermostat" as the two devices it calculates that will receive the possibility that the audio data will contain hotwords. The computing device 108 has a device group 120 that lists "phone" and "thermostat" as two devices that it calculates that the audio data will receive the possibility of containing hotwords. The computing device 110 has a device group 122 that lists "phone" and "tablet" as two devices that receive the possibility that the audio data contains hotwords, which it calculates.
0019When the user 102 makes an utterance 104, i.e. speaks "OK computer", each computing device with a microphone near the user 102 detects and processes the utterance 104. Each computing device detects utterance 104 via an audio input device such as a microphone. Each microphone provides audio data to individual audio subsystems. Individual audio subsystems buffer, filter, and digitize audio data. Also, in some embodiments, each computing device may perform termination determination and speaker identification on the audio data. The audio subsystem provides the processed audio data to the hotword section. The hotword section compares the processed audio data with known hotword data to calculate a reliability score indicating the likelihood that the utterance 104 corresponds to the hotword. The hotword section can extract audio features such as filter bank energy or mel frequency cepstrum coefficients from the processed audio data. The hotword section may use a classification window and process these audio features, such as by using a support vector machine or neural network. Based on the processing of the audio features, the hotword section 124 calculates the reliability score 0.85, the hotword section 126 calculates the reliability score 0.6, and the hotword section 128 calculates the reliability score 0.45. .. In some embodiments, the confidence score may be normalized to the range of zero to one, where higher numbers indicate higher reliability, with utterance 104 containing hotwords.
0020Each computing device sends an individual reliability score data packet to other computing devices in the device group. Each reliability score data packet contains an individual reliability score and an individual device identifier for the computing device. For example, the computing device 106 transmits a reliability score data packet 130 containing a reliability score of 0.85 and the device identifier "phone" to a computing device in the device group 118, i.e., the computing device 108,110. The computing device 108 transmits the reliability score data packet 132 including the reliability score 0.6 and the device identifier tablet to the computing device in the device group 120, that is, the computing device 106,110. The computing device 110 transmits a reliability score data packet 134 containing the reliability score 0.45 and the device identifier "thermostat" to the computing device in the device group 122, that is, the computing devices 106,108.
0021In some embodiments, the computing device may send a reliability score data packet if the reliability score has reached the hotword score threshold. For example, if the hotword score threshold is 0.5, the computing device 110 will not send the reliability score data packet 134 to other computing devices in device group 122. The computing devices 106,108 will continue to send reliability score data packets 130,132 to the computing devices in the device groups 118,120, respectively.
0022In some embodiments, the computing device that sends the reliability score data packet may send the reliability score data packet directly to another computing device. For example, the computing device 106 may transmit reliability score data packets to the computing devices 108,110 via short-range wireless communication. The communication protocol used between the two computing devices may be Universal Plug and Play. In some embodiments, the computing device that sends the reliability score data packet may broadcast the reliability score data packet. In this example, the reliability score data packet can be received by the computing devices in the device group and by other computing devices. In some embodiments, the computing device that sends the reliability score data packet sends the reliability score data packet to the server, and then the server sends the reliability score data packet to other compute in the device group. It may be sent to the wing device. The server may be located in the local area network of the computing device or in a network accessible via the Internet. For example, the computing device 108 sends a reliability score data packet 132 and a list of computing devices in the device group 120 to the server. The server sends the reliability score data packet 132 to the computing devices 106, 110. In an example where a computing device sends a reliability score data packet to another computing device, the receiving computing device may return a confirmation notification that it has received the reliability score data packet.
0023Each computing device uses a score comparison unit to compare the hotword reliability scores it receives. For example, the computing device 106 calculates a hotword reliability score of 0.85 and receives hotword reliability scores of 0.6 and 0.45. In this example, the score comparison unit 136 compares the three scores and identifies that the score 0.85 is the highest. For computing devices 108,110, score comparison units 138,140 reach similar conclusions and identify that the score 0.85 corresponding to computing device 106 is the highest.
0024The computing device that determines that its hotword reliability score is the highest value starts voice recognition for the voice data that follows the utterance of the hotword. For example, the user may say "OK computer" and the computing device 106 may determine that it has the highest hotword reliability score. The computing device 106 will start speech recognition for the audio data received after the hotword. When the user says "call Alice," the computing device 106 will process the utterance and execute the appropriate command. In some embodiments, receiving the hotword may wake the computing device receiving the hotword from sleep. In this example, the computing device with the highest hotword reliability score stays awake, while the other computing device without the highest hotword reliability score is the voice that follows the hotword utterance. Go to sleep without processing data.
0025As shown in FIG. 1, the score comparison unit 136 identifies that the hotword reliability score corresponding to the computing device 106 is the highest. Therefore, device state 142 is "awake". The score comparison unit 138,140 also identifies that the hotword reliability score corresponding to the computing device 106 is the highest. Therefore, device states 138,140 are "asleep". In some embodiments, the boot state of the computing device may be unaffected. For example, user 102 could have a computing device 106 in his hand while watching a movie on the computing device 108. When the user 102 speaks "OK computer", the computing device 106 initiates speech recognition for the audio data following the hotword by having the highest hotword reliability score. The computing device 108 does not start speech recognition for the audio data following the hotword and continues to play the movie.
0026In some embodiments, the computing device, which it determines to have the highest hotword reliability score, waits for a certain amount of time before it begins performing speech recognition for the speech following the hotword. By doing so, the computing device that has calculated the highest hotword reliability score is allowed to start performing speech recognition for the speech following the hotword without waiting for a higher hotword reliability score. As an example, the score comparison unit 136 of the computing device 106 receives the hotword reliability scores 0.6 and 0.45 from the computing devices 108 and 110, respectively, and receives the hotword reliability score 0.85 from the hotword unit 124. The computing device 106 waits 500 milliseconds from the time when the hotword unit 124 calculates the hotword reliability score of the audio data "OK computer" before performing voice recognition for the voice following the hotword. wait. In the example where the score comparison unit receives a higher score, the computing device does not have to wait for a certain amount of time before setting the device state to "sleep". For example, the hotword section 126 of the computing device 108 calculates a hotword reliability score of 0.6 and receives hotword reliability scores of 0.85 and 0.45. The computing device 108 can set the device state 144 to "sleep" when it receives a hotword reliability score of 0.85. In this case, it is probable that the computing device 108 received the hotword reliability score 0.85 within a specific time after the hotword unit 126 calculated the hotword reliability score 0.6.
0027In some embodiments, a computing device is specified to allow time for other computing devices to receive its reliability score data packets when it has the highest hotword reliability score. The reliability score data packet may continue to be broadcast for this time. This method would be most appropriate in the case where a computing device does not return a confirmation notification when it receives a reliability score data packet from another computing device. Therefore, if the computing device 106 sends a reliability score data packet 130 to a computing device in device group 118 and receives a confirmation notification before a certain time, such as 500 milliseconds, the hot word The execution of voice recognition for the voice following is started. In an example where a computing device broadcasts its own reliability score data packet and does not expect confirmation, the computing device receives a certain amount of time, such as 500 milliseconds, or a higher hotword reliability score. You may continue to broadcast your hotword reliability score until you do, whichever comes first. For example, the computing device 110 calculates a hotword reliability score of 0.45 and initiates broadcasting of the reliability score data packet 134. After 300 milliseconds, the computing device 110 receives the reliability score data packet 130, and the hotword reliability score 0.85 from the reliability score data packet 130 is higher than its own hotword reliability score 0.45. Therefore, the broadcast of the reliability score data packet 134 is terminated. In another broadcast example, the compute device 106 has a hotword reliability score of 0. Calculate 85 and start broadcasting the reliability score data packet 130. After 500 milliseconds, the computing device 106 ends the broadcast of the reliability score data packet 130 and begins performing speech recognition for the voice following the hotword. The computing device 106 may receive the reliability score data packets 132,134 before the lapse of 500 milliseconds, but the hotword reliability score in the reliability score data packets 132,134 is lower than 0.85, so that the computing device 106 Continues to wait until 500 milliseconds have passed.
0028In some embodiments, the computing device may begin performing speech recognition for the speech following the hotword before receiving a higher hotword reliability score. The hotword section calculates the hotword reliability score, and when the hotword reliability score reaches the threshold value, the computing device performs voice recognition for the voice following the hotword. The computing device may perform speech recognition without displaying any indication to the user about speech recognition. This gives the user the impression that voice recognition-based results can be displayed to the user faster than waiting to confirm that the highest hotword score has been calculated, even when the computing device is not active. It will give, so it would be desirable to do so. As an example, the computing device 106 calculates a hotword reliability score of 0.85 and initiates speech recognition for the speech following the hotword. The computing device 106 receives the reliability score data packets 132,134 and determines that the hotword reliability score 0.85 is the highest value. The computing device 106 continues to perform speech recognition for the speech following the hotword and presents the result to the user. For the computing device 108, the hotword unit 126 calculates a hotword reliability score of 0.6, and the computing device 108 starts performing speech recognition for the voice following the hotword without displaying the data to the user. To do. Upon receiving the reliability score data packet 130 including the hotword reliability score 0.85, the computing device 108 ends the execution of speech recognition. No data is displayed to the user, giving the user the impression that the computing device 108 remains in a "sleep" state.
0029In some embodiments, in order to avoid any waiting time after the hotword is issued, the score is placed prior to the end of the hotword, eg, for a partial hotword, in the hotword section. May be notified from. For example, if the user says "OK computer", the computing device will calculate a partial hotword reliability score when the user finishes saying "OK comp". Good. The computing device may then share the partial hotword reliability score with other computing devices. The computing device with the highest partial hotword reliability score can continue to process the user's conversation.
0030In some embodiments, the computing device emits an audible or inaudible sound, for example, of a particular frequency or frequency pattern, when it determines that the hotword reliability score has reached a threshold. It's okay. The sound will inform other computing devices that the computing device will continue to process the audio data following the hotword. Other computing devices will receive this sound and will not process the audio data. For example, the user says "OK computer". One of multiple computing devices calculates a hotword reliability score that exceeds or is equal to a threshold. The computing device emits a sound of 18 kHz when it determines that the hotword reliability score exceeds or is equal to the threshold. Other computing devices near the user may be in the process of calculating the hotword reliability score and may be in the middle of calculating the hotword reliability score when the sound is received. Other computing devices stop processing the user's conversation when they receive the sound. In some embodiments, the computing device may encode a hotword reliability score into audible or inaudible sounds. For example, if the hotword reliability score is 0.5, the computing device may produce audible or inaudible sounds that include a frequency pattern that encodes a score of 0.5.
0031In some embodiments, the computing device may use different audio measurements to select a computing device to continue processing the user's conversation. For example, a computing device may use loudness to determine which computing device continues to process a user's conversation. The computing device that detects the loudest conversation may continue to process the user's conversation. In another example, a computing device currently in use or with an active display may notify other computing devices that it will continue to process the user's conversation when it detects a hotword.
0032In some embodiments, each computing device near the user when the user is speaking receives audio data and sends the audio data to a server to improve speech recognition. Each computing device can receive audio data corresponding to the user's conversation. Each computing device may send audio data to the server, even though it appears to the user that only one computing device handles its own conversation. The server can then use the audio data received from each computing device to compare different audio samples that correspond to the same utterance, thus improving speech recognition. For example, the user says, "OK computer, remind me to buy milk (OK computer, remind me to buy milk) ". When the user finishes speaking "OK computer", nearby computing devices will probably determine which computing device has the highest hotword reliability score, and that computing device will allow the user to "milk". Remind me to buy a computer, "and it will process those words and return a response. Other computing devices also receive the utterance "Remind me to buy milk." Other computing devices may send audio data to the server that corresponds to "Remind me to buy milk" even if it does not respond to the "Remind me to buy milk" utterance. A computing device that responds to "Remind me to buy milk" may also send its audio data to the server. The server can improve speech recognition because it has different audio samples from different computing devices that process the audio data and respond to the same "remind me to buy milk" utterance. ..
0033FIG. 2 is a diagram showing an example of the process 200 for hotword detection. Process 200 may be performed by a computing device such as computing device 108 in FIG. Process 200 calculates a value that corresponds to the likelihood that the speech contains a hot word, compares that value to another value calculated by another computing device, and puts it in the part of the speech after the hot word. On the other hand, it is determined whether or not to execute voice recognition.
0034The computing device receives the audio data corresponding to the utterance (210). The user speaks and the microphone of the computing device receives the audio data of the utterance. The computing device processes the audio data by buffering, filtering, endpointing, and digitizing the audio data. As an example, the user speaks "OK computer" and the microphone of the computing device receives the audio data corresponding to "OK computer". The audio subsystem of a computing device samples, buffers, filters, and terminates its audio data for further processing by the computing device.
0035The computing device determines the first value corresponding to the possibility that the utterance contains a hot word (220). A computing device can be referred to as a hotword reliability score by comparing the audio data of an utterance with a group of audio samples containing hotwords, or by analyzing the audio characteristics of the audio data of an utterance. Determine the value of. The first value may be normalized to the range of zero to one, where 1 indicates the highest likelihood that the utterance will contain a hot word. In some embodiments, the computing device identifies a second computing device so that the second computing device is configured to respond to utterances, including hotwords, and responds to hotwords. Determine that it has been set by the user. The user may log in to both the computing device and the second computing device. Both the computing device and the second computing device may be configured to respond to the user's voice. The computing device and the second computing device may be connected to the same local area network. Both the computing device and the second computing device can be located within a specific distance from each other, such as 10 meters, as determined by GPS or signal strength. For example, computing devices may communicate by short-range wireless communication. The computing device may detect the strength of the signal transmitted by the second computing device, such as 5 dBm, and convert it to a corresponding distance, such as 5 meters.
0036The computing device receives a second value corresponding to the possibility that the utterance contains a hotword, the second value being determined by the second computing device (230). The second computing device receives the utterance through its own second microphone. The second computing device processes the received audio data corresponding to the utterance to determine a second value or a second hotword reliability score. The second hotword reliability score represents the possibility that the utterance contains hotwords, as calculated by the second computing device. In some embodiments, the computing device uses one or more of the following techniques to transmit a first value to a second computing device. That is, the computing device sets the first value via a server accessible over the Internet, via a server located on the local area network, or directly via the local area network or short-range wireless communication. Can be sent to 2 computing devices. The computing device may send the first value only to the second computing device, or the computing device may receive the first value so that other computing devices can also receive the first value. May be broadcast. The computing device may receive the second value from the second computing device using the same or different technology as the computing device that transmitted the first value.
0037In some embodiments, the computing device may calculate a volume score for the utterance or a signal-to-noise ratio for the utterance. The computing device may combine the volume score, signal-to-noise ratio, and hotword reliability score to determine new values for comparison with similar values from other computing devices. For example, a computing device may calculate a hotword reliability score and a signal-to-noise ratio. The computing device may then combine these two scores and compare them to scores similarly calculated for other computing devices. In some embodiments, the computing device may calculate different scores and send each score to another computing device for comparison. For example, a computing device may calculate a volume score and a hotword reliability score for an utterance. The computing device may then send these scores to other computing device for comparison.
0038In some embodiments, the computing device may transmit a first identifier along with a first value. The identifier may be based on one or more of the addresses of the computing device, the name of the computing device given by the user, or the location of the computing device. For example, the identifier may be "69.123.132.43" or "phone". Similarly, the second computing device may transmit a second identifier along with a second value. In some embodiments, the computing device may transmit a first identifier to a particular computing device, which is a computing device previously identified as being configured to respond to hotwords. .. For example, a computing device is because the second computing device is capable of responding to hotwords and the same user is logged in to the second computing device in the same way as the computing device. The second computing device may have previously been identified as being configured to respond to hotwords.
0039The computing device compares the first value with the second value (240). The computing device then initiates speech recognition processing on the audio data based on the results of the comparison (250). In some embodiments, for example, the computing device initiates speech recognition when the first value is greater than or equal to the second value. If the user says "OK computer, call Carol", the computing device will say "The first value is greater than or equal to the second value." By performing voice recognition for "Call Carol", the "Call Carol" process will be initiated. In some embodiments, the computing device sets the boot state. In an example where the first value is greater than or equal to the second value, the computing device sets the boot state as active or "awake". In the "awake" state, the computing device displays the result of speech recognition.
0040In some embodiments, the computing device compares the first value with the second value and determines that the first value is less than the second value. The computing device sets the boot state as inactive or "sleep" based on determining that the first value is less than the second value. In the "sleep" state, the computing device does not appear to the user to be activated or to process audio data.
0041In some embodiments, when the computing device determines that the first value is greater than or equal to the second value, it may wait a certain amount of time before setting the boot state to active. .. A computing device may wait for a certain amount of time to increase the likelihood that it will not receive higher values from other computing devices. The particular time may be fixed or may vary depending on the technology by which the computing device sends and receives values. In some embodiments, the computing device may continue to transmit the first value for a particular period of time when it determines that the first value is greater than or equal to the second value. By continuing to transmit the first value for a particular amount of time, the computing device increases the probability that the first value will be received by another computing device. In an example where the computing device determines that the first value is less than the second value, the computing device may stop transmitting the first value.
0042In some embodiments, the computing device may consider additional information in determining whether to execute the command following the hotword. An example of additional information may be the part of the utterance that follows the hot word. Typically, the audio data that follows the hotword is "call Sally," "play Halloween Movie," or "set the temperature to 70 degrees Fahrenheit." Please (set heat to 70) Corresponds to commands for computing devices such as degrees). A computing device can identify a typical device that can handle or handle that type of request. Typically, a request to call someone will be handled by the telephone device based on typical pre-programmed usage or based on the user's usage pattern of the device. If the user routinely watches a movie on a tablet, the tablet can handle the request to play the movie. If the thermostat has a temperature control function, the thermostat can handle temperature control.
0043In order for the computing device to consider the part of the utterance that follows the hotword, the computing device would have to start speech recognition for the audio data when it would have identified the hotword. The computing device may classify the command portion of the utterance and calculate the frequency of commands in such classification. The computing device may transmit its frequency to other computing devices along with the hotword reliability score. Each computing device may use the frequency and the hotword reliability score to determine whether to execute the command following the hotword.
0044For example, on a phone where the user says "OK computer, play Michael Jackson" and the computing device is used by the user to listen to music with a 20% chance. In some cases, the computing device may send that information along with a hotword reliability score. A computing device, such as a tablet, that is used by a user with a 5 percent chance of listening to music may send that information to other computing devices along with a hotword reliability score. The computing device may use the combination of the hotword reliability score and the music playback probability to determine whether or not to execute the command.
0045FIG. 3 shows an example of a computing device 300 and a mobile computing device 350 that can be used to implement the techniques described herein. Computing device 300 is intended to refer to various forms of digital computers such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. .. Mobile computing device 350 is intended to refer to various forms of mobile devices, such as personal digital assistants, mobile phones, smartphones, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are for illustrative purposes only and are not intended to be limiting.
0046The computing device 300 includes a processor 302, a memory 304, a storage device 306, a high-speed interface 308 connecting the memory 304 and a plurality of high-speed expansion ports 310, and a low-speed interface 312 connecting the low-speed expansion port 314 and the storage device 306. Including. Each of the processor 302, memory 304, storage device 306, high-speed interface 308, high-speed expansion port 310, and low-speed interface 312 are interconnected using various buses, on a common motherboard or otherwise appropriately. May be mounted. Processor 302 processes instructions for execution within the computing device 300, including instructions stored in memory 304 or on storage device 306, and external inputs such as display 316 connected to high-speed interface 308. / Can display graphical information for the GUI on the output device. In another embodiment, a plurality of processors and / or a plurality of buses may be appropriately used together with a plurality of memories and a plurality of types of memories. Also, multiple computing devices may be connected to each device that provides the required parts of operation (eg, as a server bank, a group of blade servers, or a multiprocessor system).
0047The memory 304 stores information in the computing device 300. In some embodiments, the memory 304 is one or more volatile memory units. In some embodiments, the memory 304 is one or more non-volatile memory units. Further, the memory 304 may be another form of computer-readable medium such as magnetic or optical disc.
0048Storage device 306 can provide high capacity storage for computing device 300. In some embodiments, the storage device 306 includes a storage area network or other device in the configuration, including a floppy® disk device, hard disk device, optical disk device, or tape device, flash memory or other similar device. It may be a computer readable medium such as a solid state memory device, or an array of devices, or may include them. The instructions may be stored in the information carrier. When the instruction is executed by one or more processing devices (eg, processor 302), it executes one or more methods, such as those described above. The instructions may also be stored by one or more storage devices, such as a computer-readable or machine-readable medium (eg, memory 304, storage device 306, or memory on processor 302).
0049The high-speed interface 308 manages the bandwidth-intensive operation for the computing device 300, while the low-speed interface 312 manages the lower-side band-intensive operation. The allocation of such functions is just one example. In some embodiments, the high-speed interface 308 is in memory 304, in display 316 (eg, via a graphics processor or accelerator), and in high-speed expansion port 310, which can accept various expansion cards (not shown). Be connected. In this embodiment, the slow interface 312 is connected to the storage device 306 and the slow expansion port 314. Slow expansion port 314, which may include various communication ports (eg USB, Bluetooth®, Ethernet®, wireless Ethernet), is a switch or switch via a keyboard, pointing device, scanner, or eg network adapter. It may be connected to one or more input / output devices, such as networking devices such as routers.
0050The computing device 300 can be implemented in a number of different forms, as shown in the figure. For example, the computing device 300 may be implemented as including a standard server 320 or a plurality of groups of such servers. Further, the computing device 300 may be implemented as a personal computer such as a laptop computer 322. The computing device 300 may also be implemented as part of the rack server system 324. Alternatively, the components in the computing device 300 may be combined with other components in the mobile device (not shown), such as the mobile computing device 350. Each such device may include one or more of the computing device 300 and the mobile computing device 350, and the entire system may consist of a plurality of computing devices communicating with each other.
0051The mobile computing device 350 includes, among other components, a processor 352, a memory 364, an input / output device such as a display 354, a communication interface 366, and a transmitter / receiver 368. The mobile computing device 350 may also be provided with a storage device, such as a microdrive or other device, to provide additional storage. Each of the processor 352, memory 364, display 354, communication interface 366, and transmitter / receiver 368 is interconnected using various buses, some of which are on a common motherboard or properly others. It may be mounted by the method.
0052Processor 352 can execute instructions in mobile computing device 350, including instructions stored in memory 364. Processor 352 may be implemented as a chipset of chips containing separate or multiple analog and digital processors. Processor 352 may be provided for coordination of other components of the mobile computing device 350, such as controlling the user interface, executing applications with the mobile computing device 350, and wireless communication with the mobile computing device 350. ..
0053The processor 352 may interact with the user via the control interface 358 and the display interface 356 connected to the display 354. The display 354 is, for example, a TFT LCD (Thin-Film-Transistor Liquid Crystal Display) display or an OLED (Organic Light Emitting). It may be a Diode) display, or other suitable display technology. The display interface 356 may include suitable circuits for driving the display 354 and presenting graphical and other information to the user. The control interface 358 may receive a command from the user and translate it for submission to the processor 352. In addition, the external interface 362 may provide communication with the processor 352 to allow close area communication between the mobile computing device 350 and other devices. The external interface 362 may be provided, for example, for wired communication in some embodiments or for wireless communication in some embodiments, and a plurality of interfaces may be used.
0054The memory 364 stores information in the mobile computing device 350. The memory 364 may be implemented as one or more computer-readable media, one or more volatile memory units, or one or more of one or more non-volatile memory units. Further, the extended memory 374 may be provided and connected to the mobile computing device 350 via the extended interface 372. The expansion interface 372 is, for example, SIMM (Single In Line Memory). Module) Card interface may be included. The extended memory 374 may provide a separate storage space for the mobile computing device 350, or may store applications or other information for the mobile computing device 350. In particular, the extended memory 374 may include instructions for executing or augmenting the process described above, and may include secure information. Thus, for example, the extended memory 374 may be provided as a security module for the mobile computing device 350 and may be programmed with an instruction permitting the secure use of the mobile computing device 350. In addition, secure applications may be provided via the SIMM card, along with additional information such as identification information placed on the SIMM card in a non-hackable manner.
0055The memory may include, for example, flash memory and / or NVRAM memory (Non-Volatile Random Access Memory) as described below. In some embodiments, the instructions are stored in the information carrier. When the instruction is executed by one or more processing devices (eg, processor 352), it executes one or more methods, such as those described above. Instructions are also stored by one or more storage devices, such as one or more computer-readable or machine-readable media (eg, memory 364, extended memory 374, or memory on processor 352). Good. In some embodiments, the instruction may be received, for example, as a signal propagated over the transmitter / receiver 368 or the external interface 362.
0056The mobile computing device 350 may perform wireless communication via the communication interface 366 and may optionally include a digital signal processing circuit. The communication interface 366 particularly includes GSM (Registered Trademark) (Global System for Mobile communications) voice calls, SMS (Short Message Service), EMS (Enhanced Messaging Service), or MMS (Multimedia Messaging Service) messaging, and CDMA (Code Division Multiple). Access), TDMA (Time Division Multiple Access), PDC (Personal Digital Cellular), WCDMA (Registered Trademark) (Wideband Code Division Multiple Access), CDMA2000, or GPRS (General Packet Radio) It can provide communication under various modes or protocols such as Service). Such communication may be performed, for example, via a transmitter / receiver 368 that uses radio frequencies. Further, short-range communication may be performed by using Bluetooth (registered trademark), Wi-Fi, or other above-mentioned transmitter / receiver (not shown). In addition, the Global Positioning System (GPS) receiver module 370 may provide the mobile computing device 350 with wireless data regarding additional navigation and positioning, such data on the mobile computing device 350. Can be used properly by applications running on.
0057In addition, the mobile computing device 350 may audibly communicate using the audio codec 360, and may receive information emitted by the user and convert it into digital information suitable for use. The audio codec 360 may generate audible sound to the user, such as through a speaker, eg, the speaker of the handset of the mobile computing device 350. Such sounds may include sounds from voice phone calls, may include recorded sounds (eg, voice messages, music files, etc.) and may be generated by an application running on the mobile computing device 350. It may include the sound to be played.
0058Mobile Computing Device Device 350 may be implemented in a number of different forms, as illustrated. For example, it may be implemented as a mobile phone device 380. It may also be implemented as a smartphone 382, PDA, or other similar mobile device.
0059Various embodiments of the systems and techniques described herein include digital electronic circuits, integrated circuits, specially designed application specific integrated circuits (ASICs), computer hardware, firmware, software, and / or theirs. It may be realized as a combination. These various embodiments may include implementation as one or more computer programs executable and / or interpretable on a programmable system that includes at least one programmable processor. The programmable system may be dedicated or general purpose, to receive data and instructions from the storage system, at least one input device, and at least one output device, and to send data and instructions to them. Is connected to.
0060These computer programs (also known as programs, software, software applications, or code) include machine language instructions for programmable processors, high-level procedural and / or object-oriented programming languages, and / or assemblies / It may be implemented in a machine language. As used herein, the terms machine-readable and computer-readable media are machine language for programmable processors, including machine-readable media that receive machine language instructions as machine-readable signals. Refers to any computer program product, device, and / or device used to provide instructions and / or data (eg, magnetic disks, optical disks, memory, programmable logic devices (PLDs)). The term machine-readable signal refers to any signal used to provide machine language instructions and / or data to a programmable processor.
0061To provide user interaction, the systems and techniques described herein include display devices for displaying information to the user (eg, a cathode line tube (CRT) or liquid crystal display (LCD) monitor). It may be performed on a computer equipped with a keyboard and a pointing device (eg, a mouse or trackball) that allows the user to provide input to the computer. Other types of devices may be used to provide interaction with the user, for example, the feedback provided to the user may be any form of sensory feedback (eg, visual feedback, auditory feedback, or tactile feedback). Feedback), and the input from the user may be received in any form, including acoustic, audible, or tactile input.
0062The systems and technologies described herein include back-end components (eg, data servers), or include middleware components (eg, application servers), or front-end components (eg, users describe in the specification). A computing system that includes (a client computer with a graphical user interface or web browser that can interact with embodiments of the system and technology), or any combination of such back-end, middleware, or front-end components. May be carried out as. The components of the system may be interconnected by digital data communication (eg, a communication network) of any form or medium. Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
0063The computing system may include a client and a server. Clients and servers are generally located apart from each other and typically interact over a communication network. The client-server relationship arises from the work of computer programs that run on their own computers and have a client-server relationship with each other.
0064Although some embodiments have been described in detail above, other modifications can be considered. For example, a client application has been described as accessing a (s) representative stations, but in another embodiment, the (s) representative stations may run on one or more servers. It may be utilized by other applications running on one or more processors. Moreover, the illustrated logical flows do not require the order or order as described to obtain the desired results. In addition, other actions may be provided for the indicated flow, or actions may be removed from the indicated flow. In addition, other components may be added to the indicated system, or components may be removed from the indicated system. Therefore, other embodiments are within the scope of the appended claims.
0065100 systems 102 users 104 utterances 106,108,110 Computing device 112,114,116 Device identifier (ID) 118,120,122 device groups 124,126,128 Hot word section 130,132,134 Confidence score data packet 136,138,140 Score comparison section 142,144,146 Device status 300 computing device 302 processor 304 memory 306 storage device 308 High speed interface 310 fast expansion port 312 low speed interface 314 slow expansion port 316 display 320 server 324 rack server system 350 mobile computing device 352 processor 354 display 356 display interface 358 Control interface 360 audio codec 362 External interface 364 memory 366 Communication interface 368 transmitter / receiver 370 GPS receiver module 372 Extended interface 374 extended memory 380 mobile phone 382 smartphone
4 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US11908479B2 | Cited by | United States of America | Applicant |
| JP2020003774A | Cited by | Japan | Search report |
| US11087765B2 | Cited by | United States of America | Applicant |
| US11244686B2 | Cited by | United States of America | Applicant |
| JP2022541207A | Cited by | Japan | Search report |
| US11355107B2 | Cited by | United States of America | Applicant |
| JP2019204103A | Cited by | Japan | Search report |
| JP2022542113A | Cited by | Japan | Search report |
| US12198691B2 | Cited by | United States of America | Applicant |
| US2021201915A1 | Cited by | United States of America | Applicant |
| JP2019204103A | Cited by | Japan | Search report |
| US11227600B2 | Cited by | United States of America | Applicant |
| JP2020505626A | Cited by | Japan | Search report |
| JP2021513693A | Cited by | Japan | Search report |
| US12315512B2 | Cited by | United States of America | Applicant |
| US11270705B2 | Cited by | United States of America | Applicant |
| US12387722B2 | Cited by | United States of America | Applicant |
| US11380331B1 | Cited by | United States of America | Applicant |
| JP2000310999A | Cites | Japan | Search report |
| JP2006227634A | Cites | Japan | Search report |
| US6023676A | Cites | United States of America | Search report |
| US8340975B1 | Cites | United States of America | Search report |
| JPH11231896A | Cites | Japan | Search report |
| JPH1152976A | Cites | Japan | Search report |
| JPS59180599A | Cites | Japan | Search report |
65 members in 7 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 62061830 | United States of America | – | |
| 201462061830 | United States of America | P | |
| 14675932 | United States of America | – | |
| 201514675932 | United States of America | A | |
| 2015052860 | United States of America | W |
Members65
| Document | Office | Kind | |
|---|---|---|---|
| US2016104480A1 | United States of America | A1 | |
| WO2016057268A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US9318107B1 | United States of America | B1 | |
| US2016217790A1 | United States of America | A1 | |
| KR20160101198A | Republic of Korea | A | |
| CN106030699A | China | A | |
| EP3084759A1 | European Patent Office (EPO) | A1 | |
| US9514752B2 | United States of America | B2 | |
| KR20170004956A | Republic of Korea | A | |
| US2017025124A1 | United States of America | A1 | |
| EP3139378A1 | European Patent Office (EPO) | A1 | |
| US2017084277A1 | United States of America | A1 | |
| JP2017072857A | Japan | A | |
| EP3171359A1 | European Patent Office (EPO) | A1 | |
| KR101752119B1 | Republic of Korea | B1 | |
| JP2017520008AThis record | Japan | A | |
| JP6208376B2 | Japan | B2 | |
| US9812128B2 | United States of America | B2 | |
| JP2017227912A | Japan | A | |
| US2018040322A1 | United States of America | A1 | |
| KR101832648B1 | Republic of Korea | B1 | |
| WO2018067528A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US10102857B2 | United States of America | B2 | |
| US10134398B2 | United States of America | B2 | |
| JP6427549B2 | Japan | B2 | |
| US2019051303A1 | United States of America | A1 | |
| US2019130914A1 | United States of America | A1 | |
| JP6530023B2 | Japan | B2 | |
| JP2019133198A | Japan | A | |
| EP3084759B1 | European Patent Office (EPO) | B1 | |
| EP3139378B1 | European Patent Office (EPO) | B1 | |
| CN106030699B | China | B | |
| US10559306B2 | United States of America | B2 | |
| US2020058306A1 | United States of America | A1 | |
| JP2020034952A | Japan | A | |
| US10593330B2 | United States of America | B2 | |
| EP3627503A1 | European Patent Office (EPO) | A1 | |
| CN111028826A | China | A | |
| EP3171359B1 | European Patent Office (EPO) | B1 | |
| US2020211556A1 | United States of America | A1 | |
| US10909987B2 | United States of America | B2 | |
| US2021118448A1 | United States of America | A1 | |
| US11024311B2 | United States of America | B2 | |
| JP6893951B2 | Japan | B2 | |
| US2021249015A1 | United States of America | A1 | |
| JP2022017569A | Japan | A | |
| JP7022733B2 | Japan | B2 | |
| US11557299B2 | United States of America | B2 | |
| DE202015010012U1 | Germany | U1 | |
| US2023147222A1 | United States of America | A1 | |
| US11670297B2 | United States of America | B2 | |
| US2023274741A1 | United States of America | A1 | |
| JP7354210B2 | Japan | B2 | |
| EP4280210A2 | European Patent Office (EPO) | A2 | |
| JP2023174674A | Japan | A | |
| EP3627503B1 | European Patent Office (EPO) | B1 | |
| EP4280210A3 | European Patent Office (EPO) | A3 | |
| CN111028826B | China | B | |
| US11915706B2 | United States of America | B2 | |
| US2024169992A1 | United States of America | A1 | |
| US12046241B2 | United States of America | B2 | |
| US2024363113A1 | United States of America | A1 | |
| US12254884B2 | United States of America | B2 | |
| JP7664335B2 | Japan | B2 | |
| US2025191590A1 | United States of America | A1 |
18 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Written notification of registration of transferJAPANESE INTERMEDIATE CODE: R350R350 | R350 | |
| Written request for registration of change of nameJAPANESE INTERMEDIATE CODE: R313533S533 | S533 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on accelerated examinationJAPANESE INTERMEDIATE CODE: A971005A975 | A975 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 | |
| Explanation of circumstances concerning accelerated examinationJAPANESE INTERMEDIATE CODE: A871A871 | A871 |
Numbers
- Publication
- 2017520008
- Application
- 2016551250
Titles2
- Japanese
- 複数のデバイス上でのホットワード検出
- English
- Hotword detection on multiple devices
Classification
- CPC, 9
- G10L15/22
- G10L15/285
- G10L2015/223
- G10L2015/088
- G06F3/167
- G10L15/32
- G10L15/08
- G10L17/22
- G10L15/01
- IPC, 6
- G10L15 10
- G10L15 00
- H04M1 00
- H04M11 00
- G06F3 16
- G10L15 28
Designated states5
- Regional, 4
- Zimbabwe
- Turkmenistan
- Türkiye
- Togo
- National, 1
- United States of America