Visual confirmation for a recognized voice-initiated action
Abstract
The techniques described herein provide computing devices that are configured to provide an indication that the computing device has recognized a voice-initiated action. In one example, a method for outputting, by the computing device, a voice recognition graphical user interface (GUI) having at least one element in a first visual format for display is provided. The method further includes receiving audio data by the computing device and determining, by the computing device, a voice-initiated action based on the audio data. The method further includes outputting an updated voice recognition GUI for display while receiving additional audio data and before performing the voice initiation action based on the audio data, where the updated voice recognition GUI is A second visual format different from the first visual format is used to display the at least one element to indicate that the voice-initiated action has been recognized.

Term
7.7 yearsto projected expiry
Projected expiry 19 June 2034, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
14 claims: 5 independent, 9 dependent
- 1一种方法,包括: 由计算设备输出具有以第一视觉格式的至少一个元素的话音识别图形用户界面(GUI) 以用于显示; 由所述计算设备接收音频数据; 由所述计算设备基于所述音频数据来确定语音发起动作;以及 在接收到附加音频数据的同时并且在基于所述音频数据来执行所述语音发起动作之 前,输出已更新的话音识别GUI以用于显示,在所述已更新的话音识别GUI中以不同于所述 第一视觉格式的第二视觉格式来显示所述至少一个元素,以指示所述语音发起动作已被识 别。
- 2根据权利要求1所述的方法,进一步包括: 由所述计算设备基于所述音频数据来确定转录; 识别与所述语音发起动作相关联的所述转录的一个或多个单词,其中,所述至少一个 元素包括所述一个或多个单词的至少一部分;以及 由所述计算设备在输出所述已更新的话音识别GUI之前输出所述转录的不包括所述一 个或多个单词的一部分以用于以所述第一视觉格式显示。
- 3根据权利要求1至2中的任一项所述的方法, 其中,在图像、色彩、字体、大小、突出显示、风格、以及位置中的一个或多个方面,所述 第二视觉格式不同于所述第一视觉格式。
- 4根据权利要求1至3中的任一项所述的方法,进一步包括: 由所述计算设备来确定所述音频数据的转录,其中: 输出所述话音识别GUI进一步包括输出所述转录的至少一部分,以及 输出所述已更新的话音识别GUI进一步包括裁剪被输出的所述转录的所述至少一部 分,使得所述转录的与所述语音发起动作相关的所述一个或多个单词被显示。
- 5根据权利要求1至4中的任一项所述的方法, 其中,所述至少一个元素的所述第一视觉格式包括表示所述计算设备的话音识别模式 的图像,以及其中,所述至少一个元素的所述第二视觉格式包括表示所述语音发起动作的 图像。
- 6根据权利要求5所述的方法, 其中,表示所述话音识别模式的所述图像响应于基于所述音频数据确定所述语音发起 动作而变体成表示所述语音发起动作的所述图像。
- 7根据权利要求1至6中的任一项所述的方法,进一步包括 响应于基于所述音频数据而确定所述语音发起动作,由所述计算设备来执行所述语音 发起动作。
- 8根据权利要求7所述的方法, 其中,执行所述语音发起动作进一步响应于由所述计算设备接收到确认所述语音发起 动作正确的指示。
- 9根据权利要求1至8中的任一项所述的方法,进一步包括,由所述计算设备并且至少 部分地基于所述音频数据来确定所述语音发起动作。
- 10根据权利要求9所述的方法,其中,确定所述语音发起动作进一步包括至少部分地 基于以所述音频数据为基础的所述转录的单词或短语与动作的预配置的集合的比较来确 定所述语音发起动作。
- 11根据权利要求9至10中的任一项所述的方法,其中,确定所述语音发起动作进一步 包括: 由所述计算设备来识别所述转录中的至少一个动词;以及 将所述至少一个动词与来自动词集合的一个或多个动词相比较,所述动词集合中的每 个动词与来自多个动作的至少一个动作相对应。
- 12根据权利要求9至11中的任一项所述的方法,其中,确定所述语音发起动作进一步 包括以下各项中的至少一个: 由所述计算设备至少部分地基于来自所述计算设备的数据来确定场境;以及 由所述计算设备至少部分地基于所述场境来确定所述语音发起动作;或者 响应于接收到取消输入的指示,由所述计算设备来输出所述至少一个元素以便以所述 第一视觉格式来显示。 13 .一种计算机可读存储介质,所述计算机可读存储介质被用指令编码以使得一个或 多个可编程处理器执行权利要求1至12中的任一项所述的方法。
- 1314. 一种计算设备,包括: 显示设备;以及 一个或多个处理器,所述一个或多个处理器可操作以: 输出具有以第一视觉格式的至少一个元素的话音识别图形用户界面(GUI)以用于在所 述显示设备处显示; 接收音频数据; 基于所述音频数据来确定语音发起动作;以及 在接收到附加音频数据的同时且在基于所述音频数据来执行所述语音发起动作之前, 输出已更新的话音识别GUI以用于显示,在所述已更新的话音识别GUI中以不同于所述第一 视觉格式的第二视觉格式来显示所述至少一个元素,以指示所述语音发起动作已被识别。
- 1415. 根据权利要求14所述的计算设备,进一步包括用于执行权利要求1至12中的任一项 所述的方法的装置。
Independent claims14
152 paragraphs, as filed
Background technology for visual confirmation of actions initiated by recognized speech
[0001] Certain computing devices (eg, mobile phones, tablet computers, personal digital assistants, etc.) may be voice activated. The voice activated computing device can be controlled by means of audio data such as human voice. Such computing devices provide functions to detect voice, determine the action indicated by the detected voice, and perform the indicated action. For example, the computing device may receive audio input corresponding to voice commands such as "search", "navigation", "play", "pause", "call", and so on. In this case, the computing device can use voice recognition technology to analyze the audio input to determine the command and then perform actions associated with the command (e.g., provide search options, execute a map application, start playing media files, stop playing media files, Make a call, etc.). In this way, the voice activated computing device can provide the user with the ability to operate some features of the computing device without using the user's hands.
Summary of the invention
[0002] In one example, the present disclosure is directed to a method for outputting, by a computing device, a voice recognition graphical user interface (GUI) having at least one element in a first visual format for display. The method also includes receiving audio data by the computing device. The method also includes determining, by the computing device, a voice-initiated action based on the audio data. The method further includes outputting an updated voice recognition GUI for display while receiving additional audio data and before performing the voice initiation action based on the audio data, where the updated voice recognition GUI is different from The second visual format of the first visual format is used to display the at least one element to indicate that the voice-initiated action has been recognized.
[0003] In another example, the present disclosure is directed to a computing device including a display device and one or more processors. The one or more processors are operable to output a voice recognition graphical user interface (GUI) having at least one element in a first visual format for display at the display device. The one or more processors are operable to receive the audio data and determine a voice initiation action based on the audio data. The one or more processors are also configured to output an updated voice recognition GUI for display while receiving additional audio data and before performing the voice initiation action based on the audio data. Update the voice recognition GUI to display the at least one element in a second visual format different from the first visual format to indicate that the voice-initiated action has been recognized.
[0004] In another example, the present disclosure is directed to a computer-readable storage medium encoded with instructions that, when executed by one or more processors of a computing device, cause the one or more processors to output A voice recognition graphical user interface (GUI) with at least one element in a first visual format for display. The instructions also cause the one or more processors to receive audio data and determine the voice-initiated action based on the audio data. The instructions also cause the one or more processors to output the updated voice recognition GUI for display while receiving additional audio data and before performing the voice initiation action based on the audio data, The updated voice recognition GUI displays the at least one element in a second visual format different from the first visual format to indicate that the voice-initiated action has been recognized.
[0005] The details of one or more examples are set forth in the drawings and the following description. Other features, objects, and advantages of the present disclosure will become apparent from the description, the drawings, and the claims.
Description of the drawings
[0006] FIG. 1 is a conceptual diagram illustrating an exemplary computing device configured to provide a graphical user interface that provides a visual indication of a recognized voice-initiated action in accordance with one or more aspects of the present disclosure.
[0007] FIG. 2 is a block diagram illustrating an example computing device for providing a graphical user interface that includes a visual indication of a recognized voice-initiated action in accordance with one or more aspects of the present disclosure.
[0008] FIG. 3 is a block diagram illustrating an example computing device that outputs graphical content for display at a remote device in accordance with one or more techniques of the present disclosure.
[0009] FIGS. 4A to 4D are screenshots illustrating an example graphical user interface (GUI) of a computing device for navigating an example in accordance with one or more techniques of the present disclosure.
[0010] FIGS. 5A to 5B are screenshots illustrating an example GUI of a computing device for a media playback example in accordance with one or more techniques of the present disclosure.
[0011] FIG. 6 is a conceptual diagram illustrating a series of example visual formats in which elements can be morphed into actions based on different voices according to one or more techniques of the present disclosure.
[0012] FIG. 7 is a flowchart illustrating an example process for a computing device to visually confirm that a voice-initiated action has been recognized in accordance with one or more techniques of the present disclosure.
Detailed ways
[0013] In general, the present disclosure is directed to techniques that computing devices can use to provide visual confirmation of voice-initiated actions determined based on received audio data. For example, in some embodiments, the computing device may receive audio data from an audio input device (e.g., a microphone), transcribe the audio data (e.g., voice), determine whether the audio data includes an indication of a voice-initiated action, and provide if so Visual confirmation of the indicated action. By outputting the visual confirmation of the voice-initiated action, the computing device may therefore enable the user to more easily and quickly determine whether the computing device has correctly recognized and is about to perform the voice-initiated action.
[0014] In some embodiments, the computing device may provide visual confirmation that the voice-initiated action has been recognized by changing the visual format of the element corresponding to the voice-initiated action. For example, the computing device may output the element in the first visual format. In response to determining at least one of the one or more words of the transcription of the received audio data corresponding to the specific voice-initiated action, the computing device may update the visual format of the element to a second visual format that is different from the first visual format. format. Therefore, the observable differences between these visual formats can provide a mechanism that the user can use to visually confirm that the voice-initiated action has been recognized by the computing device and that the computing device will perform the voice-initiated action. The element may be, for example, one or more graphical icons, images, words of text (based on, for example, a transcription of received audio data), or any combination thereof. In some examples, the elements are interactive user interface elements. Therefore, a computing device configured in accordance with the techniques described herein can change the visual appearance of an output element to indicate that the computing device has recognized a voice-initiated action associated with audio data received by the computing device.
[0015] FIG. 1 is a conceptual diagram illustrating an example computing device 2 configured to provide a graphical user interface 16 that provides a visual view of recognized voice-initiated actions in accordance with one or more aspects of the present disclosure. Instructions. The computing device 2 may be a mobile device or a fixed device. For example, in the example of FIG. 1, the computing device 2 is illustrated as a mobile phone such as a smart phone. However, in other examples, the computing device 2 may be a desktop computer, a host computer, a tablet computer, a personal digital assistant (PDA), a laptop computer, a portable game device, a portable media player, a global positioning system (GPS)
Device, e-book reader, glasses, watch, TV platform, car navigation system, wearable computing platform, or another type of computing device.
[0016] As shown in FIG. 1, the computing device 2 includes a user interface device (UID) 4<sub>O</sub>The UID 4 of the computing device 2 can serve as an input device or an output device for the computing device 2. Various techniques can be used to implement UID 4. For example, UID 4 can act as an input device using a presence-sensitive input display, such as a resistive touch screen, surface acoustic wave touch screen, capacitive touch screen, projected capacitive touch screen, pressure sensitive screen, acoustic pulse recognition touch screen, or another presence-sensitive display technology. UID 4 can act as an output (for example, display) device using any one or more display devices, such as liquid crystal displays (LCD), dot matrix displays, light emitting diode (LED) displays, organic light emitting diodes ( OLED) display, electronic ink, or similar monochrome or color display capable of outputting visible information to the user of the computing device 2.
[0017] The UID 4 of the computing device 2 may include a presence-sensitive display, which may receive tactile input from, for example, a user of the computing device 2. The UID 4 may receive an indication of tactile input by detecting one or more gestures from the user of the computing device 2 (for example, the user touches or points to one or more locations of the UID 4 with a finger or a stylus pen). UID 4 may present output to the user, for example, where a sensitive display is present. UID 4 can present the output as a graphical user interface (eg, user interface 16) that can be associated with the functions provided by computing device 2. For example, UID 4 may present various user interfaces of applications (eg, electronic messaging applications, navigation applications, Internet browser applications, media player applications, etc.) that are executed at the computing device 2 or can be accessed by the computing device 2. The user can interact with the corresponding user interface of the application to cause the computing device 2 to perform function-related operations.
[0018] The example of the computing device 2 shown in FIG. 1 also includes a microphone 12. The microphone 12 may be one of one or more input devices of the computing device 2. The microphone 12 is a device for receiving auditory input such as audio data. The microphone 12 can receive audio data including voice from the user. The microphone 12 detects the audio and provides relevant audio data to other components of the computing device 2 for processing. In addition to the microphone 12, the computing device 2 may also include other input devices.
[0019] For example, changing a part of the transcribed text corresponding to a voice command (for example, "voice-initiated action") so that the visual appearance of the part of the transcribed text corresponding to the voice command is different from that of The visual appearance of the transcribed text corresponding to the voice command. For example, the computing device 2 receives audio data at the microphone 12. The voice recognition module 8 can transcribe the voice included in the audio data, which can be in real time or near real time with the received audio data. The computing device 2 outputs the non-command text 20 corresponding to the transcribed voice for display. In response to determining a portion of the transcribed voice corresponding to the command, computing device 2 may provide at least one indication that the voice portion is recognized as a voice command. In some examples, computing device 2 may perform the actions identified in the voice-initiated actions. "Voice command" as used herein may also be referred to as "voice-initiated action" ο
[0020] In order to instruct the computing device 2 to recognize a voice-initiated action within the audio data, the computing device 2 may change the visual format of a portion of the transcribed text (for example, the command text 22) corresponding to the voice command. In some examples, the computing device 2 may change the visual appearance of the portion of the transcribed text corresponding to the voice command so that the visual appearance is different from the visual appearance of the transcribed text that does not correspond to the voice command. For simplicity, any text that is associated with or recognized as a voice-initiated action is referred to herein as "command text." Likewise, any text that is not associated with or recognized as a voice-initiated action is referred to herein as "non-command text."
[0021] The text associated with the voice-initiated action (eg, command text 22) may have a different font, color, size, or other visual characteristics than text associated with non-command speech (eg, non-command text 20). In another example, the command text
The present 22 may be highlighted in some way, while the non-command text 20 is not highlighted. The UI device 4 can change any other characteristics of the visual format of the text, so that the transcribed command text 22 is visually different from the transcribed non-command text 20. In other examples, the computing device 2 may use any combination of changes or alterations to the visual appearance of the command text 22 described herein to visually distinguish the command text 22 from the non-command text 20.
[0022] In another example, the computing device 2 may serve as a substitute for the transcribed text or output graphical elements such as icons 24 or other images for display in addition to the transcribed text. The term "graphic element" as used herein refers to any visual element displayed in a graphical user interface, and may also be referred to as a "user interface element". The graphic element may be an icon indicating that the action computing device 2 is currently executing or executable. In this example, when the computing device 2 recognizes the voice-initiated action, the user interface ("UI") device module 6 causes the graphical element 24 to change from the first visual format to the second visual format, which indicates that the computing device 2 has recognized Voice initiates the action. The image of the graphical element 24 in the second visual format may correspond to the voice-initiated action. For example, the UI device 4 may display the graphical element 24 in the first visual format, while the computing device 2 is receiving audio data. The first visual format may be, for example, an icon 24 with an image of a microphone. In response to determining that the audio data contains a voice-initiated action requesting route guidance to a specific address, for example, the computing device 2 causes the icon 24 to change from a first visual format (for example, an image of a microphone) to a second visual format (for example, a compass arrow image).
[0023] In some examples, in response to recognizing the voice-initiated action, the computing device 2 outputs a new graphical element corresponding to the voice-initiated action. For example, rather than automatically taking an action associated with a voice-initiated action, the techniques described herein may enable the computing device 2 to first provide an indication of the voice-initiated action. In some examples, in accordance with various techniques of the present disclosure, the computing device 2 may be configured to update the graphical user interface 16 so that the elements are presented in different visual formats based on audio data including recognized instructions for voice-initiated actions.
[0024] In addition to the UI device module 6, the computing device 2 may also include a voice recognition module 8 and a voice activation module 10. Modules 6, 8, and 10 may use software, hardware, firmware or a mixture of hardware, software, and firmware resident in and executed on the computing device 2 to perform the described actions. The computing device 2 can execute modules 6, 8, and 10 with multiple processors. The computing device 2 can execute the modules 6, 8, and 10 as virtual machines executed on the underlying hardware. Modules 6, 8, and 10 can be executed as one or more services of the operating system and computing platform. Modules 6, 8, and 10 may be executed as one or more remote computing services such as one or more services provided by cloud and/or cluster-based computing systems. Modules 6, 8, and 10 can be executed as one or more executable programs at the application layer of the computing platform.
[0025] The voice recognition module 8 of the computing device 2 may receive one or more indications of audio data from, for example, the microphone 12. Using voice recognition technology, the voice recognition module 8 can analyze and transcribe the voice included in the audio data. The voice recognition module 8 may provide the transcribed voice to the UI device module 6° The UI device module 6 may instruct the UID 4 to output text related to the transcribed voice such as the non-command text 20 of the GUI 16 for display.
[0026] The voice activation module 10 of the computing device 2 may receive the transcribed voice text characters from the audio data detected at the microphone 12 from, for example, the voice recognition module 8. The voice activation module 10 may analyze the transcribed text to determine whether it includes keywords or phrases that activate voice-initiated actions. Once the voice activation module 10 recognizes the word or phrase corresponding to the voice-initiated action, the voice activation module 10 causes the UID 4 to display graphical elements in a second, different visual format within the user interface 16 to indicate that the voice-initiated action has been successfully Recognition. For example, when the voice activation module 10 determines a word in the transcribed text corresponding to the voice-initiated action, UID 4 outputs the word from the first visual format (which may be the same as the rest of the transcribed non-command text 20 The same visual format) becomes a second, different visual format. For example, the visual characteristic style of the keyword or phrase corresponding to the voice-initiated action is different from other words that do not correspond to the voice-initiated action, so as to indicate that the computing device 2 recognizes the voice-initiated action. In another example, when the voice activation module
10 When a voice-initiated action is recognized, the icon or other image included in the GUI 16 changes from one visual format to another visual format.
[0027] The ui device module 6 can cause the UID 4 to present the user interface 16. The user interface 16 includes graphical indications (eg, elements) displayed at various locations of the UID 4. FIG. 1 illustrates the icon 24 within the user interface 16 as an example graphical indication. FIG. 1 also illustrates graphical elements 26, 28, and 40 within the user interface 16 as examples of graphical indications for selecting options or performing additional functions related to the application executed at the computing device 2. The UI module 6 may receive information identifying the graphic element displayed in the first visual format at the user interface 16 as corresponding to or associated with a voice-initiated action as input from the voice activation module 10. In response to the computing device 2 identifying the graphical element as being associated with the voice-initiated action, the UI module 6 may update the user interface 16 to change the graphical element from the first visual format to the second visual format.
[0028] The UI device module 6 may act as an intermediary between various components of the computing device 2 to make determinations based on the input detected by the UID 4 and generate the output presented by the UID 4. For example, the UI module 6 receives transcribed text characters of audio data as input from the voice recognition module 8. The UI module 6 causes the UID 4 to display the transcribed text characters in the first visual format at the user interface 16. The UI module 6 receives information that recognizes at least a part of the text characters as corresponding to the command text from the voice activation command 10. Based on the identification information, the UI module 6 displays the text associated with the voice command or another graphic element in a second visual format that is different from the first visual format originally used to display the command text or graphic element .
[0029] For example, the UI module 6 receives information that recognizes a part of the transcribed text characters as corresponding to a voice-initiated action as an input from the voice activation module 10. In response to the voice activation module 10 determining that the transcribed text portion corresponds to the voice-initiated action, the UI module 6 changes the visual format of a portion of the transcribed text character. That is, the UI module 6 changes the graphic element from the first visual format to the second visual format in response to recognizing the graphic element as being associated with the voice-initiated action. The UI module 6 can cause the UID 4 to present the updated user interface 16. For example, GUI 16 includes text related to voice commands, command text 22 (ie, "listen"). In response to the voice activation module 10 determining that "listen" corresponds to the command, the UI device 4 updates the GUI 16 to display the command text 22 in a second format, which is different from the format of the rest of the non-command text 20.
[0030] In the example of FIG. 1, the user interface 16 is branched into two areas: an editing area 18-A and an action area 18-B. The editing area 18-A and the action area 18-B may include graphic elements such as transcribed text, images, objects, hyperlinks, characters of the text, menus, fields, virtual buttons, virtual keys, etc. Any of the graphical elements listed above as used herein may be user interface elements. FIG. 1 shows only one example layout for the user interface 16. There may be other examples in which the user interface 16 differs in one or more of layout, number of regions, appearance, format, version, color scheme, or other visual characteristics.
[0031] The editing area 18-A may be an area of the UI device 4 configured to receive input or output information. For example, the computing device 2 may receive a voice input recognized by the voice recognition module 8 as a voice, and the editing area 18-A outputs information related to the voice input. For example, as shown in FIG. 1, the user interface 16 displays non-command text 20 in the editing area 18-A. In other examples, the editing area 18-A may update information displayed based on touch-based or gesture-based input.
[0032] The action area 18-B may be configured to accept input from a user or provide an indication of an action that the computing device 2 has taken, is currently taking, or will take in the past. In some examples, action zone 18-B includes a graphical keyboard, which includes graphical elements displayed as keys. In some examples, while the computing device 2 is in the voice recognition mode, the action zone 18-B will not include a graphical keyboard.
[0033] In the example of FIG. 1, the computing device 2 outputs a user interface 16 for display, the user interface 16 including at least one graphical element that can be displayed in a visual format indicating that the computing device 2 has recognized a voice-initiated action. For example, the UI design
The backup module 6 can generate the user interface 16 and include graphical elements 22 and 24° in the user interface 16. The UI device module 6 can send information to the UID 4, including the information used to display the user interface 16 at the presence-sensitive display 5 of the UID 4 Instructions. UID 4 can receive this information and cause the presence-sensitive display 5 of UID 4 to present a user interface 16, which includes graphical elements that can change the visual format to provide an indication that a voice-initiated action has been recognized.
[0034] The user interface 16 includes one or more graphical elements displayed at various positions of the UID 4. As shown in the example of FIG. 1, many graphic elements are displayed in the editing area 18-A and the action area 18-B. In this example, the computing device 2 is in the speech recognition mode, which means that the microphone 12 is turned on to receive the first frequency input and the speech recognition module 8 is activated. The initial activation module 10 may also be active in the voice recognition mode in order to detect voice-initiated actions. When the computing device 2 is not in the voice recognition mode, the voice recognition module 8 and the voice passive module 10 may not be active. In order to indicate that the computing device 2 is in the voice recognition module and is listening, the words "listening..." may be displayed in the area 18-B. As shown in Figure 1, the icon 24 is in the image of the microphone.
[0035] The icon 24 indicates that the computing device 2 is in a voice recognition mode (for example, it can receive audio data, such as spoken words)<sub>O</sub>The UID 4 is displayed in the action area 18-B of the GUI 16 to enable selection of the language element 26 that the user is speaking, so that the voice recognition module 8 can transcribe the user's speech in the correct language. The GUI 16 includes a drop-down menu 28 to provide options to change the language used by the voice recognition module 8 to transcribe audio data. The GUI 16 also includes a virtual button 30 to provide an option to cancel the voice recognition mode of the computing device 2. As shown in FIG. 1, the visual button 30 includes the word "done" to indicate its purpose to end the voice recognition mode. Both the drop-down menu 28 and the virtual button 30 may be user interactive graphic elements such as touch targets, which may be triggered, converted or otherwise interact with the UI device 4 based on the input received at the UI device 4. For example, when the user is speaking, the user can tap the user interface 16 at or near the area of the virtual button 30 to switch the computing device 2 from the voice recognition mode.
[0036] The voice recognition module 8 can transcribe words spoken by the user or input into the computing device 2 in other ways. In one example, the user says "I want to listen to jazz...". Directly or indirectly, the microphone 12 can provide information related to audio data containing words spoken to the voice recognition module 8. The voice recognition module 8 may apply a language model corresponding to the selected language (for example, English, as shown in the language element 26) to transcribe audio data. The voice recognition module 8 can provide information related to transcription to the UI device 4, which in turn can output characters other than the command text 20 at the user interface 16 in the editing area 18-A.
[0037] The voice recognition module 8 may provide the transcribed text to the voice activation module 10. The voice activation module 10 can review the transcribed text of the voice-initiated action. In one example, the voice activation module 10 may determine that the word "listen" in the phrase "I want to listen to jazz" indicates or describes a voice-initiated action. This word corresponds to listening to something, and the voice activation module 10 can determine it as meaning listening to the audio file. Based on the context of the sentence, the voice activation module 10 determines that the user wants to listen to jazz. Therefore, the voice activation module 10 can trigger actions including opening the media player and causing the media player to play jazz. For example, the computing device 2 may play an album identified as a jazz genre stored on a memory device accessible by the computing device 2.
[0038] In response to recognizing that the word "listen" indicates a voice-initiated action, the voice activation module 10 directly or indirectly provides UID 4 with information that recognizes "listen" as corresponding to the voice-initiated action. The UID 4 then changes the visual format of at least one graphical element displayed at the user interface 16 to indicate that the voice-initiated action has been recognized. As shown in the example of Figure 1, the spoken word "listen" has been recognized as a voice command.
[0039] FIG. 1 illustrates graphical elements related to the text "listen" in a different visual format for the words "I want" and "jazz". FIG. 1 illustrates the display of the transcribed text characters 20 and the voice-initiated action text 22 (also referred to herein as "commands
Text 22") of the editing area 18-A. The command text 22 is a graphic element corresponding to the voice-initiated action transcribed by the voice recognition module 8 and recognized as a voice command by the voice activation module 10. The command text 22 may be visually different The non-command text in the text character 20. For example, Figure 1 illustrates the command text 22 (for example, "LISTEN TO") as written in uppercase letters and underlined, while the non-command text 20 is generally lowercase and not Underlined (for example, "I want" and "Jazz").
[0040] In another example, the visual format of the icon 24 may change when a voice-initiated action is detected. In Fig. 1, the icon 24 is in the image of the microphone. The icon 24 may initially have this image because the computing device 2 is in the voice recognition mode. In response to the voice activation module 10 determining that the audio data contains a voice-initiated action, the UID 4 may change the visual format of the icon 24. For example, UID 4 may change the icon 24 to have a visual format related to the action requested by the voice-initiated action. In this example, the icon 24 may be changed from a first visual format (for example, a microphone) to a visual format related to a voice-initiated action (for example, a play icon for playing a media file). In some examples, the icon 24 may undergo an animation change between the two visual formats.
[0041] In this way, the technology of the present disclosure may enable the computing device 2 to update the voice recognition graphical user interface 16, in which the command text 22 and the command text 22 and One or both of icons 24. The techniques of the present disclosure may enable the computing device 2 to provide an indication that a voice-initiated action has been recognized and is about to be or is being taken. The present technology may further enable the user to verify or confirm that the action to be taken is the action that the user wants the computing device 2 to take with his voice command, or cancel the action if the action is incorrect or for any other reason. A computing device 2 configured with these features can provide users with increased confidence that a voice-initiated action is or can be achieved. This can improve the user's overall satisfaction with the computing device 2 and its voice recognition features. The technology can improve user experience by using voice control of computing devices configured according to various technologies of the present disclosure.
[0042] FIG. 2 is a block diagram illustrating an example computing device 2 for providing a graphical user interface including a visual indication of a recognized voice-initiated action in accordance with one or more aspects of the present disclosure. The computing device 2 of FIG. 2 is described below in the context of FIG. 1. Figure 2 illustrates only one specific example of computing device 2, and many other examples of computing device 2 may be used in other situations. Other examples of the computing device 2 may include a subset of the components included in the exemplary computing device 2 or may include additional components not shown in FIG. 2.
[0043] As shown in the example of FIG. 2, the computing device 2 includes a user interface device 4 ("UID 4"), one or more processors 40, one or more input devices 42, one or more microphones 12, One or more communication units 44, one or more output devices 46, and one or more storage devices 48. The storage device 48 of the computing device 2 also includes a UID module 6, a voice recognition module 8, a voice activation module 10, application modules 14A-14N (collectively referred to as "application modules 14 "), a language database 56, and an action data storage 58. One or more communication channels 50 may interconnect each of the components 4, 40, 42, 44, 46, and 48 for inter-component communication (physically, communicatively, and/or operationally) . In some examples, the communication channel 50 may include a system bus, a network connection, an inter-process communication data structure, or any other technology for transferring data.
[0044] One or more input devices 42 of the computing device 2 may receive input. Examples of inputs are tactile, motion, audio, and video inputs. The input device 42 of the computing device 2 includes, in one example, a presence-sensitive display 5, a touch-sensitive screen, a mouse, a keyboard, a voice response system, a camera, a microphone (such as the microphone 12) or any other for detecting input from a human or machine Type of equipment.
[0045] One or more output devices 46 of the computing device 2 may generate output. Examples of output are tactile, audio, electromagnetic, and video output. In one example, the output device 46 of the computing device 2 includes presence-sensitive displays, speakers, cathode ray tube (CRT) monitors, liquid crystal displays (LCD), motors, actuators, electromagnets, piezoelectric sensors, or To humans
Or any other type of device that produces output from a machine. The output device 46 may utilize one or more of a sound card or a horizontal graphics adapter card to produce audible or visual output, respectively.
[0046] The one or more communication units 44 of the computing device 2 may communicate with external devices via one or more networks by transmitting and/or receiving network signals on one or more networks. The communication unit 44 can be connected to any public or private communication network. For example, the computing device 2 may use the communication unit 44 to transmit and/or receive radio signals on a radio network such as a cellular radio network. Likewise, the communication unit 44 may transmit and/or receive satellite signals on a global navigation satellite system (GNNS) such as a global positioning system (GPS). Examples of the communication unit 44 include a network interface card (for example, an Ethernet card), an optical transceiver, a radio frequency transceiver, a GPS receiver, or any other type of device that can send or receive information. Other examples of the communication unit 44 may include shortwave radios, cellular data radios, wireless Ethernet radios, and universal serial bus (USB) controllers.
[0047] In some examples, the UID 4 of the computing device 2 may include the functions of the input device 42 and/or the output device 46. In the example of FIG. 2, UID 4 may be or may include presence sensitive display 5. In some examples, the presence-sensitive display 5 may detect objects at and/or near the presence-sensitive display 5. As an example range, the presence sensitive display 5 can detect an object such as a finger or a stylus pen within six centimeters or less of the presence of the sensitive display 5. The presence-sensitive display 5 can determine the position (for example, (x,y) coordinates) of the presence-sensitive display 5 where the object is detected. In another example range, the presence-sensitive display 5 can detect objects that are fifteen centimeters or less from the presence-sensitive display 5, and other ranges are also possible. The presence of the sensitive display 5 can determine the position of the screen selected by the user's finger using capacitive, inductive, and/or optical recognition technology. In some examples, the presence sensitive display 5 uses tactile, audio, or video stimuli as described with respect to the output device 46 to provide output to the user. In the example of FIG. 2, UID 4 presents a user interface (such as user interface 16 of FIG. 1) at the presence-sensitive display 5 of UID 4.
[0048] Although illustrated as an internal component of the computing device 2, UID 4 also represents an external component that shares a data path with the computing device 2 for transmitting and/or receiving input and output. For example, in one example, UID 4 represents a built-in component (for example, a screen on a mobile phone) of the computing device 2 that is located in the outer package of the computing device 2 and is physically connected to the outer package. In another example, UID 4 represents an external component of the computing device 2 that is located outside the packaging of the computing device 2 and is physically separated therefrom (for example, a monitor, a projector that shares a wired and/or wireless data path with a tablet computer) Wait).
[0049] One or more storage devices 48 within the computing device 2 may store information for processing during the operation of the computing device 2 (for example, the computing device 2 may store data in the voice recognition module during execution at the computing device 2). 8 and the voice database 56 and the action data storage 58 accessed by the voice activation module 10). In some examples, the storage device 48 acts as temporary storage, meaning that the storage device 48 is not used for long-term storage. The storage device 48 on the computing device 2 can be configured as a volatile memory for short-term storage of information, and therefore does not retain the stored content if the power is cut off. Examples of volatile memory include random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), and other forms of volatile memory known in the art.
[0050] In some examples, the storage device 48 also includes one or more computer-readable storage media. The storage device 48 may be configured to store a larger amount of information than volatile memory. The storage device 48 may be further configured as a non-volatile memory space for long-term information storage, and to retain information after power-on/power-off cycles. Examples of non-volatile memory include magnetic hard disks, optical disks, floppy disks, flash memory, or various forms of electrically programmable memory (EPROM) or electrically erasable programmable (EEPROM) memory. The storage device 48 may store program instructions and/or data associated with the modules 6, 8, 10, and 14.
[0051] One or more processors 40 may implement functions and/or execute instructions in the computing device 2. For example, computing device 2
The upper processor 40 can receive and execute the instructions stored by the storage device 60 to execute the functions of the UID module 6, the voice recognition module 8, the voice activation module 10, and the application module 14. These instructions executed by the processor 40 may cause the computing device 2 to store information in the storage device 48 during the execution of the program. The processor 40 can execute the instructions in the modules 6, 8, and 10 to cause the UID 4 to display a user interface 16 with graphical elements when the computing device 2 recognizes a voice-initiated action, which has a visual format different from the previous visual format . That is, modules 6, 8, and 10 can be operated by the processor 40 to perform various actions, including transcribing and receiving audio data, analyzing the audio data for voice-initiated actions, and updating the presence-sensitive display 5 of UID 4 to change The visual format of the graphical elements associated with the voice-initiated action. In addition, the UID module 6 can be operated by the processor 40 to perform various actions, including receiving gesture instructions at the location of the presence-sensitive display 5 of the UID 4 and causing the UID 4 to present the user interface 14 at the presence-sensitive display 5 of the UID 4. .
[0052] According to aspects of the present disclosure, the computing device 2 of FIG. 2 may output a voice recognition GUI having at least one element in a first visual format at the user interface device 4. The microphone 12 of the computing device 2 receives audio data. Before performing a voice initiation action based on audio data and while receiving additional audio data, UID 4 outputs an updated voice recognition GUI in which the at least one element is presented in a second visual format different from the first visual format Provide an indication that a voice-initiated action has been recognized.
[0053] The voice recognition module 8 of the computing device 2 may receive from the microphone 12, for example, one or more indications of audio data detected at the microphone 12. Generally, the microphone 12 can provide received audio data or an indication of audio data, and the voice recognition module 8 can receive audio data from the microphone 12. The voice recognition module 8 may determine whether the information corresponding to the audio data received from the microphone 12 includes voice. Using voice recognition technology, the voice recognition module 8 can transcribe audio data. If the audio data does include speech, the speech recognition module 8 can use the language database 6 to transcribe the audio data.
[0054] The voice recognition module 8 may also determine whether the audio data includes the voice of a specific user. In some examples, if the audio data corresponds to a human voice, the voice recognition module 8 determines whether the voice belongs to the previous user of the computing device 2. If the voice in the audio data does belong to the previous user, the voice recognition module 8 can modify the voice recognition technology based on certain characteristics of the user's voice. These characteristics may include pitch, accent, rhythm, fluency, consonants, fixed pitch, resonance, or other characteristics of speech. Taking into account the known characteristics about the user's voice, the voice recognition module 8 can improve the result of transcribing audio data for the user.
[0055] In an example where the computing device 2 has more than one user using voice recognition, the computing device 2 may have a profile for each user. The voice recognition module 8 may update the user's profile in response to receiving additional voice input from the user in order to improve the voice recognition for the user in the future. In other words, the voice recognition module 8 can be adapted to the specific characteristics of each user of the computing device 2. The voice recognition module 8 can be adapted to each user by using machine learning technology. These voice recognition features of the voice recognition module 8 may be optional for each user of the computing device 2.
[0056] In some examples, the voice recognition module 8 transcribes the voice in the audio data received from the microphone 12 directly or indirectly by the voice recognition module 8. The voice recognition module 8 can provide the UI device 4 with text data related to the transcribed voice. For example, the voice recognition module 8 provides the characters of the transcribed text to the UI device 4. The UI device 4 may output the text related to the transcribed voice recognized in the information related to the transcribed voice at the user interface 16 for display.
[0057] The voice activation module 10 of the computing device 2 may receive the transcribed voice text characters from the audio data detected at the microphone 12 from, for example, the voice recognition module 8. The voice activation module 10 may analyze the transcribed text or audio data to determine whether it includes keywords or phrases that activate voice-initiated actions. In some examples, the voice activation module 10 compares words or phrases from audio data with a list of actions that can be triggered by voice activation. For example, the action list may be a list of verbs such as run, play, close, open, start, email, etc. Voice activation module 10
The action data store 58 can be used to determine whether a word or phrase corresponds to an action. That is, the voice activation module 10 can compare words or phrases from the audio data with the action data storage 58. The action data store 58 may contain data of words or phrases associated with actions.
[0058] Once the voice initiation module 10 recognizes the word or phrase that activates the voice-initiated action, the voice activation module 10 causes the UID 4 to display graphical elements in the user interface 16 in a second, different visual format to indicate that the voice-initiated action has been Successfully identified. For example, when the voice activation module 10 determines a word in the transcribed text corresponding to the voice-initiated action, UID 4 outputs the word from the first visual format (which may be the same as the rest of the transcribed text). Visual format) becomes a second, different visual format. For example, keywords or phrases related to a voice-initiated action immediately or approximately immediately adopt a different style in the transcribed display to indicate that the computing device 2 recognizes the voice-initiated action. In another example, when the computing device 2 recognizes a voice-initiated action, the icon or other image is transformed from one visual format to another visual format, which can initiate the action based on the recognized voice.
[0059] The computing device 2 may further include one or more application modules 14-A to 14-N. In addition to other modules specifically described in this disclosure, the application module 14 may also include any other applications executable by the computing device 2. For example, the application module 14 may include a web browser, a media player, a file system, a map program, or any other number of applications or features that the computing device 2 may include.
[0060] The techniques described herein may enable the computing device 2 to improve the user's experience when using voice commands to control the computing device 2. For example, the technology of the present disclosure may enable the computing device 2 to output a visual indication that a voice-initiated action has been accurately recognized. For example, the computing device 2 outputs the graphical element associated with the voice-initiated action in a visual format that is different from the visual format of the similar graphical element that is not associated with the voice-initiated action. In addition, the computing device 2 indicates that the voice-initiated action has been recognized, which can provide the user with increased confidence that the computing device 2 can achieve or is achieving the correct voice-initiated action. The computing device 2 that outputs the graphic element in the second visual format can improve the user's overall satisfaction with the computing device 2 and its voice recognition features.
[0061] The techniques described herein may further enable the computing device 2 to provide the user with an option to confirm whether the computing device 2 correctly used the audio data to determine the action. In some examples, if the computing device 2 receives an indication that it did not correctly determine the action, it can cancel the action. In another example, the computing device 2 performs the voice-initiated action only when it receives an indication that the computing device 2 correctly determined the action. The techniques described herein can improve the performance and overall ease of use of the computing device 2.
[0062] FIG. 3 is a block diagram illustrating an example computing device 100 that outputs graphical content for display at a remote device in accordance with one or more techniques of the present disclosure. Graphical content may generally include any visual information that can be output for display, such as text, images, a set of moving images, and so on. The example shown in FIG. 3 includes a computing device 100, a presence-sensitive display 101, a communication unit 110, a projector 120, a projector screen 122, a mobile device 126, and a visual display device 130. Although shown in FIGS. 1 and 2 as a stand-alone computing device 2 for illustrative purposes, a computing device such as the computing device 100 may generally be any component or system that includes a processor for executing software instructions or other suitable computing The environment, and does not need to include, for example, the presence of sensitive displays.
[0063] As shown in the example of FIG. 3, the computing device 100 may be a processor including functions as described with respect to the processor 40 in FIG. In such examples, the computing device 100 may be operatively coupled to the presence-sensitive display 101 by a communication channel 102A, which may be a system bus or other suitable connection. As described further below, the computing device 100 may also be operatively coupled to the communication unit 110 by a communication channel 102B, which may also be a system bus or other suitable connection. Although shown separately in FIG. 3 as an example, the computing device 100 may be operatively coupled by any number of one or more communication channels
Until there is a sensitive display 101 and a communication unit 110.
[0064] In other examples such as previously illustrated with computing device 2 in FIGS. 1 to 2, computing device may refer to portable or mobile devices such as mobile phones (including smart phones), laptop computers, and the like. In some examples, the computing device may be a desktop computer, a tablet computer, a smart TV platform, a camera, a personal digital assistant (PDA), a server, a host, etc.
[0065] The presence-sensitive display 101 (such as the example of the user interface device 4 shown in FIG. 1) may include a display device 103 and a presence-sensitive input device 105. The display device 103 may, for example, receive data from the computing device 100 and display graphical content associated with the data. In some examples, the presence-sensitive input device 105 may use capacitive, inductive, and/or optical recognition techniques to determine the presence of one or more user inputs at the sensitive display 101 (eg, continuous gestures, multi-touch gestures, single-point Touch gestures, etc.), and use the communication channel 102 to send such user-input instructions to the computing device 100. In some examples, the presence of a sensitive input device 105 may be physically located on top of the display device 103, such that when the user positions the input unit on a graphical element displayed by the display device 103, there is a sensitive input device 105 there. Corresponds to the position of the display device 103 where the graphic element is displayed. In other examples, the presence of the sensitive input device 105 can be physically located separately from the display device 103, and the location of the presence of the sensitive input device 105 can correspond to the location of the display device 103, so that input can be performed at the presence of the sensitive input device 105 It is used to interact with the graphic element displayed at the corresponding position of the display device 103.
[0066] As shown in FIG. 3, the computing device 100 may also include and/or be operatively coupled with a communication unit 110. The communication unit 110 may include the functions of one or more communication units 44 as described in FIG. 2. Examples of the communication unit 110 may include a network interface card, an Ethernet card, an optical transceiver, a radio frequency transceiver, or any other type of device that can send and receive information. Other examples of such communication units may include Bluetooth, 3G, and Wi-Fi radios, universal serial bus (USB) interfaces, and so on. The computing device 100 may also include and/or be operatively coupled with one or more other devices, such as input devices, output devices such as those shown in FIGS. 1 and 2 Equipment, memory, storage device, etc.
[0067] FIG. 3 also illustrates the projector 120 and the projector screen 122. Other such examples of projection devices may include electronic whiteboards, holographic display devices, and any other suitable devices for displaying graphic content. The projector 120 and the projector screen 122 may include one or more communication units that enable each device to communicate with the computing device 100. In some examples, one or more communication units may enable communication between the projector 120 and the projector screen 122. The projector 120 may receive data including graphic content from the computing device 100. The projector 120 may project graphic content onto the projector screen 122 in response to receiving the data. In some examples, the projector 120 may use optical recognition or other appropriate technology to determine one or more user inputs at the projector screen (eg, continuous gestures, multi-touch gestures, single-touch gestures, etc.), and use One or more communication units send such user-input instructions to the computing device 100. In such examples, the projector screen 122 may be unnecessary, and the projector 120 may project graphical content on any suitable medium and use optical recognition or other such suitable technology to detect one or more user inputs.
[0068] In some examples, the projector screen 122 may include a presence-sensitive display 124. The presence sensitive display 124 may include a subset or all of the functions of the UI device 4 as described in this disclosure. In some examples, the presence sensitive display 124 may include additional functionality. The projector screen 122 (for example, an electronic whiteboard) may receive data from the computing device 100 and display graphical content. In some examples, the presence-sensitive display 124 may use capacitive, inductive, and/or optical recognition techniques to determine one or more user inputs at the projector screen 122 (eg, continuous gestures, multi-touch gestures, single-point touch Gestures, etc.), and use one or more communication units to send such user-input instructions to the computing device 100.
[0069] FIG. 3 also illustrates the mobile device 126 and the visual display device 130. The mobile device 126 and the visual display device 130 may each include computing and connectivity capabilities. Examples of mobile devices 126 may include e-reader devices, convertible notebook devices, hybrid tablet devices, and so on. Examples of the visual display device 130 may include other semi-fixed devices such as televisions, computer monitors, and the like. As shown in FIG. 3, the mobile device 126 may include a presence-sensitive display 128. The visual display device 130 may include a presence-sensitive display 132. The presence sensitive display 128, 132 may include a subset or all of the functions of the presence sensitive display 4 as described in this disclosure. In some examples, the presence sensitive display 128, 132 may include additional functionality. In any case, the presence-sensitive display 132 can receive data from the computing device 100 and display graphical content, for example. In some examples, the presence-sensitive display 132 may use capacitive, inductive, and/or optical recognition techniques to determine one or more user inputs at the projector screen (eg, continuous gestures, multi-touch gestures, single-touch gestures Etc.), and use one or more communication units to send such user-input instructions to the computing device 100.
[0070] As described above, in some examples, the computing device 100 may output graphical content for display at the presence-sensitive display 101 coupled to the computing device 100 by a system bus or other suitable communication channel. The computing device 100 may also output graphical content for display at one or more remote devices such as the projector 120, the projector screen 122, the mobile device 126, and the visual display device 130. For example, according to the technology of the present disclosure, the computing device 100 may execute one or more instructions to generate and/or modify graphical content. The computing device 100 may output data including graphic content to a communication unit (such as the communication unit 110) of the computing device 100. The communication unit 110 may send data to one or more of remote devices such as the projector 120, the projector screen 122, the mobile device 126, and/or the visual display device 130. In this way, the computing device 100 can output graphical content for display at one or more of the remote devices. In some examples, one or more of the remote devices may output graphical content at a presence-sensitive display included in and/or operatively coupled to the corresponding remote device.
[0071] In some examples, the computing device 100 may not output graphical content at the presence-sensitive display 101 that is operatively coupled to the computing device 100. In other examples, the computing device 100 may output graphical content for display at both the presence-sensitive display 101 and one or more remote devices coupled to the computing device 100 by the communication channel 102A. In such examples, the graphical content may be displayed at each corresponding device substantially simultaneously. For example, the communication delay of sending data including graphic content to a remote device may introduce a certain delay. In some examples, the graphical content generated by the computing device 100 and output for display at the presence-sensitive display 101 may be different from the graphical content output for display at one or more remote devices.
[0072] The computing device 100 may use any suitable communication technology to send and receive data. For example, the computing device 100 may be operatively coupled to the external network 114 using the network link 112A. Each remote device illustrated in FIG. 3 may be operatively coupled to the network external network 114 by one of the respective network links 112B, 112C, and 112D. The external network 114 may include a network hub, a network switch, a network router, etc., which are operatively coupled to each other to provide information exchange between the computing device 100 and the remote device illustrated in FIG. 3. In some examples, the network links 112A to 112D may be Ethernet, ATM, or other network connections. Such connections can be wireless and/or wired connections.
[0073] In some examples, direct device communication 118 may be used to operatively couple the computing device 1000 to one or more of the remote devices included in FIG. 3. Direct device communication 118 may include communication through which the computing device 100 directly sends and receives data with a remote device using wired or wireless communication. That is, in certain examples of direct device communication 118, the data sent by the computing device 100 may not be forwarded by one or more additional devices before being received at the remote device, and vice versa. Examples of direct device communication 118 may include Bluetooth, near field communication, universal serial bus, WiFi, infrared, and the like. The communication links 116A to 116D can connect one or more of the remote devices illustrated in FIG. 3 to the computing device.
The device 100 is operatively coupled. In some examples, the communication links 116A to 116D may be connections using Bluetooth, near field communication, universal serial bus, infrared, or the like. Such connections can be wireless and/or wired connections.
[0074] According to the technology of the present disclosure, the computing device 100 may be operatively coupled to the visual display device 130 using the external network 114. The computing device 100 may output a graphical keyboard for display at the presence-sensitive display 132. For example, the computing device 100 may transmit data including a representation of a graphical keyboard to the communication unit 110. The communication unit 110 may use the external network 114 to transmit data including a representation of a graphical keyboard to the visual display device 130. The visual display device 130 may cause the presence-sensitive display 132 to output a graphical keyboard in response to receiving data using the external network 114. In response to the user performing a gesture at the presence-sensitive display 132 (for example, at an area where the graphical keyboard is outputted from the presence-sensitive display 132), the visual display device 130 may use the external network 114 to send an indication of the gesture to the computing device 100. The communication unit 110 may receive an instruction of the gesture and send the instruction to the computing device 100.
[0075] In response to receiving the voice included in the audio data, the computing device 100 may transcribe the voice into text. The computing device 100 may cause one of the display devices such as the presence-sensitive input display 105, the projector 120, the presence-sensitive display 128, or the presence-sensitive display 132 to output a graphic element in a first visual format, which may include at least the transcribed text Part. The computing device 100 may determine that the voice includes a voice-initiated action, and cause one of the display devices 105, 120, 128, or 132 to output a graphical element related to the voice-initiated action. The graphical element may be output in a second visual format different from the first visual format to indicate that the computing device 100 has detected a voice-initiated action. The computing device 100 may perform a voice-initiated action.
[0076] FIGS. 4A to 4D are screenshots illustrating an example graphical user interface (GUI) of a computing device for navigating an example in accordance with one or more techniques of the present disclosure. The computing device 200 of FIGS. 4A to 4D may be any computing device including a mobile computing device as discussed above in relation to FIGS. 1 to 3. Furthermore, the computing device 200 may be configured to include any subset of the features and techniques described herein as well as additional features and techniques. 4A to 4D include graphic elements 204-A to 204-C (collectively referred to as "graphic elements 204") that can have different visual formats.
[0077] FIG. 4A depicts a computing device 200 having a graphical user interface (GUI) 202 and operating a state in which the computing device 200 can receive audio data. For example, a microphone such as the microphone 12 of FIGS. 1 and 2 may be initialized and capable of detecting audio data including voice. GUI 202 may be a voice recognition GUI. GUI 202 includes graphical elements 202 and 204-A. The graphical element 202 is text and expresses "Speak Now", which may indicate that the computing device 200 is capable of receiving audio data. The graphic element 204-A is an icon representing a microphone. Therefore, the graphical element 204-A may indicate that the computing device 200 is capable of performing the action of recording audio data.
[0078] FIG. 4B illustrates the computing device 200 outputting the GUI 206 in response to receiving audio data in FIG. 4A<sub>o</sub>GUI 206 includes graphical elements 204-A, 208, and 210. In this example, the computing device 200 has used, for example, the voice recognition module 8 and the language database 56 to transcribe the received audio data. As indicated by the microphone icon 204-A, the computing device 200 may still be receiving additional audio data. The transcribed audio data is output as text in the graphic element 208 and includes the word "I want to navigate to". The graphical element 210 may further indicate that the computing device 200 may still be receiving additional audio data or the voice recognition module 8 may still be transcribing the received audio data.
[0079] GUI 206 includes graphical elements 208 in a first visual format. That is, the graphic element 208 includes text with a specific font, size, color, position, and so on. The word "navigate to" is included as part of the graphical element 208 and presented in the first visual format. Likewise, GUI 206 includes graphical elements 204-A in a first visual format. The first visual format of the graphic element 204-A is an icon including an image of a microphone. The graphical element 204-A may indicate an action that the computing device 200 is or will perform.
[0080] FIG. 4C depicts the computing device 200 outputting the updated GUI 212. The updated GUI 212 includes graphical elements 204-B, 208, 210, and 214. In this example, the voice activation module 10 may have analyzed the transcribed audio data and recognized the voice-initiated action. For example, the voice activation module 10 may compare one or more words or phrases in the transcribed text shown in the graphic element 208 with the action data store 58. In this example, the voice activation module 10 determines that the phrase "navigate to" corresponds to the voice-initiated action instruction. In response to detecting the action instruction, the voice activation module 10 may have instructed the UID module 6 to output the updated GUI 212 where, for example, the sensitive display 5 is present.
[0081] The updated GUI 212 includes an updated graphical element 204-B in a second visual format. The graphical element 204-B is an icon depicting an image of an arrow, which may be associated with the navigation feature of the computing device 200. Conversely, the graphic element 204-A is an icon depicting a microphone. Therefore, the graphic element 204-B has the second visual format, and the graphic element 204-A has the first visual format. The icon of the graphical element 204-B indicates that the computing device 200 may perform a voice-initiated action, such as performing a navigation function.
[0082] Similarly, the updated GUI 202 also includes an updated graphical element 214. The graphical element 214 includes the word "navigate to" in a second visual format other than in the GUI 206. In the GUI 202, the second visual format of the graphical element 214 includes highlighting provided by colored or shaded shapes around the word and bolding of the word. In other examples, other characteristics or visual aspects, including size, color, font, style, location, etc., that "navigate to" can be changed from the first visual format to the second visual format. The graphical element 214 provides that the computing device 200 has recognized The voice in the audio data indicates that the action is initiated. In some examples, the GUI 212 provides additional graphical elements that indicate that the computing device 2 requires confirmation before performing the voice-initiated action.
[0083] In FIG. 4D, the computing device 200 has continued to receive and transcribe audio data since the GUI 212 was displayed. The computing device 200 outputs the updated GUI 216<sub>O</sub>GUI 216 includes graphic elements 204-(3, 208, 214, 218, 220, and 222. Graphic element 204-C has re-taken the first visual format, that is, the image of the microphone, because the computing device 200 has performed a voice-initiated action and is Continue to detect audio data.
[0084] The computing device 200 receives and transcribes the additional word Starbucks in FIG. 4D. In summary, in this example, the computing device 200 has detected and transcribed the sentence "I want to navigate to Starbucks." The voice activation module 10 may have determined that Starbucks is the place to which the speaker (for example, the user) wishes to navigate. The computing device 200 has performed a voice-initiated action recognition action and navigated to Starbuckso. Therefore, the computing device 200 has executed a navigation application and performed a search for Starbucks. In one example, the computing device 200 uses the context information to determine what the voice-initiated action is and how to perform the action. For example, the computing device 200 may have used the current location of the computing device 200 to centrally search for the local Starbucks location based on it.
[0085] The graphic element 208 may include only a part of the transcribed text, so that a graphic element representing a voice-initiated action may be included in the GUI 216, that is, the graphic element 214. The GUI 216 includes a map graphic element 220 showing the location of Starbucks. Graphical element 22 may include an interactive list of Starbucks locations.
[0086] In this manner, graphical elements 204-B and 214 may be updated to indicate that the computing device 200 has recognized a voice-initiated action and can perform a voice-initiated action. The computing device 200 configured in accordance with the techniques described herein can provide a user with an improved experience of interacting with the computing device 200 via voice commands.
[0087] FIGS. 5A to 5B are screenshots illustrating an example GUI of a computing device for a media playback example in accordance with one or more techniques of the present disclosure. The computing device 200 of FIGS. 5A and 5B may be any computing device including a mobile computing device as discussed above with respect to FIGS. 1 to 4D. Furthermore, the computing device 200 may be configured to include any subset of the features and techniques described herein as well as additional features and techniques.
[0088] FIG. 5A illustrates the computing device 200 outputting a GUI 240 including graphical elements 242, 244, 246, and 248. The graphic element 244 corresponds to the text "I want..." transcribed by the voice recognition module 8 and is presented in the first visual format. The graphic element 246 is the text of the phrase of the voice activation module 10 recognized as the voice-initiated action "listening", and is presented in a second visual format that is different from the first visual format of the graphic element 244. The voice-initiated action may, for example, be playing a media file. The graphic element 242-A is an icon that can represent a voice-initiated action, such as having the appearance of a play button. Graphical element 242-A represents a play button because the voice activation module 10 has determined that the computing device 200 has received a voice instruction to play media including audio components. The graphical element 248 provides an indication that the computing device 200 may still be receiving, transcribing, or analyzing audio data.
[0089] FIG. 5B illustrates the computing device 200 outputting a GUI 250 including graphical elements 242-B, 244, 246, and 248. The graphic element 242-B has a visual format corresponding to the image of the microphone to indicate that the computing device 200 is capable of receiving audio data. The graphic element 242-B no longer has a visual format corresponding to the voice-initiated action, that is, the image of the play button, because the computing device 200 has performed an action related to the voice-initiated action, which may be a voice-initiated action.
[0090] The voice activation module 10 has determined that the voice-initiated action "listening" is applicable to the word "killer", which may be a band. The computing device 200 may have determined an application, such as a video or audio player, to play media files that include audio components. The computing device 200 may also have determined media files that meet the requirements (satisfy the "killer" requirements), such as music stored on a local storage device (such as the storage device 48 of FIG. 2) accessible through a network such as the Internet file. The computing device 200 has performed the task of executing applications to play such files. The application may be, for example, a media player application, which instructs UID 4 to output a GUI 250, which includes graphical elements 252 related to a playlist for the media player application.
[0091] FIG. 6 is a conceptual diagram illustrating a series of example visual formats in which elements may vary based on different voice-initiated actions according to one or more techniques of the present disclosure. The element may be a graphic element such as the graphic elements 204 and 242 of FIGS. 4A to 4D, 5A, and 5B. This element can change the visual format represented by the images 300-1 to 300-4, 302-1 to 302-5, 304-1 to 304-5, and 306-1 to 306-5.
[0092] The image 300-1 represents a microphone, and may be the first visual format of the user interface element. When the element has the visual format of image 300-1, a computing device such as computing device 2 may be able to receive audio data from an input device such as microphone 12. In response to the computing device 200 determining that a voice-initiated action corresponding to the command to play the media file has been received, the visual format of the element may be changed from image 300-1 to image 302-1. In some examples, image 300 The -1 transforms into an image 302-1, which may be an animation. For example, the image 300-1 becomes the image 302-1, and in doing so, the element takes the intermediate figures 300-2, 300-3, and 300-4.
[0093] Similarly, in response to the computing device 2 determining that it has received a voice-initiated action to stop playing the media file after it starts playing, the computing device 2 may change the visual format of the element from image 302-1 to image 304- 1, the image corresponding to the stop. Image 302-1 may adopt intermediate images 302-2, 302-3, 302-4, and 302-5 as it is transformed into image 304-1.
[0094] Similarly, in response to the computing device 2 determining that it has received a voice-initiated action to pause the playback of the media file after it starts playing, the computing device 2 may change the visual format of the element from image 304-1 to image 306- 1, the image corresponding to the pause. The image 304-1 can take the intermediate images 304-2, 304-3, 304-4, and 304-5 as it is transformed into the image 306-1.
[0095] In addition, in response to the computing device 2 determining that the additional voice-initiated action has not been received for a predetermined period of time, the computing device 2 may change the visual format of the element from image 306-1 back to image 300-1, that is, with audio recording The corresponding image. Figure
Image 306-1 may adopt intermediate images 306-2, 306-3, 306-4, and 306-5 as it is transformed into image 300-1<sub>o</sub>In other examples, the element may be deformed or changed into other visual formats with different images.
[0096] FIG. 7 is a flowchart illustrating an example process 500 for a computing device to visually confirm a recognized voice-initiated action in accordance with one or more techniques of the present disclosure. The process 500 will be discussed in terms of the computing device 2 performing the process 500 of FIGS. 1 and 2. However, any computing device, such as the computing device 100 or 200 of FIGS. 3, 4A to 4D, 5A, and 5D, may perform the process 500.
[0097] Process 500 includes outputting, by computing device 2, a voice recognition graphical user interface (GUI) (such as GUI 16 or 202) having at least one element in a first visual format for display (510). This element can be, for example, an icon or text. The first visual format may be a first image (such as microphone image 300-1) or one or more words (such as non-command text 208).
[0098] The process 500 further includes receiving audio data by the computing device 2 (520). For example, the microphone 12 detects environmental noise. Process 500 may further include determining, by the computing device, a voice-initiated action based on the audio data (530). For example, the voice recognition module 8 can determine the voice-initiated action based on the audio data. Examples of voice-initiated actions can include sending text messages, listening to music, getting directions, calling businesses, calling contacts, sending emails, watching maps, going to websites, writing notes, replaying the last number, opening apps, calling voice mailboxes , Read appointments, check phone status, search the web, check signal strength, check network, check battery, or any other actions.
[0099] The process 500 may further include the computing device 2 transcribing audio data and outputting an updated voice recognition GUI for display while receiving additional audio data and before performing a voice initiation action based on the audio data. In the voice recognition GUI, the at least one element is displayed in a second visual format different from the first visual format to indicate that the voice-initiated action has been recognized, such as the graphical element 214 (540) shown in FIG. 4C.
[0100] In some examples, outputting the voice recognition GUI further includes outputting a portion of the transcribed audio data, and wherein outputting the updated voice recognition GUI further includes cropping at least a portion of the transcribed audio data so that the voice is initiated One or more words of the transcribed audio data related to the action are displayed. In some examples where the computing device 2 has a relatively small screen, the displayed transcribed text may focus more on words corresponding to the voice-initiated action.
[0101] The process 500 further includes outputting an updated voice recognition GUI such as the GUI 212 before performing a voice initiation action based on the audio data and while receiving additional audio data, in a second visual format different from the first visual format To present the at least one element to provide an indication that a voice-initiated action has been recognized. In some examples, the second visual format is different from the first visual format in terms of image, color, font, size, highlighting, style, and location.
[0102] Process 500 may also include computing device 2 analyzing audio data to determine a voice-initiated action. The computing device 2 may analyze the transcription of audio data to determine a voice-initiated action based at least in part on a comparison of words or phrases of the transcribed audio data with a database of actions. The computing device can search for keywords in the transcribed audio data. For example, the computing device 2 may detect at least one verb in the transcription of audio data, and compare the at least one verb with a verb set, where each verb in the verb set corresponds to a voice-initiated action. For example, the set of verbs may include listen and "play", both of which may be related to voice-initiated actions to play media files with audio components.
[0103] In some examples, the computing device 2 determines the context of the computing device 2, such as the current location of the computing device 2, what application is currently or recently being executed by the computing device 2, daytime, the identity of the user issuing the voice command , Or any other contextual information. The computing device 2 may use the context information to at least partially determine the voice-initiated action. In some examples
, The computing device 2 captures more audio data before determining that the voice initiates the action. If subsequent words change the meaning of the voice-initiated action, computing device 2 may update the visual format of the element to reflect the new meaning. In some examples, the computing device 2 may use the context to make subsequent decisions, such as which location in a chain hotel to obtain route guidance.
[0104] In some examples, the first visual format of the at least one element has an image representing a voice recognition mode, and wherein the second visual format of the at least one element has an image representing a voice-initiated action. For example, the element represented in FIG. 6 may have a first visual format 300-1 representing a voice recognition mode (for example, a microphone) and a second visual format 302-1 representing a voice-initiated action (for example, playing a media file). In some examples, the image representing the voice recognition mode is transformed into an image representing a voice-initiated action. In other examples, any element having the first visual format can be transformed into the second visual format.
[0105] The computing device 2 may actually perform a voice-initiated action based on audio data. That is, in response to the computing device 2 determining that the voice-initiated action will obtain route guidance to the address, the computing device 2 performs tasks such as executing a map application and searching for route guidance. The computing device 2 can determine the confidence threshold value at which the recognized voice-initiated action is correct. If the confidence level for a specific voice-initiated action is below the confidence threshold, the computing device 2 may request the user's confirmation before proceeding to perform the voice-initiated action.
[0106] In some examples, the computing device 2 performs the voice-initiated action only in response to receiving an indication confirming that the voice-initiated action is correct. For example, the computing device 2 may output a prompt for display before the computing device 2 performs an action, the prompt requesting feedback that the recognized voice initiates the correct action. In some cases, the computing device 2 updates the voice recognition GUI so that it is presented in the first visual format in response to receiving an instruction to cancel the input or in response to not receiving feedback that the recognized voice initiates the correct action within a predetermined period of time. The element. In some examples, the voice recognition GUI includes interactive graphical elements for canceling voice-initiated actions.
[0107] Clause 1. A method comprising: outputting, by a computing device, a voice recognition graphical user interface (GUI) having at least one element in a first visual format for display; receiving audio data by the computing device; by The computing device determines a voice-initiated action based on the audio data; and while receiving additional audio data and before performing the voice-initiated action based on the audio data, output an updated voice recognition GUI for use Displaying, displaying the at least one element in a second visual format different from the first visual format in the updated voice recognition GUI to indicate that the voice-initiated action has been recognized.
[0108] Clause 2. The method of Clause 1, further comprising: determining, by the computing device, a transcription based on the audio data; identifying one or more words of the transcription associated with the speech-initiated action , Wherein the at least one element includes at least a part of the one or more words; and the computing device outputs a transcription that does not include the one or more words before outputting the updated voice recognition GUI A part is used for display in the first visual format.
[0109] Clause 3. The method according to any one of Clauses 1 to 2, wherein, in one or more aspects of image, color, font, size, highlighting, style, and position, the first The second visual format is different from the first visual format.
[0110] Clause 4. The method of any one of clauses 1 to 3, wherein the computing device determines the transcription of the audio data, wherein: outputting the voice recognition GUI further includes outputting the transcription At least a part, and outputting the updated voice recognition GUI further includes cropping the at least part of the transcription that is output so that the one or more words of the transcription related to the voice initiation action are displayed .
[0111] Clause 5. The method according to any one of clauses 1 to 3, wherein the first of the at least one element
A visual format includes an image representing the voice recognition mode of the computing device, and wherein the second visual format of the at least one element includes an image representing the voice-initiated action.
[0112] Clause 6. The method of Clause 5, wherein the image representing the voice recognition mode is morphed into representing the voice-initiated action in response to determining the speech-initiated action based on the audio data image.
[0113] Clause 7. The method according to any one of clauses 1 to 6, further comprising: in response to determining the voice-initiated action based on the audio data, performing the voice-initiated action by the computing device.
[0114] Clause 8. The method of clause 7, wherein performing the voice-initiated action is further responsive to receipt by the computing device confirming that the voice-initiated action is correct.
[0115] Clause 9. The method of any one of clauses 1 to 8, further comprising determining, by the computing device and based at least in part on the audio data, the voice-initiated action.
[0116] Clause 10. The method of clause 9, wherein determining the voice-initiated action further comprises based at least in part on a comparison of a transcribed word or phrase based on the audio data with a pre-configured set of actions To determine the voice-initiated action.
[0117] Clause 11. The method according to any one of clauses 9 to 10, wherein determining the voice-initiated action further comprises: identifying, by the computing device, at least one verb in the transcription; and The at least one verb is compared with one or more verbs from a set of verbs, and each verb in the set of verbs corresponds to at least one action from a plurality of actions.
[0118] Clause 12. The method of any one of clauses 9 to 11, wherein determining the voice-initiated action further comprises: determining by the computing device based at least in part on data from the computing device Context; and the computing device determines the voice-initiated action based at least in part on the context.
[0119] Clause 13. The method according to any one of clauses 9 to 12, further comprising: in response to receiving an instruction to cancel the input, outputting, by the computing device, the at least one element for use Describe the first visual format to display zjs ο
[0120] Clause 14. A computing device comprising: a display device; and one or more processors operable to: output a voice recognition having at least one element in a first visual format A graphical user interface (GUI) for displaying at the display device; receiving audio data; determining a voice-initiated action based on the audio data; and performing while receiving additional audio data and based on the audio data Before the voice initiates the action, output an updated voice recognition GUI for display, and display the at least one element in a second visual format different from the first visual format in the updated voice recognition GUI To indicate that the voice-initiated action has been recognized.
[0121] Clause 15. The computing device of clause 14, wherein the one or more processors are further operable to: determine a transcription based on the audio data; identify the transcription and the voice-initiated action Associated one or more words, wherein the at least one element includes at least a part of the one or more words; and outputting the transcription before outputting the updated voice recognition GUI does not include the one Or a part of a plurality of words for display in the first visual format.
[0122] Clause 16. The computing device of any one of clauses 14 to 15, wherein the first visual format of the at least one element has an image representing the voice recognition mode of the computing device , And wherein the second visual format of the at least one element has an image that represents the voice-initiated action, and wherein the image that represents the voice recognition mode is transformed into all that represents the voice-initiated actionNarrated picture.
[0123] Clause 17. The computing device of any one of clauses 14 to 16, wherein the one or more processors are further operable to respond to determining the voice-initiated action based on the audio data Perform the voice initiation action.
[0124] Clause 18. A computer-readable storage medium encoded with instructions that, when executed by one or more processors of a computing device, cause the one or more processors: A voice recognition graphical user interface (GUI) of at least one element of a visual format for display; receiving audio data; determining the voice-initiated action based on the audio data; and while receiving additional audio data and based on all Before the audio data is used to perform the voice initiation action, output an updated voice recognition GUI for display, and display in a second visual format different from the first visual format in the updated voice recognition GUI The at least one element to indicate that the voice-initiated action has been recognized.
[0125] Clause 19. The computer-readable storage medium according to Clause 18, wherein the instructions further cause the one or more processors to determine a transcription based on the audio data; identify the origin of the transcription and the voice One or more words associated with an action, wherein the at least one element includes at least a part of the one or more words; and the output of the transcription before outputting the updated voice recognition GUI does not include the A part of one or more words is used for display in the first visual format.
[0126] Clause 20. The computer-readable storage medium of any one of clauses 18 to 19, wherein the first visual format of the at least one element includes the voice recognition that represents the computing device Mode image, wherein the second visual format of the at least one element includes an image representing the voice-initiated action, and wherein the image representing the voice recognition mode responds to the determination based on the audio data The voice-initiated action is transformed into the image representing the voice-initiated action.
[0127] Clause 21. A computing device comprising at least one processor and at least one module, the at least one module being operable by the at least one processor to perform any of the methods of clauses 1 to 13 One method.
[0128] Clause 22. A computing device comprising means for performing any of the methods of clauses 1 to 13.
[0129] Clause 23. A computer-readable storage medium including instructions that, when executed by at least one processor of a computing device, configure the computing device to perform any of the methods of clauses 1 to 13 .
[0130] In one or more examples, the described functions may be implemented by hardware, software, firmware, or any combination thereof. If implemented by software, the function can be stored as one or more instructions or codes, or transmitted on or through a computer-readable medium, and executed by a hardware-based processing unit. The computer-readable medium may include a computer-readable storage medium, which corresponds to a tangible medium such as a data storage medium or a communication medium including any medium that facilitates the transfer of a computer program from one place to another place, for example, according to a communication protocol. In this way, the computer-readable medium may generally correspond to: (1) a tangible computer-readable storage medium, which is non-transitory, or (2) a communication medium, such as a signal or carrier wave. A data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, codes, and/or data structures for implementing the techniques described in this disclosure. The computer program product may include a computer readable medium.
[0131] By way of example and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CDROM or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or may be used Any other medium that can store desired program codes in the form of instructions or data structures and can be accessed by a computer. And, any connection is properly termed a computer-readable medium. For example, if you use coaxial cable, fiber optic cable, twisted pair, digital
Subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave transmit commands from a website, server, or other remote source, and the definition of the medium includes coaxial cable, fiber optic cable, twisted pair, DSL, or Wireless technologies such as infrared, radio, and microwave. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but are instead directed to non-transitory, tangible storage media. Discs and discs as used herein include compact discs (CD), laser discs, optical discs, digital versatile discs (DVD), floppy discs, and blu-ray discs, where discs usually reproduce data magnetically, while discs usually use lasers and optical To reproduce the data. Combinations of the above should also be included in the scope of computer-readable media.
[0132] The instructions are executed by one or more processors, such as one or more digital signal processors (DSP), general-purpose microprocessors, application-specific integrated circuits (ASIC), field programmable logic arrays (FPGA), or others. Price integrated or discrete logic circuit. Therefore, the term "processor" as used herein may refer to any of the foregoing structure or any other structure suitable for implementing the techniques described herein. In addition, in some aspects, the functions described herein may be provided in dedicated hardware and/or software modules. Moreover, the technology can be completely implemented by one or more circuits or logic elements.
[0133] The technology of the present disclosure can be implemented in a variety of devices or devices, including wireless handsets, integrated circuits (ICs), or sets of ICs (eg, chipsets). Various components, modules, or units are described in the present disclosure to emphasize the functional aspects of devices configured to implement the disclosed technology, but they do not necessarily require different hardware units to be implemented. Conversely, as described above, various units can be combined in a hardware unit, or provided by many interoperable hardware units. The hardware unit includes one or more processors as described above, and appropriate software and/or firmware. Combine.
[0134] Various embodiments have been described in this disclosure. These and other embodiments are within the scope of the following claims.
11 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11
Every citation, both ways
| Document | Relation | Office | Category | Cited during | Relevant claims |
|---|---|---|---|---|---|
| CN108683937A | Cited by | China | – | Search report | – |
| US11893203B2 | Cited by | United States of America | – | Applicant | – |
| US12425511B2 | Cited by | United States of America | – | Applicant | – |
| US12417596B2 | Cited by | United States of America | – | Applicant | – |
| CN110720085A | Cited by | China | – | Search report | – |
| WO2020135811A1 | Cited by | World Intellectual Property Organization (WIPO) | – | International search | – |
| CN105933073A | Cited by | China | – | Search report | – |
| CN113516979A | Cited by | China | – | Search report | – |
| US11972678B2 | Cited by | United States of America | – | Applicant | – |
| US11144176B2 | Cited by | United States of America | – | Applicant | – |
| CN107832036A | Cited by | China | – | Search report | – |
| US10999426B2 | Cited by | United States of America | – | Applicant | – |
| US12204734B2 | Cited by | United States of America | – | Applicant | – |
| CN112634883A | Cited by | China | – | Search report | – |
| US12386482B2 | Cited by | United States of America | – | Applicant | – |
| US11765114B2 | Cited by | United States of America | – | Applicant | – |
| US11521469B2 | Cited by | United States of America | – | Applicant | – |
| US10971145B2 | Cited by | United States of America | – | Applicant | – |
| US11693529B2 | Cited by | United States of America | – | Applicant | – |
| CN112714896A | Cited by | China | – | Search report | – |
| CN1198203C | Cites | China | A | Search report | 1-12 |
| US2007061148A1 | Cites | United States of America | A | Search report | 1-12 |
| US2010312547A1 | Cites | United States of America | A | Search report | 1-12 |
| US2013080178A1 | Cites | United States of America | A | Search report | 1-12 |
| US8275617B1 | Cites | United States of America | A | Search report | 1-12 |
12 members in 6 offices
Priority claims14
| Document | Office | Kind | Date |
|---|---|---|---|
| 201361860679 | United States of America | P | |
| 201361860679 | United States of America | P | |
| 61860679 | United States of America | – | |
| 14109660 | United States of America | – | |
| 201314109660 | United States of America | A | |
| 201314109660 | United States of America | A | |
| 2014043236 | United States of America | W | |
| 2014043236 | United States of America | W | |
| 14109660 | – | – | – |
| 61860679 | – | – | – |
| PCTUS2014043236 | – | – | – |
| US201314109660 | – | – | – |
| US201361860679P | – | – | – |
| WO2014US43236 | – | – | – |
Members12
| Document | Office | Kind | |
|---|---|---|---|
| US2015040012A1 | United States of America | A1 | |
| WO2015017043A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU2014296734A1 | Australia | A1 | |
| CN105453025AThis record | China | A | |
| KR20160039244A | Republic of Korea | A | |
| EP3028136A1 | European Patent Office (EPO) | A1 | |
| AU2014296734B2 | Australia | B2 | |
| KR101703911B1 | Republic of Korea | B1 | |
| US9575720B2 | United States of America | B2 | |
| US2017116990A1 | United States of America | A1 | |
| CN105453025B | China | B | |
| EP3028136B1 | European Patent Office (EPO) | B1 |
5 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Termination of patent right due to non-payment of annual feeCF01 | CF01 | |
| Patent grantGrantedGR01 | GR01 | |
| Change of applicant informationCB02 | CB02 | |
| Entry into substantive examinationC10 | C10 | |
| PublicationC06 | C06 |
Numbers
- Publication
- 105453025
- Publication, DOCDB
- 105453025
- Publication, EPODOC
- CN105453025
- Application
- 800429369
- Application, DOCDB
- 201480042936
- Application, EPODOC
- CN2014842936
Titles2
- Chinese
- 用于已识别语音发起动作的视觉确认
- English
- Used for visual confirmation of actions initiated by recognized voices
Classification
- CPC, 7
- G10L15/22
- G01C21/3608
- G06F3/04817
- G06F3/167
- G10L15/1815
- G10L2015/223
- G10L2015/228
- IPC, 2
- G06F3 16
- G10L15 22