Nova Patents
EP4535348A2

Multi-modal inputs for voice commands

Abstract

Systems and processes for operating an intelligent automated assistant are provided. In one example process, a first input including activation of an affordance is received. A domain associated with the affordance is determined. A second input including user speech is received, where a user intent is determined based on the domain and the user speech. A determination is made whether the user intent includes a command associated with the affordance. In accordance with a determination that the user intent includes a command associated with the affordance, a task in furtherance of the command is performed.

EP4535348A2, drawing sheet 1
Sheet 1 of 55

Term

12.9 yearsto projected expiry

Projected expiry 30 August 2039, counted from filing; an application has no term until it is granted.

  1. Priority
  2. Filed
  3. Published
  4. Today
  5. Projected expiry

15 claims: 12 independent, 3 dependent

  1. 1
    A computer-implemented method, comprising:at an electronic device with one or more processors and memory: detecting a user gaze including activation of a text input area;determining whether the activation of the text input area is associated with a predetermined input type;in accordance with a determination that the activation of the text input area is associated with a predetermined input type, sampling audio data;determining a text representation based on the sampled audio data;and providing, within the text input area, an output including the text representation.
  2. 3
    The method of any one of claims 1-2, wherein detecting the user gaze includes eye tracking.
  3. 4
    The method any one of claims 1-3, wherein determining whether the activation of the text input area is associated with a predetermined input type comprises:in accordance with detecting the user gaze including activation of the text input area, displaying an affordance proximate to the text input area;and in accordance with receiving a user input including activation of the affordance, determining that the activation of the text input area is associated with a predetermined input type.
  4. 5
    The method of any one of claims 1-4, wherein detecting the user gaze includes communicating with a secondary device.
  5. 6
    The method of any one of claims 1-5, comprising:in accordance with a determination that the activation of the text input area is associated with a predetermined input type, receiving audio data from a secondary device, wherein sampling audio data includes sampling the audio data received from the secondary device.
  6. 7
    The method of any one of claims 1-6, wherein the text input area corresponds to at least one of an instant messaging text field, an e-mail text field, a calendar event text field, a notes field, a social media text field, or a search text field.
  7. 8
    The method of any one of claims 1-7, comprising:in accordance with a determination that the activation of the text input area is associated with a predetermined input type, providing an output to indicate an initiation of speech recognition, wherein the output includes at least one of an audible output, a visual output, or a haptic output.
  8. 9
    The method of any one of claims 1-8, wherein providing, within the text input area, an output including the text representation comprises:while determining the text representation based on the sampled audio data, displaying text within the text input area, wherein the displayed text corresponds to a determined portion of the text representation.
  9. 10
    The method of any one of claims 1-9, comprising:receiving audio data;storing a predetermined duration of the received audio data;in accordance with detecting the user gaze including activation of the text input area, sampling the predetermined duration of audio data;and determining at least a portion of the text representation based on the sampled predetermined duration of audio data.
  10. 12
    The method of any one of claims 1-11, comprising:while sampling the audio data, determining whether detection of the user gaze including activation of the text input area is maintained;and in accordance with a determination that the input including activation of the text input area is maintained, continue sampling the audio data;in accordance with a determination that the input including activation of the text input area is not maintained, discontinue sampling the audio data.
  11. 13
    The method of any one of claims 1-12, comprising:determining whether a predetermined duration of audio includes user speech;in accordance with a determination that the predetermined duration of audio includes user speech, continue sampling the audio data;and in accordance with a determination that the predetermined duration of audio does not include user speech, discontinue sampling the audio data.
  12. 14
    A computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system that is in communication with one or more display generation components and one or more input devices, the one or more programs including instructions for performing the method of any of claims 1-13.
  13. 15
    An electronic device, comprising:one or more processors;a memory;and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing the methods of any one of claims 1-13.