Intelligent digital assistant in a multi-tasking environment
Abstract
Systems and processes for operating a digital assistant are provided. In one example, the method includes receiving a first speech input from a user. The method further includes identifying context information and determining user intent based on the first speech input and context information. The method further includes determining whether the user intent is to perform the task using a search process or an object management process. The retrieval process is configured to retrieve data, and the object management process is configured to manage objects. The method includes, in accordance with a determination that the user intent is to perform the task using the search process, performing the task using the search process; and performing the task using the object management process according to a determination that the user intent is to perform the task using the object management process.

Term
Projected expiry 1 November 2036.
- Priority
- Filed
- Published
- Today
- Projected expiry
239 claims: 31 independent, 208 dependent
- 1디지털 어시스턴트 서비스를 제공하기 위한 방법으로서, 하나 이상의 프로세서 및 메모리를 구비한 사용자 디바이스에서:사용자로부터 제1 스피치 입력을 수신하는 단계;상기 사용자 디바이스와 연관된 컨텍스트 정보를 식별하는 단계;상기 제1 스피치 입력 및 상기 컨텍스트 정보에 기초하여 사용자 의도를 결정하는 단계;상기 사용자 의도가 검색 프로세스를 이용하여 태스크를 수행하는 것인지 아니면 객체 관리 프로세스를 이용하여 태스크를 수행하는 것인지 결정하는 단계 - 상기 검색 프로세스는 상기 사용자 디바이스에 내부적으로 또는 외부적으로 저장된 데이터를 검색하도록 구성되고, 상기 객체 관리 프로세스는 상기 사용자 디바이스와 연관된 객체들을 관리하도록 구성됨 -;상기 사용자 의도가 상기 검색 프로세스를 이용하여 상기 태스크를 수행하는 것이라는 결정에 따라, 상기 검색 프로세스를 이용하여 상기 태스크를 수행하는 단계;및 상기 사용자 의도가 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하는 것이라는 상기 결정에 따라, 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하는 단계를 포함하는, 방법.
- 2제1항에 있어서, 상기 제1 스피치 입력을 수신하기 이전에:상기 사용자 디바이스와 연관된 디스플레이 상에, 상기 디지털 어시스턴트 서비스를 호출하기 위한 어포던스를 디스플레이하는 단계를 추가로 포함하는, 방법.
- 3제2항에 있어서, 사전결정된 구절을 수신하는 것에 응답하여 상기 디지털 어시스턴트를 호출하는 단계를 추가로 포함하는, 방법.
- 4제2항에 있어서, 상기 어포던스의 선택을 수신하는 것에 응답하여 상기 디지털 어시스턴트를 호출하는 단계를 추가로 포함하는, 방법.
- 5제1항에 있어서, 상기 사용자 의도를 결정하는 단계는:하나 이상의 행동가능한 의도를 결정하는 단계;및 상기 행동가능한 의도와 연관된 하나 이상의 파라미터를 결정하는 단계를 포함하는, 방법.
- 6제1항에 있어서, 상기 컨텍스트 정보는 사용자 특정 데이터, 하나 이상의 객체와 연관된 메타데이터, 센서 데이터, 및 사용자 디바이스 구성 데이터 중 적어도 하나를 포함하는, 방법.
- 7제1항에 있어서, 상기 사용자 의도가 상기 검색 프로세스를 이용하여 상기 태스크를 수행하는 것인지 아니면 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하는 것인지 결정하는 단계는:상기 스피치 입력이 상기 검색 프로세스를 나타내는 하나 이상의 키워드를 포함하는지 아니면 상기 객체 관리 프로세스를 나타내는 하나 이상의 키워드를 포함하는지 결정하는 단계를 포함하는, 방법.
- 8제1항에 있어서, 상기 사용자 의도가 상기 검색 프로세스를 이용하여 상기 태스크를 수행하는 것인지 아니면 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하는 것인지 결정하는 단계는:상기 태스크가 검색과 연관되는지 결정하는 단계;상기 태스크가 검색과 연관된다는 결정에 따라, 상기 태스크를 수행하는 것이 상기 검색 프로세스를 필요로 하는지 결정하는 단계;및 상기 태스크가 검색과 연관되지 않는다는 결정에 따라, 상기 태스크가 적어도 하나의 객체를 관리하는 것과 연관되는지 결정하는 단계를 포함하는, 방법.
- 9제8항에 있어서, 상기 태스크는 검색과 연관되고, 상기 태스크를 수행하는 것이 상기 검색 프로세스를 필요로 하지 않는다는 결정에 따라, 상기 검색 프로세스 또는 상기 객체 관리 프로세스를 선택하기 위한 음성 요청을 출력하는 단계, 및 상기 사용자로부터, 상기 검색 프로세스 또는 상기 객체 관리 프로세스의 상기 선택을 나타내는 제2 스피치 입력을 수신하는 단계를 추가로 포함하는, 방법.
- 10제8항에 있어서, 상기 태스크는 검색과 연관되고, 상기 태스크를 수행하는 것이 상기 검색 프로세스를 필요로 하지 않는다는 결정에 따라, 사전결정된 구성에 기초하여, 상기 태스크가 상기 검색 프로세스를 이용하여 수행될 것인지 아니면 상기 객체 관리 프로세스를 이용하여 수행될 것인지 결정하는 단계를 추가로 포함하는, 방법.
- 11제8항에 있어서, 상기 태스크는 검색과 연관되지 않고, 상기 태스크가 상기 적어도 하나의 객체를 관리하는 것과 연관되지 않는다는 결정에 따라:상기 태스크가 상기 사용자 디바이스에 이용가능한 제4 프로세스를 이용하여 수행될 수 있는지 결정하는 단계;및 상기 사용자와 대화를 개시하는 단계들 중 적어도 하나를 수행하는 단계를 수행하는 단계를 추가로 포함하는, 방법.
- 12제1항에 있어서, 상기 검색 프로세스를 이용하여 상기 태스크를 수행하는 단계는:상기 검색 프로세스를 이용하여 적어도 하나의 객체를 검색하는 단계를 포함하는, 방법.
- 13제12항에 있어서, 상기 적어도 하나의 객체는 적어도 하나의 폴더 또는 파일을 포함하는, 방법.
- 14제13항에 있어서, 상기 파일은 적어도 하나의 사진, 오디오, 또는 비디오를 포함하는, 방법.
- 15제13항에 있어서, 상기 파일은 상기 사용자 디바이스에 내부적으로 또는 외부적으로 저장된, 방법.
- 16제13항에 있어서, 적어도 하나의 상기 폴더 또는 상기 파일을 검색하는 단계는 상기 폴더 또는 상기 파일과 연관된 메타데이터에 기초하는, 방법.
- 17제12항에 있어서, 상기 적어도 하나의 객체는 통신을 포함하는, 방법.
- 18제17항에 있어서, 상기 통신은 적어도 하나의 이메일, 메시지, 통지, 또는 음성메일을 포함하는, 방법.
- 19제17항에 있어서, 상기 통신과 연관된 메타데이터를 검색하는 단계를 추가로 포함하는, 방법.
- 20제12항에 있어서, 상기 적어도 하나의 객체는 적어도 하나의 연락처 또는 캘린더를 포함하는, 방법.
- 21제12항에 있어서, 상기 적어도 하나의 객체는 애플리케이션을 포함하는, 방법.
- 22제12항에 있어서, 상기 적어도 하나의 객체는 온라인 정보제공 소스를 포함하는, 방법.
- 23제1항에 있어서, 상기 태스크는 검색과 연관되고, 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하는 단계는:상기 객체 관리 프로세스를 이용하여 상기 적어도 하나의 객체를 검색하는 단계를 포함하는, 방법.
- 24제23항에 있어서, 상기 적어도 하나의 객체는 적어도 하나의 폴더 또는 파일을 포함하는, 방법.
- 25제24항에 있어서, 상기 파일은 적어도 하나의 사진, 오디오, 또는 비디오를 포함하는, 방법.
- 26제24항에 있어서, 상기 파일은 상기 사용자 디바이스에 내부적으로 또는 외부적으로 저장된, 방법.
- 27제24항에 있어서, 적어도 하나의 상기 폴더 또는 상기 파일을 검색하는 단계는 상기 폴더 또는 상기 파일과 연관된 메타데이터에 기초하는, 방법.
- 28제1항에 있어서, 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하는 단계는 상기 객체 관리 프로세스를 즉시실행하는 단계를 포함하고, 상기 객체 관리 프로세스를 즉시실행하는 단계는 상기 객체 관리 프로세스를 호출하는 단계, 상기 객체 관리 프로세스의 새로운 인스턴스를 생성하는 단계, 또는 상기 객체 관리 프로세스의 기존 인스턴스를 실행하는 단계를 포함하는, 방법.
- 29제1항에 있어서, 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하는 단계는 상기 적어도 하나의 객체를 생성하는 단계를 포함하는, 방법.
- 30제1항에 있어서, 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하는 단계는 상기 적어도 하나의 객체를 저장하는 단계를 포함하는, 방법.
- 31제1항에 있어서, 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하는 단계는 상기 적어도 하나의 객체를 압축하는 단계를 포함하는, 방법.
- 32제1항에 있어서, 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하는 단계는 상기 적어도 하나의 객체를 제1 물리적 또는 가상 저장장치로부터 제2 물리적 또는 가상 저장장치로 이동하는 단계를 포함하는, 방법.
- 33제1항에 있어서, 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하는 단계는 상기 적어도 하나의 객체를 제1 물리적 또는 가상 저장장치로부터 제2 물리적 또는 가상 저장장치로 복사하는 단계를 포함하는, 방법.
- 34제1항에 있어서, 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하는 단계는 물리적 또는 가상 저장장치에 저장된 상기 적어도 하나의 객체를 삭제하는 단계를 포함하는, 방법.
- 35제1항에 있어서, 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하는 단계는 물리적 또는 가상 저장장치에 저장된 적어도 하나의 객체를 복구하는 단계를 포함하는, 방법.
- 36제1항에 있어서, 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하는 단계는 상기 적어도 하나의 객체를 마킹하는 단계를 포함하고, 상기 적어도 하나의 객체의 마킹은 표시되거나 또는 상기 적어도 하나의 객체의 메타데이터와 연관되는 것 중 적어도 하나인, 방법.
- 37제1항에 있어서, 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하는 단계는 백업을 위한 사전결정된 기간에 따라 상기 적어도 하나의 객체를 백업하는 단계를 포함하는, 방법.
- 38제1항에 있어서, 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하는 단계는 상기 사용자 디바이스에 통신가능하게 연결된 하나 이상의 전자 디바이스 사이에 상기 적어도 하나의 객체를 공유하는 단계를 포함하는, 방법.
- 39제1항에 있어서, 상기 검색 프로세스 또는 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행한 결과에 기초하여 응답을 제공하는 단계를 추가로 포함하는, 방법.
- 40제39항에 있어서, 상기 검색 프로세스 또는 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행한 상기 결과에 기초하여 상기 응답을 제공하는 단계는:상기 검색 프로세스 또는 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행한 상기 결과를 제공하는 제1 사용자 인터페이스를 디스플레이하는 단계를 포함하는, 방법.
- 41제39항에 있어서, 상기 검색 프로세스 또는 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행한 상기 결과에 기초하여 상기 응답을 제공하는 단계는:상기 검색 프로세스를 이용하여 상기 태스크를 수행한 상기 결과와 연관된 링크를 제공하는 단계를 포함하는, 방법.
- 42제39항에 있어서, 상기 검색 프로세스 또는 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행한 상기 결과에 기초하여 상기 응답을 제공하는 단계는:상기 검색 프로세스 또는 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행한 상기 결과에 따른 음성 출력을 제공하는 단계를 포함하는, 방법.
- 43제39항에 있어서, 상기 검색 프로세스 또는 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행한 상기 결과에 기초하여 상기 응답을 제공하는 단계는:상기 사용자가 상기 검색 프로세스 또는 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행한 상기 결과를 조작하도록 인에이블하는 어포던스를 제공하는 단계를 포함하는, 방법.
- 44제39항에 있어서, 상기 검색 프로세스 또는 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행한 상기 결과에 기초하여 상기 응답을 제공하는 단계는:상기 태스크를 수행한 상기 결과를 이용하여 동작하는 제3 프로세스를 즉시실행하는 단계를 포함하는, 방법.
- 45제39항에 있어서, 검색 프로세스 또는 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행한 상기 결과에 기초하여 상기 응답을 제공하는 단계는:신뢰 수준을 결정하는 단계;및 상기 신뢰 수준의 상기 결정에 따른 상기 응답을 제공하는 단계를 포함하는, 방법.
- 46제45항에 있어서, 상기 신뢰 수준은 상기 제1 스피치 입력 및 상기 사용자 디바이스와 연관된 컨텍스트 정보에 기초하여 상기 사용자 의도를 결정하는 정확성을 나타내는, 방법.
- 47제45항에 있어서, 상기 신뢰 수준은 상기 사용자 의도가 상기 검색 프로세스를 이용하여 상기 태스크를 수행하는 것인지 아니면 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하는 것인지 결정하는 정확성을 나타내는, 방법.
- 48제45항에 있어서, 상기 신뢰 수준은 상기 검색 프로세스 또는 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하는 정확성을 나타내는, 방법.
- 49제45 내지 제48항 중 어느 한 항에 있어서, 상기 신뢰 수준의 상기 결정에 따른 상기 응답을 제공하는 단계는:상기 신뢰 수준이 임계 신뢰 수준 이상인지 결정하는 단계;상기 신뢰 수준이 상기 임계 신뢰 수준 이상이라는 결정에 따라, 제1 응답을 제공하는 단계;및 상기 신뢰 수준이 임계 신뢰 수준 미만이라는 결정에 따라, 제2 응답을 제공하는 단계를 포함하는, 방법.
- 50하나 이상의 프로그램을 저장하는 비일시적 컴퓨터 판독가능 저장 매체로서, 상기 하나 이상의 프로그램은, 전자 디바이스의 하나 이상의 프로세서에 의해 실행될 때, 상기 전자 디바이스로 하여금:사용자로부터 제1 스피치 입력을 수신하고;상기 사용자 디바이스와 연관된 컨텍스트 정보를 식별하고;상기 제1 스피치 입력 및 상기 컨텍스트 정보에 기초하여 사용자 의도를 결정하고;상기 사용자 의도가 검색 프로세스를 이용하여 태스크를 수행하는 것인지 아니면 객체 관리 프로세스를 이용하여 태스크를 수행하는 것인지 결정하고 - 상기 검색 프로세스는 상기 사용자 디바이스에 내부적으로 또는 외부적으로 저장된 데이터를 검색하도록 구성되고, 상기 객체 관리 프로세스는 상기 사용자 디바이스와 연관된 객체들을 관리하도록 구성됨 -;상기 사용자 의도가 상기 검색 프로세스를 이용하여 상기 태스크를 수행하는 것이라는 결정에 따라, 상기 검색 프로세스를 이용하여 상기 태스크를 수행하고;상기 사용자 의도가 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하는 것이라는 상기 결정에 따라, 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하게 하는 명령어들을 포함하는, 비일시적 컴퓨터 판독가능 저장 매체.
- 51전자 디바이스로서, 하나 이상의 프로세서; 메모리; 및 메모리에 저장된 하나 이상의 프로그램을 포함하고, 상기 하나 이상의 프로그램은:사용자로부터 제1 스피치 입력을 수신하고;상기 사용자 디바이스와 연관된 컨텍스트 정보를 식별하고;상기 제1 스피치 입력 및 상기 컨텍스트 정보에 기초하여 사용자 의도를 결정하고;상기 사용자 의도가 검색 프로세스를 이용하여 태스크를 수행하는 것인지 아니면 객체 관리 프로세스를 이용하여 태스크를 수행하는 것인지 결정하고 - 상기 검색 프로세스는 상기 사용자 디바이스에 내부적으로 또는 외부적으로 저장된 데이터를 검색하도록 구성되고, 상기 객체 관리 프로세스는 상기 사용자 디바이스와 연관된 객체들을 관리하도록 구성됨 -;상기 사용자 의도가 상기 검색 프로세스를 이용하여 상기 태스크를 수행하는 것이라는 결정에 따라, 상기 검색 프로세스를 이용하여 상기 태스크를 수행하고;상기 사용자 의도가 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하는 것이라는 상기 결정에 따라, 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하기 위한 명령어들을 포함하는, 전자 디바이스.
- 52전자 디바이스로서, 사용자로부터 제1 스피치 입력을 수신하기 위한 수단;상기 사용자 디바이스와 연관된 컨텍스트 정보를 식별하기 위한 수단;상기 제1 스피치 입력 및 상기 컨텍스트 정보에 기초하여 사용자 의도를 결정하기 위한 수단;상기 사용자 의도가 검색 프로세스를 이용하여 태스크를 수행하는 것인지 아니면 객체 관리 프로세스를 이용하여 태스크를 수행하는 것인지 결정하기 위한 수단 - 상기 검색 프로세스는 상기 사용자 디바이스에 내부적으로 또는 외부적으로 저장된 데이터를 검색하도록 구성되고, 상기 객체 관리 프로세스는 상기 사용자 디바이스와 연관된 객체들을 관리하도록 구성됨 -;상기 사용자 의도가 상기 검색 프로세스를 이용하여 상기 태스크를 수행하는 것이라는 결정에 따라, 상기 검색 프로세스를 이용하여 상기 태스크를 수행하기 위한 수단;및 상기 사용자 의도가 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하는 것이라는 상기 결정에 따라, 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하기 위한 수단을 포함하는, 전자 디바이스.
- 53전자 디바이스로서, 하나 이상의 프로세서;메모리;및 메모리에 저장된 하나 이상의 프로그램을 포함하고, 상기 하나 이상의 프로그램은 제1항 내지 제49항 중 어느 한 항의 방법을 수행하기 위한 명령어들을 포함하는, 전자 디바이스.
- 54전자 디바이스로서, 제1항 내지 제49항 중 어느 한 항의 방법을 수행하기 위한 수단을 포함하는, 전자 디바이스.
- 55비일시적 컴퓨터 판독가능 저장 매체로서, 전자 디바이스의 하나 이상의 프로세서에 의한 실행을 위한 하나 이상의 프로그램을 포함하고, 상기 하나 이상의 프로그램은, 상기 하나 이상의 프로세서에 의해 실행될 때, 상기 전자 디바이스로 하여금 제1항 내지 제49항 중 어느 한 항의 방법을 수행하게 하는 명령어들을 포함하는, 비일시적 컴퓨터 판독가능 저장 매체.
- 56디지털 어시스턴트를 동작시키기 위한 시스템으로서, 제1항 내지 제49항 중 어느 한 항의 방법을 수행하기 위한 수단을 포함하는, 시스템.
- 57전자 디바이스로서, 사용자로부터 제1 스피치 입력을 수신하도록 구성된 수신 유닛;및 프로세싱 유닛을 포함하며, 상기 프로세싱 유닛은, 상기 사용자 디바이스와 연관된 컨텍스트 정보를 식별하고;상기 제1 스피치 입력 및 상기 컨텍스트 정보에 기초하여 사용자 의도를 결정하고;상기 사용자 의도가 검색 프로세스를 이용하여 태스크를 수행하는 것인지 아니면 객체 관리 프로세스를 이용하여 태스크를 수행하는 것인지 결정하고 - 상기 검색 프로세스는 상기 사용자 디바이스에 내부적으로 또는 외부적으로 저장된 데이터를 검색하도록 구성되고, 상기 객체 관리 프로세스는 상기 사용자 디바이스와 연관된 객체들을 관리하도록 구성됨 -;상기 사용자 의도가 상기 검색 프로세스를 이용하여 상기 태스크를 수행하는 것이라는 결정에 따라, 상기 검색 프로세스를 이용하여 상기 태스크를 수행하고;상기 사용자 의도가 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하는 것이라는 상기 결정에 따라, 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하도록 구성된, 전자 디바이스.
- 58제57항에 있어서, 상기 제1 스피치 입력을 수신하기 이전에:상기 사용자 디바이스와 연관된 디스플레이 상에, 상기 디지털 어시스턴트 서비스를 호출하기 위한 어포던스를 디스플레이하는 것을 추가로 포함하는, 전자 디바이스.
- 59제58항에 있어서, 사전결정된 구절을 수신하는 것에 응답하여 상기 디지털 어시스턴트를 즉시실행하는 것을 추가로 포함하는, 전자 디바이스.
- 60제58항에 있어서, 상기 어포던스의 선택을 수신하는 것에 응답하여 상기 디지털 어시스턴트를 즉시실행하는 것을 추가로 포함하는, 전자 디바이스.
- 61제57항에 있어서, 상기 사용자 의도를 결정하는 것은:하나 이상의 행동가능한 의도를 결정하는 것;및 상기 행동가능한 의도와 연관된 하나 이상의 파라미터를 결정하는 것을 포함하는, 전자 디바이스.
- 62제57항에 있어서, 상기 컨텍스트 정보는 사용자 특정 데이터, 하나 이상의 객체와 연관된 메타데이터, 센서 데이터, 및 사용자 디바이스 구성 데이터 중 적어도 하나를 포함하는, 전자 디바이스.
- 63제57항에 있어서, 상기 사용자 의도가 상기 검색 프로세스를 이용하여 상기 태스크를 수행하는 것인지 아니면 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하는 것인지 결정하는 것은:상기 스피치 입력이 상기 검색 프로세스를 나타내는 하나 이상의 키워드를 포함하는지 아니면 상기 객체 관리 프로세스를 나타내는 하나 이상의 키워드를 포함하는지 결정하는 것을 포함하는, 전자 디바이스.
- 64제57항에 있어서, 상기 사용자 의도가 상기 검색 프로세스를 이용하여 상기 태스크를 수행하는 것인지 아니면 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하는 것인지 결정하는 것은:상기 태스크가 검색과 연관되는지 결정하는 것;상기 태스크가 검색과 연관된다는 결정에 따라, 상기 태스크를 수행하는 것이 상기 검색 프로세스를 필요로 하는지 결정하는 것;및 상기 태스크가 검색과 연관되지 않는다는 결정에 따라, 상기 태스크가 적어도 하나의 객체를 관리하는 것과 연관되는지 결정하는 것을 포함하는, 전자 디바이스.
- 65제64항에 있어서, 상기 태스크는 검색과 연관되고, 상기 태스크를 수행하는 것이 상기 검색 프로세스를 필요로 하지 않는다는 결정에 따라, 상기 검색 프로세스 또는 상기 객체 관리 프로세스를 선택하기 위한 음성 요청을 출력하는 것, 및 상기 사용자로부터, 상기 검색 프로세스 또는 상기 객체 관리 프로세스의 상기 선택을 나타내는 제2 스피치 입력을 수신하는 것을 추가로 포함하는, 전자 디바이스.
- 66제64항에 있어서, 상기 태스크는 검색과 연관되고, 상기 태스크를 수행하는 것이 상기 검색 프로세스를 필요로 하지 않는다는 결정에 따라, 사전결정된 구성에 기초하여, 상기 태스크가 상기 검색 프로세스를 이용하여 수행될 것인지 아니면 상기 객체 관리 프로세스를 이용하여 수행될 것인지 결정하는 것을 추가로 포함하는, 전자 디바이스.
- 67제64항에 있어서, 상기 태스크는 검색과 연관되지 않고, 상기 태스크가 상기 적어도 하나의 객체를 관리하는 것과 연관되지 않는다는 결정에 따라:상기 태스크가 상기 사용자 디바이스에 이용가능한 제4 프로세스를 이용하여 수행될 수 있는지 결정하는 것;및 상기 사용자와 대화를 개시하는 것 중 적어도 하나를 수행하는 것을 추가로 포함하는, 전자 디바이스.
- 68제57항에 있어서, 상기 검색 프로세스를 이용하여 상기 태스크를 수행하는 것은:상기 검색 프로세스를 이용하여 적어도 하나의 객체를 검색하는 것을 포함하는, 전자 디바이스.
- 69제68항에 있어서, 상기 적어도 하나의 객체는 적어도 하나의 폴더 또는 파일을 포함하는, 전자 디바이스.
- 70제69항에 있어서, 상기 파일은 적어도 하나의 사진, 오디오, 또는 비디오를 포함하는, 전자 디바이스.
- 71제69항에 있어서, 상기 파일은 상기 사용자 디바이스에 내부적으로 또는 외부적으로 저장된, 전자 디바이스.
- 72제69항에 있어서, 적어도 하나의 상기 폴더 또는 상기 파일을 검색하는 것은 상기 폴더 또는 상기 파일과 연관된 메타데이터에 기초하는, 전자 디바이스.
- 73제68항에 있어서, 상기 적어도 하나의 객체는 통신을 포함하는, 전자 디바이스.
- 74제73항에 있어서, 상기 통신은 적어도 하나의 이메일, 메시지, 통지, 또는 음성메일을 포함하는, 전자 디바이스.
- 75제73항에 있어서, 상기 통신과 연관된 메타데이터를 검색하는 것을 추가로 포함하는, 전자 디바이스.
- 76제68항에 있어서, 상기 적어도 하나의 객체는 적어도 하나의 연락처 또는 캘린더를 포함하는, 전자 디바이스.
- 77제68항에 있어서, 상기 적어도 하나의 객체는 애플리케이션을 포함하는, 전자 디바이스.
- 78제68항에 있어서, 상기 적어도 하나의 객체는 온라인 정보제공 소스를 포함하는, 전자 디바이스.
- 79제57항에 있어서, 상기 태스크는 검색과 연관되고, 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하는 것은:상기 객체 관리 프로세스를 이용하여 상기 적어도 하나의 객체를 검색하는 것을 포함하는, 전자 디바이스.
- 80제79항에 있어서, 상기 적어도 하나의 객체는 적어도 하나의 폴더 또는 파일을 포함하는, 전자 디바이스.
- 81제80항에 있어서, 상기 파일은 적어도 하나의 사진, 오디오, 또는 비디오를 포함하는, 전자 디바이스.
- 82제79항에 있어서, 상기 파일은 상기 사용자 디바이스에 내부적으로 또는 외부적으로 저장된, 전자 디바이스.
- 83제79항에 있어서, 적어도 하나의 상기 폴더 또는 상기 파일을 검색하는 것은 상기 폴더 또는 상기 파일과 연관된 메타데이터에 기초하는, 전자 디바이스.
- 84제57항에 있어서, 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하는 것은 상기 객체 관리 프로세스를 즉시실행하는 것을 포함하고, 상기 객체 관리 프로세스를 즉시실행하는 것은 상기 객체 관리 프로세스를 호출하는 것, 상기 객체 관리 프로세스의 새로운 인스턴스를 생성하는 것, 또는 상기 객체 관리 프로세스의 기존 인스턴스를 실행하는 것을 포함하는, 전자 디바이스.
- 85제57항에 있어서, 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하는 것은 상기 적어도 하나의 객체를 생성하는 것을 포함하는, 전자 디바이스.
- 86제57항에 있어서, 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하는 것은 상기 적어도 하나의 객체를 저장하는 것을 포함하는, 전자 디바이스.
- 87제57항에 있어서, 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하는 것은 상기 적어도 하나의 객체를 압축하는 것을 포함하는, 전자 디바이스.
- 88제57항에 있어서, 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하는 것은 상기 적어도 하나의 객체를 제1 물리적 또는 가상 저장장치로부터 제2 물리적 또는 가상 저장장치로 이동하는 것을 포함하는, 전자 디바이스.
- 89제57항에 있어서, 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하는 것은 상기 적어도 하나의 객체를 제1 물리적 또는 가상 저장장치로부터 제2 물리적 또는 가상 저장장치로 복사하는 것을 포함하는, 전자 디바이스.
- 90제57항에 있어서, 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하는 것은 물리적 또는 가상 저장장치에 저장된 상기 적어도 하나의 객체를 삭제하는 것을 포함하는, 전자 디바이스.
- 91제57항에 있어서, 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하는 것은 물리적 또는 가상 저장장치에 저장된 적어도 하나의 객체를 복구하는 것을 포함하는, 전자 디바이스.
- 92제57항에 있어서, 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하는 것은 상기 적어도 하나의 객체를 마킹하는 것을 포함하고, 상기 적어도 하나의 객체의 마킹은 표시되거나 또는 상기 적어도 하나의 객체의 메타데이터와 연관되는 것 중 적어도 하나인, 전자 디바이스.
- 93제57항에 있어서, 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하는 것은 백업을 위한 사전결정된 기간에 따라 상기 적어도 하나의 객체를 백업하는 것을 포함하는, 전자 디바이스.
- 94제57항에 있어서, 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하는 것은 상기 사용자 디바이스에 통신가능하게 연결된 하나 이상의 전자 디바이스 사이에 상기 적어도 하나의 객체를 공유하는 것을 포함하는, 전자 디바이스.
- 95제57항에 있어서, 상기 검색 프로세스 또는 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행한 결과에 기초하여 응답을 제공하는 것을 추가로 포함하는, 전자 디바이스.
- 96제95항에 있어서, 상기 검색 프로세스 또는 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행한 상기 결과에 기초하여 상기 응답을 제공하는 것은:상기 검색 프로세스 또는 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행한 상기 결과를 제공하는 제1 사용자 인터페이스를 디스플레이하는 것을 포함하는, 전자 디바이스.
- 97제95항에 있어서, 상기 검색 프로세스 또는 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행한 상기 결과에 기초하여 상기 응답을 제공하는 것은:상기 검색 프로세스를 이용하여 상기 태스크를 수행한 상기 결과와 연관된 링크를 제공하는 것을 포함하는, 전자 디바이스.
- 98제95항에 있어서, 상기 검색 프로세스 또는 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행한 상기 결과에 기초하여 상기 응답을 제공하는 것은:상기 검색 프로세스 또는 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행한 상기 결과에 따른 음성 출력을 제공하는 것을 포함하는, 전자 디바이스.
- 99제95항에 있어서, 상기 검색 프로세스 또는 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행한 상기 결과에 기초하여 상기 응답을 제공하는 것은:상기 사용자가 상기 검색 프로세스 또는 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행한 상기 결과를 조작하도록 인에이블하는 어포던스를 제공하는 것을 포함하는, 전자 디바이스.
- 100제95항에 있어서, 상기 검색 프로세스 또는 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행한 상기 결과에 기초하여 상기 응답을 제공하는 것은:상기 태스크를 수행한 상기 결과를 이용하여 동작하는 제3 프로세스를 즉시실행하는 것을 포함하는, 전자 디바이스.
- 101제95항에 있어서, 검색 프로세스 또는 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행한 상기 결과에 기초하여 상기 응답을 제공하는 것은:신뢰 수준을 결정하는 것;및 상기 신뢰 수준의 상기 결정에 따른 상기 응답을 제공하는 것을 포함하는, 전자 디바이스.
- 102제101항에 있어서, 상기 신뢰 수준은 상기 제1 스피치 입력 및 상기 사용자 디바이스와 연관된 컨텍스트 정보에 기초하여 상기 사용자 의도를 결정하는 정확성을 나타내는, 전자 디바이스.
- 103제101항에 있어서, 상기 신뢰 수준은 상기 사용자 의도가 상기 검색 프로세스를 이용하여 상기 태스크를 수행하는 것인지 아니면 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하는 것인지 결정하는 정확성을 나타내는, 전자 디바이스.
- 104제101항에 있어서, 상기 신뢰 수준은 상기 검색 프로세스 또는 상기 객체 관리 프로세스를 이용하여 상기 태스크를 수행하는 정확성을 나타내는, 전자 디바이스.
- 105제101항 내지 제104항 중 어느 한 항에 있어서, 상기 신뢰 수준의 상기 결정에 따른 상기 응답을 제공하는 것은:상기 신뢰 수준이 임계 신뢰 수준 이상인지 결정하는 것;상기 신뢰 수준이 상기 임계 신뢰 수준 이상이라는 결정에 따라, 제1 응답을 제공하는 것;및 상기 신뢰 수준이 임계 신뢰 수준 미만이라는 결정에 따라, 제2 응답을 제공하는 것을 포함하는, 전자 디바이스.
- 106디지털 어시스턴트 서비스를 제공하기 위한 방법으로서, 하나 이상의 프로세서 및 메모리를 구비한 사용자 디바이스에서:사용자로부터 태스크를 수행하기 위한 스피치 입력을 수신하는 단계;상기 사용자 디바이스와 연관된 컨텍스트 정보를 식별하는 단계;상기 스피치 입력 및 상기 사용자 디바이스와 연관된 컨텍스트 정보에 기초하여 사용자 의도를 결정하는 단계;사용자 의도에 따라, 상기 태스크가 상기 사용자 디바이스에서 수행될 것인지 아니면 상기 사용자 디바이스에 통신가능하게 연결된 제1 전자 디바이스에서 수행될 것인지 결정하는 단계;상기 태스크는 상기 사용자 디바이스에서 수행될 것이고 상기 태스크를 수행하기 위한 콘텐츠가 다른 곳에 위치한다는 결정에 따라, 상기 태스크를 수행하기 위한 상기 콘텐츠를 수신하는 단계;및 상기 태스크는 상기 제1 전자 디바이스에서 수행될 것이고 상기 태스크를 수행하기 위한 상기 콘텐츠는 상기 제1 전자 디바이스와 다른 곳에 위치한다는 결정에 따라, 상기 태스크를 수행하기 위한 상기 콘텐츠를 상기 제1 전자 디바이스에 제공하는 단계를 포함하는, 방법.
- 107제106항에 있어서, 상기 사용자 디바이스는 복수의 사용자 인터페이스를 제공하도록 구성된, 방법.
- 108제106항에 있어서, 상기 사용자 디바이스는 랩톱 컴퓨터, 데스크톱 컴퓨터, 또는 서버를 포함하는, 방법.
- 109제106항에 있어서, 상기 제1 전자 디바이스는 랩톱 컴퓨터, 데스크톱 컴퓨터, 서버, 스마트폰, 태블릿, 셋톱 박스, 또는 시계를 포함하는, 방법.
- 110제106항에 있어서, 상기 스피치 입력을 수신하기 이전에:상기 사용자 디바이스의 디스플레이 상에, 상기 디지털 어시스턴트를 호출하기 위한 어포던스를 디스플레이하는 단계를 추가로 포함하는, 방법.
- 111제110항에 있어서, 사전결정된 구절을 수신하는 것에 응답하여 상기 디지털 어시스턴트를 즉시실행하는 단계를 추가로 포함하는, 방법.
- 112제110항에 있어서, 상기 어포던스의 선택을 수신하는 것에 응답하여 상기 디지털 어시스턴트를 즉시실행하는 단계를 추가로 포함하는, 방법.
- 113제106항에 있어서, 상기 사용자 의도를 결정하는 단계는:하나 이상의 행동가능한 의도를 결정하는 단계;및 상기 행동가능한 의도와 연관된 하나 이상의 파라미터를 결정하는 단계를 포함하는, 방법.
- 114제106항에 있어서, 상기 컨텍스트 정보는 사용자 특정 데이터, 센서 데이터, 및 사용자 디바이스 구성 데이터 중 적어도 하나를 포함하는, 방법.
- 115제106항에 있어서, 상기 태스크가 상기 사용자 디바이스에서 수행될 것인지 아니면 상기 제1 전자 디바이스에서 수행될 것인지 결정하는 단계는 상기 스피치 입력에 포함되는 하나 이상의 키워드에 기초하는, 방법.
- 116제106항에 있어서, 상기 태스크가 상기 사용자 디바이스에서 수행될 것인지 아니면 상기 제1 전자 디바이스에서 수행될 것인지 결정하는 단계는:상기 사용자 디바이스에서 상기 태스크를 수행하는 것이 수행 기준을 충족하는지 결정하는 단계;및 상기 사용자 디바이스에서 상기 태스크를 수행하는 것이 상기 수행 기준을 충족한다는 결정에 따라, 상기 태스크는 상기 사용자 디바이스에서 수행될 것이라고 결정하는 단계를 포함하는, 방법.
- 117제116항에 있어서, 상기 사용자 디바이스에서 상기 태스크를 수행하는 것이 상기 수행 기준을 충족하지 않는다는 결정에 따라:상기 제1 전자 디바이스에서 상기 태스크를 수행하는 것이 상기 수행 기준을 충족하는지 결정하는 단계를 추가로 포함하는, 방법.
- 118제117항에 있어서, 상기 제1 전자 디바이스에서 상기 태스크를 수행하는 것이 상기 수행 기준을 충족한다는 결정에 따라, 상기 태스크는 상기 제1 전자 디바이스에서 수행될 것이라고 결정하는 단계; 및 상기 제1 전자 디바이스에서 상기 태스크를 수행하는 것이 상기 수행 기준을 충족하지 않는다는 결정에 따라:상기 제2 전자 디바이스에서 상기 태스크를 수행하는 것이 상기 수행 기준을 충족하는지 결정하는 단계를 추가로 포함하는, 방법.
- 119제116항에 있어서, 상기 수행 기준은 하나 이상의 사용자 선호도에 기초하여 결정되는, 방법.
- 120제116항에 있어서, 상기 수행 기준은 상기 디바이스 구성 데이터에 기초하여 결정되는, 방법.
- 121제116항에 있어서, 상기 수행 기준은 역동적으로 업데이트되는, 방법.
- 122제106항에 있어서, 상기 태스크는 상기 사용자 디바이스에서 수행될 것이고 상기 태스크를 수행하기 위한 콘텐츠가 다른 곳에 위치한다는 결정에 따라, 상기 태스크를 수행하기 위한 상기 콘텐츠를 수신하는 단계는:상기 제1 전자 디바이스로부터 상기 콘텐츠의 적어도 일부분을 수신하는 단계를 포함하며, 상기 콘텐츠의 적어도 일부분은 상기 제1 전자 디바이스에 저장되어 있는, 방법.
- 123제106항에 있어서, 상기 태스크는 상기 사용자 디바이스에서 수행될 것이고 상기 태스크를 수행하기 위한 콘텐츠가 다른 곳에 위치한다는 결정에 따라, 상기 태스크를 수행하기 위한 상기 콘텐츠를 수신하는 단계는:제3 전자 디바이스로부터 상기 콘텐츠의 적어도 일부분을 수신하는 단계를 포함하는, 방법.
- 124제106항에 있어서, 상기 태스크는 상기 제1 전자 디바이스에서 수행될 것이고 상기 태스크를 수행하기 위한 상기 콘텐츠는 상기 제1 전자 디바이스와 다른 곳에 위치한다는 결정에 따라, 상기 태스크를 수행하기 위한 상기 콘텐츠를 상기 제1 전자 디바이스에 제공하는 단계는:상기 콘텐츠의 적어도 일부분을 상기 사용자 디바이스로부터 상기 제1 전자 디바이스로 제공하는 단계를 포함하고, 상기 콘텐츠의 적어도 일부분은 상기 사용자 디바이스에 저장되어 있는, 방법.
- 125제106항에 있어서, 상기 태스크는 상기 제1 전자 디바이스에서 수행될 것이고 상기 태스크를 수행하기 위한 상기 콘텐츠는 상기 제1 전자 디바이스와 다른 곳에 위치한다는 결정에 따라, 상기 태스크를 수행하기 위한 상기 콘텐츠를 상기 제1 전자 디바이스에 제공하는 단계는:상기 콘텐츠의 적어도 일부분이 제4 전자 디바이스로부터 상기 제1 전자 디바이스로 제공되게 하는 단계를 포함하고, 상기 콘텐츠의 적어도 일부분은 상기 제4 전자 디바이스에 저장되어 있는, 방법.
- 126제106항에 있어서, 상기 태스크는 상기 사용자 디바이스에서 수행될 것이고, 상기 사용자 디바이스에서 상기 수신된 콘텐츠를 이용하여 제1 응답을 제공하는 단계를 추가로 포함하는, 방법.
- 127제126항에 있어서, 상기 사용자 디바이스에서 상기 제1 응답을 제공하는 단계는:상기 사용자 디바이스에서 상기 태스크를 수행하는 단계를 포함하는, 방법.
- 128제127항에 있어서, 상기 사용자 디바이스에서 상기 태스크를 수행하는 단계는 상기 사용자 디바이스와 다른 곳에서 부분적으로 수행된 태스크의 연속인, 방법.
- 129제126항에 있어서, 상기 사용자 디바이스에서 상기 제1 응답을 제공하는 단계는:상기 사용자 디바이스에서 수행될 상기 태스크와 연관된 제1 사용자 인터페이스를 디스플레이하는 단계를 포함하는, 방법.
- 130제126항에 있어서, 상기 사용자 디바이스에서 상기 제1 응답을 제공하는 단계는:상기 사용자 디바이스에서 수행될 상기 태스크와 연관된 링크를 제공하는 단계를 포함하는, 방법.
- 131제126항에 있어서, 상기 사용자 디바이스에서 상기 제1 응답을 제공하는 단계는:상기 사용자 디바이스에서 수행될 상기 태스크에 따른 음성 출력을 제공하는 단계를 포함하는, 방법.
- 132제106항에 있어서, 상기 태스크는 상기 제1 전자 디바이스에서 수행될 것이고, 상기 사용자 디바이스에서 제2 응답을 제공하는 단계를 추가로 포함하는, 방법.
- 133제132항에 있어서, 상기 사용자 디바이스에서 상기 제2 응답을 제공하는 단계는:상기 태스크가 상기 제1 전자 디바이스에서 수행되게 하는 단계를 포함하는, 방법.
- 134제133항에 있어서, 상기 제1 전자 디바이스에서 수행될 상기 태스크는 상기 제1 전자 디바이스와 다른 곳에서 수행된 태스크의 연속인, 방법.
- 135제132항에 있어서, 상기 사용자 디바이스에서 상기 제2 응답을 제공하는 단계는:상기 제1 전자 디바이스에서 수행될 상기 태스크에 따른 음성 출력을 제공하는 단계를 포함하는, 방법.
- 136제132항 내지 제135항 중 어느 한 항에 있어서, 상기 사용자 디바이스에서 상기 제2 응답을 제공하는 단계는:상기 사용자가 상기 태스크의 수행을 위하여 다른 전자 디바이스를 선택하도록 인에이블하는 어포던스를 제공하는 단계를 포함하는, 방법.
- 137하나 이상의 프로그램을 저장하는 비일시적 컴퓨터 판독가능 저장 매체로서, 상기 하나 이상의 프로그램은, 전자 디바이스의 하나 이상의 프로세서에 의해 실행될 때, 상기 전자 디바이스로 하여금:사용자로부터 태스크를 수행하기 위한 스피치 입력을 수신하고;상기 사용자 디바이스와 연관된 컨텍스트 정보를 식별하고;상기 스피치 입력 및 상기 사용자 디바이스와 연관된 컨텍스트 정보에 기초하여 사용자 의도를 결정하고;사용자 의도에 따라, 상기 태스크가 상기 사용자 디바이스에서 수행될 것인지 아니면 상기 사용자 디바이스에 통신가능하게 연결된 제1 전자 디바이스에서 수행될 것인지 결정하고;상기 태스크는 상기 사용자 디바이스에서 수행될 것이고 상기 태스크를 수행하기 위한 콘텐츠가 다른 곳에 위치한다는 결정에 따라, 상기 태스크를 수행하기 위한 상기 콘텐츠를 수신하고;상기 태스크는 상기 제1 전자 디바이스에서 수행될 것이고 상기 태스크를 수행하기 위한 상기 콘텐츠는 상기 제1 전자 디바이스와 다른 곳에 위치한다는 결정에 따라, 상기 태스크를 수행하기 위한 상기 콘텐츠를 상기 제1 전자 디바이스에 제공하게 하는 명령어들을 포함하는, 비일시적 컴퓨터 판독가능 저장 매체.
- 138사용자 디바이스로서, 하나 이상의 프로세서; 메모리; 및 메모리에 저장된 하나 이상의 프로그램을 포함하고, 상기 하나 이상의 프로그램은:사용자로부터 태스크를 수행하기 위한 스피치 입력을 수신하고;상기 사용자 디바이스와 연관된 컨텍스트 정보를 식별하고;상기 스피치 입력 및 상기 사용자 디바이스와 연관된 컨텍스트 정보에 기초하여 사용자 의도를 결정하고;사용자 의도에 따라, 상기 태스크가 상기 사용자 디바이스에서 수행될 것인지 아니면 상기 사용자 디바이스에 통신가능하게 연결된 제1 전자 디바이스에서 수행될 것인지 결정하고;상기 태스크는 상기 사용자 디바이스에서 수행될 것이고 상기 태스크를 수행하기 위한 콘텐츠가 다른 곳에 위치한다는 결정에 따라, 상기 태스크를 수행하기 위한 상기 콘텐츠를 수신하고;상기 태스크는 상기 제1 전자 디바이스에서 수행될 것이고 상기 태스크를 수행하기 위한 상기 콘텐츠는 상기 제1 전자 디바이스와 다른 곳에 위치한다는 결정에 따라, 상기 태스크를 수행하기 위한 상기 콘텐츠를 상기 제1 전자 디바이스에 제공하기 위한 명령어들을 포함하는, 사용자 디바이스.
- 139사용자 디바이스로서, 사용자로부터 태스크를 수행하기 위한 스피치 입력을 수신하기 위한 수단;상기 사용자 디바이스와 연관된 컨텍스트 정보를 식별하기 위한 수단;상기 스피치 입력 및 상기 사용자 디바이스와 연관된 컨텍스트 정보에 기초하여 사용자 의도를 결정하기 위한 수단;사용자 의도에 따라, 상기 태스크가 상기 사용자 디바이스에서 수행될 것인지 아니면 상기 사용자 디바이스에 통신가능하게 연결된 제1 전자 디바이스에서 수행될 것인지 결정하기 위한 수단;상기 태스크는 상기 사용자 디바이스에서 수행될 것이고 상기 태스크를 수행하기 위한 콘텐츠가 다른 곳에 위치한다는 결정에 따라, 상기 태스크를 수행하기 위한 상기 콘텐츠를 수신하기 위한 수단;및 상기 태스크는 상기 제1 전자 디바이스에서 수행될 것이고 상기 태스크를 수행하기 위한 상기 콘텐츠는 상기 제1 전자 디바이스와 다른 곳에 위치한다는 결정에 따라, 상기 태스크를 수행하기 위한 상기 콘텐츠를 상기 제1 전자 디바이스에 제공하기 위한 수단을 포함하는, 사용자 디바이스.
- 140사용자 디바이스로서, 하나 이상의 프로세서;메모리;및 메모리에 저장된 하나 이상의 프로그램을 포함하고, 상기 하나 이상의 프로그램은 제106항 내지 제136항 중 어느 한 항의 방법을 수행하기 위한 명령어들을 포함하는, 사용자 디바이스.
- 141사용자 디바이스로서, 제106항 내지 제136항 중 어느 한 항의 방법을 수행하기 위한 수단을 포함하는, 사용자 디바이스.
- 142비일시적 컴퓨터 판독가능 저장 매체로서, 전자 디바이스의 하나 이상의 프로세서에 의한 실행을 위한 하나 이상의 프로그램을 포함하고, 상기 하나 이상의 프로그램은, 상기 하나 이상의 프로세서에 의해 실행될 때, 상기 전자 디바이스로 하여금 제106항 내지 제136항 중 어느 한 항의 방법을 수행하게 하는 명령어들을 포함하는, 비일시적 컴퓨터 판독가능 저장 매체.
- 143디지털 어시스턴트를 동작시키기 위한 시스템으로서, 제106항 내지 제136항 중 어느 한 항의 방법을 수행하기 위한 수단을 포함하는, 시스템.
- 144사용자 디바이스로서, 사용자로부터 태스크를 수행하기 위한 스피치 입력을 수신하도록 구성된 수신 유닛;및 프로세싱 유닛을 포함하며, 상기 프로세싱 유닛은, 상기 사용자 디바이스와 연관된 컨텍스트 정보를 식별하고;상기 스피치 입력 및 상기 사용자 디바이스와 연관된 컨텍스트 정보에 기초하여 사용자 의도를 결정하고;사용자 의도에 따라, 상기 태스크가 상기 사용자 디바이스에서 수행될 것인지 아니면 상기 사용자 디바이스에 통신가능하게 연결된 제1 전자 디바이스에서 수행될 것인지 결정하고;상기 태스크는 상기 사용자 디바이스에서 수행될 것이고 상기 태스크를 수행하기 위한 콘텐츠가 다른 곳에 위치한다는 결정에 따라, 상기 태스크를 수행하기 위한 상기 콘텐츠를 수신하고;상기 태스크는 상기 제1 전자 디바이스에서 수행될 것이고 상기 태스크를 수행하기 위한 상기 콘텐츠는 상기 제1 전자 디바이스와 다른 곳에 위치한다는 결정에 따라, 상기 태스크를 수행하기 위한 상기 콘텐츠를 상기 제1 전자 디바이스에 제공하도록 구성된, 사용자 디바이스.
- 145제144항에 있어서, 상기 사용자 디바이스는 복수의 사용자 인터페이스를 제공하도록 구성된, 사용자 디바이스.
- 146제144항에 있어서, 상기 사용자 디바이스는 랩톱 컴퓨터, 데스크톱 컴퓨터, 또는 서버를 포함하는, 사용자 디바이스.
- 147제144항에 있어서, 상기 제1 전자 디바이스는 랩톱 컴퓨터, 데스크톱 컴퓨터, 서버, 스마트폰, 태블릿, 셋톱 박스, 또는 시계를 포함하는, 사용자 디바이스.
- 148제144항에 있어서, 상기 스피치 입력을 수신하기 이전에:상기 사용자 디바이스의 디스플레이 상에, 상기 디지털 어시스턴트를 호출하기 위한 어포던스를 디스플레이하는 것을 추가로 포함하는, 사용자 디바이스.
- 149제148항에 있어서, 사전결정된 구절을 수신하는 것에 응답하여 상기 디지털 어시스턴트를 즉시실행하는 것을 추가로 포함하는, 사용자 디바이스.
- 150제148항에 있어서, 상기 어포던스의 선택을 수신하는 것에 응답하여 상기 디지털 어시스턴트를 즉시실행하는 것을 추가로 포함하는, 사용자 디바이스.
- 151제144항에 있어서, 상기 사용자 의도를 결정하는 것은:하나 이상의 행동가능한 의도를 결정하는 것;및 상기 행동가능한 의도와 연관된 하나 이상의 파라미터를 결정하는 것을 포함하는, 사용자 디바이스.
- 152제144항에 있어서, 상기 컨텍스트 정보는 사용자 특정 데이터, 센서 데이터, 및 사용자 디바이스 구성 데이터 중 적어도 하나를 포함하는, 사용자 디바이스.
- 153제144항에 있어서, 상기 태스크가 상기 사용자 디바이스에서 수행될 것인지 아니면 상기 제1 전자 디바이스에서 수행될 것인지 결정하는 것은 상기 스피치 입력에 포함되는 하나 이상의 키워드에 기초하는, 사용자 디바이스.
- 154제144항에 있어서, 상기 태스크가 상기 사용자 디바이스에서 수행될 것인지 아니면 상기 제1 전자 디바이스에서 수행될 것인지 결정하는 것은:상기 사용자 디바이스에서 상기 태스크를 수행하는 것이 수행 기준을 충족하는지 결정하는 것;및 상기 사용자 디바이스에서 상기 태스크를 수행하는 것이 상기 수행 기준을 충족한다는 결정에 따라, 상기 태스크는 상기 사용자 디바이스에서 수행될 것이라고 결정하는 것을 포함하는, 사용자 디바이스.
- 155제154항에 있어서, 상기 사용자 디바이스에서 상기 태스크를 수행하는 것이 상기 수행 기준을 충족하지 않는다는 결정에 따라:상기 제1 전자 디바이스에서 상기 태스크를 수행하는 것이 상기 수행 기준을 충족하는지 결정하는 것을 추가로 포함하는, 사용자 디바이스.
- 156제155항에 있어서, 상기 제1 전자 디바이스에서 상기 태스크를 수행하는 것이 상기 수행 기준을 충족한다는 결정에 따라, 상기 태스크는 상기 제1 전자 디바이스에서 수행될 것이라고 결정하는 것; 및 상기 제1 전자 디바이스에서 상기 태스크를 수행하는 것이 상기 수행 기준을 충족하지 않는다는 결정에 따라:상기 제2 전자 디바이스에서 상기 태스크를 수행하는 것이 상기 수행 기준을 충족하는지 결정하는 것을 추가로 포함하는, 사용자 디바이스.
- 157제154항에 있어서, 상기 수행 기준은 하나 이상의 사용자 선호도에 기초하여 결정되는, 사용자 디바이스.
- 158제154항에 있어서, 상기 수행 기준은 상기 디바이스 구성 데이터에 기초하여 결정되는, 사용자 디바이스.
- 159제154항에 있어서, 상기 수행 기준은 역동적으로 업데이트되는, 사용자 디바이스.
- 160제144항에 있어서, 상기 태스크는 상기 사용자 디바이스에서 수행될 것이고 상기 태스크를 수행하기 위한 콘텐츠가 다른 곳에 위치한다는 결정에 따라, 상기 태스크를 수행하기 위한 상기 콘텐츠를 수신하는 것은:상기 제1 전자 디바이스로부터 상기 콘텐츠의 적어도 일부분을 수신하는 것을 포함하고, 상기 콘텐츠의 적어도 일부분은 상기 제1 전자 디바이스에 저장되어 있는, 사용자 디바이스.
- 161제144항에 있어서, 상기 태스크는 상기 사용자 디바이스에서 수행될 것이고 상기 태스크를 수행하기 위한 콘텐츠가 다른 곳에 위치한다는 결정에 따라, 상기 태스크를 수행하기 위한 상기 콘텐츠를 수신하는 것은:제3 전자 디바이스로부터 상기 콘텐츠의 적어도 일부분을 수신하는 것을 포함하는, 사용자 디바이스.
- 162제144항에 있어서, 상기 태스크는 상기 제1 전자 디바이스에서 수행될 것이고 상기 태스크를 수행하기 위한 상기 콘텐츠는 상기 제1 전자 디바이스와 다른 곳에 위치한다는 결정에 따라, 상기 태스크를 수행하기 위한 상기 콘텐츠를 상기 제1 전자 디바이스에 제공하는 것은:상기 콘텐츠의 적어도 일부분을 상기 사용자 디바이스로부터 상기 제1 전자 디바이스로 제공하는 것을 포함하며, 상기 콘텐츠의 적어도 일부분은 상기 사용자 디바이스에 저장되어 있는, 사용자 디바이스.
- 163제144항에 있어서, 상기 태스크는 상기 제1 전자 디바이스에서 수행될 것이고 상기 태스크를 수행하기 위한 상기 콘텐츠는 상기 제1 전자 디바이스와 다른 곳에 위치한다는 결정에 따라, 상기 태스크를 수행하기 위한 상기 콘텐츠를 상기 제1 전자 디바이스에 제공하는 것은:상기 콘텐츠의 적어도 일부분이 제4 전자 디바이스로부터 상기 제1 전자 디바이스로 제공되게 하는 것을 포함하며, 상기 콘텐츠의 적어도 일부분은 상기 제4 전자 디바이스에 저장되어 있는, 사용자 디바이스.
- 164제144항에 있어서, 상기 태스크는 상기 사용자 디바이스에서 수행될 것이고, 상기 사용자 디바이스에서 상기 수신된 콘텐츠를 이용하여 제1 응답을 제공하는 것을 추가로 포함하는, 사용자 디바이스.
- 165제164항에 있어서, 상기 사용자 디바이스에서 상기 제1 응답을 제공하는 것은:상기 사용자 디바이스에서 상기 태스크를 수행하는 것을 포함하는, 사용자 디바이스.
- 166제165항에 있어서, 상기 사용자 디바이스에서 상기 태스크를 수행하는 것은 상기 사용자 디바이스와 다른 곳에서 부분적으로 수행된 태스크의 연속인, 사용자 디바이스.
- 167제164항에 있어서, 상기 사용자 디바이스에서 상기 제1 응답을 제공하는 것은:상기 사용자 디바이스에서 수행될 상기 태스크와 연관된 제1 사용자 인터페이스를 디스플레이하는 것을 포함하는, 사용자 디바이스.
- 168제164항에 있어서, 상기 사용자 디바이스에서 상기 제1 응답을 제공하는 것은:상기 사용자 디바이스에서 수행될 상기 태스크와 연관된 링크를 제공하는 것을 포함하는, 사용자 디바이스.
- 169제164항에 있어서, 상기 사용자 디바이스에서 상기 제1 응답을 제공하는 것은:상기 사용자 디바이스에서 수행될 상기 태스크에 따른 음성 출력을 제공하는 것을 포함하는, 사용자 디바이스.
- 170제144항에 있어서, 상기 태스크는 상기 제1 전자 디바이스에서 수행될 것이고, 상기 사용자 디바이스에서 제2 응답을 제공하는 것을 추가로 포함하는, 사용자 디바이스.
- 171제170항에 있어서, 상기 사용자 디바이스에서 상기 제2 응답을 제공하는 것은:상기 태스크가 상기 제1 전자 디바이스에서 수행되게 하는 것을 포함하는, 사용자 디바이스.
- 172제171항에 있어서, 상기 제1 전자 디바이스에서 수행될 상기 태스크는 상기 제1 전자 디바이스와 다른 곳에서 수행된 태스크의 연속인, 사용자 디바이스.
- 173제170항에 있어서, 상기 사용자 디바이스에서 상기 제2 응답을 제공하는 것은:상기 제1 전자 디바이스에서 수행될 상기 태스크에 따른 음성 출력을 제공하는 것을 포함하는, 사용자 디바이스.
- 174제170항 내지 제173항 중 어느 한 항에 있어서, 상기 사용자 디바이스에서 상기 제2 응답을 제공하는 것은:상기 사용자가 상기 태스크의 수행을 위하여 다른 전자 디바이스를 선택하도록 인에이블하는 어포던스를 제공하는 것을 포함하는, 사용자 디바이스.
- 175디지털 어시스턴트 서비스를 제공하기 위한 방법으로서, 하나 이상의 프로세서 및 메모리를 구비한 사용자 디바이스에서:사용자로부터 상기 사용자 디바이스의 하나 이상의 시스템 구성을 관리하기 위한 스피치 입력을 수신하는 단계 - 상기 사용자 디바이스는 동시에 복수의 사용자 인터페이스를 제공하도록 구성됨 -;상기 사용자 디바이스와 연관된 컨텍스트 정보를 식별하는 단계;스피치 입력 및 컨텍스트 정보에 기초하여 사용자 의도를 결정하는 단계;상기 사용자 의도가 정보제공 요청을 나타내는지 아니면 태스크를 수행하기 위한 요청을 나타내는지 결정하는 단계;상기 사용자 의도가 정보제공 요청을 나타낸다는 결정에 따라, 상기 정보제공 요청에 대한 음성 응답을 제공하는 단계;및 상기 사용자 의도가 태스크를 수행하기 위한 요청을 나타낸다는 결정에 따라, 상기 태스크를 수행하기 위하여 상기 사용자 디바이스와 연관된 프로세스를 즉시실행하는 단계를 포함하는, 방법.
- 176제175항에 있어서, 상기 스피치 입력을 수신하기 이전에:상기 사용자 디바이스의 디스플레이 상에, 상기 디지털 어시스턴트를 호출하기 위한 어포던스를 디스플레이하는 단계를 추가로 포함하는, 방법.
- 177제176항에 있어서, 사전결정된 구절을 수신하는 것에 응답하여 상기 디지털 어시스턴트 서비스를 즉시실행하는 단계를 추가로 포함하는, 방법.
- 178제176항에 있어서, 상기 어포던스의 선택을 수신하는 것에 응답하여 상기 디지털 어시스턴트 서비스를 즉시실행하는 단계를 추가로 포함하는, 방법.
- 179제175항에 있어서, 상기 사용자 디바이스의 상기 하나 이상의 시스템 구성은 오디오 구성들을 포함하는, 방법.
- 180제175항에 있어서, 상기 사용자 디바이스의 상기 하나 이상의 시스템 구성은 날짜 및 시간 구성들을 포함하는, 방법.
- 181제175항에 있어서, 상기 사용자 디바이스의 상기 하나 이상의 시스템 구성은 구술 구성들을 포함하는, 방법.
- 182제175항에 있어서, 상기 사용자 디바이스의 상기 하나 이상의 시스템 구성은 디스플레이 구성들을 포함하는, 방법.
- 183제175항에 있어서, 상기 사용자 디바이스의 상기 하나 이상의 시스템 구성은 입력 디바이스 구성들을 포함하는, 방법.
- 184제175항에 있어서, 상기 사용자 디바이스의 상기 하나 이상의 시스템 구성은 네트워크 구성들을 포함하는, 방법.
- 185제175항에 있어서, 상기 사용자 디바이스의 상기 하나 이상의 시스템 구성은 통지 구성들을 포함하는, 방법.
- 186제175항에 있어서, 상기 사용자 디바이스의 상기 하나 이상의 시스템 구성은 프린터 구성들을 포함하는, 방법.
- 187제175항에 있어서, 상기 사용자 디바이스의 상기 하나 이상의 시스템 구성은 보안 구성들을 포함하는, 방법.
- 188제175항에 있어서, 상기 사용자 디바이스의 상기 하나 이상의 시스템 구성은 백업 구성들을 포함하는, 방법.
- 189제175항에 있어서, 상기 사용자 디바이스의 상기 하나 이상의 시스템 구성은 애플리케이션 구성들을 포함하는, 방법.
- 190제175항에 있어서, 상기 사용자 디바이스의 상기 하나 이상의 시스템 구성은 사용자 인터페이스 구성들을 포함하는, 방법.
- 191제175항에 있어서, 상기 사용자 의도를 결정하는 단계는:하나 이상의 행동가능한 의도를 결정하는 단계;및 상기 행동가능한 의도와 연관된 하나 이상의 파라미터를 결정하는 단계를 포함하는, 방법.
- 192제175항에 있어서, 상기 컨텍스트 정보는 사용자 특정 데이터, 디바이스 구성 데이터, 및 센서 데이터 중 적어도 하나를 포함하는, 방법.
- 193제175항에 있어서, 상기 사용자 의도가 정보제공 요청을 나타내는지 아니면 태스크를 수행하기 위한 요청을 나타내는지 결정하는 단계는:상기 사용자 의도가 시스템 구성을 변경하는 것인지 결정하는 단계를 포함하는, 방법.
- 194제175항에 있어서, 상기 사용자 의도가 정보제공 요청을 나타낸다는 결정에 따라, 상기 정보제공 요청에 대한 상기 음성 응답을 제공하는 단계는:상기 정보제공 요청에 따른 하나 이상의 시스템 구성의 상태를 획득하는 단계;및 상기 하나 이상의 시스템 구성의 상태에 따른 상기 음성 응답을 제공하는 단계를 포함하는, 방법.
- 195제175항에 있어서, 상기 사용자 의도가 정보제공 요청을 나타낸다는 결정에 따라, 상기 정보제공 요청에 대한 상기 음성 응답을 제공하는 것에 추가하여:상기 하나 이상의 시스템 구성의 상기 상태에 따른 정보를 제공하는 제1 사용자 인터페이스를 디스플레이하는 단계를 추가로 포함하는, 방법.
- 196제175항에 있어서, 상기 사용자 의도가 정보제공 요청을 나타낸다는 결정에 따라, 상기 정보제공 요청에 대한 상기 음성 응답을 제공하는 것에 추가하여:상기 정보제공 요청과 연관된 링크를 제공하는 단계를 추가로 포함하는, 방법.
- 197제175항에 있어서, 상기 사용자 의도가 태스크를 수행하기 위한 요청을 나타낸다는 결정에 따라, 상기 태스크를 수행하기 위하여 상기 사용자 디바이스와 연관된 상기 프로세스를 즉시실행하는 단계는:상기 프로세스를 이용하여 상기 태스크를 수행하는 단계를 포함하는, 방법.
- 198제197항에 있어서, 상기 태스크를 수행한 결과에 따른 제1 음성 출력을 제공하는 단계를 추가로 포함하는, 방법.
- 199제197항에 있어서, 상기 사용자가 상기 태스크를 수행한 결과를 조작하도록 인에이블하는 제2 사용자 인터페이스를 제공하는 단계를 추가로 포함하는, 방법.
- 200제199항에 있어서, 상기 제2 사용자 인터페이스는 상기 태스크를 수행한 상기 결과와 연관된 링크를 포함하는, 방법.
- 201제175항에 있어서, 상기 사용자 의도가 태스크를 수행하기 위한 요청을 나타낸다는 결정에 따라, 상기 태스크를 수행하기 위하여 상기 사용자 디바이스와 연관된 상기 프로세스를 즉시실행하는 단계는:상기 사용자가 상기 태스크를 수행하도록 인에이블하는 제3 사용자 인터페이스를 제공하는 단계를 포함하는, 방법.
- 202제201항에 있어서, 상기 제3 사용자 인터페이스는 상기 사용자가 상기 태스크를 수행하도록 인에이블하는 링크를 포함하는, 방법.
- 203제201항 또는 제202항에 있어서, 상기 제3 사용자 인터페이스와 연관된 제2 음성 출력을 제공하는 단계를 추가로 포함하는, 방법.
- 204하나 이상의 프로그램을 저장하는 비일시적 컴퓨터 판독가능 저장 매체로서, 상기 하나 이상의 프로그램은, 전자 디바이스의 하나 이상의 프로세서에 의해 실행될 때, 상기 전자 디바이스로 하여금:사용자로부터 상기 사용자 디바이스의 하나 이상의 시스템 구성을 관리하기 위한 스피치 입력을 수신하고 - 상기 사용자 디바이스는 동시에 복수의 사용자 인터페이스를 제공하도록 구성됨 -;상기 사용자 디바이스와 연관된 컨텍스트 정보를 식별하고;상기 스피치 입력 및 컨텍스트 정보에 기초하여 사용자 의도를 결정하고;상기 사용자 의도가 정보제공 요청을 나타내는지 아니면 태스크를 수행하기 위한 요청을 나타내는지 결정하고;상기 사용자 의도가 정보제공 요청을 나타낸다는 결정에 따라, 상기 정보제공 요청에 대한 음성 응답을 제공하고;상기 사용자 의도가 태스크를 수행하기 위한 요청을 나타낸다는 결정에 따라, 상기 태스크를 수행하기 위하여 상기 사용자 디바이스와 연관된 프로세스를 즉시실행하게 하는 명령어들을 포함하는, 비일시적 컴퓨터 판독가능 저장 매체.
- 205전자 디바이스로서, 하나 이상의 프로세서;메모리;및 메모리에 저장된 하나 이상의 프로그램들을 포함하고, 상기 하나 이상의 프로그램들은, 사용자로부터 상기 사용자 디바이스의 하나 이상의 시스템 구성을 관리하기 위한 스피치 입력을 수신하고 - 상기 사용자 디바이스는 동시에 복수의 사용자 인터페이스를 제공하도록 구성됨 -;상기 사용자 디바이스와 연관된 컨텍스트 정보를 식별하고;상기 스피치 입력 및 컨텍스트 정보에 기초하여 사용자 의도를 결정하고;상기 사용자 의도가 정보제공 요청을 나타내는지 아니면 태스크를 수행하기 위한 요청을 나타내는지 결정하고;상기 사용자 의도가 정보제공 요청을 나타낸다는 결정에 따라, 상기 정보제공 요청에 대한 음성 응답을 제공하고;상기 사용자 의도가 태스크를 수행하기 위한 요청을 나타낸다는 결정에 따라, 상기 태스크를 수행하기 위하여 상기 사용자 디바이스와 연관된 프로세스를 즉시실행하기 위한 명령어들을 포함하는, 전자 디바이스.
- 206전자 디바이스로서, 사용자로부터 상기 사용자 디바이스의 하나 이상의 시스템 구성을 관리하기 위한 스피치 입력을 수신하기 위한 수단 - 상기 사용자 디바이스는 동시에 복수의 사용자 인터페이스를 제공하도록 구성됨 -;상기 사용자 디바이스와 연관된 컨텍스트 정보를 식별하기 위한 수단;상기 스피치 입력 및 컨텍스트 정보에 기초하여 사용자 의도를 결정하기 위한 수단;상기 사용자 의도가 정보제공 요청을 나타내는지 아니면 태스크를 수행하기 위한 요청을 나타내는지 결정하기 위한 수단;상기 사용자 의도가 정보제공 요청을 나타낸다는 결정에 따라, 상기 정보제공 요청에 대한 음성 응답을 제공하기 위한 수단;및 상기 사용자 의도가 태스크를 수행하기 위한 요청을 나타낸다는 결정에 따라, 상기 태스크를 수행하기 위하여 상기 사용자 디바이스와 연관된 프로세스를 즉시실행하기 위한 수단을 포함하는, 전자 디바이스.
- 207전자 디바이스로서, 하나 이상의 프로세서;메모리;및 메모리에 저장된 하나 이상의 프로그램을 포함하고, 상기 하나 이상의 프로그램은 제175항 내지 제203항 중 어느 한 항의 방법을 수행하기 위한 명령어들을 포함하는, 전자 디바이스.
- 208전자 디바이스로서, 제175항 내지 제203항 중 어느 한 항의 방법을 수행하기 위한 수단을 포함하는, 전자 디바이스.
- 209비일시적 컴퓨터 판독가능 저장 매체로서, 전자 디바이스의 하나 이상의 프로세서에 의한 실행을 위한 하나 이상의 프로그램을 포함하고, 상기 하나 이상의 프로그램은, 상기 하나 이상의 프로세서에 의해 실행될 때, 상기 전자 디바이스로 하여금 제175항 내지 제203항 중 어느 한 항의 방법을 수행하게 하는 명령어들을 포함하는, 비일시적 컴퓨터 판독가능 저장 매체.
- 210디지털 어시스턴트를 동작시키기 위한 시스템으로서, 제175항 내지 제203항 중 어느 한 항의 방법을 수행하기 위한 수단을 포함하는, 시스템.
- 211전자 디바이스로서, 사용자로부터 상기 사용자 디바이스의 하나 이상의 시스템 구성을 관리하기 위한 스피치 입력을 수신하도록 구성된 수신 유닛 - 상기 사용자 디바이스는 동시에 복수의 사용자 인터페이스를 제공하도록 구성됨 -;및 프로세싱 유닛을 포함하며, 상기 프로세싱 유닛은, 상기 사용자 디바이스와 연관된 컨텍스트 정보를 식별하고;상기 스피치 입력 및 컨텍스트 정보에 기초하여 사용자 의도를 결정하고;상기 사용자 의도가 정보제공 요청을 나타내는지 아니면 태스크를 수행하기 위한 요청을 나타내는지 결정하고;상기 사용자 의도가 정보제공 요청을 나타낸다는 결정에 따라, 상기 정보제공 요청에 대한 음성 응답을 제공하고;상기 사용자 의도가 태스크를 수행하기 위한 요청을 나타낸다는 결정에 따라, 상기 태스크를 수행하기 위하여 상기 사용자 디바이스와 연관된 프로세스를 즉시실행하도록 구성된, 전자 디바이스.
- 212제211항에 있어서, 스피치 입력을 수신하기 이전에:상기 사용자 디바이스의 디스플레이 상에, 상기 디지털 어시스턴트를 호출하기 위한 어포던스를 디스플레이하는 것을 추가로 포함하는, 전자 디바이스.
- 213제212항에 있어서, 사전결정된 구절을 수신하는 것에 응답하여 상기 디지털 어시스턴트 서비스를 즉시실행하는 것을 추가로 포함하는, 전자 디바이스.
- 214제212항에 있어서, 상기 어포던스의 선택을 수신하는 것에 응답하여 상기 디지털 어시스턴트 서비스를 즉시실행하는 것을 추가로 포함하는, 전자 디바이스.
- 215제211항에 있어서, 상기 사용자 디바이스의 상기 하나 이상의 시스템 구성은 오디오 구성들을 포함하는, 전자 디바이스.
- 216제211항에 있어서, 상기 사용자 디바이스의 상기 하나 이상의 시스템 구성은 날짜 및 시간 구성들을 포함하는, 전자 디바이스.
- 217제211항에 있어서, 상기 사용자 디바이스의 상기 하나 이상의 시스템 구성은 구술 구성들을 포함하는, 전자 디바이스.
- 218제211항에 있어서, 상기 사용자 디바이스의 상기 하나 이상의 시스템 구성은 디스플레이 구성들을 포함하는, 전자 디바이스.
- 219제211항에 있어서, 상기 사용자 디바이스의 상기 하나 이상의 시스템 구성은 입력 디바이스 구성들을 포함하는, 전자 디바이스.
- 220제211항에 있어서, 상기 사용자 디바이스의 상기 하나 이상의 시스템 구성은 네트워크 구성들을 포함하는, 전자 디바이스.
- 221제211항에 있어서, 상기 사용자 디바이스의 상기 하나 이상의 시스템 구성은 통지 구성들을 포함하는, 전자 디바이스.
- 222제211항에 있어서, 상기 사용자 디바이스의 상기 하나 이상의 시스템 구성은 프린터 구성들을 포함하는, 전자 디바이스.
- 223제211항에 있어서, 상기 사용자 디바이스의 상기 하나 이상의 시스템 구성은 보안 구성들을 포함하는, 전자 디바이스.
- 224제211항에 있어서, 상기 사용자 디바이스의 상기 하나 이상의 시스템 구성은 백업 구성들을 포함하는, 전자 디바이스.
- 225제211항에 있어서, 상기 사용자 디바이스의 상기 하나 이상의 시스템 구성은 애플리케이션 구성들을 포함하는, 전자 디바이스.
- 226제211항에 있어서, 상기 사용자 디바이스의 상기 하나 이상의 시스템 구성은 사용자 인터페이스 구성들을 포함하는, 전자 디바이스.
- 227제211항에 있어서, 상기 사용자 의도를 결정하는 것은:하나 이상의 행동가능한 의도를 결정하는 것;및 상기 행동가능한 의도와 연관된 하나 이상의 파라미터를 결정하는 것을 포함하는, 전자 디바이스.
- 228제211항에 있어서, 상기 컨텍스트 정보는 사용자 특정 데이터, 디바이스 구성 데이터, 및 센서 데이터 중 적어도 하나를 포함하는, 전자 디바이스.
- 229제211항에 있어서, 상기 사용자 의도가 정보제공 요청을 나타내는지 아니면 태스크를 수행하기 위한 요청을 나타내는지 결정하는 것은:상기 사용자 의도가 시스템 구성을 변경하는 것인지 결정하는 것을 포함하는, 전자 디바이스.
- 230제211항에 있어서, 상기 사용자 의도가 정보제공 요청을 나타낸다는 결정에 따라, 상기 정보제공 요청에 대한 상기 음성 응답을 제공하는 것은:상기 정보제공 요청에 따른 하나 이상의 시스템 구성의 상태를 획득하는 것;및 상기 하나 이상의 시스템 구성의 상태에 따른 상기 음성 응답을 제공하는 것을 포함하는, 전자 디바이스.
- 231제211항에 있어서, 상기 사용자 의도가 정보제공 요청을 나타낸다는 결정에 따라, 상기 정보제공 요청에 대한 상기 음성 응답을 제공하는 것에 추가하여:상기 하나 이상의 시스템 구성의 상기 상태에 따른 정보를 제공하는 제1 사용자 인터페이스를 디스플레이하는 것을 추가로 포함하는, 전자 디바이스.
- 232제211항에 있어서, 상기 사용자 의도가 정보제공 요청을 나타낸다는 결정에 따라, 상기 정보제공 요청에 대한 상기 음성 응답을 제공하는 것에 추가하여:상기 정보제공 요청과 연관된 링크를 제공하는 것을 추가로 포함하는, 전자 디바이스.
- 233제211항에 있어서, 상기 사용자 의도가 태스크를 수행하기 위한 요청을 나타낸다는 결정에 따라, 상기 태스크를 수행하기 위하여 상기 사용자 디바이스와 연관된 상기 프로세스를 즉시실행하는 것은:상기 프로세스를 이용하여 상기 태스크를 수행하는 것을 포함하는, 전자 디바이스.
- 234제233항에 있어서, 상기 태스크를 수행한 결과에 따른 제1 음성 출력을 제공하는 것을 추가로 포함하는, 전자 디바이스.
- 235제233항에 있어서, 상기 사용자가 상기 태스크를 수행한 결과를 조작하도록 인에이블하는 제2 사용자 인터페이스를 제공하는 것을 추가로 포함하는, 전자 디바이스.
- 236제235항에 있어서, 상기 제2 사용자 인터페이스는 상기 태스크를 수행한 상기 결과와 연관된 링크를 포함하는, 전자 디바이스.
- 237제211항에 있어서, 상기 사용자 의도가 태스크를 수행하기 위한 요청을 나타낸다는 결정에 따라, 상기 태스크를 수행하기 위하여 상기 사용자 디바이스와 연관된 상기 프로세스를 즉시실행하는 것은:상기 사용자가 상기 태스크를 수행하도록 인에이블하는 제3 사용자 인터페이스를 제공하는 것을 포함하는, 전자 디바이스.
- 238제237항에 있어서, 상기 제3 사용자 인터페이스는 상기 사용자가 상기 태스크를 수행하도록 인에이블하는 링크를 포함하는, 전자 디바이스.
- 239제237항 또는 제238항에 있어서, 상기 제3 사용자 인터페이스와 연관된 제2 음성 출력을 제공하는 것을 추가로 포함하는, 전자 디바이스.
Independent claims239
1,022 paragraphs in 1 section, as filed
Intelligent digital assistant in a multitasking environment
CROSS-REFERENCE TO RELATED APPLICATIONS
This application is filed on June 10, 2016, and is entitled "INTELLIGENT DIGITAL ASSISTANT IN A MULTI-TASKING ENVIRONMENT," US Provisional Patent Application Nos. 62/348,728; and U.S. Regular Patent Application Serial No. 15/271,766, filed September 21, 2016, entitled "INTELLIGENT DIGITAL ASSISTANT IN A MULTI-TASKING ENVIRONMENT;" The contents of each of these are incorporated herein by reference in their entirety.
technical field
BACKGROUND This disclosure relates generally to digital assistants, and more particularly, to digital assistants that interact with users to perform tasks in a multitasking environment.
Digital assistants are growing in popularity. In a desktop or tablet environment, users frequently perform multitasking tasks, such as searching for files or information, managing files or folders, playing movies or songs, editing documents, adjusting system configurations, and sending emails. carry out It is often cumbersome and inconvenient for a user to manually parallel multiple tasks and frequently switch tasks. Accordingly, it may be desirable for a digital assistant to have the ability to help a user perform some tasks in a multitasking environment based on the user's voice input.
Some existing techniques for helping a user perform a task in a multitasking environment may include, for example, dictation. Typically, a user may be required to manually perform many different tasks in a multitasking environment. For example, a user may have been working on a presentation on his desktop computer yesterday and would like to continue working on the presentation. Users typically have to manually locate the presentation on their desktop computer, open the presentation, and continue editing the presentation.
As another example, the user may be booking a flight on his smartphone while the user is away from his desktop computer. A user may wish to continue booking flights when a desktop computer is available. In existing technologies, the user has to open a web browser on the user's desktop computer and start the flight reservation process. In other words, the previous flight booking process the user made on the smartphone may not continue on the user's desktop computer.
As another example, while editing a document on his desktop computer, a user may wish to change the system configuration, such as changing the brightness level of the screen or turning on a Bluetooth connection. In existing technologies, a user may need to stop editing a document, find and launch a brightness configuration application, and manually change settings. In a multi-tasking environment, some existing techniques may not be able to perform tasks as described in the examples above based on the user's speech input. Therefore, it may be desirable and advantageous to provide a voice-assisted digital assistant in a multi-tasking environment.
Systems and processes for operating a digital assistant are provided. According to one or more examples, a method includes receiving, at a user device having one or more processors and memory, a first speech input from a user. The method further includes identifying context information associated with the user device and determining user intent based on the first speech input and the context information. The method further includes determining whether the user intent is to perform the task using a search process or an object management process. The retrieval process is configured to retrieve data stored internally or externally to the user device, and the object management process is configured to manage objects associated with the user device. The method further includes performing the task using the search process in accordance with a determination that the user intent is to perform the task using the search process. The method further includes performing the task using the object management process in accordance with a determination that the user intent is to perform the task using the object management process.
In accordance with one or more examples, a method includes, at a user device having one or more processors and memory, receiving speech input from a user to perform a task. The method further includes identifying context information associated with the user device and determining user intent based on the speech input and context information associated with the user device. The method further includes, according to the user intent, determining whether the task is to be performed at the user device or at a first electronic device communicatively coupled to the user device. The method further includes receiving the content for performing the task in accordance with a determination that the task will be performed at the user device and the content for performing the task is located elsewhere. The method further includes providing content for performing the task to the first electronic device according to a determination that the task will be performed at the first electronic device and content for performing the task is located in a different location than the first electronic device include as
According to one or more examples, a method includes receiving, at a user device having one or more processors and memory, a speech input from a user to manage one or more system configurations of the user device. The user device is configured to simultaneously provide a plurality of user interfaces. The method further includes identifying context information associated with the user device and determining user intent based on the speech input and the context information. The method further includes determining whether the user intent indicates a request for information or a request to perform a task. The method further includes providing a voice response to the informational request in accordance with a determination that the user intent indicates the informational request. The method further includes, in accordance with a determination that the user intent indicates a request to perform the task, instantiating a process associated with the user device to perform the task.
Executable instructions for performing these functions are, optionally, included in a non-transitory computer-readable storage medium or other computer program product configured for execution by one or more processors. Executable instructions for performing these functions are optionally included in a transitory computer readable storage medium or other computer program product configured for execution by one or more processors.
For a better understanding of the variously described embodiments, reference should be made to the specific details for carrying out the following invention in connection with the following drawings in which like reference numerals indicate corresponding parts throughout the drawings. 1 is a block diagram illustrating a system and environment for implementing a digital assistant, according to various examples. 2A is a block diagram illustrating a portable multifunction device implementing a client-side portion of a digital assistant, in accordance with some embodiments. 2B is a block diagram illustrating example components for event processing, according to various examples. 3 illustrates a portable multifunction device implementing a client-side portion of a digital assistant, according to various examples. 4 is a block diagram of an exemplary multifunction device having a display and a touch-sensitive surface, in accordance with various examples. 5A illustrates an example user interface for a menu of applications on a portable multifunction device, according to various examples. 5B illustrates an example user interface for a multifunction device having a touch-sensitive surface separate from a display, according to various examples. 6A illustrates a personal electronic device according to various examples. 6B is a block diagram illustrating a personal electronic device according to various examples. 7A is a block diagram illustrating a digital assistant system, or a server portion thereof, in accordance with various examples. 7B illustrates the functions of the digital assistant shown in FIG. 7A according to various examples. 7C shows a portion of an ontology according to various examples. 8A-8F illustrate functions for performing a task using a search process or an object management process by a digital assistant according to various examples. 9A-9H illustrate functions for performing a task using a search process by a digital assistant according to various examples. 10A and 10B illustrate functions for performing a task using an object management process by a digital assistant according to various examples. 11A-11D illustrate functions for performing a task using a search process by a digital assistant according to various examples. 12A-12D illustrate functions for performing a task using a search process or an object management process by a digital assistant according to various examples. 13A-13C illustrate functions for performing a task using an object management process by a digital assistant according to various examples. 14A-14D illustrate functions for performing a task on a user device using content located elsewhere by a digital assistant according to various examples; 15A to 15D illustrate functions of performing a task in a first electronic device using content located elsewhere by a digital assistant according to various examples; 16A to 16C illustrate functions of performing a task in a first electronic device using content located elsewhere by a digital assistant according to various examples; 17A-17E illustrate functions of performing a task on a user device using content located elsewhere by a digital assistant according to various examples; 18A to 18F illustrate functions of providing system configuration information in response to a user's request for information provision by a digital assistant according to various examples. 19A-19D illustrate functions for performing a task in response to a user request by a digital assistant according to various examples. 20A-20G show a flow diagram of an example process for operating a digital assistant according to various examples. 21A-21E show a flow diagram of an example process for operating a digital assistant in accordance with various examples. 22A-22D show a flow diagram of an example process for operating a digital assistant in accordance with various examples. 23 shows a block diagram of an electronic device in accordance with various examples.
In the following description of the invention and embodiments, reference is made to the accompanying drawings, wherein specific embodiments in which they may be practiced are shown by way of example in the drawings. It should be understood that other embodiments and examples may be practiced and changes may be made without departing from the scope of the invention.
Techniques for providing digital assistants in a multitasking environment are desirable. As described herein, a technique for providing a digital assistant in a multi-tasking environment reduces the hassle of searching for objects or information, enables efficient object management, and enables tasks performed on user devices and other electronic devices. It is desirable for a variety of purposes, such as maintaining continuity between the two and reducing the user's effort to adjust the system configuration. These techniques advantageously allow a user to operate a digital assistant to perform various tasks using speech inputs in a multitasking environment. In addition, these techniques eliminate the hassle or inconvenience associated with performing various tasks in a multi-tasking environment. Also, by having users perform tasks verbally, they can keep both hands on the keyboard or mouse while performing tasks that may require a context switch - effectively, digitally like the user's "third hand." Let your assistant perform tasks. As will be appreciated, by having the user perform tasks verbally, the user can more efficiently complete tasks that may require multiple interactions with multiple applications. For example, searching for images and emailing them to individuals opens a search interface, enters a search term, selects one or more results, opens an email to compose, copies or moves result files to an open email, and email You may request to address and send to . Tasks like this can be accomplished more efficiently by voice commands such as "Find photos from Date X and send them to my family." Similar requests to move files, search for information on the Internet, and compose messages can all be made more effectively using voice, while allowing users to perform other tasks using their hands.
Although the following description uses terms such as "first", "second", etc. to describe various elements, these elements should not be limited by the terms. These terms are only used to distinguish one element from another. For example, a first storage device may be referred to as a second storage device, and similarly, a second storage device may be referred to as a first storage device, without departing from the scope of the variously described examples. The first storage device and the second storage device may both be storage devices, and in some cases may be separate and different storage devices.
The terminology used in the description of the variously described examples herein is for the purpose of describing the specific examples only, and is not intended to be limiting. As used in the description of the variously recited examples and in the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. It is intended It will also be understood that the term "and/or" as used herein indicates and encompasses any and all possible combinations of one or more of the listed associated items. The terms "include," "including," "comprise," and/or "comprising," as used herein, refer to the recited features, integers Specifies the presence of steps, steps, actions, elements, and/or components, but the presence of one or more other features, integers, steps, actions, elements, components, and/or groups thereof. Or it will be further understood that the addition is not excluded.
The term "if" refers to "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. It can be interpreted as meaning "in response to detecting". Similarly, the phrases "if it is determined that" or "if [the stated condition or event] is detected" means "on determining that" or "in response to determining that" or "[the stated condition or event]," depending on the context. or upon detecting [the event]" or "in response to detecting [the stated condition or event]".
One. system and environment
1 shows a block diagram of a system 100 according to various examples. In some examples, system 100 may implement a digital assistant. The terms "digital assistant", "virtual assistant", "intelligent automated assistant", or "automatic digital assistant" refer to interpreting natural language input in spoken and/or written form to infer user intent, and the inferred user It may refer to any information processing system that performs actions based on intent. For example, to operate according to an inferred user intent, the system may perform one or more of the following: identify a task flow using steps and parameters designed to achieve the inferred user intent; inputting specific requirements from inferred user intent into the task flow; executing a task flow by invoking programs, methods, services, APIs, etc.; and generating output responses to the user in audible (eg, speech) and/or visual form.
Specifically, the digital assistant may accept user requests, at least in part, in the form of natural language commands, requests, statements, statements, and/or questions. Typically, a user request may either seek an informed answer or performance of a task by a digital assistant. A satisfactory response to a user request may be the provision of the requested informational answer, the performance of the requested task, or a combination of the two. For example, the user may ask the digital assistant a question such as "Where am I now?" Based on the user's current location, the digital assistant may reply, "You are in Central Park near West Gate." The user can also request the performance of a task, for example "invite my friends to my girlfriend's birthday party next week". In response, the digital assistant may acknowledge the request by saying "yes, I'll do it right away" and then send an appropriate calendar invitation on behalf of the user to each of the user's friends listed in the user's electronic address book. invite) can be sent. During performance of the requested task, the digital assistant may interact with the user in a continuous conversation involving multiple exchanges of information, sometimes over an extended period of time. There are numerous other ways of interacting with a digital assistant to request information or performance of various tasks. In addition to providing verbal responses and taking programmed actions, the digital assistant may also provide responses in other visual or audible forms, such as text, notifications, music, videos, animations, and the like.
1 , in some examples, the digital assistant may be implemented according to a client-server model. The digital assistant includes a client-side portion 102 running on the user device 104 (hereinafter, "DA Client 102"), and a server-side portion 106 running on a server system 108 (hereinafter "DA Server"). (106)"). The DA client 102 may communicate with the DA server 106 via one or more networks 110 . The DA client 102 may provide client-side functions such as user-facing input and output processing, and communication with the DA server 106 . The DA server 106 may provide server-side functions to any number of DA clients 102 , each residing on an individual user device 104 .
In some examples, the DA server 106 has a client-facing I/O interface 112 , one or more processing modules 114 , data and models 116 , and an I/O interface 118 for external services. may include. The client-facing I/O interface 112 may facilitate client-facing input and output processing to the DA server 106 . One or more processing modules 114 may process speech input using data and models 116 and determine a user's intent based on natural language input. Further, the one or more processing modules 114 perform task execution based on the inferred user intent. In some examples, the DA server 106 may communicate with external services 120 via the network(s) 110 to complete a task or obtain information. An I/O interface 118 to external services may facilitate such communications.
User device 104 may be any suitable electronic device. For example, user devices may be a portable multifunction device (eg, device 200 described below with reference to FIG. 2A ), a multifunction device (eg, device 400 described below with reference to FIG. 4 ), or a personal electronic device (such as , a device 600 described later with reference to FIGS. 6A and 6B ). The portable multifunction device may be, for example, a mobile phone that also includes other functions such as PDA and/or music player functions. Specific examples of portable multifunction devices may include iPhone®, iPod Touch®, and iPad® devices from Apple Inc. of Cupertino, CA. have. Other examples of portable multifunction devices may include, without limitation, laptop or tablet computers. Also, in some examples, user device 104 may be a non-portable multifunction device. In particular, the user device 104 may be a desktop computer, a game console, a television, or a television set-top box. In some examples, user device 104 may operate in a multitasking environment. The multi-tasking environment allows a user to act as the device 104 to perform multiple tasks in parallel. For example, the multi-tasking environment may be a desktop or laptop environment, wherein the device 104 may perform a task in response to user input received from a physical user-interface device, and at the same time, the user's voice input You can perform other tasks in response. In some examples, user device 104 may include a touch-sensitive surface (eg, touch screen displays and/or touchpads). In addition, user device 104 may optionally include one or more other physical user interface devices, such as a physical keyboard, mouse, and/or joystick. Various examples of electronic devices, such as multifunction devices, are described in greater detail below.
Examples of communication network(s) 110 may include a local area network (LAN) and a wide area network (WAN), such as the Internet. Communication network(s) 110 is, for example, Ethernet (Ethernet), Universal Serial Bus (USB), FireWire (FIREWIRE), GSM (Global System for Mobile Communications), EDGE (Enhanced Data GSM Environment) ), code division multiple access (CDMA), time division multiple access (TDMA), Bluetooth, Wi-Fi, voice over Internet Protocol (VoIP), Wi-MAX, or It may be implemented using any known network protocol including a variety of wired or wireless protocols such as any other suitable communication protocol.
The server system 108 may be implemented on one or more standalone data processing devices or distributed computer networks. In some examples, the server system 108 may be configured with various virtualizations of third-party service providers (eg, third-party cloud service providers) to provide the underlying computing resources and/or infrastructure resources of the server system 108 . Devices and/or services may also be employed.
In some examples, user device 104 can communicate with DA server 106 via second user device 122 . The second user device 122 may be similar or identical to the user device 104 . For example, the second user device 122 may be similar to the devices 200 , 400 , or 600 described below with reference to FIGS. 2A , 4 , 6A and 6B . The user device 104 may be configured to be communicatively coupled to the second user device 122 via a direct communication connection, eg, Bluetooth, NFC, BTLE, etc., or via a wired or wireless network, eg, a local Wi-Fi network. can In some examples, the second user device 122 may be configured to act as a proxy between the user device 104 and the DA server 106 . For example, the DA client 102 of the user device 104 is configured to transmit information (eg, a user request received at the user device 104 ) to the DA server 106 via the second user device 122 . can be The DA server 106 may process the information and return relevant data (eg, data content in response to the user request) to the user device 104 via the second user device 122 .
In some examples, the user device 104 may be configured to communicate abbreviated requests for data to the second user device 122 to reduce the amount of information transmitted from the user device 104 . The second user device 122 may be configured to determine supplemental information to add to the shortened request to generate a complete request to send to the DA server 106 . Such a system architecture advantageously allows a user device 104 (eg, a watch or similar small electronic device) with limited communication capabilities and/or limited battery power to be configured with a second device with greater communication capabilities and/or battery power. 2 By using the user device 122 (eg, mobile phone, laptop computer, tablet computer, etc.) as a proxy to the DA server 106 , it can access services provided by the DA server 106 . Although only two user devices 104 , 122 are shown in FIG. 1 , system 100 may include any number and type of user devices configured with such a proxy configuration to communicate with DA server system 106 . It should be understood that there can be
Although the digital assistant shown in FIG. 1 may include both a client-side portion (eg, DA client 102 ) and a server-side portion (eg, DA server 106 ), in some examples, the functions of the digital assistant include: It can be implemented as a standalone application installed on a user device. Moreover, the division of functions between the client-side portion and the server-side portion of the digital assistant may be different in different implementations. For example, in some examples, the DA client may be a thin-client that provides only user-facing input and output processing functions and delegates all other functions of the digital assistant to a backend server.
2. electronic device
Attention is now directed to embodiments of electronic devices for implementing the client-side portion of the digital assistant. 2A is a block diagram illustrating a portable multifunction device 200 having a touch-sensitive display system 212 in accordance with some embodiments. Touch-sensitive display 212 is sometimes referred to as a "touch screen" for convenience, and is sometimes known or referred to as a "touch-sensitive display system". Device 200 includes memory 202 (optionally including one or more computer-readable storage media), memory controller 222 , one or more processing units (CPUs) 220 , peripherals interface 218 , RF Circuitry 208 , audio circuitry 210 , speaker 211 , microphone 213 , input/output (I/O) subsystem 206 , other input control devices 216 , and external port 224 . includes Device 200 optionally includes one or more optical sensors 264 . Device 200 may optionally include one or more contact intensity sensors for detecting intensity of contacts on device 200 (eg, a touch-sensitive surface, such as touch-sensitive display system 212 of device 200 ). 265). Device 200 may optionally generate tactile outputs on device 200 (eg, a touch such as touch-sensitive display system 212 of device 200 or touchpad 455 of device 400 ) one or more tactile output generators 267 for generating tactile outputs on the responsive surface. These components optionally communicate via one or more communication buses or signal lines 203 .
As used in the specification and claims, the term "strength" of a contact on a touch-sensitive surface refers to the force or pressure (force per unit area) of a contact (eg, finger contact) on the touch-sensitive surface, or touch-sensitive. Refers to a substitute (proxy) for the force or pressure of contact on a surface. The intensity of the contact has a range of values including at least four distinct values and more typically hundreds (eg, at least 256) distinct values. The intensity of the contact is optionally determined (or measured) using various approaches, and various sensors or combinations of sensors. For example, one or more force sensors below or adjacent to the touch-sensitive surface are optionally used to measure force at various points on the touch-sensitive surface. In some implementations, force measurements from multiple force sensors are combined (eg, weighted average) to determine an estimated force of contact. Similarly, a pressure-sensitive tip of the stylus is optionally used to determine the pressure of the stylus on the touch-sensitive surface. Alternatively, the size and/or changes thereto of the contact area detected on the touch-sensitive surface, the capacitance and/or changes thereto of the touch-sensitive surface in the vicinity of the contact, and/or the touch-sensitive surface in the vicinity of the contact The resistance and/or changes thereto are optionally used as a substitute for the force or pressure of contact on the touch-sensitive surface. In some implementations, the surrogate measurements for contact force or pressure are used directly to determine whether an intensity threshold has been exceeded (eg, the intensity threshold is described in units corresponding to the surrogate measurements). In some implementations, alternative measurements for contact force or pressure are converted to an estimated force or pressure, and the estimated force or pressure is used to determine whether an intensity threshold has been exceeded (eg, the intensity threshold is pressure threshold measured in units of pressure). Using the intensity of the contact as an attribute of the user input may not otherwise display affordances (eg, on a touch-sensitive display) and/or (eg, a touch-sensitive display, a touch-sensitive surface, or Enables user access to additional device functions that may not be accessible by the user on reduced sized devices with limited footprint to receive user input (via physical/mechanical controls such as knobs or buttons) make it
As used in the specification and claims, the term "tactile output" refers to the physical displacement of a device relative to its previous position, a component of the device (eg, a touch-sensitive surface) relative to another component of the device (eg, a housing). ), or the displacement of a component with respect to the center of mass of the device to be detected by the user using the user's tactile sense. For example, in a situation where the device or component of the device is in contact with a touch-sensitive surface of a user (eg, a finger, palm, or other part of the user's hand), the tactile output generated by the physical displacement may affect the user to be interpreted as a tactile sensation corresponding to a perceived change in the physical properties of a device or component of a device. For example, movement of a touch-sensitive surface (eg, a touch-sensitive display or trackpad) may optionally be directed to the user as a "down click" or "up click" of a physical actuator button. interpreted by In some cases, a user may experience a tactile sense, such as a "down click" or "up click," even in the absence of movement of a physical actuator button associated with the touch-sensitive surface that is physically depressed (eg, displaced) by movement of the user. will feel As another example, movement of the touch-sensitive surface is optionally interpreted or sensed as "roughness" of the touch-sensitive surface by the user, even when there is no change in smoothness of the touch-sensitive surface. . Although these interpretations of touch by the user will be influenced by the user's individualized sensory perception, there are many touch sensory perceptions that are common to the majority of users. Thus, when tactile output is described as corresponding to a user's specific sensory perception (eg, "up click", "down click", "roughness"), unless otherwise stated, the tactile output generated is typical (or average) corresponding to the physical displacement of the device or component thereof that will produce the described sensory perception for the user.
Device 200 is only one example of a portable multifunction device, and device 200 may optionally have more or fewer components than shown, optionally combine two or more components, or optionally a component It should be understood that they have a different configuration or arrangement of the The various components shown in FIG. 2A are implemented in hardware, software, or a combination of both hardware and software, including one or more signal processing and/or application-specific integrated circuits (ASICs).
Memory 202 may include one or more computer-readable storage media. Computer-readable storage media can be tangible and non-transitory. Memory 202 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic disk storage devices, flash memory devices, or other non-volatile solid state memory devices. Memory controller 222 may control access to memory 202 by other components of device 200 .
In some examples, the non-transitory computer-readable storage medium of memory 202 fetches instructions from an instruction execution system, apparatus, or device, such as a computer-based system, processor-containing system, or instruction execution system, apparatus, or device. may be used to store instructions for use by or in connection with another system capable of fetching and executing instructions (eg, to perform aspects of process 1200 described below). In other examples, instructions (eg, for performing aspects of process 1200 described below) may be stored on a non-transitory computer-readable storage medium (not shown) of server system 108 or memory 202 ) and the non-transitory computer-readable storage medium of the server system 108 . In the context of this specification, a "non-transitory computer-readable storage medium" may be any medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.
Peripherals interface 218 may be used to couple input and output peripherals of the device to CPU 220 and memory 202 . One or more processors 220 run or execute various software programs and/or sets of instructions stored in memory 202 to process data and to perform various functions for device 200 . In some embodiments, peripherals interface 218 , CPU 220 and memory controller 222 may be implemented on a single chip, such as chip 204 . In some other embodiments, they may be implemented on separate chips.
Radio frequency (RF) circuitry 208 receives and transmits RF signals, also referred to as electromagnetic signals. RF circuitry 208 converts electrical signals to/from electromagnetic signals and communicates with communication networks and other communication devices via the electromagnetic signals. The RF circuitry 208 may optionally include, but is not limited to, an antenna system, an RF transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a CODEC chipset, a subscriber identity module (SIM) card, memory, and the like. However, it includes well-known circuitry for performing these functions. The RF circuitry 208 may optionally include networks such as the Internet, an intranet, and/or a wireless network, also referred to as the World Wide Web (WWW), such as a cellular telephone network, a wireless local area network (LAN) and/or a MAN. (metropolitan area network), and communicate by wireless communication with other devices. The RF circuitry 208 optionally includes well-known circuitry for detecting near field communication (NFC) fields, for example by means of a short-range communication radio. Wireless communication is optionally, Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), high-speed downlink packet access (HSDPA), high-speed uplink packet access (HSUPA), Evolution (EV-DO) , Data-Only), HSPA, HSPA+, DC-HSPDA (Dual-Cell HSPA), LTE (long term evolution), NFC (near field communication), W-CDMA (wideband code division multiple access), CDMA (code division multiple) access), time division multiple access (TDMA), Bluetooth, Bluetooth Low Energy (BTLE), Wireless Fidelity (Wi-Fi) (eg, Protocols for IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, IEEE 802.11n and/or IEEE 802.11ac), voice over Internet Protocol (VoIP), Wi-MAX, e-mail (eg, Internet message access protocol (IMAP) and/or or post office protocol (POP)), instant messaging (eg, extensible messaging and presence protocol (XMPP), Session Initiation Protocol for Instant Messaging and Presence Leveraging Extensions (SIMPLE), Instant Messaging and Presence Service (IMPS)), and/or including Short Message Service (SMS), or communication protocols not yet developed as of the filing date of this document; It uses any of a plurality of communication standards, protocols and technologies including, but not limited to, any other suitable communication protocol.
The audio circuitry 210 , the speaker 211 , and the microphone 213 provide an audio interface between the user and the device 200 . The audio circuit unit 210 receives audio data from the peripheral device interface 218 , converts the audio data into an electrical signal, and transmits the electrical signal to the speaker 211 . The speaker 211 converts the electrical signal into sound waves that humans can hear. The audio circuitry 210 also receives an electrical signal converted from sound waves by the microphone 213 . The audio circuitry 210 converts the electrical signal into audio data and transmits the audio data to the peripherals interface 218 for processing. Audio data may be retrieved from and/or transmitted to memory 202 and/or RF circuitry 208 by peripherals interface 218 . In some embodiments, audio circuitry 210 also includes a headset jack (eg, 312 in FIG. 3 ). The headset jack is a connection between audio circuitry 210 and detachable audio input/output peripherals, such as output-only headphones, or a headset that has both an output (eg, one or both ear headphones) and an input (eg, a microphone). Provides an interface.
I/O subsystem 206 couples input/output peripherals on device 200 , such as touch screen 212 and other input control devices 216 , to peripherals interface 218 . I/O subsystem 206 optionally includes one or more input controllers for display controller 256 , light sensor controller 258 , intensity sensor controller 259 , haptic feedback controller 261 , and other input or control devices. (260). One or more input controllers 260 receive/transmit electrical signals to/from other input control devices 216 . Other input control devices 216 optionally include physical buttons (eg, a push button, a rocker button, etc.), a dial, a slider switch, a joystick, a click wheel, and the like. In some alternative embodiments, the input controller(s) 260 are optionally coupled to (or not coupled to) any of a keyboard, an infrared port, a USB port, and a pointer device such as a mouse. One or more buttons (eg, 308 in FIG. 3 ) optionally include up/down buttons for volume control of speaker 211 and/or microphone 213 . The one or more buttons optionally include a push button (eg, 306 in FIG. 3 ).
A quick press of the push button may initiate the process of unlocking the touch screen 212 or using gestures on the touch screen to unlock the device, which was filed on December 23, 2005. As described in U.S. Patent Application Serial No. 11/322,549 (U.S. Patent No. 7,657,849), "Unlocking a Device by Performing Gestures on an Unlock Image," which is incorporated herein by reference in its entirety. A longer press of the push button (eg, 306 ) may power on or off the device 200 . A user may customize the function of one or more of the buttons. The touch screen 212 is used to implement virtual or soft buttons and one or more soft keyboards.
The touch-sensitive display 212 provides an input interface and an output interface between the device and the user. Display controller 256 receives and/or transmits electrical signals to/from touch screen 212 . Touch screen 212 displays a visual output to the user. Visual output may include graphics, text, icons, video, and any combination thereof (collectively "graphics"). In some embodiments, some or all of the visual output may correspond to user interface objects.
Touch screen 212 has a touch-sensitive surface, sensor, or set of sensors that accepts input from a user based on haptic and/or tactile contact. Touch screen 212 and display controller 256 (along with any associated modules and/or sets of instructions within memory 202) contact (and any movement or interruption of) contact on touch screen 212 . , and converts the detected contact into interaction with user interface objects (eg, one or more soft keys, icons, webpages, or images) displayed on touch screen 212 . In an exemplary embodiment, the point of contact between the touch screen 212 and the user corresponds to the user's finger.
The touch screen 212 may use liquid crystal display (LCD) technology, light emitting polymer display (LPD) technology, or light emitting diode (LED) technology, although other display technologies may be used in other embodiments. . The touch screen 212 and display controller 256 are configured to determine one or more points of contact with the touch screen 212 , or other proximity sensor arrays, as well as capacitive, resistive, infrared, and surface acoustic wave technologies. Any of a plurality of currently known or later developed touch sensing technologies, including but not limited to other elements, may be used to detect a contact and any movement or disruption thereof. In an exemplary embodiment, a projected mutual capacitance sensing technology such as that found in the iPhone® and iPod Touch® from Apple Inc. of Cupertino, CA is used.
The touch-sensitive display in some embodiments of the touch screen 212 is described in US Pat. It may be similar to the multi-touch sensitive touchpads described in US Patent Publication No. 2002/0015024A1, each of which is incorporated herein by reference in its entirety. However, touch screen 212 displays visual output from device 200 , whereas touch-sensitive touchpads do not provide visual output.
The touch-sensitive display in some embodiments of the touch screen 212 may be as described in the following applications: (1) U.S. Patent Application Serial No. 11/381,313, filed May 2, 2006, "Multipoint Touch Surface Controller"; (2) US Patent Application Serial No. 10/840,862, "Multipoint Touchscreen", filed May 6, 2004; (3) U.S. Patent Application Serial No. 10/903,964, "Gestures For Touch Sensitive Input Devices", filed July 30, 2004; (4) US Patent Application Serial No. 11/048,264, "Gestures For Touch Sensitive Input Devices," filed on January 31, 2005; (5) US Patent Application Serial No. 11/038,590, "Mode-Based Graphical User Interfaces For Touch Sensitive Input Devices," filed on January 18, 2005; (6) US Patent Application Serial No. 11/228,758, "Virtual Input Device Placement On A Touch Screen User Interface," filed September 16, 2005; (7) US Patent Application Serial No. 11/228,700, "Operation Of A Computer With A Touch Screen Interface," filed September 16, 2005; (8) US Patent Application Serial No. 11/228,737, "Activating Virtual Keys Of A Touch-Screen Virtual Keyboard," filed September 16, 2005; and (9) US Patent Application Serial No. 11/367,749, "Multi-Functional Hand-Held Device", filed March 3, 2006. All of these applications are incorporated herein by reference in their entirety.
The touch screen 212 may have a video resolution greater than 100 dpi. In some embodiments, the touch screen has a video resolution of approximately 160 dpi. A user may contact the touch screen 212 using any suitable object or appendage, such as a stylus, finger, or the like. In some embodiments, the user interface is primarily designed to operate using finger-based contacts and gestures, which may be less precise than stylus-based input due to the larger contact area of the finger on the touch screen. In some embodiments, the device converts coarse finger-based input into precise pointer/cursor positions or commands for performing actions desired by the user.
In some embodiments, in addition to a touch screen, device 200 may include a touchpad (not shown) for activating or deactivating certain functions. In some embodiments, the touchpad is a touch-sensitive area of the device that, unlike the touch screen, does not display visual output. The touchpad may be a touch-sensitive surface separate from the touch screen 212 or an extension of the touch-sensitive surface formed by the touch screen.
Device 200 also includes a power system 262 for powering various components. Power system 262 may include a power management system, one or more power sources (eg, a battery or alternating current (AC)), a recharge system, a power fault detection circuit, a power converter or inverter, a power status indicator (eg, a light emitting diode). diode), and any other components associated with the generation, management, and distribution of power within portable devices.
Device 200 may also include one or more optical sensors 264 . 2A shows a light sensor coupled to a light sensor controller 258 in the I/O subsystem 206 . The optical sensor 264 may include a charge-coupled device (CCD) or complementary metal-oxide semiconductor (CMOS) phototransistors. The light sensor 264 receives light from the surrounding environment, projected through one or more lenses, and converts the light into data representing an image. In conjunction with imaging module 243 (also referred to as a camera module), light sensor 264 can capture still images or video. In some embodiments, the optical sensor is located on the back of the device 200 opposite the touch screen display 212 on the front of the device so that the touch screen display acts as a viewfinder for still and/or video image acquisition. make it usable In some embodiments, an optical sensor is positioned on the front of the device to allow an image of the user to be acquired for the videoconference while the user views other videoconference participants on the touch screen display. In some embodiments, the position of the optical sensor 264 can be changed by the user (eg, by rotating a lens and sensor within the device housing) so that a single optical sensor 264 can video conferencing and stationary and /or be used for both video image acquisition.
Device 200 also optionally includes one or more contact intensity sensors 265 . 2A shows a contact intensity sensor coupled to an intensity sensor controller 259 in the I/O subsystem 206 . Contact intensity sensor 265 may optionally include one or more piezoresistive strain gauges, capacitive force sensors, electrical force sensors, piezoelectric force sensors, optical force sensors, capacitive touch-sensitive surfaces, or other intensity sensors (eg, , sensors used to measure the force (or pressure) of a contact on a touch-sensitive surface. The contact intensity sensor 265 receives contact intensity information (eg, pressure information or a surrogate for pressure information) from the environment. In some embodiments, at least one contact intensity sensor is collocated with or proximate to a touch-sensitive surface (eg, touch-sensitive display system 212 ). In some embodiments, the at least one contact intensity sensor is located on the back side of the device 200 opposite the touch screen display 212 located on the front side of the device 200 .
Device 200 may also include one or more proximity sensors 266 . 2A shows a proximity sensor 266 coupled to a peripherals interface 218 . Alternatively, the proximity sensor 266 may be coupled to the input controller 260 within the I/O subsystem 206 . Proximity sensor 266 is described in US Patent Applications 11/241,839, "Proximity Detector In Handheld Device;" 11/240,788, "Proximity Detector In Handheld Device"; 11/620,702, "Using Ambient Light Sensor To Augment Proximity Sensor Output"; 11/586,862, "Automated Response To And Sensing Of User Activity In Portable Devices"; and 11/638,251, "Methods And Systems For Automatic Configuration Of Peripherals," which are incorporated herein by reference in their entirety. In some embodiments, the proximity sensor turns off and disables the touch screen 212 when the multifunction device is placed near the user's ear (eg, when the user is on a phone call).
Device 200 also optionally includes one or more tactile output generators 267 . 2A shows a tactile output generator coupled to a haptic feedback controller 261 within the I/O subsystem 206 . Tactile output generator 267 optionally generates one or more electroacoustic devices and/or motors, solenoids, electroactive polymers, piezoelectric actuators, electrostatic actuators, or other tactile output devices, such as speakers or other audio components. Electromechanical devices that convert energy into linear motion, such as a component (eg, a component that converts electrical signals into tactile outputs on the device). The contact intensity sensor 265 receives tactile feedback generation instructions from the haptic feedback module 233 to generate tactile outputs on the device 200 that can be sensed by a user of the device 200 . In some embodiments, the at least one tactile output generator is positioned with or proximate to a touch-sensitive surface (eg, touch-sensitive display system 212 ), and optionally, vertically (eg, vertically) the touch-sensitive surface. , generate a tactile output by moving it in/out of the surface of the device 200 ) or laterally (eg, back and forth in the same plane as the surface of the device 200 ). In some embodiments, the at least one tactile output generator sensor is located on the back side of the device 200 opposite the touch screen display 212 located on the front side of the device 200 .
Device 200 may also include one or more accelerometers 268 . 2A shows an accelerometer 268 coupled to a peripherals interface 218 . Alternatively, the accelerometer 268 may be coupled to the input controller 260 within the I/O subsystem 206 . Accelerometer 268 is described in US Patent Publication No. 20050190059, "Acceleration-based Theft Detection System for Portable Electronic Devices," and US Patent Publication 2006017692, "Methods And Apparatuses For Operating A Portable Device Based On An Accelerometer." as described above, both of which are incorporated herein by reference in their entirety. In some embodiments, information is displayed on the touch screen display in a portrait view or a landscape view based on analysis of data received from one or more accelerometers. Device 200 optionally includes, in addition to accelerometer(s) 268 , a magnetometer (not shown), and a GPS (or GLONASS or other global navigation system) receiver (not shown).
In some embodiments, software components stored in memory 202 include operating system 226 , communication module (or set of instructions) 228 , contact/motion module (or set of instructions) 230 , graphics module ( or set of instructions) 232 , text input module (or set of instructions) 234 , GPS module (or set of instructions) 235 , digital assistant client module 229 , and applications (or set of instructions) s) 236 . The memory 202 may also store data and models, such as user data and models 231 . Moreover, in some embodiments, the memory ( 202 in FIG. 2A or 470 in FIG. 4 ) stores the device/global internal state 257 as shown in FIGS. 2A and 4 . The device/global internal state 257, if present, includes an active application state indicating which applications are currently active; a display state indicating which applications, views, or other information occupy the various areas of the touch screen display 212 ; sensor status, including information obtained from various sensors and input control devices 216 of the device; and location information regarding the location and/or posture of the device.
Operating system 226 (eg, an embedded operating system such as Darwin, RTXC, LINUX, UNIX, OS X, iOS, WINDOWS, or VxWorks) is responsible for general system tasks (eg, memory management, storage device control, power management, etc.) It includes various software components and/or drivers for controlling and managing the software, and facilitates communication between various hardware and software components.
The communication module 228 facilitates communication with other devices via one or more external ports 224 , and also provides various functions for processing data received by the RF circuitry 208 and/or external ports 224 . Includes software components. External port 224 (eg, Universal Serial Bus (USB), FIREWIRE, etc.) is configured to couple to other devices either directly or indirectly through a network (eg, Internet, wireless LAN, etc.). In some embodiments, the external port is a multi-pin (eg, 30-pin) connector that is the same as, similar to, and/or compatible with the 30-pin connector used in iPod® (trademark of Apple Inc.) devices.
Contact/motion module 230 optionally detects contact with touch screen 212 (in conjunction with display controller 256 ), and other touch-sensitive devices (eg, a touchpad or physical click wheel). do. The contact/motion module 230 is configured to determine whether a contact has occurred (eg, to detect a finger-down event), the strength of the contact (eg, the force or pressure of the contact, or the determining a substitute for force or pressure), determining if there is movement of a contact and tracking movement across a touch-sensitive surface (e.g., detecting one or more finger-dragging events); ), and various software components for performing various operations related to detection of a contact, such as determining whether the contact has ceased (eg, detecting a finger-up event or contact cessation). include those The contact/motion module 230 receives contact data from the touch-sensitive surface. Determining the movement of the point of contact represented by the set of contact data optionally determines a speed (magnitude), velocity (magnitude and direction), and/or acceleration (change in magnitude and/or direction) of the point of contact. includes doing These operations are, optionally, applied to single contacts (eg, one finger contacts) or to multiple simultaneous contacts (eg, "multi-touch"/multiple finger contacts). In some embodiments, contact/motion module 230 and display controller 256 detect a contact on the touchpad.
In some embodiments, the contact/motion module 230 is a set of one or more intensity thresholds to determine whether an action was performed by the user (eg, to determine whether the user "clicked" on an icon). use the In some embodiments, at least a subset of the intensity thresholds are determined according to software parameters (eg, the intensity thresholds are not determined by activation thresholds of specific physical actuators, changing the physical hardware of the device 200 ) can be adjusted without). For example, the mouse "click" threshold of a trackpad or touch screen display can be set to any of a wide range of predefined threshold values without changing the trackpad or touch screen display hardware. Additionally, in some implementations, a user of the device (eg, by adjusting individual intensity thresholds and/or by adjusting a plurality of intensity thresholds at once with a system level click "Intensity" parameter) can select one of a set of intensity thresholds. Software settings are provided to adjust for anomalies.
The contact/motion module 230 optionally detects a gesture input by the user. Different gestures on the touch-sensitive surface have different contact patterns (eg, different motions, timings, and/or intensities of detected contacts). Thus, the gesture is selectively detected by detecting a specific contact pattern. For example, detecting a finger tap gesture may include detecting a finger-down event followed by a finger-down event at the same location (or substantially the same location) as the finger-down event (eg, at the location of the icon). and detecting an up (liftoff) event. As another example, detecting a finger swipe gesture on a touch-sensitive surface may detect a finger-down event followed by one or more finger-dragging events followed by a finger-up (lift-off) event. ) to detect the event.
The graphics module 232 renders and renders graphics on the touch screen 212 or other display, including components for changing the visual effect (eg, brightness, transparency, saturation, contrast, or other visual properties) of the displayed graphics. It includes various well-known software components for display. As used herein, the term "graphic" includes, without limitation, text, web pages, icons (eg, user interface objects including soft keys), digital images, videos, animations, and the like. , including any object that can be displayed to the user.
In some embodiments, graphics module 232 stores data representing graphics to be used. Each graphic is optionally assigned a corresponding code. Graphics module 232 receives one or more codes from applications, etc. that specify a graphic to be displayed, along with coordinate data and other graphic attribute data, if necessary, and then generates screen image data for output to display controller 256 . do.
The haptic feedback module 233 is an instruction used by the tactile output generator(s) 267 to generate tactile outputs at one or more locations on the device 200 in response to user interactions with the device 200 . It includes various software components for creating
The text input module 234 , which may be a component of the graphics module 232 , may be used for various applications (eg, contacts 237 , email 240 , IM 241 , browser 247 , and text input It provides soft keyboards for entering text in any other application).
The GPS module 235 determines the location of the device and uses this information for use in various applications (eg, to the phone 238 for use in location-based dialing; the camera 243 as photo/video metadata). and to applications that provide location-based services such as weather widgets, local yellow page widgets, and map/navigation widgets).
Digital assistant client module 229 may include various client-side digital assistant instructions for providing client-side functions of the digital assistant. For example, digital assistant client module 229 may be configured to interface with various user interfaces of portable multifunction device 200 (eg, microphone 213 , accelerometer(s) 268 , touch-sensitive display system 212 , light sensor). may accept voice input (eg, speech input), text input, touch input, and/or gesture input via (s) 264 , other input control devices 216 , etc.). Digital assistant client module 229 may also provide various output interfaces of portable multifunction device 200 (eg, speaker 211 , touch-sensitive display system 212 , tactile output generator(s) 267 , etc.) may provide output in audible (eg, speech output), visual, and/or tactile forms. For example, the output may be provided as voice, sound, notification, text message, menu, graphic, video, animation, vibration, and/or combinations of two or more of these. During operation, digital assistant client module 229 may communicate with DA server 106 using RF circuitry 208 .
User data and models 231 may be derived from various data associated with the user (eg, user-specific vocabulary data, user preference data, user-specific name pronunciations, the user's electronic address book) to provide client-side functions of the digital assistant. data, to-do lists, shopping lists, etc.). In addition, user data and models 231 may include various models (eg, speech recognition models, statistical language models, natural language processing models, ontology, task flow models) to process user input and determine user intent. , service models, etc.).
In some examples, digital assistant client module 229 is configured to use various sensors, subsystems, and peripheral devices of portable multifunction device 200 to establish a context associated with the user, current user interaction, and/or current user input. They may be used to collect additional information from the surrounding environment of the portable multifunction device 200 . In some examples, digital assistant client module 229 can provide contextual information or a subset thereof along with user input to DA server 106 to help infer the user's intent. In some examples, the digital assistant may also use the context information to determine how to prepare and deliver outputs to the user. Context information may be referred to as context data.
In some examples, contextual information accompanying user input may include sensor information, such as lighting, ambient noise, ambient temperature, images or videos of the ambient environment, and the like. In some examples, the context information may also include the physical state of the device, such as device orientation, device location, device temperature, power level, speed, acceleration, motion patterns, cellular signal strength, and the like. In some examples, information related to the software state of the DA server 106 , such as running processes, installed programs, past and present network activities, background services, error logs, resource usage, etc., and portable The information of the multifunction device 200 may be provided to the DA server 106 as context information associated with a user input.
In some examples, digital assistant client module 229 can optionally provide information stored on portable multifunction device 200 (eg, user data 231 ) in response to requests from DA server 106 . have. In some examples, digital assistant client module 229 may also elicit additional input from the user through natural language conversations or other user interfaces upon request by DA server 106 . The digital assistant client module 229 may pass additional inputs to the DA server 106 to aid the DA server 106 in intent inference and/or fulfillment of the user's intent expressed in the user request.
A more detailed description of the digital assistant is described below with reference to FIGS. 7A-7C . It should be appreciated that digital assistant client module 229 may include any number of submodules of digital assistant module 726 described below.
Applications 236 may include the following modules (or sets of instructions), or a subset or superset thereof:
<img file="KR20180098409A_D0001.tif" />contacts module 237 (sometimes referred to as an address book or contact list);
<img file="KR20180098409A_D0002.tif" />phone module 238;
<img file="KR20180098409A_D0003.tif" />video conferencing module 239;
<img file="KR20180098409A_D0004.tif" />email client module 240;
<img file="KR20180098409A_D0005.tif" />an instant messaging (IM) module 241 ;
<img file="KR20180098409A_D0006.tif" />exercise support module 242;
<img file="KR20180098409A_D0007.tif" />a camera module 243 for still and/or video images;
<img file="KR20180098409A_D0008.tif" />image management module 244;
<img file="KR20180098409A_D0009.tif" />video player module;
<img file="KR20180098409A_D0010.tif" />music player module;
<img file="KR20180098409A_D0011.tif" />browser module 247;
<img file="KR20180098409A_D0012.tif" />calendar module 248;
<img file="KR20180098409A_D0013.tif" />Weather widget 249-1, stock widget 249-2, calculator widget 249-3, alarm clock widget 249-4, dictionary widget 249-5, and other widgets obtained by the user as well as widget modules 249, which may include one or more of user-created widgets 249-6;
<img file="KR20180098409A_D0014.tif" />a widget generator module 250 for creating user-created widgets 249-6;
<img file="KR20180098409A_D0015.tif" />search module 251;
<img file="KR20180098409A_D0016.tif" />a video and music player module 252 incorporating a video player module and a music player module;
<img file="KR20180098409A_D0017.tif" />memo module 253;
<img file="KR20180098409A_D0018.tif" />map module 254; and/or
<img file="KR20180098409A_D0019.tif" />Online video module (255).
Examples of other applications 236 that may be stored in memory 202 are other word processing applications, other image editing applications, drawing applications, presentation applications, JAVA-enabled applications, encryption, digital rights. management, voice recognition, and voice replication.
Along with touch screen 212 , display controller 256 , contact/motion module 230 , graphics module 232 , and text input module 234 , contacts module 237 may include an address book or contact list (eg, (stored in application internal state 292 of contact module 237 in memory 202 or memory 470), including: adding name(s) to an address book; ; deleting name(s) from the address book; associating a name with a phone number(s), email address(es), physical address(es), or other information; associating an image with a name; sorting and sorting names; providing phone numbers or email addresses to initiate and/or facilitate communication by phone 238 , video conferencing module 239 , email 240 , or IM 241 , and the like.
RF circuitry 208 , audio circuitry 210 , speaker 211 , microphone 213 , touch screen 212 , display controller 256 , contact/motion module 230 , graphics module 232 , and text In conjunction with the input module 234 , the phone module 238 enters a sequence of characters corresponding to a phone number, accesses one or more phone numbers in the contacts module 237 , modifies the entered phone number, and It can be used to dial the phone number of As noted above, wireless communication may utilize any of a plurality of communication standards, protocols, and technologies.
RF circuitry 208 , audio circuitry 210 , speaker 211 , microphone 213 , touch screen 212 , display controller 256 , optical sensor 264 , optical sensor controller 258 , contact/motion Along with module 230 , graphics module 232 , text input module 234 , contacts module 237 , and phone module 238 , video conferencing module 239 can communicate with the user according to user instructions and one or more other users. and executable instructions for initiating, conducting, and terminating a video conference between participants.
Along with RF circuitry 208 , touch screen 212 , display controller 256 , contact/motion module 230 , graphics module 232 and text input module 234 , email client module 240 provides user instruction executable instructions for composing, sending, receiving, and managing email in response to In conjunction with the image management module 244 , the email client module 240 greatly facilitates creating and sending emails with still or video images taken with the camera module 243 .
The instant messaging module 241, along with the RF circuitry 208, the touch screen 212, the display controller 256, the contact/motion module 230, the graphics module 232, and the text input module 234, are Enter a sequence of characters corresponding to a message, modify previously entered characters (eg, Short Message Service (SMS) for phone-based instant messages) or Multimedia Message Service (Multimedia Message Service) ; MMS) protocol, or using XMPP, SIMPLE or IMPS for Internet-based instant messages), including executable instructions to send, receive instant messages, and view received instant messages, respectively. . In some embodiments, the instant messages sent and/or received are graphics, pictures, audio files, video files, and/or as supported in MMS and/or Enhanced Messaging Service (EMS). Other attachments may be included. As used herein, "instant messaging" refers to phone-based messages (eg, messages sent using SMS or MMS) and Internet-based messages (eg, messages sent using XMPP, SIMPLE, or IMPS). ) refers to both.
RF circuitry 208, touch screen 212, display controller 256, contact/motion module 230, graphics module 232, text input module 234, GPS module 235, map module 254 , and in conjunction with the music player module, the exercise support module 242 devises exercises (eg, along with time, distance, and/or calorie expenditure goals); communicate with motion sensors (sports devices); receive motion sensor data; calibrate sensors used to monitor movement; select and play music for exercise; executable instructions to display, store, and transmit athletic data.
with touch screen 212 , display controller 256 , optical sensor(s) 264 , optical sensor controller 258 , contact/motion module 230 , graphics module 232 and image management module 244 . , camera module 243 captures still images or video (including video streams) and stores them in memory 202 , modifies characteristics of the still image or video, Contains executable instructions to delete a video.
Along with the touch screen 212 , the display controller 256 , the contact/motion module 230 , the graphics module 232 , the text input module 234 and the camera module 243 , the image management module 244 is configured to perform static and and/or executable instructions for arranging, modifying (eg, editing), or otherwise manipulating, labeling, deleting, presenting (eg, in a digital slide show or album), and storing video images .
Along with RF circuitry 208 , touch screen 212 , display controller 256 , contact/motion module 230 , graphics module 232 and text input module 234 , browser module 247 is configured to: executable instructions to browse the Internet according to user instructions, including retrieving, linking to, receiving, and displaying files or portions thereof, as well as attachments and other files linked to web pages include those
RF circuitry 208 , touch screen 212 , display controller 256 , contact/motion module 230 , graphics module 232 , text input module 234 , email client module 240 , and browser module ( In conjunction with 247 ), calendar module 248 is configured to generate, display, modify, and store calendars and data associated with calendars (eg, calendar entries, to-do lists, etc.) according to user instructions. It contains executable instructions.
Widget modules, along with RF circuitry 208 , touch screen 212 , display controller 256 , contact/motion module 230 , graphics module 232 , text input module 234 and browser module 247 . 249 may be downloaded and used by a user (eg, a weather widget 249-1, a stock widget 249-2, a calculator widget 249-3, an alarm clock widget 249-4, and a dictionary widget). 249-5), or mini-applications that may be created by the user (eg, user-created widgets 249-6). In some embodiments, the widget includes a Hypertext Markup Language (HTML) file, a Cascading Style Sheets (CSS) file, and a JavaScript file. In some embodiments, a widget includes an Extensible Markup Language (XML) file and a JavaScript file (eg, Yahoo! widgets).
A widget generator module, along with RF circuitry 208 , touch screen 212 , display controller 256 , contact/motion module 230 , graphics module 232 , text input module 234 and browser module 247 . 250 may be used by a user to create widgets (eg, change a user-specific portion of a web page into a widget).
With touch screen 212 , display controller 256 , contact/motion module 230 , graphics module 232 , and text input module 234 , search module 251 may be configured to perform one or more search criteria according to user instructions. executable instructions to search for text, music, sound, image, video, and/or other files in memory 202 that match (eg, one or more user specific search terms).
With touch screen 212 , display controller 256 , contact/motion module 230 , graphics module 232 , audio circuitry 210 , speaker 211 , RF circuitry 208 and browser module 247 . , video and music player module 252 provides executable instructions that enable a user to download and play recorded music and other sound files stored in one or more file formats, such as MP3 or AAC files, and videos (eg, , to display, show, or otherwise play on the touch screen 212 or on an externally connected display through the external port 224). In some embodiments, device 200 optionally includes the functionality of an MP3 player, such as an iPod (trademark of Apple Inc.).
Along with the touch screen 212 , the display controller 256 , the contact/motion module 230 , the graphics module 232 and the text input module 234 , the memo module 253 is configured to take notes, to do, according to user instructions. executable instructions to create and manage to-do lists and the like.
RF circuitry 208 , touch screen 212 , display controller 256 , contact/motion module 230 , graphics module 232 , text input module 234 , GPS module 235 , and browser module 247 . ), the map module 254 provides maps and data associated with the maps according to user instructions (eg, driving directions; data regarding shops and other points of interest at or near a particular location; and other location-based data), and may be used to receive, display, modify, and store.
Touch screen 212, display controller 256, contact/motion module 230, graphics module 232, audio circuitry 210, speaker 211, RF circuitry 208, text input module 234, Along with the email client module 240 and browser module 247 , the online video module 255 allows a user to access, browse online videos in one or more file formats, such as H.264, (eg, streaming and receive (by download), play (eg, on a touch screen or on an externally connected display via external port 224), send an email with a link to a particular online video, and otherwise manage contains commands. In some embodiments, instant messaging module 241 rather than email client module 240 is used to send a link to a particular online video. For further description of online video applications, see U.S. Provisional Patent Application Serial No. 60/936,562, "Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos," filed June 20, 2007, and December 31, 2007. may be found in U.S. Patent Application Serial No. 11/968,067, filed on the date of "Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos," the contents of which are hereby incorporated by reference in their entirety. .
Each of the modules and applications identified above includes executable instructions for performing one or more of the functions described above and the methods described herein (eg, the computer-implemented methods and other information processing methods described herein). corresponding to a set of These modules (eg, sets of instructions) need not be implemented as separate software programs, procedures, or modules, and thus various subsets of these modules may be combined or otherwise rearranged in various embodiments. . For example, a video player module may be combined with a music player module into a single module (eg, video and music player module 252 in FIG. 2A ). In some embodiments, memory 202 may store a subset of the modules and data structures identified above. In addition, the memory 202 may store additional modules and data structures not described above.
In some embodiments, device 200 is a device in which operation of a predefined set of functions on the device is performed exclusively via a touch screen and/or touchpad. By using a touch screen and/or touchpad as the primary input control device for operation of device 200 , the number of physical input control devices (such as push buttons, dials, etc.) on device 200 may be reduced.
The predefined set of functions performed exclusively via the touch screen and/or touchpad optionally include navigation between user interfaces. In some embodiments, the touchpad, when touched by a user, navigates the device 200 from any user interface displayed on the device 200 to a main, home, or root menu. In these embodiments, a "menu button" is implemented using a touchpad. In some other embodiments, the menu button is a physical push button or other physical input control device instead of a touchpad.
2B is a block diagram illustrating example components for event processing, in accordance with some embodiments. In some embodiments, memory ( 202 in FIG. 2A or 470 in FIG. 4 ) includes event classifier 270 (eg, in operating system 226 ) and individual application 236 - 1 (eg, the application described above). (any of 237-251, 255, 480-490).
The event classifier 270 receives the event information, and determines an application 236 - 1 to which to forward the event information, and an application view 291 of the application 236 - 1 . The event sorter 270 includes an event monitor 271 and an event dispatcher module 274 . In some embodiments, application 236 - 1 includes application internal state 292 representing the current application view(s) displayed on touch-sensitive display 212 when the application is active or running. In some embodiments, device/global internal state 257 is used by event classifier 270 to determine which application(s) is currently active, and application internal state 292 is passed to event classifier 270 . It is used to determine the application views 291 to which the event information will be forwarded.
In some embodiments, application internal state 292 is a user indicating resume information to be used when application 236 - 1 resumes execution, information being displayed by application 236 - 1 or ready to be displayed. One of interface state information, a state queue to enable the user to return to a previous state or view of the application 236 - 1 , and a redo/undo queue of previous actions taken by the user. Includes additional information as above.
The event monitor 271 receives event information from the peripheral interface 218 . The event information includes information about a subevent (eg, a user touch on the touch-sensitive display 212 as part of a multi-touch gesture). Peripheral interface 218 receives from I/O subsystem 206 or sensors, such as proximity sensor 266 , accelerometer(s) 268 , and/or microphone 213 (via audio circuitry 210 ). send information to Information that peripherals interface 218 receives from I/O subsystem 206 includes information from touch-sensitive display 212 or a touch-sensitive surface.
In some embodiments, event monitor 271 sends requests to peripherals interface 218 at predetermined intervals. In response, the peripherals interface 218 transmits event information. In other embodiments, peripherals interface 218 transmits event information only when there is a significant event (eg, receiving input exceeding a predetermined noise threshold and/or input for exceeding a predetermined duration). .
In some embodiments, the event classifier 270 also includes a hit view determination module 272 and/or an active event recognizer determination module 273 .
The hit view determination module 272 provides software procedures for determining where a subevent occurred within the one or more views when the touch-sensitive display 212 displays more than one view. Views consist of controls and other elements that the user can see on the display.
Another aspect of a user interface associated with an application is a set of views, sometimes referred to herein as application views or user interface windows, in which information is displayed and touch-based gestures occur. The application views (of an individual application) for which a touch is detected may correspond to program levels in the application's program or view hierarchy. For example, the lowest-level view in which a touch is detected may be referred to as a hit view, and the set of events recognized as appropriate inputs may be determined based, at least in part, on the hit view of the initial touch that initiates the touch-based gesture.
The hit view determination module 272 receives information related to sub-events of the touch-based gesture. If the application has multiple views organized in a hierarchy, hit view determination module 272 identifies the hit view as the lowest view in the hierarchy that should handle the subevent. In most situations, the hit view is the lowest level view in which the initiated subevent occurs (eg, the first subevent in the sequence of subevents forming the event or potential event). Once a hit view is identified by the hit view determination module 272 , the hit view typically receives all subevents related to the same touch or input source that caused it to be identified as the hit view.
The active event recognizer determination module 273 determines which view or views within the view hierarchy should receive a particular sequence of subevents. In some embodiments, the active event recognizer determination module 273 determines that only the hit view should receive a particular sequence of subevents. In other embodiments, the active event recognizer determination module 273 determines that all views including the physical location of the subevent are actively involved views, so that all actively involved views are specific to the subevents. It decides that it should receive the sequence. In other embodiments, even if the touch subevents are exclusively confined to an area associated with one particular view, views higher in the hierarchy will still remain as actively engaged views.
The event dispatcher module 274 dispatches event information to an event recognizer (eg, event recognizer 280 ). In embodiments that include an active event recognizer determination module 273 , the event dispatcher module 274 communicates event information to the event recognizer determined by the active event recognizer determination module 273 . In some embodiments, event dispatcher module 274 stores event information in an event queue, which event information is fetched by respective event receiver 282 .
In some embodiments, operating system 226 includes event classifier 270 . Alternatively, application 236 - 1 includes event classifier 270 . In still other embodiments, event classifier 270 is a standalone module or is part of another module stored in memory 202 , such as contact/motion module 230 .
In some embodiments, application 236 - 1 includes a plurality of event handlers 290 and one or more application views 291 , each of which handles touch events occurring within a respective view of the application's user interface. Contains instructions for processing. Each application view 291 of application 236 - 1 includes one or more event recognizers 280 . Typically, each application view 291 includes a plurality of event recognizers 280 . In other embodiments, one or more of event recognizers 280 may be part of a separate module, such as a user interface kit (not shown) or a higher-level object from which application 236 - 1 inherits methods and other properties. am. In some embodiments, the respective event handler 290 is one of the data updater 276 , the object updater 277 , the GUI updater 278 , and/or the event data 279 received from the event classifier 270 . includes more than Event handler 290 may update application internal state 292 using or calling data updater 276 , object updater 277 , or GUI updater 278 . Alternatively, one or more of the application views 291 include one or more respective event handlers 290 . Also, in some embodiments, one or more of data updater 276 , object updater 277 , and GUI updater 278 are included within respective application view 291 .
The respective event recognizer 280 receives event information (eg, event data 279 ) from the event classifier 270 and identifies the event from the event information. The event recognizer 280 includes an event receiver 282 and an event comparator 284 . In some embodiments, event recognizer 280 also includes at least a subset of metadata 283 and event delivery instructions 288 (which may include subevent delivery instructions).
The event receiver 282 receives event information from the event classifier 270 . The event information includes information about a sub-event, for example, a touch or a touch movement. Depending on the sub-event, the event information also includes additional information such as the location of the sub-event. When the sub-event relates to the motion of the touch, the event information may also include the speed and direction of the sub-event. In some embodiments, the events include rotation of the device from one orientation to another (eg, from a portrait orientation to a landscape orientation, or vice versa), and the event information includes the current orientation of the device (also referred to as device orientation). ) and corresponding information about
The event comparator 284 compares the event information to predefined event or subevent definitions, and based on the comparison, determines an event or subevent, or determines or updates the status of the event or subevent. In some embodiments, event comparator 284 includes event definitions 286 . Event definitions 286 include definitions of events (eg, predefined sequences of sub-events), eg, Event 1 287 - 1 , Event 2 287 - 2 , and the like. In some embodiments, sub-events within event 287 include, for example, touch start, touch end, touch move, touch cancel, and multi-touch. In one example, the definition for event 1 287 - 1 is a double tap on the displayed object. A double tap is, for example, a first touch (touch start) on a displayed object during a predetermined phase, a first liftoff (touch end) during a predetermined phase, a displayed object during a predetermined phase a second touch on the upper body (touch start), and a second liftoff (touch end) for a predetermined phase. In another example, the definition for event 2 287 - 2 is dragging on the displayed object. Dragging includes, for example, a touch (or contact) on the displayed object for a predetermined phase, movement of the touch across the touch-sensitive display 212 , and liftoff of the touch (touch termination). In some embodiments, the event also includes information about one or more associated event handlers 290 .
In some embodiments, event definition 287 includes a definition of an event for an individual user interface object. In some embodiments, event comparator 284 performs a hit test to determine which user interface object is associated with a subevent. For example, in an application view in which three user interface objects are displayed on the touch-sensitive display 212 , when a touch is detected on the touch-sensitive display 212 , the event comparator 284 displays the three user interface objects A hit test is performed to determine which one is associated with a touch (subevent). When each displayed object is associated with a respective event handler 290, the event comparator uses the result of the hit test to determine which event handler 290 should be activated. For example, event comparator 284 selects an event handler associated with an object and subevent that triggers the hit test.
In some embodiments, the definition of an individual event 287 also includes deferred actions of delaying the delivery of event information until after it is determined whether the sequence of subevents corresponds or does not correspond to an event type of the event recognizer. .
If the individual event recognizer 280 determines that the set of sub-events does not match any of the events in the event definitions 286 , the individual event recognizer 280 determines that the event is impossible, event failed, or event ended. A state is entered, after which the individual event recognizer ignores subsequent subevents of the touch-based gesture. In this situation, other event recognizers that remain active for the hit view, if any, continue to track and process sub-events of the touch-based gesture in progress.
In some embodiments, individual event recognizer 280 sets configurable attributes, flags, and/or lists that indicate how the event delivery system should perform subevent delivery to actively participating event recognizers. and metadata 283 with In some embodiments, metadata 283 includes configurable properties, flags, and/or lists that indicate how event recognizers can or are enabled to interact with each other. In some embodiments, metadata 283 includes configurable properties, flags, and/or lists that indicate whether subevents are delivered to various levels in the view or program hierarchy.
In some embodiments, the respective event recognizer 280 activates an event handler 290 associated with the event when one or more specific sub-events of the event are recognized. In some embodiments, individual event recognizer 280 passes event information associated with the event to event handler 290 . Activating event handler 290 is separate from sending (and delayed sending) subevents to individual hit views. In some embodiments, event recognizer 280 sends a flag associated with the recognized event, and event handler 290 associated with the flag catches the flag and performs a predefined process.
In some embodiments, event delivery instructions 288 include subevent delivery instructions that deliver event information related to a subevent without activating an event handler. Instead, subevent delivery instructions pass event information to event handlers associated with a series of subevents or to actively engaged views. Event handlers associated with a series of subevents or actively engaged views receive event information and perform a predetermined process.
In some embodiments, data updater 276 creates and updates data used in application 236 - 1 . For example, the data updater 276 updates a phone number used in the contacts module 237 or stores a video file used in the video player module. In some embodiments, object updater 277 creates and updates objects used in application 236 - 1 . For example, the object updater 277 creates a new user interface object or updates the location of the user interface object. GUI updater 278 updates the GUI. For example, GUI updater 278 prepares display information for display on a touch-sensitive display and sends it to graphics module 232 .
In some embodiments, event handler(s) 290 includes or accesses data updater 276 , object updater 277 , and GUI updater 278 . In some embodiments, data updater 276 , object updater 277 , and GUI updater 278 are included within a single module of individual application 236 - 1 or application view 291 . In other embodiments, they are included within two or more software modules.
The discussion above regarding event handling of user touches on touch-sensitive displays also applies to other forms of user inputs for operating multifunction devices 200 having input devices, but all of which are disclosed on touch screens. You have to understand that it's not going to happen. mouse movement and mouse button presses optionally coordinated with, for example, single or multiple keyboard presses or holds; contact movements, such as tap, drag, scroll, etc., on the touchpad; pen stylus inputs; movement of the device; verbal commands; detected eye movements; biometric inputs; and/or any combination thereof is optionally used as inputs corresponding to subevents defining the event to be recognized.
3 shows a portable multifunction device 200 with a touch screen 212 in accordance with some embodiments. The touch screen, optionally, displays one or more graphics within user interface (UI) 300 . In this embodiment, as well as other embodiments described below, the user may use, for example, one or more fingers 302 (not drawn to scale in the figures) or one or more styluses 303 (shown to scale in the figures). not)) to make it possible to select one or more of the graphics by making a gesture on the graphics. In some embodiments, selection of the one or more graphics occurs when the user ceases contacting the one or more graphics. In some embodiments, the gesture may optionally include one or more taps of a finger in contact with device 200 , one or more swipes (left to right, right to left, up and/or down), and/or (right) to left, left to right, up and/or down) rolling. In some implementations or situations, unintentional contact with the graphic does not select the graphic. For example, when the gesture corresponding to the selection is a tap, the swipe gesture to sweep over the application icon optionally does not select the corresponding application.
Device 200 may also include one or more physical buttons, such as a "home" or menu button 304 . As noted above, the menu button 304 may be used to navigate to any application 236 within the set of applications that may be running on the device 200 . Alternatively, in some embodiments, the menu button is implemented as a soft key in the GUI displayed on the touch screen 212 .
In one embodiment, the device 200 includes a touch screen 212 , a menu button 304 , a push button 306 to power on/off the device and lock the device, a volume control button(s) 308 , a subscriber identity module (SIM) card slot 310 , a headset jack 312 , and a docking/charging external port 224 . push button 306 selectively powers on/off the device by pressing the button and holding the button depressed for a predefined time interval; lock the device by depressing the button and depressing the button before a predefined time interval elapses; and/or to unlock the device or initiate an unlocking process. In an alternative embodiment, device 200 also accepts verbal input for activation or deactivation of some functions via microphone 213 . Device 200 may also optionally be configured to generate tactile outputs for a user of device 200 and/or one or more contact intensity sensors 265 for detecting intensity of contacts on touch screen 212 . one or more tactile output generators 267 .
4 is a block diagram of an exemplary multifunction device having a display and a touch-sensitive surface, in accordance with some embodiments. Device 400 need not be portable. In some embodiments, device 400 is a laptop computer, desktop computer, tablet computer, multimedia player device, navigation device, educational device (such as a children's learning toy), gaming system, or control device (eg, home or industrial). controller). Device 400 typically includes one or more processing units (CPUs) 410 , one or more network or other communication interfaces 460 , memory 470 , and one or more communication buses 420 for interconnecting these components. includes Communication buses 420 optionally include circuitry (sometimes referred to as a chipset) that interconnects system components and controls communication between them. Device 400 includes an input/output (I/O) interface 430 that includes a display 440 that is typically a touch screen display. I/O interface 430 may also optionally include a keyboard and/or mouse (or other pointing device) 450 and touch pad 455 , tactile for generating tactile outputs on device 400 . Output generator 457 (eg, similar to tactile output generator(s) 267 described above with reference to FIG. 2A ), and sensors 459 (eg, the contact intensity sensor described above with reference to FIG. 2A ) light sensors, acceleration sensors, proximity sensors, touch-sensitive sensors, and/or contact intensity sensors similar to s) 265 . memory 470 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid state memory devices; optionally including non-volatile memory such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid state storage devices. Memory 470 optionally includes one or more storage devices located elsewhere from CPU(s) 410 . In some embodiments, memory 470 includes programs, modules, and data structures similar to programs, modules, and data structures stored in memory 202 of portable multifunction device 200 ( FIG. 2A ). , or a subset thereof. Also, memory 470 optionally stores additional programs, modules, and data structures that are not present in memory 202 of portable multifunction device 200 . For example, the memory 470 of the device 400 may optionally include a drawing module 480 , a presentation module 482 , a word processing module 484 , a website creation module 486 , a disk authoring module ( 488 ), and/or spreadsheet module 490 , while memory 202 of portable multifunction device 200 ( FIG. 2A ), optionally, does not store these modules.
Each of the previously identified elements in FIG. 4 may be stored in one or more of the previously mentioned memory devices. Each of the modules identified above corresponds to a set of instructions for performing the functions described above. The modules or programs (eg, sets of instructions) identified above need not be implemented as separate software programs, procedures, or modules, thus, various subsets of these modules may be combined in various embodiments or otherwise can be rearranged in this way. In some embodiments, memory 470 may store a subset of the modules and data structures identified above. Also, memory 470 may store additional modules and data structures not described above.
We now look at embodiments of user interfaces that may be implemented, for example, on portable multifunction device 200 .
5A illustrates an example user interface for a menu of applications on portable multifunction device 200 , in accordance with some embodiments. Similar user interfaces may be implemented on device 400 . In some embodiments, user interface 500 includes the following elements, or a subset or superset thereof:
signal strength indicator(s) 502 for wireless communication(s), such as cellular and Wi-Fi signals;
<img file="KR20180098409A_D0020.tif" />time 504;
<img file="KR20180098409A_D0021.tif" />Bluetooth indicator 505;
<img file="KR20180098409A_D0022.tif" />battery status indicator 506 ;
<img file="KR20180098409A_D0023.tif" />Tray 508 with icons for frequently used applications, such as:
o an icon 516 for the phone module 238 labeled "Calls", optionally including an indicator 514 of the number of missed calls or voicemail messages;
o an icon 518 for the email client module 240 labeled "Mail" optionally including an indicator 510 of the number of unread emails;
o icon 520 for browser module 247 labeled "Browser"; and
o Icon 522 for video and music player module 252 labeled "iPod", also referred to as iPod (trademark of Apple Inc.) module 252; and
<img file="KR20180098409A_D0024.tif" />Icons for different applications, such as:
o icon 524 for IM module 241 labeled "Message";
o icon 526 for calendar module 248 labeled "Calendar";
o icon 528 for image management module 244 labeled "Photos";
o icon 530 for camera module 243 labeled "camera";
o icon 532 for online video module 255 labeled "Online Video";
o icon 534 for stock widget 249-2 labeled "Stocks";
o icon 536 for map module 254 labeled "Map";
o icon 538 for weather widget 249-1 labeled "weather";
o icon 540 for alarm clock widget 249-4 labeled "Clock";
o icon 542 for exercise support module 242 labeled "Workout Support";
o icon 544 for notes module 253 labeled "Notes"; and
o Icon 546 for a settings application or module, labeled "Settings," that provides access to settings for device 200 and its various applications 236 .
It should be noted that the icon labels shown in FIG. 5A are merely exemplary. For example, icon 522 for video and music player module 252 may optionally be labeled "Music" or "Music Player." Other labels are optionally used for various application icons . In some embodiments, the label for the respective application icon includes the name of the application corresponding to the respective application icon. In some embodiments, the label for a particular application icon is separate from the name of the application that corresponds to the particular application icon.
5B illustrates a device (eg, device 400 ) having a touch-sensitive surface 551 (eg, tablet or touchpad 455 , FIG. 4 ) separate from display 550 (eg, touch screen display 212 ). ), which shows an exemplary user interface on FIG. 4). Device 400 may also optionally include one or more contact intensity sensors (eg, one or more of sensors 457 ) for detecting the intensity of contacts on touch-sensitive surface 551 and/or of device 400 . one or more tactile output generators 459 for generating tactile outputs for the user.
Although some examples that follow will be provided with reference to inputs on a touch screen display 212 (a touch-sensitive surface and display combined), in some embodiments, the device is separate from the display as shown in FIG. 5B . Detect inputs on the touch-sensitive surface. In some embodiments, the touch-sensitive surface (eg, 551 in FIG. 5B ) has a major axis (eg, 552 in FIG. 5B ) that corresponds to a major axis (eg, 553 in FIG. 5B ) on the display (eg, 550 ). According to these embodiments, the device communicates with the touch-sensitive surface 551 at locations corresponding to respective locations on the display (eg, in FIG. 5B , 560 corresponds to 568 and 562 corresponds to 570 ). Detect contacts (eg, 560 and 562 in FIG. 5B ). In this way, user inputs (eg, contacts 560 , 562 and their movements) detected by the device on the touch-sensitive surface (eg, 551 in FIG. 5B ) will not affect the touch-sensitive surface (eg, 551 in FIG. 5B ) if the touch-sensitive surface is separate from the display. used by the device to manipulate the user interface on the multifunction device's display (eg, 550 in FIG. 5B ). It should be understood that similar methods are, optionally, used with other user interfaces described herein.
Additionally, while the following examples are given primarily with reference to finger inputs (eg, finger contacts, finger tap gestures, and/or finger swipe gestures), in some embodiments, one or more of the finger inputs It should be understood that is replaced with input from another input device (eg, mouse-based input or stylus input). For example, a swipe gesture is optionally replaced with a mouse click (eg, instead of a contact) followed by movement of the cursor along the path of the swipe (eg, instead of movement of the contact). As another example, the tap gesture is optionally replaced with a mouse click while the cursor is positioned over the location of the tap gesture (eg, instead of detecting the contact and detecting the subsequent contact). Similarly, when multiple user inputs are detected simultaneously, it should be understood that multiple computer mice are selectively used simultaneously, or mouse and finger contacts are selectively used concurrently.
6A shows an example personal electronic device 600 . Device 600 includes a body 602 . In some embodiments, device 600 may include some or all of the features described with respect to devices 200 , 400 (eg, FIGS. 2A-4B ). In some embodiments, device 600 has a touch-sensitive display screen 604 (hereafter, touch screen 604 ). Alternatively or in addition to the touch screen 604 , the device 600 has a display and a touch-sensitive surface. As with devices 200 and 400 , in some embodiments, touch screen 604 (or touch-sensitive surface) has one or more intensity sensors for detecting the intensity of applied contacts (eg, touches). can have One or more intensity sensors of touch screen 604 (or touch-sensitive surface) may provide output data representing the intensity of touches. The user interface of device 600 may respond to touches based on the intensity of the touches, meaning that touches of different intensities may invoke different user interface actions on device 600 .
Techniques for detecting and processing touch intensity are described, for example, in Related Applications: filed May 8, 2013 and entitled "Device, Method, and Graphical User Interface for Displaying User Interface Objects Corresponding to an Application" "In International Patent Application No. PCT/US2013/040061, and International Patent Application filed on November 11, 2013 and entitled "Device, Method, and Graphical User Interface for Transitioning Between Touch Input to Display Output Relationships" PCT/US2013/069483, each of which is incorporated herein by reference in its entirety.
In some embodiments, device 600 has one or more input mechanisms 606 , 608 . Input mechanisms 606 and 608 (if included) may be physical. Examples of physical input mechanisms include push buttons and rotatable mechanisms. In some embodiments, device 600 has one or more attachment mechanisms. These attachment mechanisms (if included) allow device 600 to use, for example, a hat, eyeglasses, earrings, necklace, shirt, jacket, bracelet, watchband, chain, pants, belt, shoe, wallet. , so that it can be attached to a backpack, etc. These attachment mechanisms may allow device 600 to be worn by a user.
6B shows an example personal electronic device 600 . In some embodiments, device 600 may include some or all of the components described with respect to FIGS. 2A , 2B , and 4 . Device 600 has a bus 612 operatively coupling I/O section 614 with one or more computer processors 616 and memory 618 . The I/O section 614 can be coupled to a display 604 , which can have a touch-sensitive component 622 , and, optionally, a touch intensity-sensitive component 624 . Moreover, the I/O section 614 communicates with the communication unit 630 to receive application and operating system data using Wi-Fi, Bluetooth, Near Field Communication (NFC), cellular and/or other wireless communication techniques. can be connected Device 600 may include input mechanisms 606 and/or 608 . The input mechanism 606 may be, for example, a rotatable input device or a depressible and rotatable input device. The input mechanism 608 may, in some examples, be a button.
The input mechanism 608 may, in some examples, be a microphone. The personal electronic device 600 may include a GPS sensor 632 , an accelerometer 634 , a direction sensor 640 (eg, a compass), a gyroscope 636 , a motion sensor 638 , and/or combinations thereof. , various sensors, all of which may be operatively coupled to the I/O section 614 .
The memory 618 of the personal electronic device 600, for example, when executed by one or more computer processors 616, causes the computer processors to: It may be a non-transitory computer-readable storage medium for storing computer-executable instructions that may cause the techniques described below to be performed. Computer-executable instructions also include an instruction execution system, apparatus, or device, such as a computer-based system, processor-containing system, or other system capable of fetching instructions from an instruction execution system, apparatus, or device to execute the instructions. may be stored on and/or transmitted over in any non-transitory computer-readable storage medium for use by or in connection therewith. Personal electronic device 600 is not limited to the components and configuration of FIG. 6B , and may include other or additional components in a number of configurations.
As used herein, the term "affordance" refers to a user interaction graphic that may be displayed on a display screen of devices 200 , 400 , and/or 600 ( FIGS. 2 , 4 , and 6 ). Refers to a user-interactive graphical user interface object. For example, an image (eg, an icon), a button, and text (eg, a link) may each constitute an affordance.
As used herein, the term "focus selector" refers to an input element that represents the current portion of a user interface with which the user is interacting. In some implementations that include a cursor or other location marker, while the cursor is over a particular user interface element (eg, button, window, slider, or other user interface element), a touch-sensitive surface (eg, figure When an input (eg, a press input) is detected on the touch pad 455 of 4 or the touch-sensitive surface 551 of FIG. 5B ), the cursor becomes "focus function as a "selector". Some implementations include a touch screen display (eg, touch-sensitive display system 212 of FIG. 2A or touch screen 212 of FIG. 5A ) that enables direct interaction with user interface elements on the touch screen display. In some cases, when an input (eg, a touch input) is detected on a touch screen display at the location of a particular user interface element (eg, a button, window, slider, or other user interface element), the particular user interface element The detected contact on the touch screen serves as a "focus selector", such that is adjusted according to the detected input. In some implementations, focus is moved to one of the user interfaces without movement of a corresponding cursor or contact or movement of a corresponding cursor on the touch screen display (eg, by using a tab key or an arrow key to move focus from one button to another). moved from one area of the user interface to another area of the user interface; In such implementations, the focus selector moves in response to movement of focus between different areas of the user interface. Irrespective of the particular form the focus selector has, the focus selector is generally used to convey the user's intended interaction with the user interface (eg, by indicating to the device an element of the user interface with which the user wants to interact). A user interface element (or contact on a touch screen display) controlled by For example, the position of a focus selector (eg, cursor, contact, or selection box) over an individual button while a press input is detected on a touch-sensitive surface (eg, a touchpad or touch screen) (eg, on the device's display) It will indicate that the user is trying to activate an individual button (unlike other user interface elements shown).
As used in the specification and claims, the term "characteristic intensity" of a contact refers to a property of a contact based on one or more intensities of the contact. In some embodiments, the feature intensity is based on multiple intensity samples. The feature intensity is optionally determined for a predefined number of intensity samples, or for a predefined event (eg, after detecting the contact, before detecting liftoff of the contact, detecting the start of movement of the contact) A predetermined period of time (eg, 0.05) before or after, before detecting the end of the contact, before or after detecting an increase in the intensity of the contact, and/or before or after detecting a decrease in the intensity of the contact , 0.1, 0.2, 0.5, 1, 2, 5, 10 sec). The characteristic intensity of the contact is, optionally, a maximum value of the intensities of the contact, a mean value of the intensities of the contact, an average value of the intensities of the contact, a top 10 percentile of the intensities of the contact. value), a value of half the maximum value of the intensities of the contact, a value of 90 percent of the maximum value of the intensities of the contact, and the like. In some embodiments, the duration of the contact is used to determine the feature intensity (eg, when the feature intensity is an average of the intensity of the contact over time). In some embodiments, the characteristic intensity is compared to a set of one or more intensity thresholds to determine whether an action was performed by the user. For example, the set of one or more intensity thresholds may include a first intensity threshold and a second intensity threshold. In this example, as a result of a contact having a characteristic intensity that does not exceed a first threshold, a first action is performed, and as a result of a contact having a characteristic intensity that exceeds a first intensity threshold but does not exceed a second intensity threshold, a second An action is performed, and as a result of the contact having a characteristic intensity above the second threshold, a third action is performed. In some embodiments, a comparison between a characteristic strength and one or more thresholds is used to determine whether to perform one or more actions (eg, individual actions) rather than being used to determine whether to perform a first action or a second action. is used to decide whether to perform an action or whether to withhold performing an individual action.
In some embodiments, a portion of a gesture is identified to determine a characteristic strength. For example, a touch-sensitive surface may receive successive swipe contacts moving from a starting position to an ending position, at which point the intensity of the contact increases. In this example, the characteristic intensity of the contact at the end location may be based on only a portion of the successive swipe contacts (eg, only the portion of the swipe contact at the end location) and not the entire swipe contact. In some embodiments, a smoothing algorithm may be applied to the intensities of the swipe contact prior to determining the characteristic intensity of the contact. For example, the smoothing algorithm may optionally include an unweighted sliding-average smoothing algorithm, a triangular smoothing algorithm, a median filter smoothing algorithm, and/or an exponential smoothing algorithm. contains more than one. In some situations, these smoothing algorithms remove narrow spikes or dips in the intensities of the swipe contact to determine the characteristic intensity.
The intensity of the contact on the touch-sensitive surface may be characterized against one or more intensity thresholds, such as a touch-detection intensity threshold, a light press intensity threshold, a deep press intensity threshold, and/or one or more other intensity thresholds. In some embodiments, the light press intensity threshold corresponds to an intensity at which the device will perform operations typically associated with clicking a button or trackpad of a physical mouse. In some embodiments, the deep press intensity threshold corresponds to an intensity at which the device will perform operations that are different from those typically associated with clicking a button or trackpad of a physical mouse. In some embodiments, when a contact is detected with a characteristic intensity below a light press intensity threshold (eg, and above a nominal contact-detection intensity threshold, below which the contact is no longer detected), the device It will move the focus selector according to movement of the contact on the touch-sensitive surface without performing an action associated with the intensity threshold or the deep press intensity threshold. In general, unless stated otherwise, these intensity thresholds are consistent between different sets of user interface drawings.
An increase in the characteristic intensity of a contact from an intensity below the light press intensity threshold to an intensity between the light press intensity threshold and the deep press intensity threshold is sometimes referred to as a "light press" input. An increase in the characteristic intensity of a contact from an intensity below the deep press intensity threshold to an intensity above the deep press intensity threshold is sometimes referred to as a "deep press" input. An increase in the characteristic intensity of a contact from an intensity below the touch detection intensity threshold to an intensity between the touch detection intensity threshold and the light press intensity threshold is sometimes referred to as detecting a contact on a touch surface. A decrease in the characteristic intensity of a contact from an intensity above the contact detection intensity threshold to an intensity below the contact detection intensity threshold is sometimes referred to as detecting liftoff of the contact from the touch surface. In some embodiments, the contact-detection intensity threshold is zero. In some embodiments, the contact-detection intensity threshold is greater than zero.
In some embodiments described herein, the one or more actions include detecting an individual press input performed with an individual contact (or a plurality of contacts) or in response to detecting a gesture comprising the individual press input. wherein the respective press input is detected based, at least in part, on detecting an increase in intensity of the contact (or plurality of contacts) above a press-input intensity threshold. In some embodiments, the respective action is performed in response to detecting an increase in intensity of the respective contact above a press-input intensity threshold (eg, a "down stroke" of the respective press input). . In some embodiments, the press input comprises an increase in the intensity of the respective contact above the press-input intensity threshold and a subsequent decrease in the intensity of the contact below the press-input intensity threshold, wherein the respective action comprises the press-input intensity threshold is performed in response to detecting a subsequent decrease in the intensity of the individual contact less than (eg, an "up stroke" of the individual press input).
In some embodiments, the device uses intensity hysteresis to avoid sudden inputs, sometimes referred to as "jitter", where the device has a hysteresis intensity threshold ( For example, the hysteresis intensity threshold is a unit of X intensity lower than the press-input intensity threshold, or the hysteresis intensity threshold is 75%, 90%, or some suitable percentage of the press-input intensity threshold). As such, in some embodiments, the press input comprises an increase in the intensity of the individual contact above the press-input intensity threshold and a subsequent decrease in the intensity of the contact below the hysteresis intensity threshold corresponding to the press-input intensity threshold, , the respective action is performed in response to detecting a subsequent decrease in the intensity of the respective contact below the hysteresis intensity threshold (eg, an "up stroke" of the respective press input). Similarly, in some embodiments, the press input causes the device to increase the intensity of the contact from an intensity below the hysteresis intensity threshold to an intensity above the press-input intensity threshold, and optionally, an intensity below the hysteresis intensity threshold. detected only if a subsequent decrease in the intensity of the contact of
For convenience of explanation, descriptions of actions performed in response to a press input associated with a press-input intensity threshold or in response to a gesture comprising a press input include, optionally, an increase in the intensity of the contact above the press-input intensity threshold; an increase in the intensity of the contact from an intensity below the hysteresis intensity threshold to an intensity above the press-input intensity threshold, a decrease in the intensity of the contact below the press-input intensity threshold, and/or a hysteresis intensity threshold corresponding to the press-input intensity threshold Triggered in response to detecting any one of a decrease in the intensity of the contact below. Further, in examples in which the action is described as being performed in response to detecting a decrease in the intensity of the contact below the press-input intensity threshold, the action is, optionally, a hysteresis intensity corresponding to and lower than the press-input intensity threshold. performed in response to detecting a decrease in the intensity of the contact below a threshold.
3. digital assistant system
7A shows a block diagram of a digital assistant system 700 according to various examples. In some examples, digital assistant system 700 may be implemented on a standalone computer system. In some examples, digital assistant system 700 may be distributed across multiple computers. In some examples, some of the modules and functions of the digital assistant may be divided into a server portion and a client portion, wherein the client portion is one or more user devices (eg, devices 104 , 122 , 200 , 400 , or 600 ). ) and communicates with a server portion (eg, server system 108 ) via one or more networks, eg, as shown in FIG. 1 . In some examples, digital assistant system 700 may be an implementation of server system 108 (and/or DA server 106 ) shown in FIG. 1 . that digital assistant system 700 is only one example of a digital assistant system, and that digital assistant system 700 may have more or fewer components than shown, or may combine two or more components; It should be noted that they may have components in different configurations or arrangements. The various components shown in FIG. 7A may be implemented in hardware, including one or more signal processing and/or application specific integrated circuits, software instructions for execution by one or more processors, firmware, or a combination thereof.
The digital assistant system 700 may include a memory 702 , one or more processors 704 , an input/output (I/O) interface 706 , and a network communication interface 708 . These components may communicate with each other via one or more communication buses or signal lines 710 .
In some examples, memory 702 may be a non-transitory computer-readable medium, such as high-speed random access memory and/or a non-volatile computer-readable storage medium (eg, one or more magnetic disk storage devices, flash memory devices, or other non-volatile solid state media). state memory device).
In some examples, I/O interface 706 connects input/output devices 716 of digital assistant system 700 , such as displays, keyboards, touch screens, and microphones, to user interface module 722 . can be combined I/O interface 706 , in conjunction with user interface module 722 , can receive user inputs (eg, voice input, keyboard inputs, touch inputs, etc.) and process them accordingly. In some examples, for example, where the digital assistant is implemented on a standalone user device, the digital assistant system 700 is a component described for devices 200 , 400 , or 600 of FIGS. 2A , 4 , 6A and 6B , respectively. and I/O communication interfaces. In some examples, digital assistant system 700 can represent a server portion of a digital assistant implementation, and a user via a client-side portion residing on a user device (eg, devices 104 , 200 , 400 , or 600 ). can interact with
In some examples, network communication interface 708 can include wired communication port(s) 712 and/or wireless transmit and receive circuitry 714 . Wired communication port(s) 712 may receive and transmit communication signals over one or more wired interfaces, such as Ethernet, Universal Serial Bus (USB), FireWire, and the like. The wireless circuitry 714 may receive and transmit RF signals and/or optical signals to/from communication networks and other communication devices. The wireless communication may utilize any of a plurality of communication standards, protocols, and technologies, such as GSM, EDGE, CDMA, TDMA, Bluetooth, Wi-Fi, VoIP, Wi-MAX, or any other suitable communication protocol. have. Network communication interface 708 may be connected to digital assistant system 700 and others using networks such as the Internet, intranets, and/or wireless networks, such as cellular telephone networks, wireless local area networks (LANs), and/or metropolitan area networks (MANs). It may enable communication between devices.
In some examples, memory 702 or computer-readable storage media of memory 702 includes operating system 718 , communication module 720 , user interface module 722 , one or more applications 724 , and digital assistant It may store programs, modules, instructions, and data structures, including all or a subset of module 726 . In particular, the memory 702 or a computer-readable storage medium of the memory 702 may store instructions for performing the process 1200 described below. One or more processors 704 may execute such programs, modules, and instructions, and may read/write data structures to/from data structures.
Operating system 718 (eg, an embedded operating system such as Darwin, RTXC, LINUX, UNIX, iOS, OS X, WINDOWS, or VxWorks) performs general system tasks (eg, memory management, storage It may include various software components and/or drivers for controlling and managing device control, power management, etc.), and may facilitate communications between various hardware, firmware, and software components.
Communication module 720 may facilitate communication between digital assistant system 700 and other devices via network communication interface 708 . For example, communication module 720 can communicate with RF circuitry 208 of electronic devices, such as devices 200 , 400 , and 600 shown in FIGS. 2A , 4 , 6A and 6B , respectively. Communication module 720 may also include various components for processing data received by wireless circuitry 714 and/or wired communication port 712 .
The user interface module 722 receives commands and/or inputs from a user (eg, from a keyboard, touch screen, pointing device, controller, and/or microphone) via the I/O interface 706 and is displayed on the display. You can create user interface objects. User interface module 722 also prepares outputs (eg, speech, sound, animation, text, icons, vibrations, haptic feedback, lighting, etc.) and via I/O interface 706 (eg, display audio channels, audio channels, speakers, and touch pads, etc.) to the user.
Applications 724 may include programs and/or modules configured to be executed by one or more processors 704 . For example, if the digital assistant system is implemented on a standalone user device, applications 724 may include user applications such as games, calendar applications, navigation applications, or email applications. When digital assistant system 700 is implemented on a server, applications 724 may include, for example, resource management applications, diagnostic applications, or scheduling applications.
Memory 702 may also store digital assistant module 726 (or a server portion of digital assistant). In some examples, digital assistant module 726 may include the following submodules, or a subset or superset thereof: input/output processing module 728 , speech-to-text (STT) ) processing module 730 , natural language processing module 732 , dialog flow processing module 734 , task flow processing module 736 , service processing module 738 , and speech synthesis module 740 . Each of these modules may have access to one or more of the following systems or data and models of digital assistant module 726, or a subset or superset thereof: ontology 760, lexical index 744, user. Data 748 , task flow models 754 , service models 756 , and ASR system 731 .
In some examples, using models, data, and processing modules implemented in digital assistant module 726 , the digital assistant may perform at least some of the following: convert speech input to text; identifying the user's intent expressed in natural language input received from the user; actively eliciting and obtaining the information necessary to fully infer the user's intentions (eg, by disambiguating words, games, intentions, etc.); determining a task flow for implementing the inferred intent; and executing the task flow to fulfill the inferred intent.
In some examples, as shown in FIG. 7B , the I/O processing module 728 interacts with a user via the I/O devices 716 of FIG. 7A or communicates with the network communication interface 708 of FIG. 7A . interact with a user device (eg, devices 104 , 200 , 400 , or 600 ) via can do. The I/O processing module 728 may optionally obtain context information associated with the user input from the user device upon receipt of the user input or immediately after receipt of the user input. Context information may include user-specific data, vocabulary, and/or preferences related to user input. In some examples, the context information also includes information related to the software and hardware state of the user device at the time the user request is received and/or the user's surrounding environment at the time the user request is received. In some examples, I/O processing module 728 may also send follow-up questions to the user regarding the user request and receive answers therefrom. When the user request is received by the I/O processing module 728 and the user request may include speech input, the I/O processing module 728 converts the speech input to the STT processing module 730 for speech-to-text conversion. ) (or a speech recognizer).
The STT processing module 730 may include one or more ASR systems. One or more ASR systems may process speech input received via I/O processing module 728 to generate recognition results. Each ASR system may include a front-end speech pre-processor. The front-end speech preprocessor may extract representative features from the speech input. For example, a front-end speech preprocessor may perform a Fourier transform on the speech input to extract spectral features characterizing the speech input as a sequence of representative multidimensional vectors. Additionally, each ASR system may include one or more speech recognition models (eg, acoustic models and/or language models) and may implement one or more speech recognition engines. Examples of speech recognition models may include hidden Markov models, Gaussian-Mixture Models, Deep Neural Network Models, n-gram language models, and other statistical models. . Examples of speech recognition engines may include a dynamic time warp based engine and a weighted finite state transformer (WFST) based engine. The one or more speech recognition models and the one or more speech recognition engines provide intermediate recognition results (eg, phonemes, phonemic strings, and subwords), and ultimately text recognition results (eg, of words, word strings, or tokens). sequence) can be used to process the extracted representative features of the front-end speech preprocessor. In some examples, the speech input may be processed at least in part by a third-party service or on the user's device (eg, device 104 , 200 , 400 , or 600 ) to generate a recognition result. When the STT processing module 730 generates a recognition result including a text string (eg, words, or a sequence of words, or a sequence of tokens), the recognition result is sent to the natural language processing module 732 for intent inference can be transmitted to
Further details of speech-text processing are set forth in U.S. Utility Model Application No. 13/236,942 for "Consolidating Speech Recognition Results," filed September 20, 2011, the entire disclosure of which is herein incorporated by reference. incorporated by reference.
In some examples, the STT processing module 730 may include and/or access a vocabulary of words recognizable via the phonetic sign conversion module 731 . Each lexical word may be associated with one or more candidate pronunciations of the word represented by the speech recognition phonetic signature. In particular, the vocabulary of recognizable words may include words associated with a plurality of candidate pronunciations. For example, the vocabulary is /<img file="KR20180098409A_D0025.tif" />/ and /<img file="KR20180098409A_D0026.tif" />It may include the word "tomato" associated with the candidate pronunciation of /. Vocabulary words may also be associated with customized candidate pronunciations based on previous speech input from the user. These customized candidate pronunciations may be stored in the STT processing module 730 and may be associated with a specific user through the user's profile on the device. In some examples, a candidate pronunciation for a word may be determined based on the spelling of the word and one or more linguistic and/or phonetic rules. In some examples, candidate pronunciations may be generated manually, for example, based on known canonical pronunciations.
In some examples, the candidate pronunciation may be ranked based on the commonality of the candidate pronunciation. For example, the candidate pronunciation /<img file="KR20180098409A_D0027.tif" />/Is /<img file="KR20180098409A_D0028.tif" />It may rank higher than /, since the previous pronunciation is the more commonly used pronunciation (e.g., among all users, for users in a particular geographic area, or any other suitable subset of users). In the case of). In some examples, the candidate pronunciation may be ranked based on whether the candidate pronunciation is a customized candidate pronunciation associated with the user. For example, the customized candidate pronunciation may be ranked higher than the canonical candidate pronunciation. This can be useful for recognizing proper nouns with distinct pronunciations that deviate from canonical pronunciation. In some examples, the candidate pronunciation may be associated with one or more speech characteristics, such as geographic origin, nationality, or ethnicity. For example, the candidate pronunciation /<img file="KR20180098409A_D0029.tif" />/ can be associated with the United States, while the candidate pronunciation /<img file="KR20180098409A_D0030.tif" />/ can be associated with the UK. Further, the ranking of candidate pronunciations may be based on one or more characteristics of the user (eg, geographic origin, nationality, ethnicity, etc.) stored in a user profile on the device. For example, it may be determined from the user's profile that the user is associated with the United States. Candidate pronunciations based on users associated with the United States /<img file="KR20180098409A_D0031.tif" />/ (associated with the United States) is a candidate pronunciation /<img file="KR20180098409A_D0032.tif" />May be ranked higher than / (associated with UK). In some examples, one of the ranked candidate pronunciations may be selected as the predicted pronunciation (eg, the most likely pronunciation).
When speech input is received, STT processing module 730 may be used (eg, using an acoustic model) to determine a phoneme corresponding to the speech input, and then (eg, using a language model). ) may attempt to determine which word matches the phoneme. For example, the STT processing module 730 may generate a sequence/sequence of phonemes corresponding to a portion of speech input.<img file="KR20180098409A_D0033.tif" />If / can first be identified, it can then determine based on lexical index 744 that this sequence corresponds to the word "tomato".
In some examples, the STT processing module 730 can determine words in the speech input using an approximate matching technique. Thus, for example, the STT processing module 730 may determine that a sequence of phonemes /<img file="KR20180098409A_D0034.tif" />It can be determined that / corresponds to the word "tomato".
The digital assistant's natural language processing module 732 ("Natural Language Processor") takes the sequence of words or tokens generated by the STT processing module 730 ("token sequence") and converts the token sequence by the digital assistant It may attempt to associate it with one or more "actionable intentions" that are recognized. An "actionable intent" may represent a task that may be performed by a digital assistant, and may have an associated task flow implemented in task flow models 754 . An associative task flow may be a series of programmed actions and steps that the digital assistant takes to perform a task. The scope of the digital assistant's capabilities may depend on the number and type of task flows implemented and stored in task flow models 754 , or in other words, the number and type of "actionable intents" that the digital assistant is aware of. However, the effectiveness of a digital assistant may also depend on the assistant's ability to infer precise "actionable intent(s)" from user requests expressed in natural language.
In some examples, in addition to the words or token sequence obtained from the STT processing module 730 , the natural language processing module 732 may also provide context information associated with the user request, eg, from the I/O processing module 728 . can receive The natural language processing module 732 may optionally use context information to disambiguate, supplement, and/or further define information included in the token sequence received from the STT processing module 730 . Context information may include, for example, user preferences, hardware and/or software states of the user device, sensor information collected before, during, or immediately after a user request, prior interactions between the digital assistant and the user (e.g., , conversation) and the like. As described herein, contextual information may be dynamic and may change with time, location, content of the conversation, and other factors.
In some examples, natural language processing may be based, for example, on ontology 760 . The ontology 760 may be a hierarchical structure including many nodes, each node having a "property" related to one or more of "actionable intent" or "actionable intent", or other Represents any one of "attributes". As noted above, "actionable intent" may represent a task that a digital assistant may perform, ie, a task that it may be "actionable to" or affect. An "attribute" may represent a parameter that is associated with an actionable intent or sub-aspect of another attribute. The linkage between the actionable intent node and the attribute node in the ontology 760 may define how the parameter represented by the attribute node relates to the task represented by the actionable intent node.
In some examples, ontology 760 may be composed of actionable intent nodes and attribute nodes. Within the ontology 760 , each actionable intent node may be connected to one or more attribute nodes either directly or via one or more intermediate attribute nodes. Similarly, each attribute node may be connected to one or more actionable intent nodes either directly or via one or more intermediate attribute nodes. For example, as shown in FIG. 7C , the ontology 760 may include a "restaurant reservation" node (ie, an actionable intent node). The attribute nodes "Restaurant", "Date/Time" (for reservation), and "People" may each be directly linked to an actionable intent node (ie, a "Restaurant Reservation" node).
Additionally, the attribute nodes "Cooking", "Price Range", "Phone Number", and "Location" may be subnodes of the attribute node "Restaurant", and via the intermediate attribute node "Restaurant", the "Restaurant Reservation" node ( That is, each can be linked to an actionable intent node). For another example, as shown in FIG. 7C , the ontology 760 may also include a "set reminder" node (ie, another actionable intent node). The attribute nodes "Date/Time" (for setting a reminder) and "Subject" (for a reminder) may be linked to the "Set reminder" node, respectively. Since the attribute "Date/Time" may relate to both the task of making a restaurant reservation and the task of setting a reminder, the attribute node "Date/Time" is a "Restaurant Reservation" node and "Set Reminder" in the ontology 760 . It can be linked to both nodes.
An actionable intent node, along with its associated concept nodes, may be described as a "domain". In this discussion, each domain may be associated with an individual actionable intent and refers to a group of nodes (and relationships between them) associated with a particular actionable intent. For example, the ontology 760 illustrated in FIG. 7C may include an example of a restaurant reservation domain 762 and an example of a reminder domain 764 in the ontology 760 . The restaurant reservation domain includes the actionable intent node "Restaurant Reservation", the attribute nodes "Restaurant", "Date/Time", and "Number of People", and the sub-attribute nodes "Cooking", "Price Range", "Phone Number", and includes "location". Reminder domain 764 may include an actionable intent node "set reminder", and attribute nodes "subject" and "date/time". In some examples, ontology 760 may be composed of many domains. Each domain may share one or more attribute nodes with one or more other domains. For example, a "date/time" attribute node may be associated with many different domains (eg, a scheduling domain, a travel reservation domain, a movie ticket domain, etc.) in addition to the restaurant reservation domain 762 and the reminder domain 764 . .
Although FIG. 7C shows two example domains within ontology 760 , other domains are, for example, "find a movie", "initiate a phone call", "find a way", "schedule a meeting" , "Send a message", "Provide an answer to a question", "Read a list", "Provide navigation commands", "Provide commands for a task", and the like. A "send message" domain may be associated with a "send message" actionable intent node, and may further include attribute nodes such as "recipient(s)", "message type", and message body" The attribute node "recipient" may be further defined by sub-attribute nodes such as, for example, "recipient name" and "message address".
In some examples, ontology 760 may include all domains (and thus actionable intents) that the digital assistant can understand and act upon. In some examples, ontology 760 may be modified, such as by adding or removing entire domains or nodes, or by modifying relationships between nodes within ontology 760 .
In some examples, nodes associated with multiple related actionable intents may be clustered under a "parent domain" within ontology 760 . For example, a "travel" parent domain may contain a cluster of attribute nodes and actionable intent nodes related to travel. Actionable intent nodes related to travel may include "air reservation", "hotel reservation", "car rental", "direction finding", "finding point of interest", and the like. Actionable intent nodes under the same parent domain (eg, a "travel" parent domain) may have many attribute nodes in common. For example, the actionable intent nodes for "air reservation", "hotel reservation", "car rental", "direction", and "find point of interest" are the attribute nodes "start location", "destination", " One or more of "departure date/time", "arrival date/time", and "person number" may be shared.
In some examples, each node in ontology 760 may be associated with a set of words and/or phrases related to the attribute or actionable intent represented by the node. The respective set of words and/or phrases associated with each node may be a so-called "vocabulary" associated with the node. A respective set of words and/or phrases associated with each node may be stored in a lexical index 744 with respect to the attribute or actionable intent represented by the node. For example, returning to Figure 7b, the vocabulary associated with the node for the attribute of "restaurant" is "food", "drink", "cooking", "hungry", "eat", "pizza", "fast food", It may include words such as "meal" and the like. For another example, the vocabulary associated with the node for the actionable intent of "initiate a phone call" is "call", "call", "dial", "ring", "call this number", "call to", etc. It may contain the same words and phrases. Vocabulary index 744 may optionally include words and phrases in different languages.
The natural language processing module 732 may receive a token sequence (eg, a text string) from the STT processing module 730 and determine which nodes are implicated by words in the token sequence. In some examples, if a word or phrase in the token sequence is found to be associated with one or more nodes in ontology 760 (via lexical index 744), the word or phrase "triggers" or "activates" these nodes. can do it Based on the amount and/or relative importance of the active nodes, natural language processing module 732 may select one of the actionable intents as the task the user intended the digital assistant to perform. In some examples, the domain with the most "triggered" nodes may be selected. In some examples, the domain with the highest confidence value (eg, based on the relative importance of its various triggered nodes) may be selected. In some examples, the domain may be selected based on a combination of the number and importance of triggered nodes. In some examples, additional factors are likewise taken into account in selecting a node, such as whether the digital assistant has previously correctly interpreted a similar request from the user.
User data 748 may store user-specific information, such as user-specific vocabulary, user preferences, user addresses, the user's preferred and second language, the user's contact list, and other short-term or long-term information about each user. may include In some examples, natural language processing module 732 may use user-specific information to supplement information included in user input to further define user intent. For example, for a user request to "invite my friends to my birthday party", the natural language processing module 732 is configured to determine who the "friends" are and when and where the "birthday party" will be held. Instead of requiring the user to provide such information explicitly in his/her request for the purpose of this, user data 748 may be accessed.
Other details of searching for an ontology based on a token string are described in U.S. Utility Patent Application No. 12/341,743 for "Method and Apparatus for Searching Using An Active Ontology," filed December 22, 2008, The entire disclosure of the application is incorporated herein by reference.
In some examples, once natural language processing module 732 identifies an actionable intent (or domain) based on a user request, natural language processing module 732 executes a structured query ( structured query) can be created. In some examples, the structured query may include parameters for one or more nodes in a domain for an actionable intent, at least some of the parameters appended with specific information and requirements specified in the user request. For example, a user may say, "Make a dinner reservation at a sushi restaurant at 7 o'clock." In this case, the natural language processing module 732 may correctly identify the actionable intent based on the user input as "reserve a restaurant". According to the ontology, the structured query for the "restaurant reservation" domain may include parameters such as {cooking}, {time}, {date}, {number of people}. In some examples, based on the speech input and text derived from the speech input using the STT processing module 730, the natural language processing module 732 can generate a partially structured query for the restaurant reservation domain, Here the partially structured query includes the parameters {cooking = "sushi"} and {time = "7pm"}. However, in this example, the user's speech input contains insufficient information to complete the structured query associated with the domain. Accordingly, other essential parameters such as {number of people} and {date} may not be specific to the structured query based on currently available information. In some examples, the natural language processing module 732 may append the received context information to some parameters of the structured query. For example, in some examples, if the user requested a sushi restaurant "near me," natural language processing module 732 may append the GPS coordinates from the user device to the {location} parameter in the structured query.
In some examples, natural language processing module 732 may pass the generated structured query (including any completed parameters) to task flow processing module 736 ("task flow processor"). The task flow processing module 736 is configured to receive the structured query from the natural language processing module 732 , complete the structured query if necessary, and perform the operations necessary to "complete" the user's ultimate request. can be configured. In some examples, various procedures necessary to complete these tasks may be provided in task flow models 754 . In some examples, task flow models 754 may include procedures for obtaining additional information from a user, and task flows for performing actions associated with an actionable intent.
As noted above, to complete a structured query, task flow processing module 736 may need to initiate additional conversations with the user to obtain additional information and/or to disambiguate potentially ambiguous speech inputs. have. When such interaction is required, the task flow processing module 736 may call the dialog flow processing module 734 to participate in the conversation with the user. In some examples, conversation flow processing module 734 may determine how (and/or when) to ask the user for additional information, and may receive and process the user response. Questions may be provided to users via I/O processing module 728 and answers may be received from them. In some examples, dialog flow processing module 734 can present dialog output to the user via audio and/or visual output, and can receive input from the user via verbal or physical (eg, clicking) responses. have. Continuing with the example above, when the task flow processing module 736 calls the dialog flow processing module 734 to determine the "headcount" and "date" information for the structured query associated with the domain "restaurant reservation", the dialog flow The processing module 734 asks "How many people?" and "when?" can be generated and delivered to the user. Once answers are received from the user, the dialog flow processing module 734 may then append the missing information to the structured query, or pass the information to the task flow processing module 736 to complete the missing information from the structured query. .
Once task flow processing module 736 has completed the structured query for the actionable intent, task flow processing module 736 may proceed to perform the ultimate task associated with the actionable intent. Accordingly, task flow processing module 736 may execute steps and instructions in the task flow model according to specific parameters included in the structured query. For example, a task flow model for the actionable intent of "reserve a restaurant" may include steps and instructions for contacting a restaurant and actually requesting a reservation for a specific number of people at a specific time. For example, using a structured query such as {restaurant reservation, restaurant = ABC cafe, date = 3/12/2012, time = 7pm, number of people = 5}, the task flow processing module 736 may (1) Logging on to a restaurant reservation system, such as OPENTABLE®, or a server at ABC Cafe; (2) inputting date, time, and number of people information in a predetermined form on the website; (3) submitting the form; and (4) creating a calendar entry for the reservation in the user's calendar.
In some examples, task flow processing module 736 is an assistant of service processing module 738 ("service processing module") to complete a task requested in user input or provide an informed answer requested in user input. can be employed For example, the service processing module 738 may be configured to make a phone call on behalf of the task flow processing module 736 , set a calendar entry, invoke a map search, invoke other user applications installed on the user device, or communicate with them. interact, and to invoke or interact with third party services (eg, restaurant reservation portals, social networking websites, banking portals, etc.). In some examples, the protocols and application programming interfaces (APIs) required by each service may be specified by a respective one of the service models 756 . The service processing module 738 may access an appropriate service model for the service and generate requests for the service according to protocols and APIs required by the service according to the service model.
For example, if a restaurant has enabled an online reservation service, the restaurant may submit a service model specifying the parameters needed to make a reservation and APIs for passing values of the parameters needed for the online reservation service. When requested by the task flow processing module 736, the service processing module 738 may establish a network connection with the online reservation service using the web address stored in the service model, and format according to the API of the online reservation service. to transmit the necessary parameters of the reservation (eg, time, date, number of people) to the online reservation interface.
In some examples, the natural language processing module 732 , the dialog flow processing module 734 , and the task flow processing module 736 process information to infer and define the user's intent, to further clarify and refine the user's intent. It can be used collectively and iteratively to generate a response (ie, output to a user, or completion of a task) to obtain, and finally fulfill the user's intent. The generated response may be a dialog response to the speech input that at least partially fulfills the user's intent. Also, in some examples, the generated response may be output as a speech output. In these examples, the generated responses may be sent to a speech synthesis module 740 (eg, a speech synthesizer), where the speech synthesis module 740 may be processed to synthesize dialogue responses in speech form. In still other examples, the generated response may be data content related to satisfying a user request in speech input.
The speech synthesis module 740 may be configured to synthesize speech output for presentation to a user. The speech synthesis module 740 synthesizes the speech output based on the text provided by the digital assistant. For example, the generated dialog response may be in the form of a text string. The speech synthesis module 740 may convert the text string into an audible speech output. Speech synthesis module 740 may use any suitable speech synthesis technique to generate speech output from text, including concatenative synthesis, unit selection synthesis, diphonic synthesis, domain-specific synthesis, format formant synthesis, articulatory synthesis, hidden Markov model (HMM) based synthesis, and sinewave synthesis. In some examples, speech synthesis module 740 may be configured to synthesize individual words based on a phoneme string corresponding to the words. For example, a phoneme string may be associated with a word in the generated dialogue response. The phoneme string may be stored in metadata associated with the word. The speech synthesis module 740 may be configured to directly process the phoneme string in the metadata to synthesize words in speech form.
In some examples, instead of using (or in addition to) speech synthesis module 740 , speech synthesis may be performed on a remote device (eg, server system 108 ), and the synthesized speech is output to a user. may be transmitted to the user device for For example, this may occur in some implementations where the output to the digital assistant is generated at a server system. And since server systems typically have more processing power or resources than user devices, it may be possible to obtain higher-quality speech output in client-side synthesis.
Additional details regarding digital assistants are provided in U.S. Utility Patent Application Serial No. 12/987,982, filed January 10, 2011, entitled "Intelligent Automated Assistant," and September 30, 2011, entitled "Intelligent Automated Assistant". This "Generating and Processing Task Items That Represent Tasks to Perform," may be found in U.S. Utility Patent Application No. 13/251,088, the entire disclosures of which are incorporated herein by reference.
4. Examples of Digital Assistant's Features - Intelligent Search and Object Management
8A-8F, 9A-9H, 10A and 10B, 11A-11D, 12A-12D, and 13A-13C using a search process or object management process by a digital assistant Shows functions that perform tasks. In some examples, a digital assistant system (eg, digital assistant system 700 ) is implemented by a user device according to various examples. In some examples, a user device, server (eg, server 108 ), or a combination thereof may implement a digital assistant system (eg, digital assistant system 700 ). A user device may be implemented using, for example, device 104 , 200 , or 400 . In some examples, the user device is a laptop computer, desktop computer, or tablet computer. The user device may operate in a multi-tasking environment such as a desktop environment.
With reference to FIGS. 8A-8F, 9A-9H, 10A-10B, 11A-11D, 12A-12D, and 13A-13C , in some examples, the user device may include various user interfaces (eg, user interfaces 810 , 910 , 1010 , 1110 , 1210 , and 1310 ). The user device provides a display associated with the user device (eg, touch-sensitive display system 212 , display 440 ). display various user interfaces on the screen, various user interfaces representing one or more affordances representing different processes (eg, affordances 820, 920, 1020, 1120, 1220, and 1320 representing search processes; and object management processes representing one or more affordances) affordances 830 , 930 , 1030 , 1130 , 1230 , and 1330 ). One or more processes may be immediately executed directly or indirectly by a user. For example, a user immediately executes one or more processes by selecting affordances using an input device such as a keyboard, mouse, joystick, finger, or the like. A user may also use speech input to immediately execute one or more processes, as described in more detail below. Immediate execution of a process involves calling the process if it is not already running. When at least one instance of the process is running, executing the process immediately includes running an existing instance of the process or creating a new instance of the process. For example, immediately executing an object management process includes invoking the object management process using an existing object management process, or creating a new instance of the object management process.
8A to 8F, 9A to 9H, 10A and 10B, 11A to 11D, 12A to 12D, and 13A to 13C , the user device includes a user interface (eg, , displaying affordances (eg, affordances 840, 940, 1040, 1140, 1240, and 1340) for immediately executing a digital assistant service on user interfaces 810, 910, 1010, 1110, 1210, and 1310). do. The affordance could be, for example, a microphone icon representing a digital assistant. The affordance may be displayed anywhere on the user interface. For example, affordances may include a dock at the bottom of a user interface (eg, docks 808, 908, 1008, 1108, 1208, and 1308), a menu bar at the top of the user interface (eg, menu bar 806, 906, 1006). , 1106, 1206, and 1306), the notification center on the right side of the user interface, and the like. The affordance may also be dynamically displayed on the user interface. For example, the user device may display the affordance near an application user interface (eg, an application window) so that the digital assistant service can conveniently be launched immediately.
In some examples, the digital assistant is executed immediately in response to receiving the predetermined phrase. For example, the digital assistant is invoked in response to receiving phrases such as "Hey, assistant", "Wake up, assistant", "Listen up, assistant", "OK, assistant", and the like. In some examples, the digital assistant is executed immediately in response to receiving the selection of affordance. For example, the user selects affordances 840 , 940 , 1040 , 1140 , 1240 , and/or 1340 using an input device such as a mouse, stylus, finger, or the like. Providing a digital assistant on a user device consumes computational resources (eg, power, network bandwidth, memory, and processor cycles). In some examples, the digital assistant is suspended or shut down until the user calls it. In some examples, the digital assistant is active for various periods of time. For example, the digital assistant may be active during the time various user interfaces are displayed, during the time the user device is on, during the time the user device is hibernating or sleeping, during the time the user logs off, or a combination thereof. You can monitor the user's speech input.
8A-8F, 9A-9H, 10A and 10B, 11A-11D, 12A-12D, and 13A-13C , the digital assistant performs speech inputs 852 from the user , 854, 855, 856, 952, 954, 1052, 1054, 1152, 1252, or 1352). A user provides various speech inputs to perform a task using, for example, a search process or an object management process. In some examples, the digital assistant receives speech inputs directly from the user at the user device or indirectly through another electronic device communicatively coupled to the user device. The digital assistant receives speech inputs directly from the user, for example, via a microphone (eg, microphone 213 ) of the user device. User devices include devices configured to operate in a multitasking environment, such as laptop computers, desktop computers, tablets, servers, and the like. The digital assistant may also receive speech inputs indirectly through one or more electronic devices, such as a headset, smartphone, tablet, or the like. For example, the user may speak into a headset (not shown). The headset receives speech input from the user and sends the speech input or a representation thereof to the digital assistant of the user device, for example, via a Bluetooth connection between the headset and the user device.
8A-8F, 9A-9H, 10A-10B, 11A-11D, 12A-12D, and 13A-13C, in some embodiments, a digital assistant (eg, , represented by affordances 840 , 940 , 1040 , 1140 , 1240 , and 1340 ) identify context information associated with the user device. Context information includes, for example, user specific data, metadata associated with one or more objects, sensor data, and user device configuration data. An object may be a target or component of a process (eg, an object management process) associated with performing a task or a graphical element currently displayed on a screen, the object or graphical element having focus (eg, currently selected) or currently may not have For example, objects may include files (eg, photos, documents), folders, communications (eg, email, messages, notifications, or voicemail), contacts, calendars, applications, online resources, and the like. In some examples, the user specific data includes log information, user preferences, a history of the user's interaction with the user device, and the like. Log information indicates recent objects (eg, presentation files) used in the process. In some examples, metadata associated with one or more objects includes the title of the object, temporal information of the object, author of the object, a summary of the object, and the like. In some examples, the sensor data includes various data collected by a sensor associated with the user device. For example, the sensor data includes location data indicative of the physical location of the user device. In some examples, the user device configuration data includes a current device configuration. For example, the device configuration indicates that the user device is communicatively coupled to one or more electronic devices, such as smartphones, tablets, and the like. As described in more detail below, the user device may use the context information to perform one or more processes.
8A-8F, 9A-9H, 10A-10B, 11A-11D, 12A-12D, and 13A-13C , in response to receiving a speech input, a digital The assistant determines user intent based on the speech input. As described above, in some examples, the digital assistant is configured with an I/O processing module (eg, I/O processing module 728 as shown in FIG. 7B ), an STT processing module (eg, as shown in FIG. 7B ). STT processing module 730), and a natural language processing module (eg, natural language processing module 732 as shown in FIG. 7B) to process the speech input. The I/O processing module passes the speech input to the STT processing module (or speech recognizer) for speech-to-text conversion. Speech-to-text conversion generates text based on speech input. As described above, the STT processing module generates a sequence of words or tokens ("token sequence") and provides the token sequence to the natural language processing module. The natural language processing module performs natural language processing of the text and determines user intent based on a result of the natural language processing. For example, the natural language processing module may attempt to associate the token sequence with one or more actionable intents recognized by the digital assistant. As described, when the natural language processing module identifies an actionable intent based on user input, it generates a structured query to represent the identified actionable intent. A structured query includes one or more parameters associated with an actionable intent. The one or more parameters are used to facilitate performance of a task based on an actionable intent.
In some embodiments, the digital assistant further determines whether the user intent is to perform the task using a search process or an object management process to perform the task. The retrieval process is configured to retrieve data stored internally or externally on the user device. The object management process is configured to manage objects associated with the user device. Various examples of determining user intent are described in more detail below with respect to FIGS. 8A-8F, 9A-9H, 10A-10B, 11A-11D, 12A-12D, and 13A-13C. is provided
Referring to FIG. 8A , in some examples, the user device receives speech input 852 from the user to immediately launch the digital assistant. Speech input 852 includes, for example, "Hey, assistant." In response to the speech input, the user device immediately launches the digital assistant indicated by affordance 840 or 841 , causing the digital assistant to actively monitor subsequent speech inputs. In some examples, the digital assistant provides an audio output 872 indicating that it is running immediately. For example, the voice output 872 includes "Speak. I'm listening." In some examples, the user device receives from the user a selection of affordance 840 or affordance 841 to immediately launch the digital assistant. The selection of affordance is performed using an input device such as a mouse, stylus, finger, or the like.
With reference to FIG. 8B , in some examples, the digital assistant receives a speech input 854 . Speech input 854 includes, for example, "Open a search process to find today's AAPL stock price." or simply "Show today's AAPL stock price." Based on the speech input 854 , the digital assistant determines the user intent. For example, to determine a user intent, the digital assistant determines that the actionable intent is to obtain online information, and that one or more parameters associated with the actionable intent include "AAPL stock price", and "today".
As described, in some examples, the digital assistant further determines whether the user intent is to perform the task using a search process or an object management process to perform the task. In some embodiments, to make the determination, the digital assistant determines whether the speech input includes one or more keywords indicative of a search process or an object management process. For example, the digital assistant determines that the speech input 854 includes a keyword or phrase such as "open a search process." indicating that the user's intent is to use the search process to perform a task. Consequently, the digital assistant determines that the user intent is to perform the task using the search process.
8B , upon determining that the user intent is to perform the task using the search process, the digital assistant performs the task using the search process. As described, the natural language processing module of the digital assistant generates a structured query based on user intent, and passes the generated structured query to a task flow processing module (eg, task flow processing module 736 ). The task flow processing module receives the structured query from the natural language processing module, completes the structured query if necessary, and performs the actions necessary to "fulfill" the user's ultimate request. Performing the task using the search process includes, for example, searching for at least one object. In some embodiments, the at least one object is a folder, file (eg, photo, audio, video), communication (eg, email, message, notification, voicemail), contact, calendar, application (eg, Keynote) ), Number, iTunes, Safari), online information providers (eg, Google, Yahoo, Bloomberg), or a combination thereof. In some examples, retrieving the object is based on metadata associated with the object. For example, a search for a file or folder may utilize metadata such as tag, date, time, author, title, file type, size, number of pages, and/or file location associated with the folder or file. In some examples, the file or folder is stored internally or externally on the user device. For example, the file or folder may be stored on a hard disk of the user device or stored on a cloud server. In some examples, retrieving the communication is based on metadata associated with the communication. For example, a search for emails uses metadata such as who sent the email, who received the email, and the date the email was sent/received.
8B , in accordance with a determination that the user intent is to obtain an AAPL stock price using a search process, the digital assistant performs a search. For example, the digital assistant immediately executes the search process represented by affordance 820 and causes the search process to retrieve today's AAPL stock price. In some examples, the digital assistant further causes the search process to cause the user interface 822 (eg, the search process) to provide text corresponding to the speech input 854 (eg, "Open the search process to find the AAPL stock price today.") snippet or window).
Referring to FIG. 8C , in some embodiments, the digital assistant provides a response based on a result of performing a task using a search process. As shown in FIG. 8C , as a result of searching for AAPL stock prices, the digital assistant displays a user interface 824 (eg, a snippet or window) that provides the results of performing a task using the search process. In some embodiments, user interface 824 is located as a separate user interface within user interface 822 . In some embodiments, user interfaces 824 and 822 are integrated with each other as a single user interface. On the user interface 824, the search result of the stock price of the AAPL is displayed. In some embodiments, user interface 824 further provides affordances 831 , 833 . Affordance 831 enables closing of user interface 824 . For example, when the digital assistant receives the user's selection of affordance 831 , user interface 824 disappears or closes from the display of the user device. The affordance 833 enables to move or share the search results displayed on the user interface 824 . For example, when the digital assistant receives the user's selection of affordance 833, it may move the user interface 824 (or its search results) or process (eg, object management process) to share with a notification application. ) is executed immediately. As shown in FIG. 8C , the digital assistant displays a user interface 826 associated with a notification application for providing search results for AAPL stock prices. In some embodiments, user interface 826 displays affordance 827 . Affordance 827 enables scrolling within user interface 826 such that a user can view the entire content (eg, multiple notifications) within user interface 826 and/or for its entire length and/or width. Indicates the relative position of the document. In some embodiments, user interface 826 displays results stored by the digital assistant and/or conversation history (eg, search results obtained from current and/or past search processes). Also, in some examples, the results of performance of the task are dynamically updated over time. For example, the AAPL stock price may be dynamically updated over time and displayed on the user interface 826 .
In some embodiments, the digital assistant also provides an audio output corresponding to the search result. For example, the digital assistant (eg, indicated by affordance 840 ) provides an audio output 874 that includes "Today's AAPL price is $100.00." In some examples, user interface 822 includes text corresponding to voice output 874 .
Referring to FIG. 8D , in some examples, the digital assistant immediately executes a process (eg, an object management process) for moving or sharing a search result displayed on user interface 824 in response to subsequent speech input. For example, the digital assistant receives a speech input 855 such as "Copy the AAPL stock price to my notepad." In response, the digital assistant immediately executes a process to move or copy the search results (eg, AAPL stock price) to the user's notepad. As shown in FIG. 8D , in some examples, the digital assistant further displays a user interface 825 providing the search results copied or moved to the user's notepad. In some examples, the digital assistant further provides an audio output 875 such as "Yes, the AAPL stock price has been copied to your notepad." In some examples, user interface 822 includes text corresponding to voice output 875 .
Referring to FIG. 8E , in some examples, the digital assistant determines that the user intent is to perform the task using the object management process, and performs the task using the object management process. For example, the digital assistant receives speech input 856 such as "Open object management process and show me all the photos I took on my Colorado trip." or simply "Show me all the photos I took on my Colorado trip." Based on the speech input 856 and the context information, the digital assistant determines the user intent. For example, the digital assistant determines that the actionable intent is to display photos, and determines one or more parameters, such as "all", and "Travel to Colorado." The digital assistant further uses the context information to determine which photos correspond to the user's trip to Colorado. As described, context information includes user specific data, metadata of one or more objects, sensor data, and/or device configuration data. For example, metadata associated with one or more files (eg, file 1 , file 2 , and file 3 displayed in user interface 832 ) may indicate that the file names are "Colorado" or a city name of Colorado (eg, "Denver"). ") indicates that it contains a word. The metadata may also indicate that the folder name contains the words "Colorado" or the name of a city in Colorado (eg, "Denver"). As another example, sensor data (eg, GPS data) indicates that the user was traveling within Colorado during a particular period of time. Consequently, any photos taken by the user during that particular time period are photos taken during the user's trip to Colorado. In addition, the photos themselves can contain embedded metadata that associates the photo with the location where it was taken. Based on the contextual information, the digital assistant indicates that the user's intent is to, for example, display photos stored in a folder with the folder name "Travel to Colorado", or to display photos taken while the user was traveling within Colorado. decide
As described, in some examples, the digital assistant determines whether the user intent is to perform the task using a search process or an object management process to perform the task. To make such a determination, the digital assistant determines whether the speech input includes one or more keywords indicative of a search process or an object management process. For example, the digital assistant determines that the speech input 856 includes a keyword or phrase such as "open an object management process", indicating that the user intent is to use the object management process to perform a task.
Upon determining that the user intent is to perform the task using the object management process, the digital assistant performs the task using the object management process. For example, the digital assistant retrieves at least one object using an object management process. In some examples, the at least one object includes at least one folder or file. A file may include at least one picture, audio (eg, song), or video (eg, movie). In some examples, retrieving the file or folder is based on metadata associated with the folder or file. For example, a search for a file or folder uses metadata such as tag, date, time, author, title, file type, size, number of pages, and/or file location associated with the folder or file. In some examples, the file or folder may be stored internally or externally on the user device. For example, the file or folder may be stored on a hard disk of the user device or stored on a cloud server.
8E , a determination that the user intent is to display photos stored in a folder, for example, with the folder name "Travel to Colorado," or to display photos taken while the user was traveling within Colorado, as shown in FIG. 8E . Accordingly, the digital assistant uses the object management process to perform the task. For example, the digital assistant immediately executes the object management process indicated by affordance 830 and causes the object management process to retrieve photos taken from the user's trip to Colorado. In some examples, the digital assistant also causes the object management process to display a snippet or window (not shown) providing the text of the user's speech input 856 .
Referring to FIG. 8F , in some embodiments, the digital assistant further provides a response based on a result of performing the task using the object management process. As shown in FIG. 8F , as a result of retrieving photos of the user's trip to Colorado, the digital assistant uses a user interface 834 (eg, a snippet or window) to provide a result of performing a task using an object management process. is displayed. For example, on the user interface 834 , a preview of the photos is displayed. In some examples, the digital assistant immediately executes a process (eg, an object management process) to perform an additional task on the photo, such as inserting photos into a document or attaching photos to an email. As described in more detail below, the digital assistant may immediately execute a process to perform additional tasks in response to the user's additional speech input. In addition, the digital assistant can perform multiple tasks in response to a single speech input, such as "email my mom pictures from my trip to Colorado." The digital assistant may also immediately execute a process to perform such additional tasks in response to a user's input using an input device (eg, mouse input to select one or more affordances or to perform a drag and drop operation). In some embodiments, the digital assistant further provides an audio output corresponding to the result. For example, the digital assistant provides an audio output 876 that includes "Here are some photos I took on my trip to Colorado."
Referring to FIG. 9A , in some examples, the user's speech input may not include one or more keywords indicating whether the user's intent is to utilize a search process or an object management process. For example, the user provides a speech input 952 such as "What's the score for today's Warriors match?" Speech input 952 does not include keywords indicating "search process" or "object management process". As a result, the keywords may not be available for the digital assistant to determine whether the user's intent is to perform the task using the search process or the object management process.
In some embodiments, to determine whether the user intent is to perform a task using a search process or an object management process to perform a task, the digital assistant determines whether the task is associated with a search based on the speech input do. In some examples, a task associated with a search may be performed by either a search process or an object management process. For example, both a search process and an object management process can search folders and files. In some examples, the search process may further search various objects including online information sources (eg, websites), communications (eg, emails), contacts, calendars, and the like. In some examples, the object management process may not be configured to search for specific objects, such as online information sources.
Upon determining that the task is associated with a search, the digital assistant further determines whether performing the task requires a search process. As noted, when a task is associated with a search, either a search process or an object management process may be used to perform the task. However, the object management process may not be configured to retrieve specific objects. Consequently, in order to determine whether the user intent is to utilize a search process or an object management process, the digital assistant further determines whether the task requires a search process. For example, as shown in FIG. 9A , based on the speech input 952 , the digital assistant determines whether the user's intention is, for example, to know the score of today's Warriors match. Depending on the user's intent, the digital assistant determines that performing a further task requires searching an online source of information, and thus is associated with a search. The digital assistant further determines whether performing the task requires a search process. As noted, in some examples, the search process may be configured to search online information sources, such as a website, while the object management process may not be configured to search such online information sources. Consequently, the digital assistant determines that searching online sources of information (eg, searching the Warrior's website for a score) requires a search process.
Referring to FIG. 9B , in some embodiments, upon determining that performing the task requires a search process, the digital assistant performs the task using the search process. For example, upon determining that retrieving the score of today's Warriors match requires a search process, the digital assistant immediately executes the search process indicated by affordance 920, causing the search process to to search for the score. In some examples, the digital assistant further causes the search process to cause a user interface 922 (eg, a snippet) to provide the text of the user speech input 952 (eg, "How will today's Warriors match score?"). or window). User interface 922 includes one or more affordances 921 , 927 . Similar to as described above, affordance 921 (eg, a close button) enables closure of user interface 922 , and affordance 927 (eg, a scroll bar) allows the user to view the entirety of user interface 922 within user interface 922 . Enables scrolling within user interface 922 to view content.
Referring to FIG. 9B , in some examples, based on the search results, the digital assistant further provides one or more responses. As shown in FIG. 9B , as a result of retrieving the score for today's Warriors match, the digital assistant uses the retrieval process to display a user interface 924 (eg, a snippet or window) that provides the results of performing a task. Display. In some embodiments, user interface 924 is located as a separate user interface within user interface 922 . In some embodiments, user interfaces 924 and 922 are integrated with each other as a single user interface. In some examples, the digital assistant presents a user interface 924 providing current search results (eg, Warriors match score) to another user interface (eg, shown in FIG. 8C ) providing previous search results (eg, AAPL stock price). displayed with the user interface 824). In some embodiments, the digital assistant displays only the user interface 924 that provides current search results and no other user interface that provides previous search results. As shown in FIG. 9B , the digital assistant only displays a user interface 924 for presenting current search results (eg, a Warriors match score). In some examples, affordance 927 (eg, a scroll bar) enables scrolling within user interface 922 so that the user can view previous search results. Also, in some examples, previous search results are dynamically updated or updated such that, for example, stock prices, sports scores, weather forecasts, etc. are updated over time.
As shown in FIG. 9B , on the user interface 924 , a search result of the score of today's Warriors match is displayed (eg, Warriors 104-89 Cavaliers). In some embodiments, user interface 924 further provides affordances 923 and 925 . Affordance 923 enables closing of user interface 924 . For example, when the digital assistant receives the user's selection of affordance 923 , user interface 924 disappears or closes from the display of the user device. Affordance 925 enables to move or share search results displayed on user interface 924 . For example, when the digital assistant receives the user's selection of affordance 925 , it moves the user interface 924 (or its search results) or shares it with a notification application. As shown in FIG. 9B , the digital assistant displays a user interface 926 associated with the notification application for providing search results of the Warriors match scores. As described, the performance results of the task are dynamically updated over time. For example, a Warriors match score may be dynamically updated over time while a match is in progress and may be updated by a user interface 924 (eg, a snippet or window) and/or a user interface 926 (eg, a notification application user). interface) can be displayed. In some embodiments, the digital assistant further provides an audio output corresponding to the search result. For example, the digital assistant represented by affordance 940 or 941 provides a voice output 972 such as "The Warriors beat the Cavaliers 104-89." In some examples, user interface 922 (eg, a snippet or window) provides text corresponding to spoken output 972 .
As described above, in some embodiments, the digital assistant determines whether a task is associated with a search, and in accordance with such determination, the digital assistant determines whether performing the task requires a search process. Referring to FIG. 9C , in some embodiments, the digital assistant determines that performing the task does not require a search process. For example, as shown in FIG. 9C , the digital assistant receives a speech input 954 such as "Show me all files called Cost." Based on the speech input 954 and contextual information, the digital assistant determines that the user's intent includes the word "cost" (or portions, variants, synonyms thereof) included in their file names, metadata, content of files, etc. You decide to display all files. Depending on the user's intent, the digital assistant determines that the task to be performed includes retrieving all files associated with the word "cost". Consequently, the digital assistant determines that performing the task is associated with the search. As described above, in some examples, the retrieval process and the object management process can both perform a retrieval of files. Consequently, the digital assistant determines that performing the task of retrieving all files associated with the word "cost" does not require a retrieval process.
Referring to FIG. 9D , in some examples, upon determining that performing the task does not require a retrieval process, the digital assistant determines whether the task will be performed using the retrieval process or object management based on a predetermined configuration Decide if it will be done using the process. For example, if both the retrieval process and the object management process are capable of performing the task, the predetermined configuration may indicate that the task is to be performed using the retrieval process. The predetermined configuration may be created and updated using contextual information such as user preferences or user specific data. For example, the digital assistant determines that, for a particular user, a retrieval process has been selected more frequently than an object management process for historical file retrieval. Consequently, the digital assistant creates or updates the predetermined configuration to indicate that the retrieval process is the default process for retrieving files. In some examples, the digital assistant creates or updates the predetermined configuration to indicate that the object management process is the default process.
9D , based on the predetermined configuration, the digital assistant determines that the task of retrieving all files associated with the word "cost" will be performed using the retrieval process. Consequently, the digital assistant performs a search of all files associated with the word "cost" using the search process. For example, the digital assistant immediately executes a search process indicated by affordance 920 displayed on user interface 910 and causes the search process to retrieve all files associated with the word "cost". In some examples, the digital assistant further provides an audio output 974 to inform the user that a task is being performed. The audio output 974 includes, for example, "yes, search all files called 'cost'." In some examples, the digital assistant further causes the search process to display a user interface 928 (eg, a snippet or window) that provides text corresponding to speech input 954 and voice output 974 .
Referring to FIG. 9E , in some embodiments, the digital assistant further provides one or more responses based on results of performing a task using a search process. As shown in FIG. 9E , as a result of searching all files associated with the word "cost", the digital assistant displays a user interface 947 (eg, a snippet or window) that provides the search results. In some embodiments, user interface 947 is located as a separate user interface within user interface 928 . In some embodiments, user interfaces 947 and 928 are integrated with each other as a single user interface. On user interface 947, a list of files associated with the word "cost" is displayed. In some embodiments, the digital assistant further provides an audio output corresponding to the search result. For example, the digital assistant indicated by affordance 940 or 941 provides a voice output 976 such as "Here are all the files called cost." In some examples, the digital assistant further provides, on user interface 928 , text corresponding to voice output 976 .
In some embodiments, the digital assistant provides one or more links associated with results of performing the task using the search process. Links use the search results to enable immediate execution of processes (eg, opening files, invoking object management processes). As shown in FIG. 9E , on user interface 947 , a list of files indicated by their file names (eg, Expense File 1 , Expense File 2 , Expense File 3 ) may be associated with a link. For example, a link is displayed on one side of each file name. For another example, file names are displayed in a specific color (eg, blue) to indicate that the file names are associated with links. In some examples, file names associated with links are displayed in the same color as other items displayed on user interface 947 .
As noted, the link uses the search results to enable immediate execution of the process. Immediate execution of a process includes calling the process if it is not already running. If at least one instance of the process is running, executing the process immediately includes running an existing instance of the process or creating a new instance of the process. For example, immediately executing an object management process includes invoking the object management process using an existing object management process, or creating a new instance of the object management process. 9E and 9F , a link displayed on user interface 947 enables management of an object (eg, file) associated with the link. For example, user interface 947 receives a user selection (eg, selection by cursor 934 ) of a link associated with a file (eg, "cost file 3"). In response, the digital assistant immediately executes the object management process indicated by affordance 930 to enable management of the file. As shown in FIG. 9F , the digital assistant displays a user interface 936 (eg, a snippet or window) that provides a folder containing a file (eg, "cost file 3") associated with the link. Using the user interface 936 , the digital assistant immediately executes the object management process to perform one or more additional tasks (eg, copy, edit, view, move, compress, etc.) on the files.
Referring again to FIG. 9E , in some examples, the link displayed on user interface 947 enables direct viewing and/or editing of the object. For example, the digital assistant receives, via user interface 947 , a selection of a link (eg, selection by cursor 934 ) associated with a file (eg, "cost file 3"). In response, the digital assistant immediately executes a process (eg, a document viewing/editing process) for viewing and/or editing the file. In some examples, the digital assistant does not immediately execute the object management process but immediately executes the process for viewing and/or editing the file. For example, the digital assistant directly executes a numbering process or an excel process to view and/or edit cost file 3 directly.
9E and 9G , in some examples, the digital assistant immediately executes a process for filtering search results (eg, a search process). As shown in FIGS. 9E and 9G , the user may wish to filter the search results displayed on the user interface 947 . For example, a user may wish to select one or more files from search results. In some examples, the digital assistant receives speech input 977 from the user, such as "only those that Kelvin sent me that I tagged as drafts." Based on the speech input 977 and the context information, the digital assistant determines that the user intent was to display only cost files that were sent from Kelvin and associated with the draft tag. Based on the user intent, the digital assistant immediately executes a process (eg, a search process) to filter the search results. For example, as shown in FIG. 9G , based on the search results, the digital assistant determines that cost file 1 and cost file 2 have been sent from Kelvin to the user and are tagged. As a result, the digital assistant continues to display these two files on the user interface 947 and removes the cost file 3 from the user interface 947 . In some examples, the digital assistant provides an audio output 978 such as "Here are only the ones Kelvin sent that you tagged as drafts." The digital assistant can further provide text corresponding to the spoken output 978 on the user interface 928 .
Referring to FIG. 9H , in some examples, the digital assistant immediately executes a process (eg, object management process) to perform an object management task (eg, copy, move, share, etc.). For example, as shown in FIG. 9H , the digital assistant receives, from the user, a speech input 984 such as "Move Expense File 1 to the Documents folder." Based on the speech input 984 and the context information, the digital assistant determines that the user intent is to copy or move cost file 1 from the current folder to the Documents folder. Depending on the user's intent, the digital assistant immediately executes a process (eg, object management process) to copy or move cost file 1 from its current folder to the documents folder. In some examples, the digital assistant provides an audio output 982 such as "Yes, we are moving expense file 1 to your Documents folder." In some examples, the digital assistant further provides text corresponding to the voice output 982 on the user interface 928 .
As noted, in some examples, the user's speech input may not include keywords indicating whether the user's intent is to perform the task using the search process or the object management process. 10A and 10B , in some embodiments, the digital assistant determines that performing the task does not require a retrieval process. In accordance with the determination, the digital assistant provides a voice output requesting the user to select a search process or an object management process. For example, as shown in FIG. 10A , the digital assistant receives, from the user, a speech input 1052 such as "Show me all files called 'Cost'." Based on the speech input 1052 and the context information, the digital assistant determines that the user intent is to display all files associated with the word "cost". Depending on the user's intent, the digital assistant further determines that the task may be performed by either the retrieval process or the object management process, and thus a retrieval process is not required. In some examples, the digital assistant provides a voice output 1072 such as "Shall I search using a search process or an object management process?" In some examples, the digital assistant receives, from the user, a speech input 1054 , such as "object management process." Speech input 1054 accordingly indicates that the user intent is to perform a task using an object management process. Optionally, for example, the digital assistant immediately executes the object management process indicated by affordance 1030 to retrieve all files associated with the word "cost". As shown in FIG. 10B , similar to those described above, as a result of the search, the digital assistant provides a user interface 1032 (eg, a snippet or window) that presents a folder containing files associated with the word "cost". ) is displayed. Similar to those described above, using the user interface 1032 , the digital assistant immediately executes the object management process to perform one or more additional tasks (eg, copy, edit, view, move, compress, etc.) on the files. run
11A and 11B , in some embodiments, the digital assistant identifies context information and determines user intent based on the context information and the user's speech input. As shown in FIG. 11A , the digital assistant represented by affordance 1140 or 1141 receives a speech input 1152 such as "Open the keynote presentation I made last night." In response to receiving the speech input 1152 , the digital assistant identifies contextual information such as a history of the user's interaction with the user device, metadata associated with files the user has recently worked with, and the like. For example, the digital assistant identifies metadata such as date, time, and types of files the user worked with yesterday from 6p.m. to 2a.m. Based on the identified context information and the speech input 1152 , the digital assistant retrieves the keynote presentation file associated with metadata indicating that the user intent indicates that the file was edited yesterday from approximately 6p.m. to 12a.m; and immediately executing the process (eg, a keynote process) to open the presentation file.
In some examples, the context information includes an application name or identity (ID). For example, the user's speech input provides "Open Keynote Presentation", "Find My Page Docs", or "Find My HotNewApp Docs." Context information includes application names (eg, keynote, page, hotnewapp) or application ID. In some examples, the context information is dynamically updated or synchronized. For example, the context information is updated in real time after the user installs a new application named HotNewApp. In some examples, the digital assistant dynamically identifies updated contextual information and determines user intent. For example, the digital assistant identifies application names, ie, keynote, page, hotnewapp or their ID, and determines user intent according to the application names/ID and speech inputs.
According to the user intent, the digital assistant further determines whether the user intent is to use the search process to perform the task or the object management process to perform the task. As described, the digital assistant makes such a determination based on one or more keywords included in the speech input, based on whether the task requires a search process, based on a predetermined configuration, and/or based on a user's selection. . 11A , speech input 1152 does not include keywords indicating whether the user intent is to utilize a search process or an object management process. Consequently, the digital assistant determines, for example, based on the predetermined configuration, that the user's intention is to use the object management process. Following the determination, the digital assistant immediately executes the object management process to retrieve the keynote presentation file associated with metadata indicating that the file was edited from approximately 6p.m. to 12a.m. yesterday. In some embodiments, the digital assistant further provides an audio output 1172 such as "Yes, retrieve the keynote presentation I made last night."
In some embodiments, context information is used to perform a task. For example, the application names and/or ID may be used to form a query to retrieve the application and/or objects (eg, files) associated with the application name/ID. In some examples, the server (eg, server 108 ) forms a query using the application names (eg, keynote, page, hotnewapp) and/or ID and sends the query to the digital assistant of the user device. Based on the query, the digital assistant immediately executes a retrieval process or object management process to retrieve one or more applications and/or objects. In some examples, the digital assistant only retrieves objects (eg, files) that correspond to the application name/ID. For example, if the query includes the application name "pages", the digital assistant searches only page files, not other files (eg, word files) that can be opened by the page application. In some examples, the digital assistant retrieves all objects associated with the application name/ID in the query.
11B and 11C , in some embodiments, the digital assistant provides one or more responses according to a confidence level associated with a result of performing the task. Inaccuracies may exist or occur while performing a task and/or determining a user intent whether the user intent is to perform a task using a search process or an object management process to perform a task. In some examples, the digital assistant determines a confidence level indicative of accuracy in determining user intent based on the speech input and contextual information, whether the user intent is to perform the task using a search process or an object management process to perform the task. is the accuracy of determining whether to perform a task, and the accuracy of performing a task using a search process or an object management process, or a combination thereof.
Continuing the above example shown in FIG. 11A , based on a speech input 1152 such as "Open the keynote presentation I made last night," the digital assistant edits the file yesterday from approximately 6p.m. to 12a.m. Immediately executes the object management process to perform a retrieval of the keynote presentation file associated with the metadata indicating that it has been created. The search results may include a single file that completely matches the search criteria. That is, a single file is a presentation file edited yesterday from approximately 6p.m. to 12a.m. Accordingly, the digital assistant determines that the accuracy of the search is high and, accordingly, determines that the level of confidence is high. As another example, the search results may include a plurality of files that partially match the search criteria. For example, there are no presentation files edited from approximately 6p.m. to 12a.m yesterday, or there are several presentation files edited from approximately 6p.m. to 12a.m yesterday. Accordingly, the digital assistant determines that the accuracy of the search is medium or low, and therefore determines that the confidence level is medium or low.
11B and 11C , the digital assistant provides a response according to the determination of the confidence level. In some examples, the digital assistant determines whether the confidence level is above a threshold confidence level. Upon determining that the confidence level is greater than or equal to the threshold confidence level, the digital assistant provides a first response. Upon determining that the confidence level is below the threshold confidence level, the digital assistant provides a second response. In some examples, the second response is different from the first response. As shown in FIG. 11B , if the digital assistant determines that the confidence level is above a threshold confidence level, the digital assistant performs a keynote displayed by a process (eg, user interface 1142 ) to enable viewing and editing of the file. process) is executed immediately. In some examples, the digital assistant provides an audio output such as "Here is a presentation I made last night," and displays the text of the audio output in the user interface 1143 . 11C , if the digital assistant determines that the confidence level is below a threshold confidence level, the digital assistant displays a user interface 1122 (eg, a snippet or window) that provides a list of candidate files. Each of the candidate files may partially satisfy the search criteria. In some embodiments, the confidence level may be predetermined and/or dynamically updated based on user preferences, historical accuracy, and the like. In some examples, the digital assistant further provides an audio output 1174 such as "Here are all the presentations I made last night," and displays text corresponding to the audio output 1174 on the user interface 1122 . .
11D , in some embodiments, the digital assistant immediately executes a process (eg, a keynote presentation process) to perform additional tasks. Continuing with the example above, as shown in FIGS. 11B and 11D , the user may wish to display the presentation file in full screen mode. The digital assistant receives, from the user, a speech input 1154 such as "Give me full screen." Based on the speech input 1154 and the context information, the digital assistant determines that the user intent is to display the presentation file in full screen mode. Depending on the user's intent, the digital assistant causes the keynote presentation process to display the slides in full screen mode. In some examples, the digital assistant provides an audio output 1176 such as "Yes, I'm showing the presentation full screen."
12A-12C , in some embodiments, the digital assistant determines, based on a single speech input or utterance, that the user intent is to perform a plurality of tasks. Depending on the user's intent, the digital assistant further immediately executes one or more processes to perform a plurality of tasks. For example, as shown in Figure 12A, the digital assistant represented by affordance 1240 or 1241 can provide a single speech input ( 1252) is received. Based on the speech input 1252 and the context information, the digital assistant determines that the user intent is to perform the first task and the second task. Similar to those described above, the first task is to display photos stored in a folder named "Travel to Colorado", or to display photos that the user took while traveling within Colorado. For the second task, the context information may indicate that a particular email address stored in the user's contacts is tagged as the user's mother. Accordingly, the second task is to send an email containing photos associated with a trip to Colorado to a specific email address.
In some examples, the digital assistant determines, for each task, whether the user intent is to perform the task using a search process or an object management process to perform the task. For example, the digital assistant determines that the first task is associated with a search and that the user intent is to perform the first task using an object management process. 12B , upon determining that the user's intention is to perform the first task using the object management process, the digital assistant immediately executes the object management process to retrieve photos associated with the user's trip to Colorado. In some examples, the digital assistant displays a user interface 1232 (eg, a snippet or window) that provides a folder containing search results (eg, Photos 1, 2, and 3). For another example, the digital assistant determines that the first task is associated with a search and that the user intent is to use the search process to perform the first task. 12C , upon determining that the user's intention is to perform the first task using the search process, the digital assistant immediately executes the search process to retrieve photos associated with the user's trip to Colorado. In some examples, the digital assistant displays a user interface 1234 (eg, a snippet or window) that provides photos and/or links associated with the search results (eg, Photos 1, 2, and 3).
For another example, the digital assistant determines that the second task (eg, sending an email including photos associated with a trip to Colorado to a specific email address) is not associated with searching or managing the object. In accordance with the determination, the digital assistant determines whether the task can be performed using a process available on the user device. For example, the digital assistant determines that the second task can be performed using an email process at the user device. In accordance with the determination, the digital assistant initiates a process to perform the second task. 12B and 12C , the digital assistant immediately executes the email process and displays user interfaces 1242 and 1244 associated with the email process. The e-mail process attaches photos associated with the user's trip to Colorado to e-mail messages. 12B and 12C , in some embodiments, the digital assistant further outputs audio outputs 1272, 1274, such as "Here are photos from a trip to Colorado. Send photos to your mom." Ready to send. Would you like to proceed?" In some examples, the digital assistant displays text corresponding to the voice output 1274 on the user interface 1244 . In response to the voice outputs 1272, 1274, the user provides a speech input, eg, "OK". Upon receiving speech input from the user, the digital assistant causes the email process to send email messages.
Techniques for performing a plurality of tasks based on a single speech input or multiple commands contained within an utterance can be found in related applications, for example: filed May 28, 2015 and entitled " There is a US Provisional Patent Application No. 14/724,623, entitled "MULTI-COMMAND SINGLE UTTERANCE INPUT METHOD," filed on May 30, 2014, and entitled "MULTI-COMMAND SINGLE UTTERANCE INPUT METHOD." 62/005,556; and U.S. Provisional Patent Application No. 62/129,851, filed March 8, 2015, entitled "MULTI-COMMAND SINGLE UTTERANCE INPUT METHOD." Each of these applications is incorporated herein by reference in its entirety.
12C and 12D , in some examples, the digital assistant causes the process to perform additional tasks based on the user's additional speech inputs. For example, in a screen of search results displayed on user interface 1234, a user may wish to send some, but not all, of the photos. The user provides a speech input 1254, such as "Send only picture 1 and picture 2". In some examples, the digital assistant receives the speech input 1254 after the user selects affordance 1235 (eg, a microphone icon displayed on user interface 1234 ). The digital assistant determines, based on the speech input 1254 and the context information, that the user intent is to send an email with only Photo 1 and Photo 2 attached. Depending on the user's intent, the digital assistant causes the email process to delete picture 3 from the email message. In some examples, the digital assistant provides an audio output 1276, such as "Yes, I have attached Photo 1 and Photo 2 to your email." ) is displayed on the
Referring to FIG. 13A , in some embodiments, upon determining that the task is not associated with a search, the digital assistant determines whether the task is associated with managing at least one object. As shown in FIG. 13A , for example, the digital assistant receives a speech input 1352 , such as "Create a new folder on the desktop called Projects." Based on the speech input 1352 and the context information, the digital assistant determines that the user intent is to create a new folder on the desktop with the folder name "Projects". The digital assistant further determines that the user intent is not associated with a search, but instead relates to managing an object (eg, a folder). Accordingly, the digital assistant determines that the user intent is to perform the task using the object management process.
In some examples, upon determining that the user intent is to perform the task using the object management process, the digital assistant performs the task using the object management process. Performing a task using an object management process includes, for example, creating at least one object (eg, creating a folder or file), storing at least one object (eg, creating a folder, file, or storing communication), and compressing at least one object (eg, compressing folders and files). Performing the task using the object management process may further include, for example, copying or moving the at least one object from the first physical or virtual storage device to the second physical or virtual storage device. For example, the digital assistant immediately executes an object management process to cut and paste files from the user to a device flash drive or cloud drive.
Performing a task using an object management process may include, for example, deleting at least one object stored on physical or virtual storage (eg, deleting a folder or file) and/or storing on physical or virtual storage. It may further include recovering the at least one object (eg, recovering a deleted folder or a deleted file). Performing the task using the object management process may further include, for example, marking the at least one object. In some examples, the marking of the object may or may not be visible. For example, the digital assistant may cause the object management process to create "like" marks for social media posts, tag emails, mark files, and the like. The marking can be seen, for example, by displaying a flag, an indicator, or the like. Marking may also be performed on the metadata of the object so that the storage (eg, memory) contents of the metadata may be changed. Metadata may or may not be visible.
Performing the task using the object management process may further include backing up, for example, according to a predetermined period for backing up the at least one object or upon a user's request. For example, the digital assistant may cause the object management process to immediately run a backup program (eg, a time machine program) to back up folders and files. The backup may be performed automatically according to a predetermined schedule (eg, once a day, once a week, once a month, etc.) or according to a user request.
Performing the task using the object management process may further include, for example, sharing the at least one object between one or more electronic devices communicatively coupled to the user device. For example, the digital assistant may cause the object management process to share photos stored on the user device with another electronic device (eg, the user's smartphone or tablet).
13B , in accordance with a determination that the user intent is to perform the task using the object management process, the digital assistant performs the task using the object management process. For example, the digital assistant immediately executes the object management process to create a folder named "Projects" on the desktop of the user interface 1310 . In some examples, the digital assistant may cause the object management process to further open the folder automatically or in response to additional user input. For example, the digital assistant provides audio output 1372, such as "Yes, I have created a folder on your desktop called Projects, would you like to open it?" The user provides a speech input 1374, eg, "yes." In response to the user's speech input 1374 , the digital assistant causes the object management process to open the projects folder and display a user interface 1332 corresponding to the projects folder.
Referring to FIG. 13C , in some embodiments, the digital assistant provides one or more affordances that enable a user to manipulate the results of performing a task using a search process or an object management process. The one or more affordances include, for example, an edit button, a cancel button, a redo button, an undo button, and the like. For example, as shown in FIG. 13C , after creating a folder on the desktop named "Projects", the digital assistant presents a user interface 1334, which provides an edit button 1336A, undo button 1336B, and a redo button 1336C. In some examples, edit button 1336A enables a user to edit one or more aspects of an object (eg, edit the name of a projects folder); undo button 1336B enables the user to undo the last task performed by the object management process (eg, delete the projects folder); Redo button 1336C enables the user to repeat the last task performed by the object management process (eg, create another folder using the object management process). It will be appreciated that the digital assistant may provide any desirable affordances that enable the user to perform any manipulation of the result of performing a task using a search process or an object management process.
As described, the digital assistant can determine whether the user intent is to perform the task using a search process or an object management process to perform the task. In some examples, the digital assistant determines that the user intent is not associated with a search process or an object management process. For example, the user provides a speech input, such as "start dictation". The digital assistant determines that the task of the dictation is not related to the search. In some examples, upon determining that the task is not associated with a search, the digital assistant further determines whether the task is associated with managing the at least one object. For example, the digital assistant determines that the task of the dictation is not also associated with managing the object, such as copying, moving, or deleting a file, folder, or email. In some examples, upon determining that the task is not associated with managing the object, the digital assistant determines whether the task can be performed using a process available to the user device. For example, the digital assistant determines that the task of dictation may be performed using a dictation process available on the user device. In some examples, the digital assistant initiates a conversation with the user about performing a task using a process available on the user device. For example, the digital assistant may output a voice, such as "Yes, start dictation." Or offer "Are you sure you want to dictate to this presentation you're working on?" After providing the audio output, the digital assistant receives a response from the user confirming, for example, that the user's intent is to dictate to the presentation the user is currently working on.
5. Exemplary Functions of Digital Assistant - Continuity
14A to 14D, 15A to 15D, 16A to 16C, and 17A to 17E illustrate functions of performing a task using content located elsewhere by a digital assistant in a user device or a first electronic device. show In some examples, a digital assistant system (e.g., digital assistant system 700) is implemented by a user device (e.g., devices 1400, 1500, 1600, 1700) according to various examples. In some examples, the user device, A server (eg, server 108 ), or a combination thereof, may implement a digital assistant system (eg, digital assistant system 700 ). A user device may, for example, device 104 , 200 , or 400 . In some examples, the user device may be a laptop computer, a desktop computer, or a tablet computer. The user device operates in a multi-tasking environment such as a desktop environment.
With respect to FIGS. 14A-14D , 15A-15D, 16A-16C, and 17A-17E , in some examples, a user device (eg, devices 1400 , 1500 , 1600 , and 1700 ) is Various user interfaces (eg, user interfaces 1410 , 1510 , 1610 , and 1710 ) are provided. Similar to those described above, the user device displays various user interfaces on the display, which enable the user to immediately execute one or more processes (eg, a movie process, a photo process, a web browsing process).
As shown in FIGS. 14A-14D , 15A-15D , 16A-16C , and 17A-17E , similar to those described above, user devices (eg, devices 1400 , 1500 , 1600 ) , and 1700) provide an affordance (e.g., affordances 1440, 1540, 1640, and 1740) for immediately launching a digital assistant service on a user interface (e.g., user interfaces 1410, 1510, 1610, and 1710). ) is displayed. Similar to those described above, in some examples, the digital assistant is executed immediately in response to receiving the predetermined phrase. In some examples, the digital assistant is executed immediately in response to receiving the selection of affordance.
14A-14D , 15A-15D , 16A-16C , and 17A-17E , in some embodiments, the digital assistant receives speech inputs 1452 , 1454 , 1456 , 1458 from the user. , 1552, 1554, 1556, 1652, 1654, 1656, 1752, and 1756). A user is, for example, a user device (eg, devices 1400 , 1500 , 1600 , and 1700 ) or a first electronic device (eg, electronic devices 1420 , 1520 , 1530 , 1522 , 1532 , 1620 , 1622 ). , 1630, 1720, and 1730)) may provide various speech inputs for performing a task using content located elsewhere. Similar to those described above, in some examples, the digital assistant may receive speech inputs directly from the user at the user device or indirectly via another electronic device communicatively coupled to the user device.
14A-14D, 15A-15D, 16A-16C, and 17A-17E , in some embodiments, the digital assistant identifies context information associated with the user device. Context information includes, for example, user specific data, sensor data, and user device configuration data. In some examples, the user specific data may include user preferences, a history of interactions of the user with the user device (eg, devices 1400 , 1500 , 1600 , and 1700 ), and/or electronic devices communicatively coupled to the user device, and the like. Contains log information indicating For example, user specific data may include that the user recently took a self-portrait picture using the electronic device 1420 (eg, a smartphone); Indicates that the user recently accessed a podcast, webcast, movie, song, audiobook, etc. In some examples, sensor data includes various data collected by a sensor associated with a user device or other electronic devices. For example, sensor data includes GPS location data indicating the physical location of the user device or electronic devices communicatively coupled to the user device at any point in time or for any period of time. For example, the sensor data indicates that the photo stored on the electronic device 1420 was taken in Hawaii. In some examples, the user device configuration data includes a current or historical device configuration. For example, the user device configuration data indicates that the user device is currently communicatively connected to some electronic devices but disconnected from other electronic devices. Electronic devices include, for example, smartphones, set-top boxes, tablets, and the like. As described in more detail below, contextual information may be used to determine user intent and/or perform one or more tasks.
14A-14D , 15A-15D , 16A-16C , and 17A-17E , similar to those described above, in response to receiving the speech input, the digital assistant responds to the speech input. based on user intent. The digital assistant determines user intent based on results of natural language processing. For example, the digital assistant identifies an actionable intent based on user input and generates a structured query to represent the identified actionable intent. A structured query includes one or more parameters associated with an actionable intent. One or more parameters may be used to facilitate performance of a task based on an actionable intent. For example, based on a speech input, such as "Show me the selfie I took," the digital assistant determines that the actionable intent is to display a photo, and the parameters include a self-portrait that the user has recently taken in the last few days. In some embodiments, the digital assistant further determines the user intent based on the speech input and context information. For example, the context information indicates that the user device is communicatively connected to the user's phone using a Bluetooth connection, and indicates that the self-portrait picture was added to the user's phone two days ago. As a result, the digital assistant determines that the user's intention is to display a picture that is a self-portrait added to the user's phone two days ago. Determining user intent based on speech input and contextual information is described in more detail in various examples below.
In some embodiments, according to the user intent, the digital assistant further determines whether the task is to be performed at the user device or at a first electronic device communicatively coupled to the user device. Various examples of crystals are provided in greater detail below with respect to FIGS. 14A-14D, 15A-15D, 16A-16C, and 17A-17E.
Referring to FIG. 14A , in some examples, the user device 1400 receives a speech input 1452 from the user to invoke the digital assistant. 14A , in some examples, the digital assistant is represented by affordances 1440 or 1441 displayed on user interface 1410 . Speech input 1452 includes, for example, "Hey, assistant." In response to speech input 1452 , user device 1400 invokes the digital assistant to have the digital assistant actively monitor subsequent speech inputs. In some examples, the digital assistant provides an audio output 1472 announcing that it is being called. For example, the voice output 1472 includes "Speak. I'm listening." 14A , in some examples, user device 1400 is communicatively coupled to one or more electronic devices, such as electronic device 1420 . The electronic device 1420 may communicate with the user device 1400 using a wired or wireless network. For example, electronic device 1420 may communicate with user device 1400 using a Bluetooth connection to allow voice and data (eg, audio and video files) to be exchanged between the two devices.
Referring to FIG. 14B , in some examples, the digital assistant receives a speech input 1454 , such as "Show this device the selfies I took with my phone." Based on the speech input 1454 and/or context information, the digital assistant determines user intent. For example, as shown in FIG. 14B , the context information may indicate that the user device 1400 is communicatively connected to the electronic device 1420 using a wired or wireless network (eg, a Bluetooth connection, a Wi-Fi connection, etc.). indicates that it has been The context information also indicates that the user recently took a self-portrait, which is stored in the electronic device 1420 with the name "selfie0001". Consequently, the digital assistant determines that the user's intent is to display a photo, named selfie0001, stored on the electronic device 1420 . Alternatively, the photo may have been tagged and identified as containing the user's face by photo recognition software.
As described, according to the user intent, the digital assistant further determines whether the task is to be performed at the user device or at a first electronic device communicatively coupled to the user device. In some embodiments, determining whether the task is to be performed at the user device or the first electronic device is based on one or more keywords included in the speech input. For example, the digital assistant determines that the speech input 1454 includes a keyword or phrase indicating that the task will be performed on the user device 1400 , such as "on this device." As a result, the digital assistant determines that displaying the photo with the name selfie0001 stored on the electronic device 1420 will be performed on the user device 1400 . User device 1400 and electronic device 1420 are different devices. For example, user device 1400 may be a laptop computer and electronic device 1420 may be a telephone.
In some embodiments, the digital assistant further determines whether content associated with performance of the task is located elsewhere. At or around the time the digital assistant determines which device will perform the task, if at least a portion of the content for performing the task is not stored on the device determined to perform the task, the content is located elsewhere. For example, as shown in FIG. 14B , at or about the time the digital assistant of user device 1400 determines that the user intent is to display a photo named selfie0001 on user device 1400 , the name The photo, which is selfie0001, is not stored in the user device 1400 , but instead is stored in the electronic device 1420 (eg, a smartphone). Accordingly, the digital assistant determines that the photo is located elsewhere than the user device 1400 .
14B , in some embodiments, in accordance with a determination that the task will be performed on the user device and the content for performing the task is located elsewhere, the digital assistant of the user device provides the content for performing the task. receive In some examples, the digital assistant of the user device 1400 receives at least a portion of the content stored on the electronic device 1420 . For example, to display a photo named selfie0001, the digital assistant of user device 1400 sends a request to electronic device 1420 to obtain a photo named selfie0001. Electronic device 1420 receives the request and, in response, sends a photo named selfie0001 to user device 1400 . The digital assistant of user device 1400 then receives a photo named selfie0001.
14B , in some embodiments, after receiving the elsewhere located content, the digital assistant provides a response at the user device. In some examples, providing the response includes performing the task using the received content. For example, the digital assistant of the user device 1400 displays a user interface 1442 (eg, a snippet or window) that provides a screen 1443 of a photo named selfie0001. The screen 1443 may be a preview (eg, thumbnail) of a photo named selfie0001, an icon, or a full screen.
In some examples, providing the response includes providing a link associated with the task to be performed at the user device. Links enable immediate execution of processes. As noted, executing a process immediately includes calling the process if the process is not yet running. If at least one instance of the process is running, executing the process immediately includes running an existing instance of the process or creating a new instance of the process. As shown in FIG. 14B , user interface 1442 may provide a link 1444 associated with screen 1443 of a photo named selfie0001. Link 1444 enables immediate execution of a photo process, for example to show a full screen of a photo or to edit a photo. For example, the link 1444 is displayed on one side of the screen 1443 . As another example, screen 1443 itself includes or incorporates a link 1444 such that selection of screen 1443 immediately launches a photo process.
In some embodiments, providing the response includes providing one or more affordances that enable the user to further manipulate a result of performance of the task. 14B , in some examples, the digital assistant provides affordances 1445 , 1446 on a user interface 1442 (eg, a snippet or window). The affordance 1445 may include a button for adding a photo to an album, and the affordance 1446 may include a button for canceling the screen 1443 of the photo. The user may select one or both of affordances 1445 , 1446 . In response to selection of affordance 1445 , for example, the photo process adds the photo associated with screen 1443 to the album. In response to selection of affordance 1446 , for example, the photo process removes screen 1443 from user interface 1442 .
In some embodiments, providing the response includes providing a voice output according to the task to be performed at the user device. As shown in FIG. 14B , the digital assistant, indicated by affordances 1440 or 1441 , provides an audio output 1474 , such as "This is the most recent selfie on your phone."
Referring to FIG. 14C , in some examples, based on a single speech input/utterance and context information, the digital assistant determines that the user intent is to perform a plurality of tasks. As shown in FIG. 14C , the digital assistant receives a speech input 1456 , such as "Show this device the selfies I took with my phone and set it as my wallpaper." Based on the speech input 1456 and the context information, the digital assistant determines that the user intent is to perform a first task of displaying a photo named selfie0001 stored on the electronic device 1420, and the photo named selfie0001 is placed in the background. The second task of setting the screen is performed. Thus, based on the single speech input 1456 , the digital assistant determines that the user's intent is to perform multiple tasks.
In some embodiments, the digital assistant determines whether a plurality of tasks are to be performed at the user device or at an electronic device communicatively coupled to the user device. For example, using the keywords "this device" included in the speech input 1456 , the digital assistant determines that a plurality of tasks will be performed on the user device 1400 . Similar to those described above, the digital assistant further determines whether content for performing at least one task is located elsewhere. For example, the digital assistant determines that content for performing at least a first task (eg, displaying a photo named selfie0001) is located elsewhere. In some embodiments, in accordance with a determination that the plurality of tasks will be performed at the user device and the content for performing the at least one task is located elsewhere, the digital assistant from the other electronic device (eg, electronic device 1420 ) Request content, receive content to perform a task, and provide a response at the user device.
In some embodiments, providing the response includes performing a plurality of tasks. For example, as shown in FIG. 14C , providing a response includes performing a first task of displaying a screen 1449 of a picture named selfie0001, and setting a picture named selfie0001 as a wallpaper. and performing a second task. In some examples, the digital assistant automatically configures the wallpaper to be a picture named selfi0001 using the wallpaper settings configuration process. In some examples, the digital assistant provides a link to set wallpaper 1450 , which enables the user to manually set a wallpaper using a photo named selfie0001. For example, a user may select a link to wallpaper settings 1450 by using an input device, such as a mouse, stylus, or finger. Upon receiving a selection of a link to wallpaper settings 1450 , the digital assistant goes through a wallpaper settings configuration process that enables the user to select a photo named selfie0001 and set it as the wallpaper of the user device 1400 . start
14C , in some examples, the digital assistant initiates a conversation with the user and facilitates composition of the wallpaper in response to receiving speech input from the user. For example, the digital assistant provides an audio output 1476, such as "This is the most recent selfie on your phone. Would you like to set it as wallpaper?" The user provides a speech input, such as "OK". Upon receiving the speech input, the digital assistant immediately executes the wallpaper settings configuration process to configure the wallpaper as a photo named selfie0001.
As described, in some examples, the digital assistant determines user intent based on the speech input and context information. Referring to FIG. 14D , in some examples, the speech input may not include sufficient information to determine user intent. For example, the speech input may not indicate the location of the content to perform the task. As shown in FIG. 14D , the digital assistant receives a speech input 1458 , such as "Show me the selfie I took." Speech input 1458 does not include one or more keywords indicating which pictures will be displayed or where selfies will be displayed. Consequently, user intent may not be determined based solely on speech input 1458 . In some examples, the digital assistant determines the user intent based on the speech input 1458 and context information. For example, based on the context information, the digital assistant determines that the user device 1400 is communicatively coupled to the electronic device 1420 . In some examples, the digital assistant immediately launches a search process to retrieve photos the user has recently taken on user device 1400 and electronic device 1420 . Based on the search results, the digital assistant determines that a photo named selfie0001 is stored on the electronic device 1420 . Accordingly, the digital assistant determines that the user intent is to display a photo named selfie0001 located on the electronic device 1420 . In some examples, if the user intent cannot be determined based on the speech input and context information, the digital assistant initiates a conversation with the user to further articulate or articulate the user intent.
14D , in some examples, the speech input may not include one or more keywords indicating whether the task will be performed at the user device or at an electronic device communicatively coupled to the user device. For example, speech input 1458 does not indicate whether the task of displaying a selfie will be performed on user device 1400 or electronic device 1420 . In some examples, the digital assistant determines whether the task is to be performed at the user device or the electronic device based on the context information. For example, the context information indicates that the digital assistant receives the speech input 1458 at the user device 1400 and not the electronic device 1420 . Consequently, the digital assistant determines that the task of displaying the selfie will be performed at the user device 1400 . For another example, the context information indicates that the photo will be displayed on the electronic device 1420 according to a user preference. Consequently, the digital assistant determines that the task of displaying the selfie will be performed at the electronic device 1420 . It is understood that the digital assistant may determine whether a task is to be performed at the user device or the electronic device based on any contextual information.
Referring to FIG. 15A , in some embodiments, the digital assistant performs a task at an electronic device (eg, electronic device 1520 and/or 1530 ) communicatively coupled to the user device (eg, user device 1500 ). decides to be, and determines that the content is located elsewhere on the electronic device. As shown in FIG. 15A , in some examples, the digital assistant receives a speech input 1552 , such as "Play this movie on my TV." As described, the digital assistant can determine user intent based on the speech input 1552 and context information. For example, the context information indicates that user interface 1542 is displaying a movie named ABC.mov. Consequently, the digital assistant determines that the user's intent is to play a movie named ABC.mov.
According to the user intent, the digital assistant further determines whether the task is to be performed at the user device or at a first electronic device communicatively coupled to the user device. In some embodiments, determining whether the task is to be performed at the user device or the first electronic device is based on one or more keywords included in the speech input. For example, speech input 1552 includes words or phrases, "on my TV." In some examples, the context information indicates that the user device 1500 is connected to the set-top box 1520 and/or the TV 1530 using, for example, a wired connection, a Bluetooth connection, or a Wi-Fi connection. . Consequently, the digital assistant determines that the task of playing the movie named ABC.mov will be performed on the set top box 1520 and/or the TV 1530 .
In some embodiments, the digital assistant further determines whether content associated with performance of the task is located elsewhere. As noted, at or around the time the digital assistant determines which device will perform the task, if at least a portion of the content for performing the task is not stored on the device determined to perform the task, the content is located elsewhere. . For example, as shown in FIG. 15A , at or about the time the digital assistant of user device 1500 determines that the movie ABC.mov will be played on set top box 1520 and/or TV 1530 , At least a portion of the movie ABC.mov is stored on the user device 1500 (eg, a laptop computer) and/or on a server (not shown) and not on the set-top box 1520 and/or the TV 1530 . Accordingly, the digital assistant determines that the movie ABC.mov is located in a different location from the set top box 1520 and/or the TV 1530 .
15B , determining that a task is to be performed on a first electronic device (eg, set-top box 1520 and/or TV 1530 ) and that content for performing the task is located elsewhere than the first electronic device Accordingly, the digital assistant of the user device provides the content to the first electronic device to perform the task. For example, to play the movie ABC.mov on set-top box 1520 and/or TV 1530 , the digital assistant of user device 1500 transfers at least a portion of movie ABC.mov to set-top box 1520 and/or TV 1530 . Alternatively, it is transmitted to the TV 1530 .
In some examples, instead of providing content from the user device, the digital assistant of the user device causes at least a portion of the content to be provided from another electronic device (eg, a server) to the first electronic device to perform the task. For example, the movie ABC.mov is stored on a server (not shown) and not stored on the user device 1500 . Consequently, the digital assistant of the user device 1500 causes at least a portion of the movie named ABC.mov to be transmitted from the server to the set top box 1520 and/or the TV 1530 . In some examples, content for performing a task is provided to the set-top box 1520 , which in turn transmits the content to the TV 1530 . In some examples, the content for performing the task is provided directly to the TV 1530 .
As shown in FIG. 15B , in some examples, after the content is provided to a first electronic device (eg, set-top box 1520 and/or TV 1530 ), the digital assistant of user device 1500 activates the user device ( 1500) provides a response. In some examples, providing the response includes causing the task to be performed using the content at the set-top box 1520 and/or the TV 1530 . For example, the digital assistant of user device 1500 sends a request to set top box 1520 and/or TV 1530 to initiate a multimedia process playing the movie ABC.mov. In response to the request, the set-top box 1520 and/or TV 1530 initiates a multimedia process to play the movie ABC.mov.
In some examples, the task to be performed at the first electronic device (eg, set top box 1520 and/or TV 1530 ) is a continuation of the task performed elsewhere than the first electronic device. For example, as shown in FIGS. 15A and 15B , the digital assistant of the user device 1500 caused a multimedia process of the user device 1500 to play a portion of the movie ABC.mov on the user device 1500 . . In accordance with a determination that the user intent is to play the movie ABC.mov on a first electronic device (eg, set-top box 1520 and/or TV 1530 ), the digital assistant of user device 1500 transfers to the first electronic device. Instead of starting the movie ABC.mov from the beginning, it causes the rest of the movie ABC.mov to continue playing. Consequently, the digital assistant of the user device 1500 enables the user to continue watching the movie.
As shown in FIG. 15B , in some embodiments, providing the response includes providing one or more affordances that enable the user to further manipulate the outcome of performance of the task. As shown in FIG. 15B , in some examples, the digital assistant provides affordances 1547 , 1548 on a user interface 1544 (eg, a snippet or window). The affordance 1547 may be a button for canceling playback of the movie ABC.mov on the first electronic device (eg, the set-top box 1520 and/or the TV 1530 ). The affordance 1548 may be a button for pausing or resuming playback of the movie ABC.mov that is being played on the first electronic device. A user may select affordance 1547 or 1548 using an input device, such as a mouse, stylus, or finger. Upon receiving selection of affordance 1547 , for example, the digital assistant causes playback of the movie ABC.mov on the first electronic device to cease. In some examples, after playback on the first electronic device is stopped, the digital assistant also causes playback of the movie ABC.mov on user device 1500 to resume. Upon receiving the selection of affordance 1548, for example, the digital assistant causes to pause or resume playback of the movie ABC.mov on the first electronic device.
In some embodiments, providing the response includes providing a voice output according to the task to be performed at the first electronic device. As shown in FIG. 15B , the digital assistant represented by affordance 1540 or 1541 provides an audio output 1572 , such as "Play your movie on TV."
As described, upon determining that the task will be performed at the first electronic device and the content for performing the task is located elsewhere than the first electronic device, the digital assistant provides the content for performing the task to the first electronic device do. Referring to FIG. 15C , content for performing a task may include, for example, a document (eg, document 1560 ) or location information. For example, the digital assistant of the user device 1500 receives a speech input 1556 , such as "Open this pdf on my tablet." The digital assistant determines that the user intent is to perform the task of displaying the document 1560 , and determines that the task will be performed on a tablet 1532 communicatively coupled to the user device 1500 . Consequently, the digital assistant presents the document 1560 to the tablet 1532 for display. For another example, the digital assistant of the user device 1500 receives a speech input 1554 , such as "Send this location to my phone." The digital assistant determines that the user intent is to perform the task of navigation using the location information, and determines that the task will be performed at a phone 1522 (eg, a smartphone) communicatively coupled to the user device 1500 . Consequently, the digital assistant provides location information (eg, 1234 Main St.) to the phone 1522 to perform the task of navigation.
As described, in some examples, after providing content for performing a task to the first electronic device, the digital assistant provides a response at the user device. In some embodiments, providing the response includes causing the task to be performed at the first electronic device. For example, as shown in FIG. 15D , the digital assistant of user device 1500 sends a request to phone 1522 to perform a task of navigating to the location of 1234 Main St. The digital assistant of the user device 1500 further sends a request to the tablet 1532 to perform the task of displaying the document 1560 . In some examples, providing the response at the user device includes providing a voice output according to the task to be performed at the first electronic device. As shown in FIG. 15D , the digital assistant has an audio output 1574 , such as "Show the pdf on your tablet." and audio output 1576, such as "Navigate from your phone to 1234 Main St."
As noted, in some examples, the speech input may not include one or more keywords indicating whether the task will be performed at the user device or at a first electronic device communicatively coupled to the user device. Referring to FIG. 16A , for example, the digital assistant receives a speech input 1652 , such as "Play this movie." Speech input 1652 may indicate whether the task of playing the movie is to be performed on user device 1600 or not on a first electronic device (eg, set-top box 1620 and/or TV 1630, phone 1622, or tablet). 1632)) does not indicate whether it will be performed.
In some embodiments, to determine whether the task is to be performed at the user device or the first electronic device, the digital assistant at the user device determines whether performing the task at the user device meets a performance criterion. Performance criteria facilitate evaluating the performance of a task. For example, as shown in FIG. 16A , the digital assistant determines that the user's intention is to perform the task of playing the movie ABC.mov. Performance criteria for playing a movie may include, for example, a quality criterion for playing a movie (eg, 480p, 720p, 1080p), a smoothness criterion for playing a movie (eg, no lag or wait), a screen size criterion (eg, minimum screen size of 48 inches), sound effect criteria (eg, stereo sound, number of speakers), and the like. Performance criteria may be preconfigured and/or dynamically updated. In some examples, the performance criterion is determined based on context information, such as user specific data (eg, user preferences), device configuration data (eg, screen resolution and size of electronic devices), and the like.
In some examples, the digital assistant at the user device 1600 determines that performing the task at the user device meets the performance criterion. For example, as shown in FIG. 16A , user device 1600 may have a screen resolution, screen size, and sound effects that meet performance criteria for playing the movie ABC.mov, which may be a low-resolution online video. In accordance with a determination that performing the task at the user device 1600 meets the performance criteria, the digital assistant determines that the task will be performed at the user device 1600 .
In some examples, the digital assistant of the user device 1600 determines that performing the task at the user device does not meet the performance criterion. For example, user device 1600 may not have a screen size, resolution, and/or sound effect that meets performance criteria for playing the movie ABC.mov, which may be a high-resolution Blu-ray video. In some examples, upon determining that performing the task at the user device does not meet the performance criterion, the digital assistant of the user device 1600 determines whether performing the task at the first electronic device meets the performance criterion. As shown in FIG. 16B , the digital assistant of user device 1600 determines that performing the task of playing the movie ABC.mov on set-top box 1620 and/or TV 1630 meets the performance criteria. For example, set top box 1620 and/or TV 1630 may have a screen size of 52 inches, may have 1080p resolution, and may have eight connected speakers. Consequently, the digital assistant determines that the task will be performed on the set top box 1620 and/or the TV 1630 .
In some examples, the digital assistant of the user device 1600 determines that performing the task at the first electronic device does not meet the performance criterion. According to the determination, the digital assistant determines whether performing the task at the second electronic device meets a performance criterion. For example, as shown in FIG. 16B , TV 1630 may have a screen resolution (eg, 720p) that does not meet performance criteria (eg, 1080p). Consequently, the digital assistant determines whether either the phone 1622 (eg, a smartphone) or the tablet 1632 meets the performance criteria.
In some examples, the digital assistant determines which device provides optimal performance of the task. For example, as shown in FIG. 16B , the digital assistant plays the movie ABC.mov on each of user device 1600 , set-top box 1620 and TV 1630 , phone 1622 , and tablet 1632 . Evaluate or estimate the performance of the regenerating task. Based on the evaluation or estimation, the digital assistant determines whether performing a task on one device (eg, user device 1600 ) is better than on another device (eg, phone 1622 ), and selects a device for optimal performance. decide
As described, in some examples, in accordance with the device's determination to perform the task, the digital assistant provides a response at the user device 1600 . In some embodiments, providing the response includes providing a voice output according to the task to be performed at the device. As shown in FIG. 16B , the digital assistant, indicated by affordances 1640 or 1641 , provides an audio output 1672 , such as "I'll play this movie on your TV, shall I proceed?" In some examples, the digital assistant receives speech input 1654 from the user, such as "OK". In response, the digital assistant causes the movie ABC.mov to be played, eg, on set-top box 1620 and TV 1630, and audio output 1674, eg, "Play your movie on your TV." to" is provided.
In some examples, providing the response includes providing one or more affordances that enable the user to select another electronic device for performance of the task. As shown in FIG. 16B , for example, the digital assistant provides affordances 1655A, 1655B (eg, a cancel button and a tablet button). Affordance 1655A enables the user to cancel playing the movie ABC.mov on set top box 1620 and TV 1630 . Affordance 1655B enables the user to select tablet 1632 to continue playing the movie ABC.mov.
Referring to FIG. 16C , in some embodiments, the digital assistant of user device 1600 initiates a conversation with the user to determine a device to perform a task. For example, the digital assistant provides an audio output 1676, such as "Play your movie on TV or tablet?" The user provides speech input 1656, such as "on my tablet." Upon receiving the speech input 1656 , the digital assistant determines that the task of playing the movie will be performed on the tablet 1632 communicatively coupled to the user device 1600 . In some examples, the digital assistant further provides audio output 1678 , such as "Play your movie on your tablet."
Referring to FIG. 17A , in some embodiments, the digital assistant of the user device 1700 continues to perform a task that was partially performed elsewhere in the first electronic device. In some embodiments, the digital assistant of the user device continues to perform the task using the content received from the third electronic device. As shown in FIG. 17A , in some examples, phone 1720 may have been performing the task of booking a flight using content from a third electronic device, such as server 1730 . For example, the user may have been booking a flight from Kayak.com using phone 1720 . Consequently, the phone 1720 receives the content transmitted from the server 1730 associated with Kayak.com. In some examples, a user may be interrupted while booking his or her flight on phone 1720 and wish to continue booking a flight using user device 1700 . In some examples, a user may wish to conveniently continue with a flight reservation as it is more convenient to use the user device 1700 . Accordingly, the user may provide a speech input 1752 , such as "Continue with the Kayak flight reservation I was using on my phone."
Referring to FIG. 17B , upon receiving speech input 1752 , the digital assistant determines that the user's intent is to perform the task of booking a flight. In some examples, the digital assistant further determines that a task will be performed at the user device 1700 based on the context information. For example, the digital assistant determines that speech input 1752 is received at user device 1700 , and thus determines that a task is to be performed at user device 1700 . In some examples, the digital assistant further determines that a task will be performed at the user device 1700 using contextual information, such as user preferences (eg, user device 1700 has been frequently used to book flights in the past). .
As shown in FIG. 17B , in accordance with a determination that a task will be performed on user device 1700 and content for performing the task is located elsewhere, the digital assistant receives the content for performing the task. In some examples, the digital assistant receives at least a portion of the content from a phone 1720 (eg, a smartphone) and/or receives at least a portion of the content from a server 1730 . For example, the digital assistant receives data from the phone 1720 indicating the status of a flight reservation so that the user device 1700 can continue with the flight reservation. In some examples, data representing the status of a flight reservation is stored on server 1730 , such as a server associated with Kayak.com. Accordingly, the digital assistant receives data from the server 1730 to proceed with the flight reservation.
17B , after receiving content from phone 1720 and/or server 1730 , the digital assistant provides a response at user device 1700 . In some examples, providing the response includes continuing at phone 1720 the task of booking a flight that was partially performed elsewhere. For example, the digital assistant displays a user interface 1742 that enables the user to continue booking a flight on Kayak.com. In some examples, providing the response includes providing a link associated with the task to be performed at the user device 1700 . For example, the digital assistant displays a user interface 1742 (eg, a snippet or window) that provides the status of a current flight reservation (eg, indicates available flights). User interface 1742 also provides a link 1744 (eg, a link to a web browser) to continue performing the task of booking a flight. In some embodiments, the digital assistant also provides an audio output 1772, such as "Booking at Kayak. Do you want to continue in your web browser?"
As shown in FIGS. 17B and 17C , for example, when the user selects link 1744, the digital assistant immediately launches a web browsing process and a user interface 1746 (e.g., , snippet or window). In some examples, in response to the voice output 1772 , the user makes a speech input 1756 , such as "OK", confirming that the user wishes to continue booking a flight using a web browser of the user device 1700 . to provide. Upon receiving the speech input 1756, the digital assistant immediately executes the web browsing process and displays a user interface 1746 (eg, a snippet or window) for continuing the flight reservation task.
Referring to FIG. 17D , in some embodiments, the digital assistant of user device 1700 continues to perform a task that was partially performed elsewhere in the first electronic device. In some embodiments, the digital assistant of the user device continues to perform the task using the content received from the first electronic device and not from the third electronic device, such as a server. As shown in FIG. 17D , in some examples, the first electronic device (eg, phone 1720 or tablet 1732 ) may have been performing a task. For example, the user may be composing an email using the phone 1720 or editing a document such as a photo using the tablet 1732 . In some examples, the user is interrupted while using the phone 1720 or tablet 1732 , and/or wishes to continue performing a task using the user device 1700 . In some examples, a user may wish to conveniently continue performing a task because it is more convenient to use the user device 1700 (eg, a larger screen). Thus, the user enters speech 1758, such as "Open the document I was editing." Alternatively, a speech input 1759 may be provided, such as "Open the email I was drafting."
Referring to FIG. 17D , upon receiving speech input 1758 or 1759 , the digital assistant determines that the user's intent is to perform the task of editing a document or composing an email. Similar to those described above, in some examples, the digital assistant further determines that a task will be performed at user device 1700 based on context information, and determines that content for performing the task is located elsewhere. Similar to as described above, in some examples, the digital assistant, based on contextual information (eg, user specific data), at a first electronic device (eg, phone 1720 or at tablet 1732 ). As shown in FIG. 17D , in accordance with a determination that the task will be performed at the user device 1700 and the content for performing the task is located elsewhere, the digital assistant receives the content for performing the task. In some examples, the digital assistant receives at least a portion of the content from a phone 1720 (eg, a smartphone) and/or receives at least a portion of the content from a tablet 1730 . After receiving content from phone 1720 and/or tablet 1732 , the digital assistant provides a response at user device 1700 , eg, displaying user interface 1748 for the user to continue editing the document. and/or displays a user interface 1749 for the user to continue composing the email. It will be appreciated that the digital assistant of the user device 1700 may also cause the first electronic device to continue performing a task that was partially performed elsewhere on the user device 1700 . For example, a user may be composing an email on user device 1700 and may need to leave. The user provides a speech input, such as "Open draft email on my phone." Based on the speech input, the digital assistant determines that the user intent is to continue performing the task on the phone 1720 , and that the content is located elsewhere on the user device 1700 . In some examples, the digital assistant provides content for performing a task to the first electronic device and causes the first electronic device to continue performing the task, similar to those described above.
Referring to FIG. 17E , in some embodiments, continuing to perform a task includes sharing between a plurality of devices, including, for example, user device 1700 and a first electronic device (eg, phone 1720 ). or based on context information being synchronized. As described, in some examples, the digital assistant determines user intent based on the speech input and context information. Context information may be stored on its own or elsewhere. For example, as shown in FIG. 17E , the user provides a speech input 1760 , such as "How's the weather in New York?" to phone 1720 . The digital assistant of phone 1720 determines user intent, performs tasks to obtain New York weather information, and displays the New York weather information on the user interface of phone 1720 . The user subsequently provides a speech input 1761 , such as "How is Los Angeles?" to the user device 1700 . In some examples, the digital assistant of user device 1700 determines user intent using contextual information stored directly on phone 1720 and/or shared by phone 1720 via a server. Context information includes, for example, historical user data associated with phone 1720 , conversation status, system status, and the like. Both historical user data and conversation status indicate that the user was inquiring about weather information. Accordingly, the digital assistant of the user device 1700 determines that the user's intention is to obtain weather information for Los Angeles. Based on the user intent, the digital assistant of the user device 1700 , for example, receives weather information from a server, and provides a user interface 1751 that displays the weather information on the user device 1710 .
6. Exemplary Functions of Digital Assistant - Voice Assistant System Configuration Management
18A-18F and 19A-19D illustrate functions for providing system configuration information or performing a task by a digital assistant in response to a user request. In some examples, a digital assistant system (eg, digital assistant system 700 ) may be implemented by a user device according to various examples. In some examples, a user device, server (eg, server 108 ), or a combination thereof may implement a digital assistant system (eg, digital assistant system 700 ). A user device is implemented using, for example, device 104 , 200 , or 400 . In some examples, the user device is a laptop computer, desktop computer, or tablet computer. The user device operates in a multi-tasking environment such as a desktop environment.
18A-18F and 19A-19D , in some examples, the user device provides various user interfaces (eg, user interfaces 1810 , 1910 ). Similar to those described above, the user device displays various user interfaces on the display, which enable the user to immediately execute one or more processes (eg, system configuration processes).
As shown in FIGS. 18A-18F and 19A-19D , similar to those described above, the user device, on a user interface (eg, user interfaces 1810 , 1910 ), provides immediate access to the digital assistant service. Displays affordances (eg affordances 1840 and 1940) to facilitate execution.
Similar to those described above, in some examples, the digital assistant is executed immediately in response to receiving the predetermined phrase. In some examples, the digital assistant is executed immediately in response to receiving the selection of affordance.
18A-18F and 19A-19D , in some embodiments, the digital assistant receives one or more speech input from the user, eg, speech inputs 1852, 1854, 1856, 1858, 1860, 1862, 1952. , 1954, 1956, and 1958). The user provides various speech inputs for the purpose of managing one or more system configurations of the user device. System configurations include audio configurations, date and time configurations, dictation configurations, display configurations, input device configurations, notification configurations, print configurations, security configurations, backup configurations, application configurations, user interface configurations, etc. may include. To manage audio configurations, the speech input is "Turn off my microphone", "Turn up the volume", "Turn up the volume by 10%." and the like. To manage date and time configurations, speech inputs include "What is my time zone?", "Change my time zone to Cupertino time zone", "Add London time zone." and the like. To manage dictation configurations, speech input can be "Turn on dictation", "Turn dictation off", "Speak Chinese", "Enable advanced commands." and the like. To manage display configurations, the speech input is "Make my screen brighter", "Increase my contrast by 20%", "Extend my screen to a second monitor", "Mirror my display" ." and the like. To manage input device configurations, Speech Input can be used for "Connect my Bluetooth keyboard", "Make my mouse pointer bigger." and the like. To manage network configurations, speech input can be "Turn on Wi-Fi", "Turn Wi-Fi off", "What Wi-Fi network are you connected to?", "Are you connected to my phone?" and the like. To manage notification configuration, speech input can be "Turn on Do Not Disturb", "Stop showing me these notifications", "Show me new emails only", "Don't notify me in text messages." and the like. To manage print configurations, the speech input includes "Is my printer enough ink?", "Is my printer connected?" and the like. To manage security configurations, speech input is "Change password for John's account", "Turn on firewall", "Disable cookies." and the like. To manage backup configurations, the speech input is "Run a backup now.", "Set up a backup at an interval of once a month.", "Restore the backup from July 4th of last year." and the like. To manage application configurations, the speech input is "Change my default web browser to Safari", "Automatically log in to the Messages application every time I sign". and the like. To manage user interface configurations, speech input is "Change the wallpaper of my desktop", "Hide the dock", "Add Evernote to the dock." and the like. Various examples of managing system configurations using speech inputs are described in more detail below.
Similar to those described above, in some examples, the digital assistant receives speech inputs directly from the user at the user device or indirectly through another electronic device communicatively coupled to the user device.
18A-18F and 19A-19D , in some embodiments, the digital assistant identifies context information associated with the user device. Context information includes, for example, user specific data, sensor data, and user device configuration data. In some examples, the user specific data includes log information indicative of user preferences, a history of the user's interaction with the user device, and the like. For example, user specific data indicates the last time the user's system was backed up; Indicates a user's preference for a particular Wi-Fi network, etc. when several Wi-Fi networks are available. In some examples, the sensor data includes various data collected by the sensor. For example, the sensor data represents a printer ink level collected by a printer ink level sensor. In some examples, the user device configuration data includes current and past device configurations. For example, the user device configuration data indicates that the user device is currently communicatively connected to one or more electronic devices using a Bluetooth connection. Electronic devices may include, for example, smartphones, set-top boxes, tablets, and the like. As described in more detail below, the user device may determine user intent and/or use the contextual information to perform one or more processes.
18A-18F and 19A-19D , similar to those described above, in response to receiving the speech input, the digital assistant determines a user intent based on the speech input. The digital assistant determines user intent based on results of natural language processing. For example, the digital assistant identifies an actionable intent based on user input and generates a structured query to represent the identified actionable intent. A structured query includes one or more parameters associated with an actionable intent. One or more parameters may be used to facilitate performance of a task based on an actionable intent. For example, based on a speech input, such as "Turn up the volume by 10%," the digital assistant determines that the actionable intent is to adjust the system volume, and the parameters set the volume to 10% higher than the current volume level. includes setting up. In some embodiments, the digital assistant also determines user intent based on the speech input and context information. For example, the context information may indicate that the current volume of the user device is 50%. Consequently, upon receiving a speech input, such as "Turn up the volume by 10%," the digital assistant determines that the user's intent is to increase the volume level to 60%. Determining user intent based on speech input and contextual information is described in more detail in various examples below.
In some embodiments, the digital assistant further determines whether the user intent indicates a request for information or a request to perform a task. Various examples of determination are provided in greater detail below with respect to FIGS. 18A-18F and 19A-19D.
Referring to FIG. 18A , in some examples, the user device displays a user interface 1832 associated with performing a task. For example, the task includes creating a meeting invitation. When creating a meeting invitation, the user may want to know the time zone of the user's device so that the meeting invitation can be created appropriately. In some examples, the user provides speech input 1852 to invoke the digital assistant indicated by affordance 1840 or 1841 . Speech input 1852 includes, for example, "Hey, assistant." The user device receives the speech input 1852 and, in response, invokes the digital assistant to have the digital assistant actively monitor subsequent speech inputs. In some examples, the digital assistant provides an audio output 1872 announcing that it is being called. For example, the voice output 1872 includes "Speak. I'm listening."
Referring to FIG. 18B , in some examples, the user provides a speech input 1854 , such as "What is my time zone?" The digital assistant determines that the user intent is to obtain the time zone of the user device. The digital assistant further determines whether the user intent represents a request for information or a request to perform a task. In some examples, determining whether the user intent represents a request for information or a request to perform a task includes determining whether the user intent is to change a system configuration. For example, based on a determination that the user intent is to obtain the time zone of the user device, the digital assistant determines that the system configuration is not changed. Consequently, the digital assistant determines that the user intent represents a request for information.
In some embodiments, upon determining that the user intent indicates a request for information, the digital assistant provides a voice response to the request for information. In some examples, the digital assistant obtains a status of one or more system components according to the request for information, and provides a voice response according to the status of the one or more system components. As shown in FIG. 18B , the digital assistant determines that the user intent is to obtain the time zone of the user device, and the user intent indicates a request for information. Accordingly, the digital assistant obtains the time zone status from the time and date configuration of the user device. The time zone state indicates, for example, that the user device is set to the Pacific time zone. Based on the time zone status, the digital assistant provides an audio output 1874, such as "Your computer is set for Pacific Standard Time." In some examples, the digital assistant further provides a link associated with the request for information. As shown in FIG. 18B , the digital assistant provides a link 1834 that enables the user to further manage data and time configurations. In some examples, the user selects the link 1834 using an input device (eg, a mouse). Upon receiving the user's selection of link 1834, the digital assistant immediately executes the date and time configuration process and displays the associated date and time configuration user interface. The user can thus further manage date and time configurations using the date and time configuration user interface.
Referring to FIG. 18C , in some examples, the user device displays a user interface 1836 associated with performing a task. For example, the task includes playing a video (eg, ABC.mov). To enhance the video viewing experience, a user may wish to use a speaker and may want to know if a Bluetooth speaker is connected. In some examples, the user provides a speech input 1856, such as "Is my Bluetooth speaker connected?" The digital assistant determines that the user's intention is to obtain the connected state of the Bluetooth speaker 1820 . The digital assistant further determines that obtaining the connection status of the Bluetooth speaker 1820 does not change any system configuration, and thus it is a request for information.
In some embodiments, upon determining that the user intent indicates a request for information, the digital assistant obtains the status of the system components according to the request for information, and provides a voice response according to the status of the system components. As shown in FIG. 18C , the digital assistant obtains the connection status from the network configuration of the user device. The connected state indicates, for example, that the user device 1800 is not connected to the Bluetooth speaker 1820 . Based on the connection status, the digital assistant provides an audio output 1876, such as "No, not connected, but you can check Bluetooth devices in network configurations." In some examples, the digital assistant further provides a link associated with the request for information. As shown in FIG. 18C , the digital assistant provides a link 1838 that enables the user to further manage network configurations. In some examples, the user selects the link 1838 using an input device (eg, a mouse). Upon receiving the user's selection of link 1838, the digital assistant immediately executes the network configuration process and displays the associated network configuration user interface. The user can thus further manage network configurations using the network configuration user interface.
Referring to FIG. 18D , in some examples, the user device displays a user interface 1842 associated with performing a task. For example, tasks include viewing and/or editing documents. A user may wish to print a document and may want to know if the ink in the printer 1830 is sufficient for the print job. In some examples, the user provides a speech input 1858 , such as "Is my printer enough ink?" The digital assistant determines that the user intent is to obtain the printer ink level status of the printer. The digital assistant further determines that obtaining the printer level status does not change any system configuration and thus is a request for information.
In some embodiments, upon determining that the user intent indicates a request for information, the digital assistant obtains the status of the system components according to the request for information, and provides a voice response according to the status of the system components. 18D , the digital assistant obtains the printer ink level status from the printing configuration of the user device. The printer ink level status indicates, for example, that the printer ink level of the printer 1830 is 50%. Based on the connection status, the digital assistant provides an audio output 1878, such as "Yes, your printer has enough ink. You can also look up the printer supply level in the printer configurations." In some examples, the digital assistant further provides a link associated with the request for information. As shown in FIG. 18D , the digital assistant provides a link 1844 that enables the user to further manage print configurations. In some examples, the user selects the link 1844 using an input device (eg, a mouse). Upon receiving the user's selection of the link, the digital assistant immediately executes the printer configuration process and displays the associated printer configuration user interface. The user can thus further manage printer configurations using the printer configuration user interface.
Referring to FIG. 18E , in some examples, the user device displays a user interface 1846 associated with performing a task. For example, the task includes browsing the Internet using a web browser (eg, Safari). To browse the Internet, a user may want to know available Wi-Fi networks and may select one Wi-Fi network to connect to. In some examples, the user provides speech input 1860 , such as "Which Wi-Fi networks are available?" The digital assistant determines that the user intent is to obtain a list of available Wi-Fi networks. The digital assistant determines that obtaining a list of additionally available Wi-Fi networks does not change any system configuration and therefore is a request for information.
In some embodiments, upon determining that the user intent indicates a request for information, the digital assistant obtains the status of the system components according to the request for information, and provides a voice response according to the status of the system components. As shown in FIG. 18E , the digital assistant obtains the status of currently available Wi-Fi networks from the network configuration of the user device. The status of currently available Wi-Fi networks indicates, for example, that Wi-Fi network 1, Wi-Fi network 2, and Wi-Fi network 3 are available. In some examples, the status further indicates a signal strength of each of the Wi-Fi networks. The digital assistant displays a user interface 1845 that provides information based on the status. For example, user interface 1845 provides a list of available Wi-Fi networks. The digital assistant also provides an audio output 1880, eg, "A list of available Wi-Fi networks." In some examples, the digital assistant further provides a link associated with the request for information. As shown in FIG. 18E , the digital assistant provides a link 1847 that enables the user to further manage network configurations. In some examples, the user selects the link 1847 using an input device (eg, a mouse). Upon receiving the user's selection of link 1847, the digital assistant immediately executes the network configuration process and displays the associated network configuration user interface. The user can thus further manage configurations using the network configuration user interface.
Referring to FIG. 18F , in some examples, the user device displays a user interface 1890 associated with performing a task. For example, the task includes preparing an agenda for a meeting. When preparing the meeting agenda, the user may wish to find the date and time of the meeting. In some examples, the user provides a speech input 1862 , such as "Find my calendar for next Tuesday's morning meeting time." The digital assistant determines that the user's intention is to find an available time slot on Tuesday morning on the user's calendar. The digital assistant determines that finding an additional time slot does not change any system configuration and therefore is a request for information.
In some embodiments, upon determining that the user intent indicates a request for information, the digital assistant obtains the status of the system components according to the request for information, and provides a voice response according to the status of the system components. As shown in FIG. 18F , the digital assistant obtains the status of the user's calendar from calendar configurations. The status of the user's calendar is, for example, Tuesday 9a.m. or 11a.m. is still available. The digital assistant displays a user interface 1891 that provides information depending on the status. For example, user interface 1891 provides the user's calendar proximate to the date and time the user requested. In some examples, the digital assistant also provides an audio output 1882 , such as "Tuesday 9a.m. or 11a.m appears to be available." In some examples, the digital assistant further provides a link associated with the request for information. As shown in FIG. 18F , the digital assistant provides a link 1849 that enables the user to further manage calendar configurations. In some examples, the user selects the link 1849 using an input device (eg, a mouse). Upon receiving the user's selection of link 1849, the digital assistant immediately executes the calendar configuration process and displays the associated calendar configuration user interface. The user can thus further manage configurations using the calendar configuration user interface.
Referring to FIG. 19A , the user device displays a user interface 1932 associated with performing a task. For example, the task includes playing a video (eg, ABC.mov). While the video is playing, the user may wish to turn up the volume. In some examples, the user provides a speech input 1952 , such as "Turn up the volume." The digital assistant determines that the user's intention is to increase the volume to its maximum level. The digital assistant further determines whether the user intent represents a request for information or a request to perform a task. For example, based on a determination that the user intent is to increase the volume of the user device, the digital assistant determines that the audio configuration is changed and thus the user intent indicates a request to perform a task.
In some embodiments, upon determining that the user intent indicates a request to perform the task, the digital assistant immediately executes a process associated with the user device to perform the task. Immediate execution of a process includes calling the process if it is not already running. If at least one instance of the process is running, executing the process immediately includes running an existing instance of the process or creating a new instance of the process. For example, immediately executing the audio composition process includes invoking the audio composition process, using an existing audio composition process, or creating a new instance of the audio composition process. In some examples, immediately executing the process includes performing a task using the process. For example, as shown in FIG. 19A , upon user intent to increase the volume to its maximum level, the digital assistant immediately executes the audio configuration process to set the volume to its maximum level. In some examples, the digital assistant additionally provides an audio output 1972, such as "Yes, the volume is up."
Referring to FIG. 19B , the user device displays a user interface 1934 associated with performing a task. For example, a task includes viewing a document or editing a document. Users may wish to lower the screen brightness to protect their eyesight. In some examples, the user provides a speech input 1954 , such as "Turn my screen brightness down 10%." The digital assistant determines user intent based on the speech input 1954 and context information. For example, the context information indicates that the current brightness configuration is at 90%. Consequently, the digital assistant determines that the user's intention is to reduce the brightness level from 90% to 80%. The digital assistant further determines whether the user intent represents a request for information or a request to perform a task. For example, based on a determination that the user intent is to change the screen brightness to 80%, the digital assistant determines that the display configuration is changed, and thus the user intent indicates a request to perform a task.
In some embodiments, upon determining that the user intent represents a request to perform the task, the digital assistant immediately executes the process to perform the task. For example, as shown in FIG. 19B , upon user intent to change the brightness level, the digital assistant immediately executes the display configuration process to reduce the brightness level to 80%. In some examples, the digital assistant additionally provides an audio output 1974, such as "Yes, I have your screen brightness set to 80%." In some examples, as shown in FIG. 19B , the digital assistant provides an affordance 1936 that enables the user to manipulate the result of performing the task. For example, affordance 1936 may be a sliding bar that allows the user to further change the brightness level.
Referring to FIG. 19C , the user device displays a user interface 1938 associated with performing a task. For example, the task includes providing one or more notifications. Notifications may include notifications such as emails, messages, reminders, and the like. In some examples, notifications are provided in user interface 1938 . Notifications may be displayed or provided to the user in real time or immediately after becoming available on the user device. For example, the notification appears on the user interface 1938 and/or the user interface 1910 immediately after the user device receives the notification. Occasionally, a user may be performing an important task (eg, editing a document) and may not want to be disturbed by notifications. In some examples, the user provides a speech input 1956, such as "Don't notify me of received email." The digital assistant determines that the user's intention is to turn off notifications of emails. Based on a determination that the user intent is to turn off notifications of received emails, the digital assistant determines that the notification configuration is changed, and thus the user intent represents a request to perform a task.
In some embodiments, upon determining that the user intent represents a request to perform the task, the digital assistant immediately executes the process to perform the task. For example, as shown in FIG. 19C , upon user intent, the digital assistant immediately executes a notification configuration process to turn off notifications of emails. In some examples, the digital assistant further provides an audio output 1976, such as "Yes, I turned off notifications for mail." In some examples, as shown in FIG. 19C , the digital assistant provides a user interface 1942 (eg, a snippet or window) that enables the user to manipulate the result of performing a task. For example, user interface 1942 provides affordance 1943 (eg, a cancel button). If the user wishes to continue to receive notification of emails, for example, the user can select affordance 1943 to turn notification of emails back on. In some examples, the user may also provide another speech input, such as "Notify me about received emails," to turn on notification of emails.
Referring to FIG. 19D , in some embodiments, the digital assistant may not be able to complete the task based on the user's speech input, and thus may also provide a user interface that enables the user to perform the task. As shown in FIG. 19D , in some examples, the user provides speech input 1958 , such as "Show my custom message in my screensaver." The digital assistant determines that the user's intent is to change the screensaver settings to show a custom message. The digital assistant further determines that the user intent changes the display configuration, and thus the user intent indicates a request to perform a task.
In some embodiments, upon determining that the user intent indicates a request to perform the task, the digital assistant immediately executes a process associated with the user device to perform the task. In some examples, when the digital assistant is unable to complete the task based on user intent, it provides a user interface that enables the user to perform the task. For example, based on the speech input 1958 , the digital assistant may not be able to determine the content of the custom message to be shown on the screen saver, and thus may not be able to complete the task of displaying the custom message. 19D , in some examples, the digital assistant immediately executes the display configuration process and displays a user interface 1946 (eg, a snippet or window) that enables the user to manually change screen saver settings. to provide. For another example, the digital assistant provides a link 1944 (eg, a link to display configurations) that enables the user to perform a task. The user selects the link 1944 using an input device, such as a mouse, finger, or stylus. Upon receiving the user's selection, the digital assistant immediately executes the display configuration process and displays a user interface 1946 that enables the user to change screensaver settings. In some examples, the digital assistant further provides an audio output 1978, such as "You can explore screensaver options in screensaver configuration."
7. The process for operating the digital assistant - intelligent search and object management.
20A-20G show a flow diagram of an example process 2000 for operating a digital assistant in accordance with some embodiments. Process 2000 may be performed using one or more devices 104 , 108 , 200 , 400 , or 600 ( FIGS. 1 , 2A , 4 , or 6A and 6B ). The operations in process 2000 are optionally combined or divided and/or the order of some operations is optionally changed.
Referring to FIG. 20A , at block 2002 , prior to receiving the first speech input, an affordance for invoking a digital assistant service is displayed on a display associated with the user device. At block 2003, the digital assistant is invoked in response to receiving the predetermined phrase. At block 2004, the digital assistant is invoked in response to receiving the selection of affordance.
At block 2006 , a first speech input is received from a user. At block 2008 , context information associated with the user device is identified. At block 2009 , the context information includes at least one of: user specific data, metadata associated with one or more objects, sensor data, and user device configuration data.
At block 2010 , a user intent is determined based on the first speech input and context information. At block 2012, one or more actionable intents are determined to determine a user intent. At block 2013, one or more parameters associated with the actionable intent are determined.
Referring to FIG. 20B , at block 2015 , it is determined whether the user intent is to perform the task using the search process or the object management process to perform the task. The retrieval process is configured to retrieve data stored internally or externally to the user device, and the object management process is configured to manage objects associated with the user device. At block 2016, it is determined whether the speech input includes one or more keywords indicating a search process or one or more keywords indicating an object management process. At block 2018, it is determined whether the task is associated with a search. At block 2020, in accordance with a determination that the task is associated with a search, it is determined whether performing the task requires a search process. At block 2021 , in accordance with a determination that performing the task does not require a search process, a voice request to select a search process or an object management process is output, and a second speech input is received from the user. The second speech input indicates selection of a search process or an object management process.
At block 2022, in accordance with a determination that performing the task does not require a retrieval process, determine whether the task is to be performed using the retrieval process or the object management process based on the predetermined configuration do.
Referring to FIG. 20C , in block 2024 , in accordance with a determination that the task is not associated with a search, it is determined whether the task is associated with managing at least one object. At block 2025 , in accordance with a determination that the task is not associated with managing the at least one object, at least one of the following is performed: determining whether the task can be performed using a fourth process available to the user device and initiating a conversation with the user.
At block 2026, in accordance with a determination that the user intent is to perform the task using the search process, the task is performed using the search process. At block 2028, at least one object is retrieved using a retrieval process. At block 2029, the at least one object includes at least one folder or file. At block 2030, the file includes at least one picture, audio, or video. At block 2031 , the file is stored internally or externally to the user device. At block 2032 , retrieving the at least one folder or file is based on metadata associated with the folder or file. At block 2034, the at least one object includes communication. At block 2035, the communication includes at least one email, message, notification, or voicemail. At block 2036, metadata associated with the communication is retrieved.
Referring to FIG. 20D , at block 2037 , the at least one object includes at least one contact or calendar. At block 2038, the at least one object includes an application. At block 2039, the at least one object includes an online informational source.
At block 2040 , in accordance with a determination that the user intent is to perform the task using the object management process, the task is performed using the object management process. At block 2042 , the task is associated with a search, and the at least one object is retrieved using the object management process. At block 2043, the at least one object includes at least one folder or file. At block 2044 , the file includes at least one picture, audio, or video. At block 2045 , the file is stored internally or externally to the user device. At block 2046 , retrieving the at least one folder or file is based on metadata associated with the folder or file.
At block 2048, the object management process is executed immediately. Immediate execution of the object management process includes invoking the object management process, creating a new instance of the object management process, or running an existing instance of the object management process.
Referring to FIG. 20E , at block 2049 , at least one object is created. At block 2050, at least one object is stored. At block 2051 , at least one object is compressed. At block 2052 , the at least one object is moved from the first physical or virtual storage device to the second physical or virtual storage device. At block 2053, the at least one object is copied from the first physical or virtual storage device to the second physical or virtual storage device. At block 2054, the at least one object stored on the physical or virtual storage device is deleted. At block 2055, at least one object stored in physical or virtual storage is recovered. At block 2056, at least one object is marked. The marking of the at least one object is at least one of being indicated or associated with metadata of the at least one object. At block 2057, the at least one object is backed up according to a predetermined time period for back up. At block 2058 , the at least one object is shared between one or more electronic devices communicatively coupled to the user device.
Referring to FIG. 20F , at block 2060 , a response is provided based on a result of performing the task using a search process or an object management process. At block 2061 , a first user interface is displayed that provides a result of performing a task using a search process or an object management process. At block 2062, a link associated with the result of performing the task using the search process is displayed. At block 2063, an audio output is provided according to a result of performing a task using a search process or an object management process.
At block 2064, an affordance is provided that enables the user to manipulate the results of performing a task using a search process or an object management process. At block 2065, a third process that operates using the result of performing the task is immediately executed.
20F, at block 2066, a confidence level is determined. At block 2067 , the confidence level indicates the accuracy of determining the user intent based on the first speech input and contextual information associated with the user device. At block 2068, the confidence level indicates the accuracy of determining whether the user intent is to perform the task using the search process or the object management process.
Referring to FIG. 20G , at block 2069 , the confidence level indicates the accuracy of performing the task using the search process or object management process.
At block 2070, a response is provided according to the determination of the confidence level. At block 2071 , it is determined whether the confidence level is greater than or equal to a threshold confidence level. At block 2072 , in accordance with a determination that the confidence level is greater than or equal to a threshold confidence level, a first response is provided. At block 2073 , in accordance with a determination that the confidence level is below a threshold confidence level, a second response is provided.
8. The process for operating the digital assistant - continuity.
21A-21E show a flow diagram of an example process 2100 for operating a digital assistant in accordance with some embodiments. Process 2100 may be performed using one or more devices 104 , 108 , 200 , 400 , 600 , 1400 , 1500 , 1600 , or 1700 ( FIGS. 1 , 2A , 4 , 6A and 1700 ). 6B, 14A-14D, 15A-15D, 16A-16C, and 17A-17E). The operations in process 2100 are optionally combined or divided and/or the order of some operations is optionally changed.
Referring to FIG. 21A , at block 2102 , prior to receiving the first speech input, an affordance for invoking a digital assistant service is displayed on a display associated with the user device. At block 2103, the digital assistant is invoked in response to receiving the predetermined phrase. At block 2104, the digital assistant is invoked in response to receiving the selection of affordance.
At block 2106 , a first speech input for performing a task is received from a user. At block 2108 , context information associated with the user device is identified. At block 2109 , the user device is configured to provide a plurality of user interfaces. At block 2110 , the user device comprises a laptop computer, a desktop computer, or a server. At block 2112 , the context information includes at least one of: user specific data, metadata associated with one or more objects, sensor data, and user device configuration data.
At block 2114 , a user intent is determined based on the speech input and context information. At block 2115 , one or more actionable intents are determined to determine a user intent. At block 2116 , one or more parameters associated with the actionable intent are determined.
Referring to FIG. 21B , at block 2118 , according to the user intent, it is determined whether the task is to be performed at the user device or at a first electronic device communicatively coupled to the user device. At block 2120 , the first electronic device includes a laptop computer, desktop computer, server, smartphone, tablet, set-top box, or watch. At block 2121 , determining whether the task is to be performed at the user device or the first electronic device is based on one or more keywords included in the speech input. At block 2122 , it is determined whether performing the task at the user device meets a performance criterion. At block 2123, a performance criterion is determined based on one or more user preferences. At block 2124 , a performance criterion is determined based on the device configuration data. At block 2125, the performance criteria are dynamically updated. At block 2126 , in accordance with a determination that performing the task at the user device meets the performance criteria, it is determined that the task is to be performed at the user device.
Referring to FIG. 21C , in block 2128 , in accordance with a determination that performing the task at the user device does not meet the performance criteria, it is determined whether performing the task at the first electronic device meets the performance criteria. At block 2130 , in accordance with a determination that performing the task at the first electronic device meets the performance criterion, it is determined that the task is to be performed at the first electronic device. At block 2132 , in accordance with a determination that performing the task at the first electronic device does not meet the performance criteria, it is determined whether performing the task at the second electronic device meets the performance criteria.
At block 2134 , content for performing the task is received in accordance with a determination that the task will be performed at the user device and the content for performing the task is located elsewhere. At block 2135 , at least a portion of the content is received from the first electronic device. At least a portion of the content is stored on the first electronic device. At block 2136 , at least a portion of the content is received from a third electronic device.
Referring to FIG. 21D , in block 2138 , in accordance with a determination that the task will be performed at the first electronic device and the content for performing the task is located elsewhere than the first electronic device, the content for performing the task is 1 is provided in an electronic device. At block 2139 , at least a portion of the content is provided from the user device to the first electronic device. At least a portion of the content is stored on the user device. At block 2140 , at least a portion of the content is provided from the fourth electronic device to the first electronic device. At least a portion of the content is stored on the fourth electronic device.
At block 2142 , the task will be performed at the user device. A first response is provided to the user device using the received content. At block 2144 , the task will be performed at the user device. At block 2145 , performing the task at the user device is a continuation of the partially performed task elsewhere than at the user device. At block 2146 , a first user interface associated with the task to be performed is displayed at the user device. At block 2148 , a link associated with the task will be performed at the user device. At block 2150 , a voice output according to the task to be performed is provided at the user device.
Referring to FIG. 21E , at block 2152 , a task will be performed at the first electronic device and a second response is provided at the user device. At block 2154 , the task will be performed at the first electronic device. At block 2156 , the task to be performed at the first electronic device is a continuation of the task performed elsewhere from the first electronic device. At block 2158 , a voice output according to the task to be performed is provided at the first electronic device. At block 2160 , a voice output according to the task to be performed is provided at the first electronic device.
9. Process for operating digital assistant - system configuration management.
22A-22D show a flow diagram of an example process 2200 for operating a digital assistant in accordance with some embodiments. Process 2200 is performed using one or more devices 104 , 108 , 200 , 400 , 600 , or 1800 ( FIGS. 1 , 2A , 4 , 6A and 6B , and 18C and 18D ). can be The operations in process 2200 are optionally combined or divided and/or the order of some operations is optionally changed.
Referring to FIG. 22A , at block 2202 , prior to receiving the speech input, an affordance for invoking a digital assistant service is displayed on a display associated with the user device. At block 2203, the digital assistant is invoked in response to receiving the predetermined phrase. At block 2204, the digital assistant is invoked in response to receiving the selection of affordance.
At block 2206 , speech input is received from a user for managing one or more system configurations of the user device. The user device is configured to simultaneously provide a plurality of user interfaces. At block 2207 , the one or more system configurations of the user device include audio configurations. At block 2208 , the one or more system configurations of the user device include date and time configurations. At block 2209 , the one or more system configurations of the user device include oral configurations. At block 2210 , the one or more system configurations of the user device include display configurations. At block 2211 , the one or more system configurations of the user device include input device configurations. At block 2212 , the one or more system configurations of the user device include network configurations. At block 2213 , the one or more system configurations of the user device include notification configurations.
22B, at block 2214, one or more system configurations of the user device include printer configurations. At block 2215 , the one or more system configurations of the user device include security configurations. At block 2216 , the one or more system configurations of the user device include backup configurations. At block 2217 , the one or more system configurations of the user device include application configurations. At block 2218 , the one or more system configurations of the user device include user interface configurations.
At block 2220 , context information associated with the user device is identified. At block 2223 , the context information includes at least one of: user specific data, device configuration data, and sensor data. At block 2224, user intent is determined based on the speech input and context information. At block 2225 , one or more actionable intents are determined. At block 2226, one or more parameters associated with the actionable intent are determined.
Referring to FIG. 22C , at block 2228 , it is determined whether the user intent represents a request for information or a request to perform a task. At block 2229 , it is determined whether the user intent is to change the system configuration.
At block 2230, in accordance with a determination that the user intent indicates an informational request, a voice response is provided to the informational request. At block 2231, the status of one or more system configurations is obtained according to the request for information. At block 2232, a voice response is provided according to the status of one or more system configurations.
At block 2234 , in addition to providing a voice response to the request for information, a first user interface for providing information according to the status of one or more system configurations is displayed. At block 2236, in addition to providing a voice response to the request for information, a link associated with the request for information is provided.
At block 2238 , in accordance with a determination that the user intent indicates a request to perform the task, the process associated with the user device is immediately executed to perform the task. At block 2239, a task is performed using the process. At block 2240 , a first audio output is provided according to a result of performing the task.
Referring to FIG. 22D , at block 2242 , a second user interface is provided that enables the user to manipulate the result of performing the task. At block 2244, the second user interface includes a link associated with a result of performing the task.
At block 2246 , a third user interface is provided that enables a user to perform a task. At block 2248, the third user interface includes a link that enables the user to perform the task. At block 2250 , a second audio output associated with a third user interface is provided.
10. Electronic Devices - Intelligent Search and Object Management
23 shows FIGS. 8A-8F, 9A-9H, 10A and 10B, 11A-11F, 12A-12D, 13A-13C, 14A-14D, 15A-15D, FIG. , of an electronic device 2300 configured according to the principles of various described examples, including those described with reference to FIGS. 16A-16C , 17A-17E, 18A-18F, and 19A-19D . A functional block diagram is shown. The functional blocks of the device may, optionally, be implemented by hardware, software, or a combination of hardware and software for carrying out the principles of the various described examples. It is understood by those of ordinary skill in the art that the functional blocks described in FIG. 23 may be optionally combined or separated into sub-blocks to implement the principles of various described examples. . Accordingly, the description herein supports, optionally, any possible combination, separation, or additional definition of the functional blocks described herein.
23 , the electronic device 2300 can include a microphone 2302 and a processing unit 2308 . In some examples, processing unit 2308 includes receiving unit 2310, identifying unit 2312, determining unit 2314, performing unit 2316, providing unit 2318, immediate unit 2320, display unit ( 2322 , output unit 2324 , initiation unit 2326 , retrieve unit 2328 , create unit 2330 , execute unit 2332 , create unit 2334 , immediate unit 2335 , storage unit 2336 ), compression unit 2338 , copy unit 2340 , delete unit 2342 , restore unit 2344 , marking unit 2346 , backup unit 2348 , sharing unit 2350 , inactive unit 2352 , and an acquiring unit 2354 .
In some examples, processing unit 2308 receives a first speech input from a user (eg, with receiving unit 2310 ); identify context information associated with the user device (eg, using identification unit 2312 ); and determine user intent based on the first speech input and context information (eg, using the determining unit 2314 ).
In some examples, processing unit 2308 is configured to determine (eg, using determining unit 2314 ) whether the user intent is to perform the task using the search process or the object management process to perform the task. do. The retrieval process is configured to retrieve data stored internally or externally to the user device, and the object management process is configured to manage objects associated with the user device.
In some examples, upon determining that the user intent is to perform the task using the search process, processing unit 2308 is configured to perform the task using the search process (eg, using performing unit 2316 ) . In some examples, upon determining that the user intent is to perform the task using the object management process, processing unit 2308 is configured to perform the task using the object management process (eg, using performing unit 2316 ). is composed
In some examples, prior to receiving the first speech input, processing unit 2308 (eg, using display unit 2322 ) determines, on a display associated with the user device, an affordance for invoking a digital assistant service. configured to display.
In some examples, processing unit 2308 is configured to invoke the digital assistant in response to receiving the predetermined phrase (eg, with call unit 2320 ).
In some examples, processing unit 2308 is configured to invoke the digital assistant in response to receiving the selection of affordance (eg, with call unit 2320 ).
In some examples, processing unit 2308 determines (eg, with determining unit 2314 ) one or more actionable intents; and determine (eg, using the determining unit 2314 ) one or more parameters associated with the actionable intent.
In some examples, the context information includes at least one of: user specific data, metadata associated with one or more objects, sensor data, and user device configuration data.
In some examples, processing unit 2308 is configured to determine (eg, with determining unit 2314 ) whether the speech input includes one or more keywords representative of a search process or one or more keywords representative of an object management process. do.
In some examples, processing unit 2308 is configured to determine (eg, with determining unit 2314 ) whether the task is associated with a search. Upon determining that the task is associated with a search, the processing unit 2308 determines (eg, with the determining unit 2314 ) whether performing the task requires a search process; and according to a determination that the task is not associated with the search, determine (eg, with the determining unit 2314 ) whether the task is associated with managing the at least one object.
In some examples, the task is associated with a search, and upon determining that performing the task does not require a search process, processing unit 2308 (eg, with output unit 2324 ) manages the search process or object output a voice request for selecting a process, and receive (eg, using the receiving unit 2310 ), from the user, a second speech input indicating selection of a search process or an object management process.
In some examples, the task is associated with a retrieval, and upon determining that performing the task does not require a retrieval process, processing unit 2308 (eg, using determining unit 2314 ), in a predetermined configuration based on whether the task is to be performed using the discovery process or the object management process.
In some examples, upon determining that the task is not associated with a search and that the task is not associated with managing the at least one object, the processing unit 2308 (eg, using the performing unit 2316 ) at least of the following: and perform one of: determining (eg, using the determining unit 2314) whether the task can be performed using a fourth process available to the user device; and initiating a conversation with the user (eg, using initiation unit 2326 ).
In some examples, processing unit 2308 is configured to retrieve the at least one object using a retrieval process (eg, with retrieval unit 2328 ).
In some examples, the at least one object includes at least one folder or file. The file includes at least one picture, audio, or video. The file is stored internally or externally on the user device.
In some examples, retrieving the at least one folder or file is based on metadata associated with the folder or file.
In some examples, the at least one object includes communication. The communication includes at least one email, message, notification, or voicemail.
In some examples, processing unit 2308 is configured to retrieve metadata associated with the communication (eg, using search unit 2328 ).
In some examples, the at least one object includes at least one of a contact or a calendar.
In some examples, the at least one object includes an application.
In some examples, the at least one object includes an online informational source.
In some examples, the task is associated with a search, and the processing unit 2308 is configured to retrieve the at least one object using the object management process (eg, using the search unit 2328 ).
In some examples, the at least one object includes at least one folder or file. The file includes at least one picture, audio, or video. The file is stored internally or externally on the user device.
In some examples, retrieving the at least one folder or file is based on metadata associated with the folder or file.
In some examples, processing unit 2308 is configured to immediately execute the object management process (eg, with immediate execution unit 2335 ). Immediate execution of the object management process includes invoking the object management process, creating a new instance of the object management process, or running an existing instance of the object management process.
In some examples, processing unit 2308 is configured to create the at least one object (eg, with creation unit 2334 ).
In some examples, processing unit 2308 is configured to store the at least one object (eg, with storage unit 2336 ).
In some examples, processing unit 2308 is configured to compress the at least one object (eg, with compression unit 2338 ).
In some examples, processing unit 2308 is configured to move (eg, using move unit 2339 ) at least one object from the first physical or virtual storage to the second physical or virtual storage.
In some examples, processing unit 2308 is configured to copy (eg, using copy unit 2340 ) at least one object from the first physical or virtual storage device to the second physical or virtual storage device.
In some examples, processing unit 2308 is configured to delete (eg, with delete unit 2342 ) at least one object stored in physical or virtual storage.
In some examples, processing unit 2308 is configured to restore (eg, using restore unit 2344 ) at least one object stored in physical or virtual storage.
In some examples, processing unit 2308 is configured to mark at least one object (eg, with marking unit 2346 ). The marking of the at least one object is at least one of being indicated or associated with metadata of the at least one object.
In some examples, processing unit 2308 is configured to back up the at least one object according to a predetermined period for backup (eg, with backup unit 2348 ).
In some examples, processing unit 2308 is configured to share at least one object between one or more electronic devices communicatively coupled to the user device (eg, using sharing unit 2350 ).
In some examples, processing unit 2308 is configured to provide a response based on a result of performing the task using a retrieval process or an object management process (eg, using the providing unit 2318 ).
In some examples, processing unit 2308 is configured to display (eg, using display unit 2322 ) a first user interface that provides a result of performing a task using a search process or an object management process.
In some examples, the processing unit 2308 is configured to provide a link associated with a result of performing the task using the search process (eg, using the providing unit 2318 ).
In some examples, processing unit 2308 is configured to provide an audio output according to a result of performing a task using a search process or an object management process (eg, using the providing unit 2318 ).
In some examples, processing unit 2308 is configured to provide an affordance that enables a user (eg, using providing unit 2318 ) to manipulate a result of performing a task using a search process or an object management process. .
In some examples, processing unit 2308 is configured to immediately execute a third process that operates using the result of performing the task (eg, with immediate execution unit 2335 ).
In some examples, processing unit 2308 determines a confidence level (eg, using determining unit 2314 ); and provide a response according to the determination of the confidence level (eg, using the providing unit 2318 ).
In some examples, the confidence level indicates the accuracy of determining a user intent based on the first speech input and contextual information associated with the user device.
In some examples, the confidence level indicates the accuracy of determining whether the user's intent is to perform the task using the search process or the object management process.
In some examples, the level of confidence indicates the correctness of performing a task using a search process or an object management process.
In some examples, processing unit 2308 is configured to determine (eg, with determining unit 2314 ) whether the confidence level is greater than or equal to a threshold confidence level. Upon determining that the confidence level is greater than or equal to the threshold confidence level, the processing unit 2308 is configured to provide a first response (eg, using the providing unit 2318 ); Upon determining that the confidence level is below the threshold confidence level, the processing unit 2308 is configured to provide a second response (eg, with the providing unit 2318 ).
11. Electronic Devices - Continuity
In some examples, processing unit 2308 receives (eg, with receiving unit 2310 ) speech input to perform a task from a user; identify context information associated with the user device (eg, using identification unit 2312 ); and determine user intent based on the speech input and context information associated with the user device (eg, with the determining unit 2314 ).
In some examples, processing unit 2308 determines whether the task is to be performed at the user device (eg, using determining unit 2314 ) or at a first electronic device communicatively coupled to the user device, depending on the user intent. to decide whether it will be
In some examples, upon determining that the task will be performed at the user device and the content for performing the task is located elsewhere, processing unit 2308 is configured to perform the task (eg, with receiving unit 2310 ). configured to receive content.
In some examples, upon determining that the task will be performed at the first electronic device and the content for performing the task is located elsewhere than the first electronic device, the processing unit 2308 (eg, using the providing unit 2318 ) ) to provide content for performing the task to the first electronic device.
In some examples, the user device is configured to provide a plurality of user interfaces.
In some examples, the user device comprises a laptop computer, desktop computer, or server.
In some examples, the first electronic device comprises a laptop computer, desktop computer, server, smartphone, tablet, set-top box, or watch.
In some examples, processing unit 2308 is configured to, prior to receiving the speech input, display (eg, using display unit 2322 ), on a display of the user device, an affordance for invoking the digital assistant. do.
In some examples, processing unit 2308 is configured to invoke the digital assistant in response to receiving the predetermined phrase (eg, with call unit 2320 ).
In some examples, processing unit 2308 is configured to invoke the digital assistant in response to receiving the selection of affordance (eg, with call unit 2320 ).
In some examples, processing unit 2308 determines (eg, with determining unit 2314 ) one or more actionable intents; and determine (eg, using the determining unit 2314 ) one or more parameters associated with the actionable intent.
In some examples, the context information includes at least one of: user specific data, sensor data, and user device configuration data.
In some examples, determining whether the task is to be performed at the user device or the first electronic device is based on one or more keywords included in the speech input.
In some examples, processing unit 2308 is configured to determine (eg, with determining unit 2314 ) whether performing the task at the user device meets a performance criterion.
In some examples, upon determining that performing the task at the user device meets the performance criterion, processing unit 2308 is configured to determine (eg, with determining unit 2314 ) that the task is to be performed at the user device do.
In some examples, upon determining that performing the task at the user device does not meet the performance criterion, the processing unit 2308 determines (eg, with the determining unit 2314 ) that performing the task at the first electronic device is is configured to determine whether performance criteria are met.
In some examples, upon determining that performing the task at the first electronic device meets the performance criterion, the processing unit 2308 determines (eg, using the determination 2314 ) that the task is to be performed at the first electronic device. configured to decide.
In some examples, upon determining that performing the task at the first electronic device does not meet the performance criterion, the processing unit 2308 performs the task at the second electronic device (eg, using the determining unit 2314 ) is configured to determine whether the performance meets performance criteria.
In some examples, the performance criterion is determined based on one or more user preferences.
In some examples, the performance criterion is determined based on device configuration data.
In some examples, the performance criterion is dynamically updated.
In some examples, upon determining that the task will be performed at the user device and the content for performing the task is located elsewhere, the processing unit 2308 (eg, using the receiving unit 2310 ) generates at least a portion of the content. configured to receive from the first electronic device, wherein at least a portion of the content is stored on the first electronic device.
In some examples, upon determining that the task will be performed at the user device and the content for performing the task is located elsewhere, the processing unit 2308 (eg, with the receiving unit 2310 ) generates at least a portion of the content. and receive from a third electronic device.
In some examples, upon determining that the task will be performed at the first electronic device and the content for performing the task is located elsewhere than the first electronic device, the processing unit 2308 (eg, using the providing unit 2318 ) ) to provide at least a portion of the content from the user device to the first electronic device, wherein the at least a portion of the content is stored in the user device.
In some examples, upon determining that the task will be performed at the first electronic device and the content for performing the task is located elsewhere than the first electronic device, the processing unit 2308 (eg, using the ministry unit 2352 ) ) to cause at least a portion of the content to be provided from the fourth electronic device to the first electronic device. At least a portion of the content is stored on the fourth electronic device.
In some examples, the task will be performed at the user device, and the processing unit 2308 is configured to provide the first response using the content received at the user device (eg, with the providing unit 2318 ).
In some examples, processing unit 2308 is configured to perform a task at the user device (eg, with performing unit 2316 ).
In some examples, performing the task at the user device is a continuation of the partially performed task elsewhere than at the user device.
In some examples, processing unit 2308 is configured to display (eg, using display unit 2322 ) a first user interface associated with the task to be performed at the user device.
In some examples, processing unit 2308 is configured to provide a link associated with a task to be performed at the user device (eg, with providing unit 2318 ).
In some examples, processing unit 2308 is configured (eg, with providing unit 2318 ) to provide a voice output according to the task to be performed at the user device.
In some examples, the task will be performed at the first electronic device, and the processing unit 2308 is configured to provide the second response at the user device (eg, with the providing unit 2318 ).
In some examples, processing unit 2308 is configured to cause a task to be performed at the first electronic device (eg, using ministry unit 2352 ).
In some examples, the task to be performed at the first electronic device is a continuation of the task performed elsewhere than the first electronic device.
In some examples, the processing unit 2308 is configured (eg, using the providing unit 2318 ) to provide an audio output according to the task to be performed at the first electronic device.
In some examples, processing unit 2308 is configured (eg, with providing unit 2318 ) to provide an affordance that enables a user to select another electronic device for performance of the task.
12. Electronic Devices - System Configuration Management
In some examples, processing unit 2308 is configured to receive (eg, with receiving unit 2310 ) speech input from a user for managing one or more system configurations of the user device. The user device is configured to simultaneously provide a plurality of user interfaces.
In some examples, processing unit 2308 identifies context information associated with the user device (eg, using identification unit 2312 ); and determine user intent based on the speech input and context information (eg, with the determining unit 2314 ).
In some examples, processing unit 2308 is configured to determine (eg, with determining unit 2314 ) whether the user intent indicates a request for information or a request to perform a task.
In some examples, upon determining that the user intent indicates a request for information, processing unit 2308 is configured to provide a voice response to the request for information (eg, with providing unit 2318 ).
In some examples, upon determining that the user intent represents a request to perform a task, processing unit 2308 executes a process associated with the user device to perform the task (eg, using immediate unit 2335 ). It is configured to run immediately.
In some examples, processing unit 2308 is configured to, prior to receiving the speech input, display (eg, using display unit 2322 ), on a display of the user device, an affordance for invoking the digital assistant. do.
In some examples, processing unit 2308 is configured to invoke the digital assistant service in response to receiving the predetermined phrase (eg, using call unit 2320 ).
In some examples, processing unit 2308 is configured to invoke the digital assistant service in response to receiving the selection of affordance (eg, using call unit 2320 ).
In some examples, the one or more system configuration of the user device includes audio configurations.
In some examples, the one or more system configuration of the user device includes date and time configurations.
In some examples, the one or more system configuration of the user device includes oral configurations.
In some examples, the one or more system configuration of the user device includes display configurations.
In some examples, the one or more system configuration of the user device includes input device configurations.
In some examples, the one or more system configuration of the user device includes network configurations.
In some examples, the one or more system configuration of the user device includes notification configurations.
In some examples, the one or more system configuration of the user device includes printer configurations.
In some examples, the one or more system configuration of the user device includes security configurations.
In some examples, the one or more system configuration of the user device includes backup configurations.
In some examples, the one or more system configuration of the user device includes application configurations.
In some examples, the one or more system configuration of the user device includes user interface configurations.
In some examples, processing unit 2308 determines (eg, with determining unit 2314 ) one or more actionable intents; and determine (eg, using the determining unit 2314 ) one or more parameters associated with the actionable intent.
In some examples, the context information includes at least one of: user specific data, device configuration data, and sensor data.
In some examples, the processing unit 2308 is configured to determine (eg, with the determining unit 2314 ) whether the user intent is to change the system configuration.
In some examples, the processing unit 2308 obtains (eg, using the obtaining unit 2354 ) the status of one or more system configurations according to the informational request; and provide a voice response according to the status of one or more system configurations (eg, using the providing unit 2318 ).
In some examples, upon determining that the user intent indicates a request for information, processing unit 2308 is configured to, in addition to providing a voice response to the request for information, (eg, using display unit 2322 ) and display a first user interface that provides information according to the status of one or more system components.
In some examples, upon determining that the user intent indicates a request for information, processing unit 2308 is configured to, in addition to providing a voice response to the request for information, (eg, using providing unit 2318 ) and provide a link associated with the request for information.
In some examples, upon determining that the user intent indicates a request to perform the task, the processing unit 2308 is configured to perform the task using the process (eg, with the performing unit 2316 ).
In some examples, processing unit 2308 is configured to provide a first audio output according to a result of performing the task (eg, with providing unit 2318 ).
In some examples, the processing unit 2308 is configured to provide a second user interface that enables the user to manipulate a result of performing the task (eg, with the providing unit 2318 ).
In some examples, the second user interface includes a link associated with a result of performing the task.
In some examples, upon determining that the user intent indicates a request to perform the task, the processing unit 2308 (eg, with the providing unit 2318 ) enables the third user to perform the task. configured to provide an interface.
In some examples, the third user interface includes a link that enables the user to perform the task.
In some examples, processing unit 2308 is configured to provide a second audio output associated with the third user interface (eg, with providing unit 2318 ).
The operation described above with respect to FIG. 23 is optionally implemented by the components shown in FIGS. 1 , 2A, 4, 6A and 6B, or FIGS. 7A and 7B . For example, receive operation 2310 , identify operation 2312 , determine operation 2314 , perform operation 2316 , and provide operation 2318 are optionally implemented by processor(s) 220 . It will be apparent to those skilled in the art how other processes may be implemented based on the components shown in FIGS. 1 , 2A, 4, 6A and 6B, or in FIGS. 7A and 7B .
Those skilled in the art will appreciate that the functional blocks described in FIG. 12 may optionally be combined or separated into sub-blocks to implement the principles of various described embodiments. Accordingly, the description herein optionally supports any possible combination or separation or additional definition of the functional blocks described herein. For example, processing unit 2308 may have an associated "controller" unit operatively coupled with processing unit 2308 to enable operation. Although such a controller unit is not separately shown in FIG. 23 , it will be within the purview of one of ordinary skill in the art designing devices having a processing unit 2308 , such as device 2300 . As another example, one or more units, such as receive unit 2310 , may be hardware units external to processing unit 2308 in some embodiments. Accordingly, the description herein optionally supports a combination, separation, and/or further definition of the functional blocks described herein.
An example method, non-transitory computer-readable storage medium, system, and electronic device are described in the following items:
<u>Activate your digital assistant - intelligent search and object management</u>
One. A method for providing a digital assistant service, comprising:
In a user device having one or more processors and memory:
receiving a first speech input from a user;
identifying context information associated with the user device;
determining user intent based on the first speech input and context information;
determining whether the user intent is to perform the task using the retrieval process or the object management process to perform the task, the retrieval process configured to retrieve data stored internally or externally on the user device, the object the management process is configured to manage objects associated with the user device;
performing the task using the search process in accordance with a determination that the user intent is to perform the task using the search process; and
In accordance with a determination that the user intent is to perform the task using the object management process, performing the task using the object management process.
2. The method of item 1, prior to receiving the first speech input:
and displaying, on a display associated with the user device, an affordance for invoking a digital assistant service.
3. Item 2,
and invoking the digital assistant in response to receiving the predetermined phrase.
4. according to items 2 or 3,
and invoking the digital assistant in response to receiving the selection of affordance.
5. The method of any one of items 1-4, wherein determining user intent comprises:
determining one or more actionable intentions; and
A method comprising determining one or more parameters associated with an actionable intent.
6. The method of any one of items 1-5, wherein the context information comprises at least one of: user specific data, metadata associated with one or more objects, sensor data, and user device configuration data.
7. The method of any one of items 1 to 6, wherein determining whether the user intent is to perform the task using the search process or the object management process comprises:
and determining whether the speech input includes one or more keywords indicative of a search process or one or more keywords indicative of an object management process.
8. The method of any one of items 1 to 7, wherein determining whether the user intent is to perform the task using the search process or the object management process comprises:
determining whether the task is associated with a search;
in accordance with a determination that the task is associated with a search, determining whether performing the task requires a search process; and
and in accordance with determining that the task is not associated with a search, determining whether the task is associated with managing at least one object.
9. The method of item 8, wherein the task is associated with a search,
In accordance with the determination that performing the task does not require a discovery process,
outputting a voice request to select a search process or an object management process; and
and receiving, from the user, a second speech input indicative of a selection of a search process or an object management process.
10. Items 8 or 9, wherein the task is associated with a search and
according to a determination that performing the task does not require a retrieval process, determining whether the task is to be performed using the retrieval process or the object management process based on the predetermined configuration How to.
11. Item 8, wherein the task is not associated with a search and
In accordance with a determination that the task is not associated with managing at least one object,
determining whether the task can be performed using a fourth process available to the user device; and
The method further comprising performing at least one of initiating a conversation with the user.
12. The method of any one of items 1-10, wherein performing the task using the search process comprises:
retrieving the at least one object using a retrieval process.
13. The method of item 12, wherein the at least one object comprises at least one folder or file.
14. The method of item 13, wherein the file comprises at least one photo, audio, or video.
15. The method according to items 13 or 14, wherein the file is stored internally or externally to the user device.
16. The method of any one of items 13-15, wherein retrieving the at least one folder or file is based on metadata associated with the folder or file.
17. The method of any one of items 12-16, wherein the at least one object comprises communication.
18. The method of item 17, wherein the communication comprises at least one email, message, notification, or voicemail.
19. The method of items 17 or 18, further comprising retrieving metadata associated with the communication.
20. The method of any one of items 12-19, wherein the at least one object comprises at least one contact or calendar.
21. The method of any one of items 12-20, wherein the at least one object comprises an application.
22. The method of any one of items 12-21, wherein the at least one object comprises an online informational source.
23. The method of any one of items 1-22, wherein the task is associated with a search and performing the task using the object management process comprises:
retrieving the at least one object using an object management process.
24. The method of item 23, wherein the at least one object comprises at least one folder or file.
25. The method of item 24, wherein the file comprises at least one photo, audio, or video.
26. The method of items 24 or 25, wherein the file is stored internally or externally on the user device.
27. The method of any one of items 24-26, wherein retrieving the at least one folder or file is based on metadata associated with the folder or file.
28. The method of item 1, wherein performing the task using the object management process comprises immediately executing the object management process, and immediately executing the object management process comprises invoking the object management process; A method comprising creating a new instance, or running an existing instance of an object management process.
29. The method of any one of items 1-28, wherein performing the task using the object management process comprises creating at least one object.
30. The method of any one of items 1-29, wherein performing the task using the object management process comprises storing the at least one object.
31. The method of any one of items 1-30, wherein performing the task using the object management process comprises compressing the at least one object.
32. The method of any one of items 1-31, wherein performing the task using the object management process comprises moving the at least one object from the first physical or virtual storage device to the second physical or virtual storage device. How to.
33. The method of any one of items 1-32, wherein performing the task using the object management process comprises copying the at least one object from the first physical or virtual storage device to the second physical or virtual storage device. How to.
34. The method of any one of items 1-33, wherein performing the task using the object management process comprises deleting at least one object stored in physical or virtual storage.
35. The method of any one of items 1-34, wherein performing the task using the object management process comprises recovering at least one object stored on physical or virtual storage.
36. The method according to any one of items 1 to 35, wherein performing the task using the object management process comprises marking at least one object, the marking of the at least one object being displayed or the at least one object at least one of associated with the metadata of
37. The method of any one of items 1-36, wherein performing the task using the object management process comprises backing up the at least one object according to a predetermined time period for backing up.
38. The method of any one of items 1-37, wherein performing the task using the object management process comprises sharing at least one object between one or more electronic devices communicatively coupled to the user device. .
39. The method of any one of items 1-38, further comprising providing a response based on a result of performing the task using a search process or an object management process.
40. The method of item 39, wherein providing a response based on a result of performing the task using a search process or an object management process comprises:
and displaying a first user interface providing a result of performing the task using the search process or the object management process.
41. The method of items 39 or 40, wherein providing the response based on a result of performing the task using the search process or the object management process comprises:
and providing a link associated with a result of performing the task using the search process.
42. The method of any one of items 39-41, wherein providing a response based on a result of performing the task using a search process or an object management process comprises:
and providing an audio output according to a result of performing a task using a search process or an object management process.
43. The method of any one of items 39-42, wherein providing the response based on a result of performing the task using the search process or the object management process comprises:
A method, comprising: providing an affordance that enables a user to manipulate a result of performing a task using a search process or an object management process.
44. The method of any one of items 39 to 43, wherein providing the response based on a result of performing the task using the search process or the object management process comprises:
and immediately executing a third process that operates using a result of performing the task.
45. The method of any one of items 39-44, wherein providing a response based on a result of performing the task using a search process or an object management process comprises:
determining a confidence level; and
and providing a response according to the determination of the confidence level.
46. The method of item 45, wherein the level of confidence indicates accuracy in determining user intent based on context information associated with the first speech input and the user device.
47. The method of items 45 or 46, wherein the confidence level indicates the accuracy of determining whether the user intent is to perform the task using the search process or the object management process.
48. The method of any one of items 45-47, wherein the confidence level indicates the accuracy of performing the task using a search process or an object management process.
49. The method of any one of items 45-48, wherein providing a response according to the determination of the confidence level comprises:
determining whether the confidence level is greater than or equal to a threshold confidence level;
in response to determining that the confidence level is greater than or equal to a threshold confidence level, providing a first response; and
and in response to determining that the confidence level is below the threshold confidence level, providing a second response.
50. A non-transitory computer-readable storage medium storing one or more programs, wherein the one or more programs, when executed by one or more processors of the electronic device, cause the electronic device to:
receive a first speech input from a user;
identify context information associated with the user device;
determine user intent based on the first speech input and context information;
determine whether the user intent is to perform the task using the retrieval process or the object management process to perform the task, the retrieval process being configured to retrieve data stored internally or externally on the user device; the process is configured to manage objects associated with the user device;
in accordance with a determination that the user intent is to perform the task using the search process, perform the task using the search process;
A non-transitory computer-readable storage medium comprising instructions to, in accordance with a determination that a user intent is to perform a task using the object management process, perform the task using the object management process.
51. An electronic device comprising:
one or more processors;
Memory; and
one or more programs stored in the memory, the one or more programs comprising:
receive a first speech input from a user;
identify context information associated with the user device;
determine user intent based on the first speech input and context information;
determine whether the user intent is to perform the task using the retrieval process or the object management process to perform the task, the retrieval process being configured to retrieve data stored internally or externally on the user device; the process is configured to manage objects associated with the user device;
in accordance with a determination that the user intent is to perform the task using the search process, perform the task using the search process;
and in accordance with a determination that a user intent is to perform the task using the object management process, the electronic device comprising instructions for performing the task using the object management process.
52. An electronic device comprising:
means for receiving a first speech input from a user;
means for identifying context information associated with the user device;
means for determining user intent based on the first speech input and context information;
means for determining whether the user intent is to perform the task using the retrieval process or the object management process, the retrieval process being configured to retrieve data stored internally or externally on the user device; the object management process is configured to manage objects associated with the user device;
means for performing the task using the search process in accordance with a determination that the user intent is to perform the task using the search process; and
means for performing the task using the object management process in accordance with a determination that the user intent is to perform the task using the object management process.
53. An electronic device comprising:
one or more processors;
Memory; and
An electronic device comprising one or more programs stored in the memory, the one or more programs comprising instructions for performing the method of any one of items 1 to 49.
54. An electronic device comprising:
An electronic device comprising means for performing the method of any one of items 1 to 49.
55. A non-transitory computer-readable storage medium comprising one or more programs for execution by one or more processors of an electronic device, the one or more programs, when executed by the one or more processors, cause the electronic device to: A non-transitory computer-readable storage medium comprising instructions for performing the method of any one of the items.
56. A system for operating a digital assistant, comprising means for performing the method of any one of items 1-49.
57. An electronic device comprising:
a receiving unit configured to receive a first speech input from a user; and
A processing unit comprising:
identify context information associated with the user device;
determine user intent based on the first speech input and context information;
determine whether the user intent is to perform the task using the retrieval process or the object management process to perform the task, the retrieval process being configured to retrieve data stored internally or externally on the user device; the process is configured to manage objects associated with the user device;
in accordance with a determination that the user intent is to perform the task using the search process, perform the task using the search process;
according to a determination that the user intent is to perform the task using the object management process, the electronic device configured to perform the task using the object management process.
58. The method of item 57, prior to receiving the first speech input:
The electronic device further comprising displaying, on a display associated with the user device, an affordance for invoking a digital assistant service.
59. Item 58,
The electronic device further comprising: immediately executing the digital assistant in response to receiving the predetermined phrase.
60. according to items 58 or 59,
and immediately executing the digital assistant in response to receiving the selection of affordance.
61. The method of any one of items 57-60, wherein determining user intent comprises:
determining one or more actionable intentions; and
An electronic device comprising determining one or more parameters associated with an actionable intent.
62. The electronic device of any one of items 57-61, wherein the context information comprises at least one of user specific data, metadata associated with one or more objects, sensor data, and user device configuration data.
63. The method of any one of items 57-62, wherein determining whether the user intent is to perform the task using the search process or the object management process comprises:
and determining whether the speech input includes one or more keywords indicative of a search process or one or more keywords indicative of an object management process.
64. The method of any one of items 57 to 63, wherein determining whether the user intent is to perform the task using the search process or the object management process comprises:
determining whether a task is associated with a search;
in accordance with a determination that the task is associated with a search, determining whether performing the task requires a search process;
and in accordance with determining that the task is not associated with a search, determining whether the task is associated with managing at least one object.
65. The method of item 64, wherein the task is associated with a search,
In accordance with the determination that performing the task does not require a discovery process,
outputting a voice request to select a search process or an object management process; and
and receiving, from the user, a second speech input indicative of a selection of a search process or an object management process.
66. The method of items 64 or 65, wherein the task is associated with a search;
further comprising, in accordance with a determination that performing the task does not require a retrieval process, determining, based on the predetermined configuration, whether the task will be performed using the retrieval process or the object management process , electronic devices.
67. The method of item 64, wherein the task is not associated with a search;
In accordance with a determination that the task is not associated with managing at least one object,
determining whether the task can be performed using a fourth process available to the user device; and
The electronic device further comprising performing at least one of initiating a conversation with the user.
68. The method of any one of items 57-66, wherein performing the task using the search process comprises:
and retrieving the at least one object using a retrieval process.
69. The electronic device of item 68, wherein the at least one object comprises at least one folder or file.
70. The electronic device of item 69, wherein the file comprises at least one photo, audio, or video.
71. The electronic device of items 69 or 70, wherein the file is stored internally or externally to the user device.
72. The electronic device of any one of items 69-71, wherein retrieving the at least one folder or file is based on metadata associated with the folder or file.
73. The electronic device of any one of items 68-72, wherein the at least one object comprises communication.
74. The electronic device of item 73, wherein the communication comprises at least one email, message, notification, or voicemail.
75. The electronic device of items 73 or 74, further comprising retrieving metadata associated with the communication.
76. The electronic device of any one of items 68-75, wherein the at least one object comprises at least one contact or calendar.
77. The electronic device of any of items 68-76, wherein the at least one object comprises an application.
78. The electronic device of any one of items 68-77, wherein the at least one object comprises an online informational source.
79. The method of any one of items 57-78, wherein the task is associated with a search and performing the task using the object management process comprises:
and retrieving the at least one object using an object management process.
80. The electronic device of item 79, wherein the at least one object comprises at least one folder or file.
81. The electronic device of item 80, wherein the file comprises at least one photo, audio, or video.
82. The electronic device of any of items 79-81, wherein the file is stored internally or externally to the user device.
83. The electronic device of any of items 79-82, wherein retrieving the at least one folder or file is based on metadata associated with the folder or file.
84. The method of item 57, wherein performing the task using the object management process comprises immediately executing the object management process, and immediately executing the object management process comprises invoking the object management process, creating a new instance of the object management process. An electronic device comprising creating, or executing an existing instance of an object management process.
85. The electronic device of any one of items 57-84, wherein performing the task using the object management process comprises creating at least one object.
86. The electronic device of any one of items 57-85, wherein performing the task using the object management process comprises storing the at least one object.
87. The electronic device of any one of items 57-86, wherein performing the task using the object management process comprises compressing the at least one object.
88. The method of any one of items 57-87, wherein performing the task using the object management process comprises moving at least one object from the first physical or virtual storage device to the second physical or virtual storage device. electronic device.
89. The method of any one of items 57-88, wherein performing the task using the object management process comprises copying at least one object from the first physical or virtual storage device to the second physical or virtual storage device. electronic device.
90. The electronic device of any one of items 57-89, wherein performing the task using the object management process comprises deleting at least one object stored in physical or virtual storage.
91. The electronic device of any one of items 57-90, wherein performing the task using the object management process comprises recovering at least one object stored in physical or virtual storage.
92. The method according to any one of items 57 to 91, wherein performing the task using the object management process comprises marking at least one object, the marking of the at least one object being displayed or metaphor of the at least one object. at least one of being associated with data.
93. The electronic device of any one of items 57-92, wherein performing the task using the object management process comprises backing up the at least one object according to a predetermined time period for backing up.
94. The electronic device of any one of items 57-93, wherein performing the task using the object management process comprises sharing at least one object between one or more electronic devices communicatively coupled to the user device.
95. The electronic device of any one of items 57-94, further comprising providing a response based on a result of performing the task using a search process or an object management process.
96. The method of item 95, wherein providing the response based on a result of performing the task using the search process or the object management process comprises:
and displaying a first user interface that provides a result of performing a task using a search process or an object management process.
97. The method of items 95 or 96, wherein providing the response based on a result of performing the task using the search process or the object management process comprises:
and providing a link associated with a result of performing a task using the search process.
98. The method of any one of items 95-97, wherein providing the response based on a result of performing the task using the search process or the object management process comprises:
An electronic device comprising: providing an audio output according to a result of performing a task using a search process or an object management process.
99. The method of any one of items 95-98, wherein providing the response based on a result of performing the task using the search process or the object management process comprises:
and providing an affordance that enables a user to manipulate a result of performing a task using a search process or an object management process.
100. The method of any one of items 95-99, wherein providing the response based on a result of performing the task using a search process or an object management process comprises:
and immediately executing a third process that operates using a result of performing the task.
101. The method of any one of items 95-100, wherein providing the response based on a result of performing the task using the search process or object management process comprises:
determining the level of confidence; and
and providing a response according to the determination of the level of trust.
102. The electronic device of item 101 , wherein the confidence level indicates accuracy in determining user intent based on the first speech input and context information associated with the user device.
103. The electronic device of items 101 or 102, wherein the confidence level indicates the accuracy of determining whether the user intent is to perform the task using the search process or the object management process.
104. The electronic device of any one of items 101 - 103, wherein the confidence level indicates the accuracy of performing a task using a search process or an object management process.
105. The method of any one of items 101-104, wherein providing a response according to the determination of the confidence level comprises:
determining whether the confidence level is above a threshold confidence level;
in response to determining that the confidence level is greater than or equal to a threshold confidence level, providing a first response; and
and in response to determining that the confidence level is below the threshold confidence level, providing a second response.
<u>continuity</u>
One. A method for providing a digital assistant service, comprising:
In a user device having one or more processors and memory:
receiving speech input for performing a task from a user;
identifying context information associated with the user device;
determining user intent based on the speech input and context information associated with the user device;
determining, according to the user intent, whether the task is to be performed at the user device or at a first electronic device communicatively coupled to the user device;
receiving content for performing the task in accordance with a determination that the task will be performed at the user device and content for performing the task is located elsewhere; and
according to a determination that the task is to be performed at the first electronic device and the content for performing the task is located elsewhere than the first electronic device, the method comprising: providing content for performing the task to the first electronic device; .
2. The method of item 1 , wherein the user device is configured to provide a plurality of user interfaces.
3. The method of items 1 or 2, wherein the user device comprises a laptop computer, a desktop computer, or a server.
4. The method of any one of items 1-3, wherein the first electronic device comprises a laptop computer, a desktop computer, a server, a smartphone, a tablet, a set top box, or a watch.
5. The method of any one of items 1-4, prior to receiving the speech input:
The method of claim 1, further comprising displaying, on a display of the user device, an affordance for calling the digital assistant.
6. Item 5, wherein
and immediately executing the digital assistant in response to receiving the predetermined phrase.
7. according to items 5 or 6,
and immediately executing the digital assistant in response to receiving the selection of affordance.
8. The method of any one of items 1 to 7, wherein determining user intent comprises:
determining one or more actionable intentions; and
A method comprising determining one or more parameters associated with an actionable intent.
9. The method of any one of items 1-8, wherein the context information comprises at least one of user specific data, sensor data, and user device configuration data.
10. The method of any of items 1-9, wherein determining whether the task is to be performed at the user device or the first electronic device is based on one or more keywords included in the speech input.
11. The method of any one of items 1 to 10, wherein determining whether the task is to be performed at the user device or the first electronic device comprises:
determining whether performing the task at the user device meets a performance criterion; and
according to determining that performing the task at the user device meets the performance criterion, determining that the task will be performed at the user device.
12. The method of item 11,
Upon determining that performing a task on the user's device does not meet the performance criteria:
The method further comprising determining whether performing the task at the first electronic device meets a performance criterion.
13. The method of item 12,
determining that performing the task at the first electronic device meets the performance criterion, determining that the task will be performed at the first electronic device; and
In accordance with the determination that performing the task at the first electronic device does not meet the performance criteria:
The method further comprising determining whether performing the task at the second electronic device meets the performance criterion.
14. The method of any one of items 11-13, wherein the performance criterion is determined based on one or more user preferences.
15. The method according to any one of items 11 to 14, wherein the performance criterion is determined based on device configuration data.
16. The method according to any one of items 11 to 15, wherein the performance criterion is dynamically updated.
17. The method of any one of items 1-11 and 14-16, wherein in accordance with a determination that the task is to be performed on the user device and the content for performing the task is located elsewhere, receiving the content for performing the task comprises: :
A method, comprising: receiving at least a portion of content from a first electronic device, wherein at least a portion of the content is stored on the first electronic device.
18. The method of any one of items 1-11 and 14-17, wherein in accordance with a determination that the task is to be performed on the user device and the content for performing the task is located elsewhere, the step of receiving the content for performing the task comprises: :
and receiving at least a portion of the content from a third electronic device.
19. The content for performing the task according to any one of items 1 to 16, wherein in accordance with a determination that the task is to be performed at the first electronic device and the content for performing the task is located in a different location than the first electronic device, the content for performing the task is provided. 1 providing to the electronic device comprises:
providing at least a portion of the content from the user device to the first electronic device, wherein the at least a portion of the content is stored on the user device.
20. The content for performing the task according to any one of items 1 to 16 and 19, in accordance with a determination that the task will be performed at the first electronic device and the content for performing the task is located elsewhere than the first electronic device providing to the first electronic device includes:
causing at least a portion of the content to be provided from the fourth electronic device to the first electronic device, wherein the at least a portion of the content is stored on the fourth electronic device.
21. The method of any one of items 1-11 and 14-18, wherein the task is to be performed at the user device and further comprising providing a first response using the content received at the user device.
22. The method of item 21, wherein providing the first response at the user device comprises:
A method comprising performing a task at a user device.
23. The method of item 22, wherein performing the task at the user device is a continuation of the partially performed task elsewhere than at the user device.
24. The method of any one of items 21-23, wherein providing the first response at the user device comprises:
and displaying a first user interface associated with the task to be performed at the user device.
25. The method of any one of items 21-24, wherein providing the first response at the user device comprises:
A method comprising providing a link associated with a task to be performed at a user device.
26. The method of any one of items 21-25, wherein providing the first response at the user device comprises:
and providing a voice output according to a task to be performed at the user device.
27. The method of any one of items 1-16 and 19 and 20, wherein the task is to be performed at the first electronic device and further comprising providing a second response at the user device.
28. The method of item 27, wherein providing the second response at the user device comprises:
causing the task to be performed at the first electronic device.
29. The method of item 28, wherein the task to be performed at the first electronic device is a continuation of the task performed elsewhere than the first electronic device.
30. The method of any one of items 27-29, wherein providing the second response at the user device comprises:
A method comprising: providing a voice output according to a task to be performed at a first electronic device.
31. The method of any one of items 27-30, wherein providing the second response at the user device comprises:
A method comprising: providing an affordance that enables a user to select another electronic device for performance of a task.
32. A non-transitory computer-readable storage medium storing one or more programs, wherein the one or more programs, when executed by one or more processors of the electronic device, cause the electronic device to:
receive speech input for performing a task from the user;
identify context information associated with the user device;
determine user intent based on the speech input and context information associated with the user device;
determine, according to the user intent, whether the task is to be performed at the user device or at a first electronic device communicatively coupled to the user device;
receive content for performing the task according to a determination that the task will be performed at the user device and content for performing the task is located elsewhere;
according to a determination that the task will be performed at the first electronic device and content for performing the task is located elsewhere than the first electronic device, comprising instructions to cause providing content for performing the task to the first electronic device. A non-transitory computer-readable storage medium.
33. As a user device:
one or more processors;
Memory; and
one or more programs stored in the memory, the one or more programs comprising:
receive speech input for performing a task from the user;
identify context information associated with the user device;
determine user intent based on the speech input and context information associated with the user device;
determine, according to the user intent, whether the task is to be performed at the user device or at a first electronic device communicatively coupled to the user device;
receive content for performing the task according to a determination that the task will be performed at the user device and content for performing the task is located elsewhere;
according to a determination that the task will be performed at the first electronic device and content for performing the task is located elsewhere than the first electronic device, comprising instructions for providing the first electronic device with content for performing the task. user device.
34. A user device comprising:
means for receiving speech input for performing a task from a user;
means for identifying context information associated with the user device;
means for determining user intent based on the speech input and context information associated with the user device;
means for determining, according to the user intent, whether the task is to be performed at the user device or at a first electronic device communicatively coupled to the user device;
means for receiving content for performing the task in accordance with a determination that the task will be performed at the user device and the content for performing the task is located elsewhere;
in accordance with a determination that the task will be performed at the first electronic device and content for performing the task is located elsewhere than the first electronic device, comprising: means for causing the first electronic device to provide content for performing the task to the first electronic device , user device.
35. A user device comprising:
one or more processors;
Memory; and
A user device comprising one or more programs stored in the memory, the one or more programs comprising instructions for performing the method of any one of items 1-31.
36. A user device comprising:
A user device comprising means for performing the method of any one of items 1-31.
37. A non-transitory computer-readable storage medium comprising one or more programs for execution by one or more processors of an electronic device, the one or more programs, when executed by the one or more processors, cause the electronic device to: A non-transitory computer-readable storage medium comprising instructions for performing the method of any one of the items.
38. A system for operating a digital assistant, comprising means for performing the method of any one of items 1-31.
39. A user device comprising:
a receiving unit configured to receive a speech input for performing a task from a user; and
A processing unit comprising:
identify context information associated with the user device;
determine user intent based on the speech input and context information associated with the user device;
determine, according to the user intent, whether the task is to be performed at the user device or at a first electronic device communicatively coupled to the user device;
receive content for performing the task according to a determination that the task will be performed at the user device and content for performing the task is located elsewhere;
A user device, configured to provide content for performing the task to the first electronic device according to a determination that the task will be performed at the first electronic device and the content for performing the task is located in a different location than the first electronic device.
40. The user device of item 39, wherein the user device is configured to provide a plurality of user interfaces.
41. The user device of items 39 or 40, wherein the user device comprises a laptop computer, a desktop computer, or a server.
42. The user device of any one of items 39-41, wherein the first electronic device comprises a laptop computer, a desktop computer, a server, a smartphone, a tablet, a set-top box, or a watch.
43. The method of any one of items 39 to 42, wherein prior to receiving the speech input:
The user device further comprising displaying, on a display of the user device, an affordance for invoking the digital assistant.
44. The method of item 43,
and immediately executing the digital assistant in response to receiving the predetermined phrase.
45. according to items 43 or 44,
and immediately executing the digital assistant in response to receiving the selection of affordance.
46. The method of any one of items 39 to 45, wherein determining user intent comprises:
determining one or more actionable intentions; and
and determining one or more parameters associated with the actionable intent.
47. The user device of any one of items 39-46, wherein the context information comprises at least one of user specific data, sensor data, and user device configuration data.
48. The user device of any one of items 39-47, wherein determining whether the task is to be performed at the user device or the first electronic device is based on one or more keywords included in the speech input.
49. The method of any one of items 39 to 48, wherein determining whether the task is to be performed at the user device or the first electronic device comprises:
determining whether performing the task at the user device meets a performance criterion; and
and in accordance with determining that performing the task at the user device meets the performance criterion, determining that the task will be performed at the user device.
50. Item 49,
Upon determining that performing a task on the user's device does not meet the performance criteria:
The user device further comprising determining whether performing the task at the first electronic device meets a performance criterion.
51. The method of item 50,
according to determining that performing the task at the first electronic device meets the performance criterion, determining that the task will be performed at the first electronic device; and
In accordance with the determination that performing the task at the first electronic device does not meet the performance criteria:
The user device further comprising determining whether performing the task at the second electronic device meets the performance criterion.
52. The user device of any one of items 49-51, wherein the performance criterion is determined based on one or more user preferences.
53. The user device of any one of items 49-52, wherein the performance criterion is determined based on device configuration data.
54. The user device of any one of items 49-53, wherein the performance criterion is dynamically updated.
55. The method of any one of items 39-49 and 52-54, wherein in accordance with a determination that the task will be performed at the user device and the content for performing the task is located elsewhere, receiving the content for performing the task comprises:
A user device comprising receiving at least a portion of content from a first electronic device, wherein at least a portion of the content is stored on the first electronic device.
56. The method of any one of items 39-49 and 52-55, wherein in accordance with a determination that the task will be performed at the user device and the content for performing the task is located elsewhere, receiving the content for performing the task comprises:
and receiving at least a portion of the content from a third electronic device.
57. The method of any one of items 39 to 54, wherein in accordance with a determination that the task is to be performed at the first electronic device and the content for performing the task is located in a different location than the first electronic device, the content for performing the task is provided. 1 The electronic device provides:
A user device comprising providing at least a portion of the content from the user device to the first electronic device, wherein the at least a portion of the content is stored on the user device.
58. The method for performing the task according to any one of items 39 to 54, and 57, wherein in accordance with a determination that the task will be performed at the first electronic device and content for performing the task is located elsewhere than the first electronic device. Providing the content to the first electronic device comprises:
causing at least a portion of the content to be provided from the fourth electronic device to the first electronic device, wherein the at least a portion of the content is stored in the fourth electronic device.
59. The user device of any one of items 39-49 and 52-56, wherein the task is to be performed at the user device and further comprising providing the first response using the content received at the user device.
60. The method of item 59, wherein providing the first response at the user device comprises:
A user device comprising performing a task at the user device.
61. The user device of item 60, wherein performing the task at the user device is a continuation of the partially performed task elsewhere than at the user device.
62. The method of any one of items 59-61, wherein providing the first response at the user device comprises:
and displaying a first user interface associated with a task to be performed at the user device.
63. The method of any one of items 59-62, wherein providing the first response at the user device comprises:
A user device comprising providing a link associated with a task to be performed at the user device.
64. The method of any one of items 59-63, wherein providing the first response at the user device comprises:
A user device comprising providing a voice output according to a task to be performed at the user device.
65. The user device of any one of items 39 to 54 and 57 and 58, wherein the task is to be performed at the first electronic device and further comprising providing a second response at the user device.
66. The method of item 65, wherein providing the second response at the user device comprises:
and causing the task to be performed at the first electronic device.
67. The user device of item 66, wherein the task to be performed at the first electronic device is a continuation of the task performed elsewhere than the first electronic device.
68. The method of any one of items 65-67, wherein providing the second response at the user device comprises:
A user device, comprising providing a voice output according to a task to be performed at the first electronic device.
69. The method of any one of items 65-68, wherein providing the second response at the user device comprises:
and providing an affordance that enables the user to select another electronic device for performance of a task.
<u>System configuration management</u>
One. A method for providing a digital assistant service, comprising:
In a user device having one or more processors and memory:
receiving, from a user, speech input for managing one or more system configurations of the user device, the user device being configured to simultaneously provide a plurality of user interfaces;
identifying context information associated with the user device;
determining user intent based on the speech input and context information;
determining whether the user intent indicates a request for information or a request to perform a task;
providing a voice response to the information provision request in accordance with a determination that the user intent indicates the information provision request; and
upon determining that the user intent indicates a request to perform the task, immediately executing a process associated with the user device to perform the task.
2. The method of item 1, prior to receiving the speech input:
The method of claim 1, further comprising displaying, on a display of the user device, an affordance for calling the digital assistant.
3. Item 2,
and immediately executing the digital assistant service in response to receiving the predetermined phrase.
4. Item 2,
and immediately executing the digital assistant service in response to receiving the selection of affordance.
5. The method according to any one of items 1 to 4, wherein the one or more system configuration of the user device comprises audio configurations.
6. The method of any one of items 1-5, wherein the one or more system configuration of the user device comprises date and time configurations.
7. The method of any one of items 1-6, wherein the one or more system configuration of the user device comprises oral configurations.
8. The method of any one of items 1-7, wherein the one or more system configuration of the user device comprises display configurations.
9. The method of any one of items 1-8, wherein the one or more system configuration of the user device comprises input device configurations.
10. The method of any one of items 1-9, wherein the one or more system configurations of the user device comprise network configurations.
11. The method of any one of items 1-10, wherein the one or more system configuration of the user device comprises notification configurations.
12. The method of any one of items 1-11, wherein the one or more system configuration of the user device comprises printer configurations.
13. The method of any one of items 1-12, wherein the one or more system configuration of the user device comprises security configurations.
14. The method of any one of items 1-13, wherein the one or more system configuration of the user device comprises backup configurations.
15. The method according to any one of items 1 to 14, wherein the one or more system configurations of the user device comprises application configurations.
16. The method of any one of items 1-15, wherein the one or more system configuration of the user device comprises user interface configurations.
17. The method of any one of items 1 to 16, wherein determining user intent comprises:
determining one or more actionable intentions; and
A method comprising determining one or more parameters associated with an actionable intent.
18. The method of any one of items 1-17, wherein the context information comprises at least one of user specific data, device configuration data, and sensor data.
19. The method of any one of items 1 to 18, wherein determining whether the user intent represents a request for information or a request to perform a task comprises:
and determining whether the user intent is to change a system configuration.
20. The method according to any one of items 1 to 19, wherein in accordance with a determination that the user intent indicates a request for information, providing a voice response to the request for information comprises:
acquiring the status of one or more system configurations according to the information provision request; and
and providing a voice response according to the status of one or more system components.
21. The method of any one of items 1 to 20, wherein, in response to a determination that the user intent indicates a request for information, in addition to providing a voice response to the request for information:
The method further comprising displaying a first user interface that provides information according to the status of one or more system configurations.
22. The method of any one of items 1 to 21, wherein, in accordance with a determination that the user intent indicates a request for information, in addition to providing a voice response to the request for information:
The method further comprising the step of providing a link associated with the request for information.
23. The method of any one of items 1-22, wherein upon determining that the user intent indicates a request to perform the task, immediately executing a process associated with the user device to perform the task comprises:
A method comprising performing a task using the process.
24. The method of item 23,
The method further comprising the step of providing a first voice output according to a result of performing the task.
25. according to items 23 or 24,
The method further comprising the step of providing a second user interface that enables the user to manipulate a result of performing the task.
26. The method of item 25, wherein the second user interface includes a link associated with a result of performing the task.
27. The method of any one of items 1-19 and 23-26, wherein upon determining that the user intent indicates a request to perform the task, immediately executing a process associated with the user device to perform the task comprises:
and providing a third user interface that enables the user to perform the task.
28. The method of item 27, wherein the third user interface comprises a link that enables a user to perform a task.
29. The method of items 27 or 28, further comprising providing a second audio output associated with a third user interface.
30. A non-transitory computer-readable storage medium storing one or more programs, wherein the one or more programs, when executed by one or more processors of the electronic device, cause the electronic device to:
receive from a user speech input for managing one or more system configurations of the user device, wherein the user device is configured to simultaneously provide a plurality of user interfaces;
identify context information associated with the user device;
determine user intent based on the speech input and context information;
determine whether the user intent indicates a request for information or a request to perform a task;
provide a voice response to the information provision request according to a determination that the user intent indicates the information provision request;
A non-transitory computer-readable storage medium comprising instructions that, upon determining that a user intent indicates a request to perform a task, immediately execute a process associated with the user device to perform the task.
31. An electronic device comprising:
one or more processors;
Memory; and
one or more programs stored in the memory, the one or more programs comprising:
receive from a user speech input for managing one or more system configurations of the user device, wherein the user device is configured to simultaneously provide a plurality of user interfaces;
identify context information associated with the user device;
determine user intent based on the speech input and context information;
determine whether the user intent indicates a request for information or a request to perform a task;
provide a voice response to the information provision request according to a determination that the user intent indicates the information provision request;
An electronic device comprising instructions for immediately executing a process associated with the user device to perform the task upon determining that the user intent indicates a request to perform the task.
32. An electronic device comprising:
means for receiving, from a user, speech input for managing one or more system configurations of the user device, wherein the user device is configured to simultaneously provide a plurality of user interfaces;
means for identifying context information associated with the user device;
means for determining user intent based on the speech input and context information;
means for determining whether the user intent indicates a request for information or a request to perform a task;
means for providing, in accordance with a determination that the user intent indicates a request for information, a voice response to the request for information; and
and means for immediately executing a process associated with the user device to perform the task upon determining that the user intent indicates a request to perform the task.
33. An electronic device comprising:
one or more processors;
Memory; and
An electronic device comprising one or more programs stored in the memory, the one or more programs comprising instructions for performing the method of any one of items 1-29.
34. An electronic device comprising:
An electronic device comprising means for performing the method of any of items 1-29.
35. A non-transitory computer-readable storage medium comprising one or more programs for execution by one or more processors of an electronic device, the one or more programs, when executed by the one or more processors, cause the electronic device to: A non-transitory computer-readable storage medium comprising instructions for performing the method of any one of the items.
36. A system for operating a digital assistant, comprising means for performing the method of any of items 1-29.
37. An electronic device comprising:
a receiving unit configured to receive, from a user, a speech input for managing one or more system configurations of the user device, wherein the user device is configured to simultaneously provide a plurality of user interfaces;
A processing unit comprising:
identify context information associated with the user device;
determine user intent based on the speech input and context information;
determine whether the user intent indicates a request for information or a request to perform a task;
provide a voice response to the information provision request according to a determination that the user intent indicates the information provision request;
and upon determining that the user intent indicates a request to perform the task, immediately execute a process associated with the user device to perform the task.
38. The method of item 37, prior to receiving the speech input:
The electronic device further comprising displaying, on a display of the user device, an affordance for calling the digital assistant.
39. The method of item 38,
and immediately executing the digital assistant service in response to receiving the predetermined phrase.
40. The method of item 38,
and immediately executing the digital assistant service in response to receiving the selection of affordance.
41. The electronic device of any one of items 37-40, wherein the one or more system configuration of the user device comprises audio configurations.
42. The electronic device of any one of items 37-41, wherein the one or more system configuration of the user device comprises date and time configurations.
43. The electronic device of any one of items 37-42, wherein the one or more system configuration of the user device comprises oral configurations.
44. The electronic device of any one of items 37-43, wherein the one or more system configuration of the user device comprises display configurations.
45. The electronic device of any one of items 37-44, wherein the one or more system configuration of the user device comprises input device configurations.
46. The electronic device of any one of items 37-45, wherein the one or more system configurations of the user device comprises network configurations.
47. The electronic device of any one of items 37-46, wherein the one or more system configuration of the user device comprises notification configurations.
48. The electronic device of any one of items 37-47, wherein the one or more system configuration of the user device comprises printer configurations.
49. The electronic device of any one of items 37-48, wherein the one or more system configuration of the user device comprises security configurations.
50. The electronic device of any one of items 37-49, wherein the one or more system configuration of the user device comprises backup configurations.
51. The electronic device of any one of items 37-50, wherein the one or more system configuration of the user device comprises application configurations.
52. The electronic device of any one of items 37-51, wherein the one or more system configuration of the user device comprises user interface configurations.
53. The method of any one of items 37-52, wherein determining user intent comprises:
determining one or more actionable intentions; and
An electronic device comprising determining one or more parameters associated with an actionable intent.
54. The electronic device of any one of items 37-53, wherein the context information comprises at least one of user specific data, device configuration data, and sensor data.
55. The method of any one of items 37-54, wherein determining whether the user intent represents a request for information or a request to perform a task comprises:
and determining whether a user intent is to change a system configuration.
56. The method according to any one of items 37 to 55, wherein in accordance with a determination that the user intent indicates a request for information, providing a voice response to the request for information comprises:
obtaining the status of one or more system configurations according to the information provision request; and
and providing a voice response according to a status of one or more system configurations.
57. The method of any one of items 37-56, in addition to providing a voice response to the request for information, upon determining that the user intent represents the request for information:
The electronic device further comprising displaying a first user interface that provides information according to a status of one or more system configurations.
58. The method of any one of items 37-57, in addition to providing a voice response to the request for information, upon determining that the user intent represents the request for information:
The electronic device further comprising providing a link associated with the request for information.
59. The method of any one of items 37-58, wherein upon determining that the user intent indicates a request to perform the task, immediately executing a process associated with the user device to perform the task comprises:
An electronic device comprising performing a task using a process.
60. The method of item 59,
The electronic device further comprising: providing a first voice output according to a result of performing the task.
61. The method of items 59 or 60,
The electronic device further comprising providing a second user interface that enables the user to manipulate a result of performing the task.
62. The electronic device of item 61, wherein the second user interface comprises a link associated with a result of performing the task.
63. The method of any one of items 37-55 and 59-62, wherein upon determining that the user intent indicates a request to perform the task, immediately executing a process associated with the user device to perform the task comprises:
and providing a third user interface that enables a user to perform a task.
64. The electronic device of item 63, wherein the third user interface comprises a link that enables a user to perform a task.
65. The electronic device of items 63 or 64, further comprising providing a second audio output associated with the third user interface.
The foregoing description, for purposes of explanation, has been described with reference to specific embodiments. However, the illustrative discussions above are not intended to be exhaustive or to limit the invention to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments have been chosen and described in order to best explain the principles of the techniques and their practical applications. Accordingly, those skilled in the art will be able to best utilize the techniques and various embodiments with various modifications as are suited to the particular use contemplated.
While the present disclosure and examples have been fully described with reference to the accompanying drawings, it should be noted that various changes and modifications will become apparent to those skilled in the art to which the present invention pertains. It is to be understood that such changes and modifications are included within the scope of the disclosure and examples as defined by the claims.
117 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41 Sheet 42 Sheet 43 Sheet 44 Sheet 45 Sheet 46 Sheet 47 Sheet 48 Sheet 49 Sheet 50 Sheet 51 Sheet 52 Sheet 53 Sheet 54 Sheet 55 Sheet 56 Sheet 57 Sheet 58 Sheet 59 Sheet 60 Sheet 61 Sheet 62 Sheet 63 Sheet 64 Sheet 65 Sheet 66 Sheet 67 Sheet 68 Sheet 69 Sheet 70 Sheet 71 Sheet 72 Sheet 73 Sheet 74 Sheet 75 Sheet 76 Sheet 77 Sheet 78 Sheet 79 Sheet 80 Sheet 81 Sheet 82 Sheet 83 Sheet 84 Sheet 85 Sheet 86 Sheet 87 Sheet 88 Sheet 89 Sheet 90 Sheet 91 Sheet 92 Sheet 93 Sheet 94 Sheet 95 Sheet 96 Sheet 97 Sheet 98 Sheet 99 Sheet 100 Sheet 101 Sheet 102 Sheet 103 Sheet 104 Sheet 105 Sheet 106 Sheet 107 Sheet 108 Sheet 109 Sheet 110 Sheet 111 Sheet 112 Sheet 113 Sheet 114 Sheet 115 Sheet 116 Sheet 117
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| KR102144370B1 | Cited by | Republic of Korea | Search report |
| US11881215B2 | Cited by | United States of America | Applicant |
| US12322386B2 | Cited by | United States of America | Applicant |
| WO2021015319A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| US12405753B1 | Cited by | United States of America | Search report |
| WO2021100975A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| WO2020076087A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| US11514669B2 | Cited by | United States of America | Search report |
57 members in 8 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 62348728 | United States of America | – | |
| 201662348728 | United States of America | P | |
| 15271766 | United States of America | – | |
| 201615271766 | United States of America | A | |
| 2016059953 | United States of America | W |
Members57
| Document | Office | Kind | |
|---|---|---|---|
| US2017358305A1 | United States of America | A1 | |
| WO2017213684A1 | World Intellectual Property Organization (WIPO) | A1 | |
| DK201770032A1 | Denmark | A1 | |
| DK201770035A1 | Denmark | A1 | |
| DK201770036A1 | Denmark | A1 | |
| AU2016409890A1 | Australia | A1 | |
| AU2016409890B2 | Australia | B2 | |
| KR20180098409AThis record | Republic of Korea | A | |
| EP3371690A1 | European Patent Office (EPO) | A1 | |
| CN108701013A | China | A | |
| AU2018241102A1 | Australia | A1 | |
| US2018308485A1 | United States of America | A1 | |
| DK179545B1 | Denmark | B1 | |
| KR20190018061A | Republic of Korea | A | |
| CN109783046A | China | A | |
| CN109814832A | China | A | |
| AU2018241102B2 | Australia | B2 | |
| EP3495943A1 | European Patent Office (EPO) | A1 | |
| EP3508973A1 | European Patent Office (EPO) | A1 | |
| EP3371690A4 | European Patent Office (EPO) | A4 | |
| JP2019522250A | Japan | A | |
| DK179876B1 | Denmark | B1 | |
| DK179883B1 | Denmark | B1 | |
| US2019259386A1 | United States of America | A1 | |
| DK179876B8 | Denmark | B8 | |
| DK179883B8 | Denmark | B8 | |
| AU2019213416A1 | Australia | A1 | |
| AU2019213416B2 | Australia | B2 | |
| JP2019204517A | Japan | A | |
| JP6641517B2 | Japan | B2 | |
| AU2020201030A1 | Australia | A1 | |
| US10586535B2 | United States of America | B2 | |
| US2020118568A1 | United States of America | A1 | |
| US10733993B2 | United States of America | B2 | |
| US10839804B2 | United States of America | B2 | |
| KR102198640B1 | Republic of Korea | B1 | |
| KR20210002133A | Republic of Korea | A | |
| CN109783046B | China | B | |
| AU2020201030B2 | Australia | B2 | |
| CN112767911A | China | A | |
| EP3495943B1 | European Patent Office (EPO) | B1 | |
| US11037565B2 | United States of America | B2 | |
| CN108701013B | China | B | |
| AU2021204695A1 | Australia | A1 | |
| US2021233532A1 | United States of America | A1 | |
| CN113393837A | China | A | |
| CN113407279A | China | A | |
| EP3508973B1 | European Patent Office (EPO) | B1 | |
| CN109814832B | China | B | |
| KR20230011484A | Republic of Korea | A | |
| US11657820B2 | United States of America | B2 | |
| EP4250140A2 | European Patent Office (EPO) | A2 | |
| US2023352016A1 | United States of America | A1 | |
| EP4250140A3 | European Patent Office (EPO) | A3 | |
| KR20240110893A | Republic of Korea | A | |
| US12175977B2 | United States of America | B2 | |
| US2025140255A1 | United States of America | A1 |
10 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Decision of rejection after re-examinationX601 | X601 | |
| AmendmentAMND | AMND | |
| Decision to refuse applicationE601 | E601 | |
| Application refused [patent]X091 | X091 | |
| AmendmentAMND | AMND | |
| Divisional application of patentA107 | A107 | |
| Notification of reason for refusalE902 | E902 | |
| Request for examinationA201 | A201 | |
| Request for accelerated examinationA302 | A302 | |
| AmendmentAMND | AMND |
Numbers
- Publication
- 10-2018-0098409
- Application
- 1020187023111
Titles4
- Korean
- 멀티 태스킹 환경에서의 지능형 디지털 어시스턴트
- English
- Intelligent digital assistant in a multitasking environment
- Unlabeled
- 멀티 태스킹 환경에서의 지능형 디지털 어시스턴트
- Unlabeled
- Intelligent digital assistant in a multitasking environment
Classification
- CPC, 19
- G06F9/453
- G10L15/22
- G06F17/30746
- G06F9/4843
- G06F17/30864
- G06F3/167
- G06F17/30967
- G10L15/1822
- G10L13/02
- G10L15/00
- G10L15/1815
- G10L15/065
- G10L15/30
- G10L2015/223
- G10L2015/228
- G06F16/9032
- G06F16/685
- G06F16/951
- G06F16/953
- IPC, 7
- G10L15 22
- G06F17 30
- G06F3 16
- G10L13 02
- G10L15 18
- G10L15 30
- G06F40 00