Adaptive text-to-speech outputs
Abstract
In some implementations, the language proficiency of a user of the client device is determined by one or more computers. The one or more computers then determine a text segment for output by the text-to-speech module based on the determined language proficiency of the user. After determining a text segment for output, one or more computers generate audio data comprising a synthesized utterance of the text segment. The audio data including the synthesized utterance of the text segment is then provided to the client device for output. An improved user interface is provided through better text-to-speech conversion.

Term
10.3 yearsto projected expiry
Projected expiry 29 December 2036, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
35 claims: 12 independent, 23 dependent
- 1하나 이상의 컴퓨터들에 의해 수행되는 방법으로서, 상기 하나 이상의 컴퓨터들에 의해, 클라이언트 디바이스의 사용자의 언어 능숙도를 결정하는 단계;상기 하나 이상의 컴퓨터들에 의해, 상기 사용자의 상기 결정된 언어 능숙도에 기초하여 텍스트-투-스피치 모듈에 의한 출력을 위한 텍스트 세그먼트를 결정하는 단계;상기 하나 이상의 컴퓨터들에 의해, 상기 텍스트 세그먼트의 합성된 발언을 포함하는 오디오 데이터를 생성하는 단계;상기 하나 이상의 컴퓨터들에 의해, 상기 텍스트 세그먼트의 상기 합성된 발언을 포함하는 상기 오디오 데이터를 상기 클라이언트 디바이스에 제공하는 단계를 포함하는 것을 특징으로 하는 방법.
- 2청구항 1에 있어서, 상기 클라이언트 디바이스는 텍스트-투-스피치 인터페이스를 사용하는 모바일 어플리케이션을 디스플레이하는 것을 특징으로 하는 방법.
- 3청구항 1 또는 2에 있어서, 상기 사용자의 상기 언어 능숙도를 결정하는 단계는 상기 사용자에 의해 제출된 이전의 쿼리들에 적어도 기초하여 상기 사용자의 언어 능숙도를 추론하는 것을 포함하는 것을 특징으로 하는 방법.
- 4임의의 선행하는 청구항에 있어서, 상기 텍스트-투-스피치 모듈에 의한 출력을 위한 상기 텍스트 세그먼트를 결정하는 단계는:다수의 텍스트 세그먼트들을 상기 사용자에 대한 텍스트-투-스피치 출력을 위한 후보들로서 식별하는 것, 상기 다수의 텍스트 세그먼트들은 상이한 레벨의 언어 복잡도를 가지며;및 상기 클라이언트 디바이스의 사용자의 상기 결정된 언어 능숙도에 적어도 기초하여 상기 다수의 텍스트 세그먼트들 중에서 선택하는 것을 포함하는 것을 특징으로 하는 방법.
- 5청구항 4에 있어서, 상기 다수의 텍스트 세그먼트들 중에서 선택하는 것은:상기 다수의 텍스트 세그먼트들 각각에 대해 언어 복잡도 스코어를 결정하는 것;및 상기 클라이언트 디바이스의 사용자의 상기 언어 능숙도를 기술하는 기준 스코어와 가장 잘 매칭되는 상기 언어 복잡도 스코어를 가지는 상기 텍스트 세그먼트를 선택하는 것을 포함하는 것을 특징으로 하는 방법.
- 6임의의 선행하는 청구항에 있어서, 상기 텍스트-투-스피치 모듈에 의한 출력을 위한 상기 텍스트 세그먼트를 결정하는 단계는:상기 사용자에 대한 텍스트-투-스피치 출력을 위한 텍스트 세그먼트를 식별하는 것;상기 텍스트-투-스피치 출력에 대한 상기 텍스트 세그먼트의 복잡도 스코어를 계산하는 것;및 상기 사용자의 상기 결정된 언어 능숙도 및 상기 텍스트-투-스피치 출력을 위한 텍스트 세그먼트의 상기 복잡도 스코어에 적어도 기초하여 상기 사용자에 대한 상기 텍스트-투-스피치 출력을 위한 상기 텍스트 세그먼트를 수정하는 것을 포함하는 것을 특징으로 하는 방법.
- 7청구항 6에 있어서, 상기 사용자에 대한 상기 텍스트-투-스피치 출력을 위한 상기 텍스트 세그먼트를 수정하는 것은:상기 사용자의 상기 결정된 언어 능숙도에 적어도 기초하여 상기 사용자에 대한 종합적(overall) 복잡도 스코어를 결정하는 것;상기 사용자에 대한 상기 텍스트-투-스피치 출력을 위한 상기 텍스트 세그먼트 내에서 개별 부분들에 대한 복잡도 스코어를 결정하는 것;상기 사용자에 대한 상기 종합적 복잡도 스코어보다 큰 복잡도 스코어들을 가지는 상기 텍스트 세그먼트 내의 하나 이상의 개별 부분들을 식별하는 것;및 복잡도 스코어들을 상기 종합적 복잡도 스코어 미만으로 감소시키기 위해 상기 텍스트 세그먼트 내에서 상기 하나 이상의 개별 부분들을 수정하는 것을 포함하는 것을 특징으로 하는 방법.
- 8청구항 6에 있어서, 상기 사용자에 대한 상기 텍스트-투-스피치 출력을 위한 상기 텍스트 세그먼트를 수정하는 것은:상기 사용자와 연관된 컨텍스트를 표시하는 데이터를 수신하는 것;상기 사용자와 연관된 상기 컨텍스트에 대한 종합적 복잡도 스코어를 결정하는 것;상기 텍스트 세그먼트의 상기 복잡도 스코어가 상기 사용자와 연관된 상기 컨텍스트에 대한 상기 종합적 복잡도 스코어를 초과함을 결정하는 것;및 상기 복잡도 스코어를 상기 사용자와 연관된 상기 컨텍스트에 대한 상기 종합적 복잡도 스코어 미만으로 감소시키기 위해 상기 텍스트 세그먼트를 수정하는 것을 포함하는 것을 특징으로 하는 방법.
- 9시스템으로서, 하나 이상의 컴퓨터들; 및 명령어들이 저장된 상기 하나 이상의 컴퓨터들에 연결된 비일시적 컴퓨터 판독가능 매체를 포함하며, 상기 명령어들은 상기 하나 이상의 컴퓨터들에 의해 실행될 때, 상기 하나 이상의 컴퓨터들로 하여금 동작들을 수행하게 하며, 상기 동작들은:상기 하나 이상의 컴퓨터들에 의해, 클라이언트 디바이스의 사용자의 언어 능숙도를 결정하는 동작;상기 하나 이상의 컴퓨터들에 의해, 상기 사용자의 상기 결정된 언어 능숙도에 기초하여 텍스트-투-스피치 모듈에 의한 출력을 위한 텍스트 세그먼트를 결정하는 동작;상기 하나 이상의 컴퓨터들에 의해, 상기 텍스트 세그먼트의 합성된 발언을 포함하는 오디오 데이터를 생성하는 동작;상기 하나 이상의 컴퓨터들에 의해, 상기 텍스트 세그먼트의 상기 합성된 발언을 포함하는 상기 오디오 데이터를 상기 클라이언트 디바이스에 제공하는 동작을 포함하는 것을 특징으로 하는 시스템.
- 10청구항 9에 있어서, 상기 클라이언트 디바이스는 텍스트-투-스피치 인터페이스를 사용하는 모바일 어플리케이션을 디스플레이하는 것을 특징으로 하는 시스템.
- 11청구항 9 또는 10에 있어서, 상기 사용자의 상기 언어 능숙도를 결정하는 동작은 상기 사용자에 의해 제출된 이전의 쿼리들에 적어도 기초하여 상기 사용자의 언어 능숙도를 추론하는 것을 포함하는 것을 특징으로 하는 시스템.
- 12청구항 9 내지 10 중 어느 한 항에 있어서, 상기 텍스트-투-스피치 모듈에 의한 출력을 위한 상기 텍스트 세그먼트를 결정하는 동작은:다수의 텍스트 세그먼트들을 상기 사용자에게 텍스트-투-스피치 출력을 위한 후보들로서 식별하는 것, 상기 다수의 텍스트 세그먼트들은 상이한 레벨의 언어 복잡도를 가지며;및 상기 클라이언트 디바이스의 사용자의 상기 결정된 언어 능숙도에 적어도 기초하여 상기 다수의 텍스트 세그먼트들 중에서 선택하는 것을 포함하는 것을 특징으로 하는 시스템.
- 13청구항 12에 있어서, 상기 다수의 텍스트 세그먼트들 중에서 선택하는 것은:상기 다수의 텍스트 세그먼트들 각각에 대해 언어 복잡도 스코어를 결정하는 것;및 상기 클라이언트 디바이스의 사용자의 상기 언어 능숙도를 기술하는 기준 스코어와 가장 잘 매칭되는 상기 언어 복잡도 스코어를 가지는 상기 텍스트 세그먼트를 선택하는 것을 포함하는 것을 특징으로 하는 시스템.
- 14청구항 9 내지 13 중 어느 한 항에 있어서, 상기 텍스트-투-스피치 모듈에 의한 출력을 위한 상기 텍스트 세그먼트를 결정하는 동작은:상기 사용자에 대한 텍스트-투-스피치 출력을 위한 텍스트 세그먼트를 식별하는 것;상기 텍스트-투-스피치 출력에 대한 상기 텍스트 세그먼트의 복잡도 스코어를 계산하는 것;및 상기 사용자의 상기 결정된 언어 능숙도 및 상기 텍스트-투-스피치 출력을 위한 텍스트 세그먼트의 상기 복잡도 스코어에 적어도 기초하여 상기 사용자에 대한 상기 텍스트-투-스피치 출력을 위한 상기 텍스트 세그먼트를 수정하는 것을 포함하는 것을 특징으로 하는 시스템.
- 15청구항 14에 있어서, 상기 사용자에 대한 상기 텍스트-투-스피치 출력을 위한 상기 텍스트 세그먼트를 수정하는 것은:상기 사용자의 상기 결정된 언어 능숙도에 적어도 기초하여 상기 사용자에 대한 전체적 복잡도 스코어를 결정하는 것;상기 사용자에 대한 상기 텍스트-투-스피치 출력을 위한 상기 텍스트 세그먼트 내에서 개별 부분들에 대한 복잡도 스코어를 결정하는 것;상기 사용자에 대한 상기 종합적 복잡도 스코어보다 큰 복잡도 스코어들을 가지는 상기 텍스트 세그먼트 내의 하나 이상의 개별 부분들을 식별하는 것;및 복잡도 스코어들을 상기 종합적 복잡도 스코어 미만으로 감소시키기 위해 상기 텍스트 세그먼트 내에서 상기 하나 이상의 개별 부분들을 수정하는 것을 포함하는 것을 특징으로 하는 시스템.
- 16하나 이상의 컴퓨터들에 의해 수행되는 방법으로서, 상기 사용자와 연관된 컨텍스트를 표시하는 데이터를 수신하는 단계;상기 사용자와 연관된 상기 컨텍스트에 대한 종합적 복잡도 스코어를 결정하는 단계;상기 사용자에 대한 텍스트-투-스피치 출력을 위한 텍스트 세그먼트를 식별하는 단계;상기 텍스트 세그먼트의 복잡도 스코어가 상기 사용자와 연관된 상기 컨텍스트에 대한 상기 종합적 복잡도 스코어를 초과함을 결정하는 단계;및 상기 복잡도 스코어를 상기 사용자와 연관된 상기 컨텍스트에 대한 상기 종합적 복잡도 스코어 미만으로 감소시키기 위해 상기 텍스트 세그먼트를 수정하는 단계를 포함하는 것을 특징으로 하는 방법.
- 17청구항 16에 있어서, 상기 사용자와 연관된 상기 컨텍스트에 대한 상기 종합적 복잡도 스코어를 결정하는 단계는:상기 사용자가 상기 컨텍스트에 있었던 것으로 결정된 경우 상기 사용자에 의해 이전에 제출된 쿼리들 내에 포함된 용어들을 식별하는 것;및 상기 식별된 용어들에 적어도 기초하여 상기 사용자와 연관된 상기 컨텍스트에 대한 종합적 복잡도 스코어를 결정하는 것을 포함하는 것을 특징으로 하는 방법.
- 18청구항 16 또는 17에 있어서, 상기 사용자와 연관된 상기 컨텍스트를 표시하는 상기 데이터는 상기 사용자에 의해 이전에 제출되었던 쿼리들을 포함하는 것을 특징으로 하는 방법.
- 19청구항 16 내지 18 중 어느 한 항에 있어서, 상기 사용자와 연관된 상기 컨텍스트를 표시하는 상기 데이터는 상기 사용자와 연관된 현재 위치를 표시하는 GPS 신호를 포함하는 것을 특징으로 하는 방법.
- 20청구항 16 내지 19 중 어느 한 항에 있어서, 상기 사용자와 연관된 상기 컨텍스트를 표시하는 상기 데이터는 상기 사용자의 모바일 디바이스로부터의 센서 데이터를 포함하는 것을 특징으로 하는 방법.
- 21하나 이상의 프로세싱 디바이스들에 의해 수행되는 방법으로서, 상기 하나 이상의 프로세싱 디바이스들에 의해, 디바이스에 대한 음성 입력의 복잡도 레벨을 결정하는 단계;상기 하나 이상의 프로세싱 디바이스들에 의해, 상기 음성 입력에 응답하여, 출력을 위한 메시지를 결정하는 단계, 상기 메시지는 상기 음성 입력과 연관된 상기 결정된 복잡도에 기초하여 결정되며;상기 하나 이상의 프로세싱 디바이스들에 의해, 상기 메시지의 합성된 발언을 포함하는 오디오 데이터를 생성하는 단계;및;상기 하나 이상의 프로세싱 디바이스들에 의해, 상기 음성 입력에 응답하여, 상기 합성된 발언을 포함하는 상기 오디오 데이터를 출력을 위해 제공하는 단계를 포함하는 것을 특징으로 하는 방법.
- 22청구항 21에 있어서, 상기 음성 입력의 상기 복잡도 레벨은 상기 음성 입력의 언어 복잡도를 포함하는 것을 특징으로 하는 방법.
- 23청구항 21 또는 22에 있어서, 상기 디바이스에 상기 음성 입력을 제출한 사용자의 언어 능숙도를 결정하는 단계를 더 포함하며;상기 메시지를 결정하는 단계는 상기 디바이스에 상기 음성 입력을 제출한 상기 사용자의 언어 능숙도에 기초하는 것을 특징으로 하는 방법.
- 24청구항 21 내지 23 중 어느 한 항에 있어서, 출력을 위한 상기 메시지를 결정하는 단계는:상기 음성 입력에 응답하여 출력을 위한 기본 메시지(baseline message)를 획득하는 것;및 상기 디바이스에 대한 상기 음성 입력에 대한 상기 결정된 복잡도 레벨에 기초하여 상기 기본 메시지의 복잡도 레벨을 증가시킴으로써 조절된 메시지를 생성하는 것을 포함하는 것을 특징으로 하는 방법.
- 25청구항 21에 있어서, 출력을 위한 상기 메시지를 결정하는 단계는:상기 음성 입력에 응답하여 출력을 위한 기본 메시지(baseline message)를 획득하는 것;및 상기 디바이스에 대한 상기 음성 입력에 대한 상기 결정된 복잡도 레벨에 기초하여 상기 기본 메시지의 복잡도 레벨을 감소시킴으로써 조절된 메시지를 생성하는 것을 포함하는 것을 특징으로 하는 방법.
- 26청구항 21에 있어서, 상기 디바이스는 텍스트-투-스피치 인터페이스를 사용하는 모바일 어플리케이션를 실행하는 것을 특징으로 하는 방법.
- 27하나 이상의 컴퓨터들에 의해 수행되는 방법으로서, 상기 하나 이상의 컴퓨터들에 의해, (i) 특정한 음성 입력이 제1 사용자에 의해 제공되었고 및 (ii) 상기 특정한 음성 입력이 상기 제1 사용자와는 상이한 제2 사용자에 의해 제공되었다는 것을 표시하는 데이터를 획득하는 단계;상기 하나 이상의 컴퓨터들에 의해, (i) 상기 제1 사용자에 대한 제1 언어 능숙도 스코어 및 상기 제2 사용자에 대한 제2 언어 능숙도 스코어를 결정하는 단계, 상기 제1 언어 능숙도 스코어가 상기 제2 언어 능숙도 스코어보다 높은 레벨의 언어 능숙도를 표시하며;상기 하나 이상의 컴퓨터들에 의해, (i) 상기 제1 언어 능숙도 스코어에 기초하여 제1 메시지의 합성된 발언을 포함하는 제1 오디오 데이터 및 (ii) 상기 제2 언어 능숙도 스코어에 기초하여 제2 메시지의 합성된 발언을 포함하는 제2 오디오 데이터를 생성하는 단계, 상기 제1 메시지는 상기 제2 메시지보다 높은 언어 복잡도를 가지며;및 상기 하나 이상의 컴퓨터들에 의해, (i) 상기 특정한 음성 입력에 응답하여 상기 제1 사용자의 클라이언트 디바이스에 상기 제1 오디오 데이터를 그리고 (ii) 상기 특정한 음성 입력에 응답하여 상기 제2 사용자의 클라이언트 디바이스에 상기 제2 오디오 데이터를 제공하는 단계를 포함하는 것을 특징으로 하는 방법.
- 28청구항 27에 있어서, 상기 제1 오디오 데이터를 생성하는 단계는 상기 제1 언어 능숙도 스코어에 기초하여 상기 제1 메시지의 텍스트를 결정하는 것을 포함하며, 상기 제2 오디오 데이터를 생성하는 단계는 상기 제2 언어 능숙도 스코어에 기초하여 상기 제2 메시지의 텍스트를 결정하는 것을 포함하는 것을 특징으로 하는 방법.
- 29청구항 27 또는 28에 있어서, 상기 제1 사용자의 상기 제1 언어 능숙도 및 상기 제2 사용자의 상기 제2 언어 능숙도 스코어를 결정하는 단계는 상기 제1 사용자 및 상기 제2 사용자에 의해 제출된 각각의 이전의 쿼리들에 적어도 기초하여 상기 제1 사용자 및 제2 사용자의 각각의 언어 능숙도를 추론하는 것을 포함하는 것을 특징으로 하는 방법.
- 30청구항 27 내지 29 중 어느 한 항에 있어서, 상기 제1 오디오 데이터를 생성하는 단계는:상기 제1 사용자에 대한 텍스트-투-스피치 출력을 위한 텍스트 세그먼트를 식별하는 것;상기 텍스트 세그먼트의 복잡도 스코어를 계산하는 것;및 상기 제1 사용자의 상기 제1 언어 능숙도 스코어 및 상기 텍스트-투-스피치 출력을 위한 텍스트 세그먼트의 상기 복잡도 스코어에 적어도 기초하여 상기 제1 사용자에 대한 상기 텍스트-투-스피치 출력을 위한 상기 텍스트 세그먼트를 수정하는 것을 포함하는 것을 특징으로 하는 방법.
- 31청구항 30에 있어서, 상기 제1 사용자에 대한 상기 텍스트-투-스피치 출력을 위한 상기 텍스트 세그먼트를 수정하는 것은:상기 제1 사용자의 상기 제1 언어 능숙도에 적어도 기초하여 상기 제1 사용자에 대한 종합적(overall) 복잡도 스코어를 결정하는 것;상기 제1 사용자에 대한 상기 텍스트-투-스피치 출력을 위한 상기 텍스트 세그먼트 내에서 개별 부분들에 대한 복잡도 스코어를 결정하는 것;상기 제1 사용자에 대한 상기 종합적 복잡도 스코어보다 큰 복잡도 스코어들을 가지는 상기 텍스트 세그먼트 내의 하나 이상의 개별 부분들을 식별하는 것;및 복잡도 스코어들을 상기 종합적 복잡도 스코어 미만으로 감소시키기 위해 상기 텍스트 세그먼트 내에서 상기 하나 이상의 개별 부분들을 수정하는 것을 포함하는 것을 특징으로 하는 방법.
- 32청구항 30 또는 31에 있어서, 상기 제1 사용자에 대한 상기 텍스트-투-스피치 출력을 위한 상기 텍스트 세그먼트를 수정하는 것은:상기 제1 사용자와 연관된 컨텍스트를 표시하는 데이터를 수신하는 것;상기 제1 사용자와 연관된 상기 컨텍스트에 대한 종합적 복잡도 스코어를 결정하는 것;상기 텍스트 세그먼트의 상기 복잡도 스코어가 상기 제1 사용자와 연관된 상기 컨텍스트에 대한 상기 종합적 복잡도 스코어를 초과함을 결정하는 것;및 상기 복잡도 스코어를 상기 제1 사용자와 연관된 상기 컨텍스트에 대한 상기 종합적 복잡도 스코어 미만으로 감소시키기 위해 상기 텍스트 세그먼트를 수정하는 것을 포함하는 것을 특징으로 하는 방법.
- 33청구항 27 내지 32 중 어느 한 항에 있어서, 상기 제1 오디오 데이터 및 상기 제2 오디오 데이터를 제공하는 단계는:상기 하나 이상의 컴퓨터들에 의해, (i) 컴퓨터 네트워크를 통해 상기 제1 오디오 데이터를 상기 제1 사용자의 상기 클라이언트 디바이스에 그리고 (ii) 컴퓨터 네트워크를 통해 상기 제2 오디오 데이터를 상기 제2 사용자의 상기 클라이언트 디바이스에 제공하는 것을 포함하는 것을 특징으로 하는 방법.
- 34하나 이상의 프로세싱 디바이스들 및 명령어들을 저장하는 하나 이상의 기계 판독가능 저장 디바이스들을 포함하는 시스템으로서, 상기 명령어들은 상기 하나 이상의 프로세싱 디바이스들에 의해 수행될 때, 상기 시스템으로 하여금 청구항 1 내지 9 또는 16 내지 33 중 어느 한 항의 방법을 수행하게 하는 것을 특징으로 하는 시스템.
- 35명령어들을 저장하는 하나 이상의 기계 판독가능 저장 디바이스들로서, 상기 명령어들은 하나 이상의 프로세싱 디바이스들에 의해 수행될 때, 상기 하나 이상의 프로세싱 디바이스들로 하여금 청구항 1 내지 9 또는 16 내지 33 중 어느 한 항의 방법을 수행하게 하는 것을 특징으로 하는 기계 판독가능 저장 디바이스.
Independent claims35
109 paragraphs in 1 section, as filed
Adaptive text-to-speech output
CROSS-REFERENCE TO RELATED APPLICATIONS
This application claims priority to US Application Serial No. 15/009,432 "Adaptive Text-to-Speech Output," filed January 28, 2016, which is incorporated herein by reference in its entirety.
technical field
This specification generally describes electronic communications.
Speech synthesis refers to the artificial production of human speech. Speech synthesizers may be implemented in software or hardware components for generating speech output corresponding to text. For example, text-to-speech (TTS) systems typically convert plain language text into speech by concatenating pieces of recorded speech stored in a database.
As much of electronic computing moves from desktops to mobile environments, speech synthesis is becoming central to the user experience. For example, the increasing use of smaller mobile devices without a display has increased the use of text-to-speech (TTS) systems for accessing and using content displayed on mobile devices.
This disclosure discloses improved user interfaces, in particular better computer-to-user communication via improved TTS.
One particular issue with existing TTS systems is that they are often incapable of adapting to the various language proficiencies of different users. This lack of flexibility often prevents users with limited language proficiencies from understanding complex text-to-speech output. For example, non-native language speakers using the TTS system have difficulty understanding text-to-speech thrust because of their limited language familiarity. Another issue with existing TTS systems is that a user's immediate ability to understand text-to-speech thrust may also vary based on a particular user context. For example, some user contexts include background noise that can make longer or more complex text-to-speech output more difficult to understand.
In some implementations, the system adjusts the text used in the text-to-speech output based on the user's language proficiency to increase the likelihood that the user can understand the text-to-speech output. For example, a user's language proficiency may be inferred from previous user activity and used to adjust the text-to-speech output to an appropriate complexity commensurate with the user's language proficiency. In some examples, the system obtains multiple candidate text segments corresponding to different levels of language proficiency. The system then selects the candidate text segment that best matches and most closely corresponds to the user's language proficiency, and provides the synthesized utterances of the selected text segment as output to the user. In another example, the system converts the text into text segments to better match the user's language proficiency before generating text-to-speech output. Various aspects of a text segment including vocabulary, sentence structure, length, and the like may be adjusted. The system then provides a synthesized utterance of the modified text segment for output to the user.
In cases where the systems discussed herein collect or use personal information about users, programs or configurations that provide users with personal information, such as the user's social network, social actions, or Opportunities will be provided to control whether and/or how to receive content from a content server that is more relevant to the user, to control whether information about activities, occupation, preferences of the user or current location of the user is collected. can Additionally, certain data is anonymized in one or more various ways before it is stored or used, so that personally identifiable information is removed. For example, the user's identity may be anonymized so that personally identifiable information about the user cannot be determined, or the user's geographic location may be generalized (at the city, zip code or state level) from where the location information was obtained. A specific location cannot be determined. Thus, the user can have control over how information about him or her is collected and used by the content server.
In one aspect, a computer-implemented method includes determining, by one or more computers, a language proficiency of a user of a client device; determining, by the one or more computers, a text segment for output by a text-to-speech module based on the determined language proficiency of the user; generating, by the one or more computers, audio data comprising a synthesized utterance of the text segment; providing, by the one or more computers, the audio data comprising the synthesized utterance of the text segment to the client device.
Other versions include computer programs and corresponding systems configured to perform the actions of the methods encoded on computer storage devices.
One or more implementations may include the following optional configurations. For example, in some implementations, the client device displays a mobile application using a text-to-speech interface.
In some implementations, determining the language proficiency of the user comprises inferring the language proficiency of the user based at least on previous queries submitted by the user.
In some implementations, determining a text segment for output by the text-to-speech module comprises identifying a plurality of text segments as candidates for text-to-speech output to the user, the plurality of Text segments have different levels of language complexity; and selecting from among the plurality of text segments based at least on the determined language proficiency of the user of the client device.
In some implementations, selecting among the plurality of text segments includes determining a language complexity score for each of the plurality of text segments; and selecting the text segment having the language complexity score that best matches a reference score describing the language proficiency of the user of the client device.
In some implementations, determining a text segment for output by the text-to-speech module includes identifying a text segment for text-to-speech output to the user; calculating a complexity score of the text segment for the text-to-speech output; and modifying the text segment for text-to-speech output for the user based at least on the determined language proficiency of the user and the complexity score of the text segment for text-to-speech output. .
In some implementations, modifying the text segment for the text-to-speech output for the user comprises determining an overall complexity score for the user based at least on the determined language proficiency of the user. thing; determining a complexity score for respective portions within the text segment for the text-to-speech output to the user; identifying one or more distinct portions within the text segment that have complexity scores greater than the overall complexity score for the user; and modifying the one or more individual portions within the text segment to reduce complexity scores below the overall complexity score.
In some implementations, modifying the text segment for the text-to-speech output to the user includes receiving data indicative of a context associated with the user; determining an overall complexity score for the context associated with the user; determining that the complexity score of the text segment exceeds the overall complexity score for the context associated with the user; and modifying the text segment to reduce the complexity score below the overall complexity score for the context associated with the user.
In another aspect, a computer program includes machine readable instructions that, when executed by a computing device, cause the computing device to perform any of the above methods.
In another general aspect, a computer-implemented method includes: receiving data indicative of a context associated with the user; determining an overall complexity score for the context associated with the user; identifying a text segment for text-to-speech output to the user; determining that the complexity score of the text segment exceeds the overall complexity score for the context associated with the user; and modifying the text segment to reduce the complexity score below the overall complexity score for the context associated with the user.
In some implementations, determining an overall complexity score for the context associated with the user includes identifying terms included in queries previously submitted by the user when it was determined that the user was in the context. ; and determining an overall complexity score for the context associated with the user based at least on the identified terms.
In some implementations, the data indicating the context associated with the user includes queries that have been previously submitted by the user.
In some implementations, the data indicative of the context associated with the user comprises a GPS signal indicative of a current location associated with the user.
The details of one or more implementations are set forth in the accompanying drawings and the description below. Other potential configurations and advantages will become apparent from the description, drawings and claims.
Other implementations of these aspects include computer programs encoded in corresponding systems, apparatuses and computer storage devices configured to perform the actions of the methods.
1 is a diagram illustrating examples of processes for generating text-to-speech output based on language proficiency. 2 is a diagram illustrating an example of a system for generating adaptive text-to-speech output based on a user context. 3 is a diagram illustrating an example of a system for modifying sentence structure within text-to-speech output. 4 is a block diagram illustrating an example of a system for generating adaptive text-to-speech output based on using clustering techniques. 5 is a flow diagram illustrating an example of a process for generating an adaptive text-to-speech output. 6 is a block diagram of computing devices in which the processes described herein, or portions thereof, may be implemented. In the drawings, like numbers represent corresponding parts throughout.
1 is a diagram illustrating examples of processes 100A and 100B for generating text-to-speech output based on language proficiency. Processes 100A and 100B are used to generate, for text query 104 , different text-to-speech outputs for user 102a with high language proficiency and user 102b with low language proficiency, respectively. . As shown, after receiving the query 104 at the user devices 106a and 106b, the process 100A generates a high complexity text-to-speech output 108a for the user 102a, while Process 100B produces a low complexity output 108b for user 102b. Additionally, TTS systems executing processes 100A and 100B may include a language proficiency predictor 110 , a text-to-speech engine 120 . Additionally, the text-to-speech engine 120 may further include a text analyzer 122 , a linguistics analyzer 124 , and a waveform generator 126 .
In general, the content of the text used to generate the text-to-speech output may be determined according to the language proficiency of the user. Additionally or alternatively, the text to be used to generate the text-to-speech output may be determined based on the user's context, eg, the user's location or activity, the presence of background noise, the user's current task, and the like. Additionally, text to be converted into an audible form may be adjusted or determined using other information, such as indications that the user has failed to complete a task or is repeating an action.
In the example, two users, user 102a and user 102b, provide the same query 104 to user devices 106a and 106b, respectively, as input to an application, webpage, or other search function. For example, query 104 may be a voice query sent to user devices 106a and 106b to determine a weather forecast for the current date. The query 104 is then sent to the text-to-speech engine 120 to generate text-to-speech output in response to the query 104 .
Language proficiency predictor 110 may be a software module within a TTS system that determines a language proficiency score associated with a particular user (eg, user 102a or user 102b) based on user data 108a. A language proficiency score may be a predictor of a user's ability to understand communication in a particular language, particularly speech in a particular language. One measure of language proficiency is a user's ability to successfully complete a voice control task. Many types of tasks, such as scheduling and finding directions, are followed by a sequence of interactions in which the user and device exchange verbal communication. The rate at which a user successfully completes the workflow of these tasks via a voice interface may be a strong indicator of a user's language proficiency. For example, a user who succeeded in 9 out of 10 voice tasks initiated by the user is highly likely to have high language proficiency. On the other hand, a user who fails to complete most of the user-initiated voice tasks may be inferred to have low language proficiency because the user did not fully understand the communication from the device or was unable to provide an appropriate verbal response. As will be discussed further below, if a user does not complete a workflow that includes standard TTS output, it results in a low language proficiency score, which TTS will increase the user's ability to understand and complete various tasks. Adapted, simplified output that can be used is available.
As shown, user data 108a may include the words used in previous text queries submitted by the user, whether English or any other language utilized by the TTS system is the user's native language, and the user's may include a set of activities and/or actions that reflect language comprehension skills. For example, as shown in FIG. 1 , a user's typing speed may be used to determine a user's language fluency in a language. Additionally, a language vocabulary complexity score or language proficiency score may be assigned to a user based on associating a predetermined complexity with words used by the user in previous text queries. In another example, the number of misrecognized words in previous queries may also be used to determine a language proficiency score. For example, a high number of unrecognized words may be used to indicate low language proficiency. In some implementations, the language proficiency score is determined by querying a stored score associated with the user that was determined for the user prior to submitting the query 104 .
Although FIG. 1 shows the language proficiency predictor 110 as a separate component from the TTS engine 120 , in some implementations, as shown in FIG. 2 , the language proficiency predictor 110 is the TTS engine 120 . It can be integrated into a software module within In the above examples, operations related to language proficiency prediction may be directly modularized by the TTS engine 120 .
In some implementations, a language proficiency score assigned to a user may be based on a particular user context predicted for the user. For example, as described in more detail with respect to FIG. 2 , a user context determination may be used to determine context specific language proficiencies that may cause a user to temporarily have limited language comprehension ability. For example, if the user context displays significant background noise, or if the user is involved in a task such as driving, the language proficiency score may be used to indicate that the user's current language comprehension ability has temporarily decreased compared to other user contexts. .
In some implementations, instead of inferring language proficiency based on previous user activity, the language proficiency score may be provided directly to the TTS engine 120 without the use of the language proficiency predictor 110 . For example, a language proficiency score may be assigned to a user based on user input during an enrollment process that specifies a user's level of language proficiency. For example, during registration, the user may provide a selection specifying the user's ability level, which may then be used to calculate an appropriate language proficiency for the user. In other examples, the user may provide other types of information, such as demographic information, education level, place of residence, etc. that may be used to specify the user's level of language proficiency.
In the examples described above, the language proficiency score may be a set of discrete values that are periodically adjusted based on recently generated user activity data or may be a continuous score initially specified during the registration process. In a first example, the value of a language proficiency score may be biased based on one or more factors indicative of the user's current language comprehension and proficiency may be weakened (eg, the user context indicates significant background noise). . In a second example, the value of the language proficiency score can be pre-fitted after the initial calculation, and is of a certain significance indicating that the user's language proficiency has increased (eg, increasing the typing rate or decreasing the correction rate for that language). It can only be adjusted after events. In another example, a combination of these two techniques can be used to vary the text-to-speech output based on a particular text input. In the example above, multiple language proficiency scores, each representing a particular aspect of the user's language ability, may be used to determine how best to adjust text-to-speech output to the user. For example, one language proficiency score may indicate the complexity of the user's vocabulary, while another language proficiency score may be used to indicate the grammatical ability of the user.
The TTS engine 120 may use the language proficiency score to generate text-to-speech output adapted to the language proficiency indicated by the user's language proficiency score. In some examples, TTS engine 120 may adapt text-to-speech output based on selecting a particular TTS string from a set of candidate TTSs for text query 104 . In the example above, the TTS engine 120 selects a particular TTS string based on using the user's language proficiency score to anticipate the likelihood that each of the candidate TTS strings will be interpreted correctly by the user. More specific descriptions related to the present techniques are provided with reference to FIG. 2 . Alternatively, in another example, the TTS engine 120 may select a baseline TTS string and adjust the structure of the TTS string based on the user's level of language proficiency indicated by the language proficiency score. In the example above, the TTS engine 120 may adjust the grammar of the base TTS string to provide alternative words and/or reduce the complexity of the sentence to produce an adapted TTS string that is more likely to be understood by the user. . More specific descriptions related to the present techniques are provided with reference to FIG. 3 .
Referring to FIG. 1 , TTS engine 120 may generate different text-to-speech outputs for users 102a and 102b because language proficiency scores for users are different. For example, in process 100A, language proficiency score 106a indicates high English-language proficiency, indicating that user 102a has a complex vocabulary, English is the native language, and relatively per minute in previous user queries. Inferred from user data 108a indicating that it has many words. Based on the value of the language proficiency score 106a, the TTS engine 120 generates a high-complexity text-to-speech output 108a comprising a complex grammatical structure. As shown, the text-to-speech output 108a includes an independent clause stating that today's weather forecast is clear, and a dependent clause additionally containing additional information about the daily maximum and minimum temperatures. (subordinate clause).
In the example of process 100B, language proficiency score 106a indicates low English-language proficiency, which indicates that user 102b has a simple vocabulary, English is a second foreign language, and provided the previous 10 incorrect queries. It is inferred from the user data 108b indicating that the In this example, the TTS engine 120 generates a text-to-speech output 108b of lower complexity comprising a simpler grammatical structure compared to the text-to-speech output 108a. For example, instead of including multiple clauses within a single sentence, text-to-speech output 108b conveys the same key information as text-to-speech output 108a, but with the best of the day and Include a simple independent clause that does not contain additional information on the minimum temperature.
Adaptation of text to TTS output can be performed by a variety of different devices and software modules. For example, the TTS engine of the server system may include the ability to adjust text based on a language proficiency score and then output audio comprising a synthesized utterance of the adjusted text. As another example, the pre-processing module of the server system may condition the text and pass the adjusted text to the TTS engine for speech synthesis. As another example, the user device may include a TTS engine or a TTS engine and a text pre-processor that may generate appropriate TTS outputs.
In some implementations, the TTS system may include software modules configured to exchange communications with a third-party mobile application or webpage of the client device. For example, the TTS functionality of the system may be made available for use by third-party mobile applications through an application package interface (API). The API may include a defined set of protocols that an application or website may use to request TTS audio from a server system running the TTS engine 120 . In some implementations, the API may cause the TTS function to run locally on the user's device. For example, an API may be made available to an application or web page through an inter-process communication (IPC), remote procedure call (RPC) or other system call or function. The TTS engine and associated language proficiency analysis or text preprocessing run locally on the user's device to determine the appropriate text for the user's language proficiency and also generate audio for the synthesized speech.
For example, a third-party application or webpage may use the API to generate a set of voice commands provided to the user based on the workflow of the voice interface of the third-party application or webpage. The API may specify that the application or webpage should provide text to be converted into speech. In some examples, other information may be provided, such as a user identifier or a language proficiency score.
In implementations where the TTS engine 120 exchanges communications with a third-party application using an API, the TTS engine 120 determines that a text segment from the third-party application generates text-to-speech output for the text. It can be used to determine if it should be adjusted. For example, an API may include computer-implemented protocols that specify conditions in a third-party application that initiate the generation of adaptive text-to-speech output .
As an example, an API allows an application to submit a number of different text segments as candidates for TTS output, wherein the different text segments correspond to different levels of language proficiency. For example, the candidates may be text segments with the same meaning but different complexity levels (eg, high complexity response, medium complexity response and low complexity response). The TTS engine 120 may then determine the level of language proficiency required to understand each candidate, determine an appropriate language proficiency score for the user, and select the candidate text that best corresponds to the language proficiency score. . Thereafter, the TTS engine 120 provides the synthesized audio for the selected text back to the application through a network using, for example, an API. In some examples, the API may be available locally to user devices 106a and 106b. In the example above, the API may be accessible via various types of inter-process communication (IPC) or via a system call. For example, since the API operates locally on user devices 106a and 106b , the output of the API on user devices 106a and 106b may be the text-to-speech output of the TTS engine 120 .
In another example, the API may cause a third-party application to provide a single text segment and a value indicating whether the TTS engine 120 is allowed to modify the text segment to create text segments with different complexity. If the application or webpage indicates that a change is permitted, the TTS system 120 reduces the complexity of the text, for example, if the language proficiency score suggests that the original text is more complex than the user could understand in the uttered response. You can make various changes to the text, such as In another example, the API may also cause a third-party application to provide user data (eg, previous user queries submitted to the third-party application) along with a text segment, so that the TTS engine 120 can display the user associated with the user. Determine a context and adjust to produce a specific text-to-speech output based on the determined user context. Similarly, the API allows an application to provide an indication of user context or context data from the user device (eg, GPS, accelerometer data, ambient noise levels, etc.) control the text-to-speech output that will ultimately be presented to the user. In some examples, a third-party application may provide the API with data that may be used to determine a user's language proficiency.
In some implementations, TTS engine 120 may adjust text-to-speech output for a user query without using the user's language proficiency or determining the context associated with the user. In this implementation, the TTS engine 120 determines that the initial text-to-speech output is based on receiving signals that the user does not understand the output (eg, multiple retries for the same query or operation). You may decide it's too complicated for the user. In response, the TTS engine 120 may reduce the complexity of a subsequent text-to-speech response to the retried query or related queries. Thus, if the user fails to successfully complete an action, the TTS engine 120 may progressively reduce the amount of detail or language proficiency required to understand the TTS output until it reaches a level that the user understands. .
2 is a diagram illustrating an example of a system 200 for adaptively generating text-to-speech output based on a user context. Briefly, system 200 includes a query analyzer 211 , a language proficiency predictor 212 , an interpolator 213 , a linguistics analyzer 214 , and a re-ranker 215 . ) and a TTS engine 210 including a waveform generator 216 . System 200 also includes a context store 220 that stores a set of context profiles 232 and a user history manager 230 that stores user history data 234 . In some examples, the TTS engine 210 corresponds to the TTS engine 120 described with respect to FIG. 1 .
In the example, the user 202 initially submits a query 204 at the user device 208 that includes a request for information related to the user's first meeting on that day. User device 208 then sends query 204 and context data 206 associated with user 202 to query analyzer 211 and language proficiency predictor 212 respectively. Other types of TTS output other than responses to queries, eg, schedule reminders, notifications, workflows, etc., may be adapted using the same techniques.
Context data 206 may include time intervals between repeated text queries, GPS data indicative of a position, speed, or movement pattern associated with user 202 , and previous text queries submitted to TTS engine 210 within a specified period of time. or other types of information related to a particular context associated with the user 202 , such as background information that may indicate user activity related to the TTS engine 210 . In some examples, the context data 206 may be transmitted to the TTS engine 210 , such as whether the query 204 is a text segment associated with a user action or an instruction sent to the TTS engine 210 to generate text-to-speech output. ) may indicate the type of query 204 submitted.
After receiving the query 204 , the query analyzer 211 parses the query 204 to identify information that is a response to the query 204 . For example, in some instances where query 204 is a voice query, query analyzer 211 initially generates a transcription of the voice query and then provides the query to a search engine, eg, Individual words or segments within the query 204 are processed by receiving the search results to determine information that is a response to the query 204 . The transcription of the query and the identified information may then be sent to a linguistic analyzer 214 .
Referring now to the language proficiency predictor 212 , after receiving the context data 206 , the language proficiency predictor 212 may be configured to generate a user based on the received context data 206 using the techniques described in connection with FIG. 1 . Calculate the language proficiency for (202). In particular, the language proficiency predictor 212 parses through the various context profiles 232 stored in the storage 220 . Context profile 232 may be an archived library that is associated with a particular user context and contains relevant types of information that may be included in text-to-speech output. The context profile 232 has a value associated with each type of information indicating whether each type of information is likely to be understood by the user 202 if the user 202 is currently within the context associated with the context profile 232 . additionally specify.
In the example shown in FIG. 2 , the context profile 232 specifies that the user 202 is currently in a context indicating that the user 202 is going to work for his/her job. Additionally, the context profile 232 also specifies values for individual words and phrases that are likely to be understood by the user 202 . For example, date or time information is associated with a value of "0.9" for "SINCE" so that user 202 can generalize associated with a meeting rather than detailed information associated with the meeting (eg, meeting attendees or location of the meeting). Indicates that they are likely to understand the information that has been received (eg, the time of the next upcoming meeting). In this example, differences in values indicate differences in the user's ability to understand a particular type of information because the user's ability to understand complex or detailed information has been reduced.
A value associated with individual words and phrases may be determined based on user activity data from previous user sessions if user 202 was previously in the context indicated by context data 206 . For example, historical user data may be sent from the user history manager 230 retrieving data stored in the query log 234 . In an example, the value for the date and time information may be increased based on determining that the user generally accesses date and time information associated with the meetings more frequently than the locations of the meetings.
After language proficiency predictor 212 selects a particular context profile 232 corresponding to the received context data 206 , language proficiency predictor 212 sends the selected context profile 232 to interpolator 213 . do. The interpolator 213 parses the selected context profile 232 and extracts the individual words and phrases included and their associated values. In some examples, interpolator 213 sends different types of information and associated values directly to linguistic analyzer 214 to generate text-to-speech output candidates 240a. In this example, interpolator 213 extracts certain types of information and associated values from the selected context profile 232 and sends them to linguistic analyzer 214 . In another example, the interpolator 213 may also send the selected context profile 232 to the re-ranker 215 .
In some examples, a set of structured data (eg, fields of a calendar event) may be provided to the TTS engine 210 . In the example above, interpolator 213 may convert the structured data into text at a level matching the user's proficiency indicated by the context profile 232 . For example, the TTS engine 210 may have access to data representing one or more grammars representing different levels of detail or complexity for representing information in the structured data, and based on the user's language proficiency score. You can choose an appropriate grammar. Similarly, the TTS engine 210 may use dictionaries to select appropriate words for a given language proficiency score.
The linguistic analyzer 214 performs processing operations such as normalization on the information contained within the query 204 . For example, the query analyzer 211 can assign phonetic transcriptions for each word or snippet included within the query 204 and convert the query 204 to phrases, clauses and It can be divided into prosodic units such as sentences. The linguistic analyzer 214 generates a list 240a including a plurality of text-to-speech output candidates identified as being in response to the or query 204 . In the example, list 240a includes multiple text-to-speech output candidates having different levels of complexity. For example, the response "Mr. John near DuPont Circle at 12:00 PM" is the most complex response because it identifies the time for the meeting, the location for the meeting, and the person with whom the meeting will take place. In comparison, the response "3 hours later" is minimally complex as it only identifies the time for the meeting.
List 240a also includes a default ranking for text-to-speech candidates based on the likelihood that each text-to-speech output candidate is likely a response to query 204 . In the example, list 240a indicates that the most complex text-to-speech output candidate is most likely a response to query 204 because it contains the largest amount of information associated with the content of query 204 . .
After the linguistic analyzer has generated the list 240a of text-to-speech output candidates, the re-ranker 215 generates a list 240b, which based on the received context data 206 is text-to-speech. - Adjust the ranking for speech output candidates. For example, the re-ranker 215 may adjust the ranking based on scores associated with a particular type of information contained within the selected context profile 232 .
In the example, the re-ranker 215 will likely understand date and time information within a text-to-speech response, given the user's current context in which the user indicates that the user is on the go, but within the text-to-speech response. Based on the context profile 232 indicating that the attendee name or location information is less likely to be understood, the user 202 may rank the simplest text-to-speech output as the highest. In this regard, the received context data 206 may include a description of a particular text-to-speech output candidate in order to increase the likelihood that the user 202 will understand the content of the text-to-speech output 204c of the TTS engine 210 . It can be used to control selection.
3 is a diagram illustrating an example of a system 300 for modifying sentence structure within text-to-speech output. Briefly, the TTS engine 310 receives a query 302 and a language proficiency profile 304 for a user (eg, user 202 ). The TTS engine 310 then performs operations 312 , 314 and 316 to generate a conditioned text-to-speech output 302c that is a response to the query 302 . In some examples, the TTS engine 310 corresponds to the TTS engine 120 described with respect to FIG. 1 or the TTS engine 210 described with respect to FIG. 2 .
In general, the TTS engine 310 may modify the sentence structure of the basic text-to-speech output 306a for the query 302 using different types of adjustment techniques. As an example, the TTS engine 310 may determine that the complexity score associated with individual words or phrases is greater than a threshold score indicated by the user's language complexity profile 304 , based on a determination that the Words or phrases may be substituted. As another example, the TTS engine 310 rearranges the individual sentence clauses so that the overall complexity of the basic text-to-speech output 306a is reduced to a satisfactory level based on the language complexity profile 304 . The TTS engine 310 may also reorder words, split or combine sentences, and make other changes to adjust the complexity of the text.
More specifically, during operation 312 , the TTS engine 310 initially generates a basic text-to-speech output 306a that is a response to the query 302 . The TTS engine 310 then parses the basic text-to-speech output 306a into segments 312a-312c. In addition, the TTS engine 310 detects punctuation marks (eg, commas, periods, semicolons, etc.) that indicate breakpoints between individual segments. The TTS engine 310 also computes a complexity score for each of the segments 312a-312c. In some examples, a complexity score may be calculated based on the frequency of a particular word within a particular language. Alternative techniques may include calculating a complexity score based on a frequency of use by a user or a frequency of occurrence in historical content (eg, news articles, webpages, etc.) accessed by the user. In each of these examples, the complexity score can be used to indicate words that are more likely to be understood by the user and other words that are less likely to be understood by the user.
In the example, segments 312a and 312b are determined to be relatively complex based on the inclusion of highly complex terms such as "forecast" and "consistent," respectively. However, segment 312c is determined to be relatively simple as the terms involved are relatively simple. This decision is represented by segments 312a and 312b having high complexity scores (eg, 0.83, 0.75) compared to the complexity score (eg, 0.41) for segment 312c.
As described above, the language proficiency profile 304 may be used to calculate a threshold complexity score that represents the maximum complexity understandable by the user. In the example, the threshold complexity score is calculated to be "0.7", causing TTS 310 to determine that segments 312a and 312b are unlikely to be understood by the user.
After identifying individual segments associated with complexity scores greater than the threshold complexity score indicated by the language proficiency profile 304 , during operation 314 , the TTS engine 310 makes the identified words more likely to be understood by the user. Replace with alternative words that are expected to be high. As shown in FIG. 3 , "forecast" may be replaced with "weather" and "schedule" may be replaced with "change". In these examples, segments 314a and 314b represent simpler alternatives having a lower complexity score below the threshold complexity score indicated by language proficiency profile 304 .
In some implementations, the TTS engine 310 uses a high-complexity trained skip-gram model that uses unsupervised techniques to determine to replace appropriately complex words with highly complex words. It can process word replacements for words. Also, in some examples, the TTS engine 310 may process word replacement for high complexity words using synonym or thesaurus data.
Referring now to operation 316 , sentence clauses of the query calculate the complexity associated with particular sentence structures and determine whether the user will be able to understand the sentence structure based on the language proficiency indicated by the language proficiency profile 304 . It can be adjusted based on
In the example, the TTS engine 310 indicates that the basic text-to-speech response 306a includes three sentence clauses (eg, "Today's forecast is sunny," "but not consistent," and "warm"). Based on determining that it does, it determines that the basic text-to-speech response 306a has a high sentence complexity. In response, TTS engine 310 may generate adjusted sentence portions 316a and 316b that combine the dependent and independent clauses into a single clause that does not contain segment punctuation. As a result, the adjusted text-to-speech response 306b contains a simpler vocabulary (eg, "weather", "changes") as well as a simpler sentence structure (eg, a separate clause None), and increases the likelihood that the user will understand the adjusted text-to-speech output 306b. The conditioned text-to-speech output 306b is then generated for output by the TTS engine 310 as output 306c.
In some implementations, the TTS engine 310 adjusts the sentence structure using a user-specific restructuring algorithm that includes adjusting the base query 302a using weighting factors to avoid particular sentence structures identified as being problematic for the user. can be performed. For example, a user-specific restructuring algorithm may specify an option to reduce the inclusion of dependent clauses or increase sentence clauses with simple subject-verb object sentences.
4 is a block diagram illustrating an example of a system 400 for adaptively generating text-to-speech outputs based on using clustering techniques. The system 400 includes a language proficiency predictor 410 , a user similarity determiner 420 , a complexity optimizer and a machine learning system 400 .
Briefly, the language proficiency predictor 410 receives data from a plurality of users 402 . The language proficiency predictor 410 then predicts a set of language complexity profiles 412 for each of the plurality of users 402 , which is then sent to the user similarity determiner 420 . User similarity determiner 420 identifies user clusters 424 of similar users. Thereafter, the complexity optimizer 430 and the machine learning system 400 analyze the language complexity profile 412 of each user in the user clusters 424 and the context data received from the plurality of users 402 to determine the complexity Create a mapping 442 .
In general, system 400 may be used to analyze the relationship between active and passive language complexity for a population of users. Active language complexity refers to detected language input provided by a user (eg, text queries, voice input, etc.). Passive language complexity refers to a user's ability to understand or understand speech signals presented to the user. In this regard, system 400 determines active language complexity and passive language complexity for multiple users to determine an appropriate passive language complexity for each individual user for which a particular user has a high probability of understanding text-to-speech output. The determined relationship between the two can be used.
The plurality of users 402 may be multiple users using an application associated with the TTS engine (eg, the TTS engine 120 ). For example, the plurality of users 402 may be a set of users using a mobile application that utilizes a TTS engine to provide text-to-speech configurations to users via a user interface of the mobile application. In the usual example, data from a plurality of users 402 (eg, previous user queries, user selections, etc.) may be tracked by a mobile application and analyzed by a language proficiency predictor 410 . can be collected for
The language proficiency predictor 410 may initially measure the passive language complexities for the plurality of users 402 using substantially similar techniques as previously described with respect to FIG. 1 . The language proficiency predictor 410 may then generate a language complexity profile 412 , which includes a respective language complexity profile for each of the plurality of users 402 . Each respective language complexity profile includes data indicative of a passive language complexity and an active language complexity for each of the plurality of users 402 .
User similarity determiner 420 identifies similar users within plurality of users 402 using language complexity data included in set of language proficiency profiles 412 . In some examples, user similarity determiner 420 can group users with similar active language complexities (eg, similar language inputs, provided speech queries, etc.). In another example, the user similarity determiner 420 may determine similar users by comparing words contained in previous user submitted queries, specific user actions on a mobile application, or user locations. The user similarity determiner 420 then clusters similar users to create a user cluster 424 .
In some implementations, the user similarity determiner 420 generates the user cluster 424 based on the stored cluster data 422 including data aggregated for users in particular clusters. For example, the cluster data 422 may be grouped by certain parameters (eg, number of incorrect query responses, etc.) indicative of the passive language complexity associated with the plurality of users 402 .
After creating the user cluster 424, the complexity optimizer 430 varies the complexity of the language output by the TTS system and displays the user's performance in understanding the language output by the TTS system to indicate user capabilities. Measure the user's manual language complexity using a set of parameters. For example, parameters may be used to indicate how well users within each cluster 424 understand a given text-to-speech output. In the example above, complexity optimizer 430 initially provides a low complexity speech signal for the user and iteratively provides additional speech signals within a range of complexity.
Further, in some implementations, complexity optimizer 430 may determine an optimal passive language complexity for various user contexts associated with each user cluster 424 . For example, after measuring a user's language proficiency using the set of parameters, the complexity optimizer 430 may classify the measured data by context data received from the plurality of users 402 , thus Allows the optimal manual language complexity to be determined for each user context.
Then, after gathering performance data for a range of passive language complexities, the machine learning system 400 determines the particular passive language complexity for which the performance parameters indicate that the user's comprehension of the language is the best. For example, the machine learning system 400 collects performance data for all users within a particular user cluster 424 to determine the relationship between active language complexity, passive language complexity, and user context.
The aggregated data for the user cluster 424 is then compared to the individual data for each user in the user cluster 424 to determine the actual language complexity score for each user in the user cluster 424 . For example, as shown in Figure 4, complexity mapping 442 calculates the relationship between active language complexity and passive language complexity to infer the actual language complexity corresponding to the active language complexity mapped to the optimal passive language complexity. can be expressed
Complexity mapping 442 represents the relationship between active language complexity, TTS complexity and passive language complexity for all user clusters within the plurality of users 402, which then calculates the appropriate TTS complexity for subsequent queries by individual users. can be used to predict. For example, as described above, user inputs (eg, queries, text messages, emails, etc.) may be used to group similar users into a user cluster 424 . For each cluster, the system provides TTS outputs that require various levels of language proficiency to understand. The system then evaluates the responses received from users, the ratio of task completion to the various TTS outputs, to determine the appropriate level of language complexity for the users in each cluster. The system stores a mapping 442 between cluster identifiers and complexity scores corresponding to the identified clusters. The system then uses the complexity mapping 442 to determine an appropriate level of complexity for the TTS output to the user. For example, the system identifies a cluster representing the user's active language proficiency, queries the mapping 442 for a corresponding TTS complexity score (eg, indicative of a level of passive language comprehension) for the cluster, and Generate a TTS output having a complexity level indicated by the retrieved TTS complexity score.
The actual language complexity determined for the user can then be used to tune the TTS system using the techniques described with respect to FIGS. 1-3 . In this regard, language complexity data collected from a group of similar users (eg, user cluster 424 ) can be used to intelligently adjust the performance of the TTS system with respect to a single user.
5 is a flow diagram illustrating an example of a process 500 for adaptively generating text-to-speech output. Briefly, process 500 includes determining 510 a language proficiency of a user of a client device, determining 520 a text segment for output by a text-to-speech module, and a synthesized utterance of the text segment. It may include generating (530) the audio data including the and providing (540) the audio data to the client device.
In more detail, process 500 includes determining 510 a language proficiency of a user of the client device. For example, as described with respect to FIG. 1 , language proficiency predictor 110 may determine a language proficiency for a user using various techniques. In some examples, language proficiency may represent an assigned score indicative of a level of language proficiency. In another example, language proficiency may represent an assigned one of a plurality of categories of language proficiency. In another example, language proficiency may be determined based on user input and/or actions indicative of the user's proficiency level.
In some implementations, language proficiency may be inferred from different user signals. For example, as described with respect to FIG. 1 , language proficiency is the lexical complexity of user inputs, the rate of data input by the user, number of unrecognized words from speech input, TTs of completed voice actions for different levels of complexity. It can be inferred from the number or level of complexity of texts viewed by the user (eg, text on a book, article webpage, etc.).
Process 500 can include determining 520 a text segment for output by the text-to-speech module. For example, the TTS engine may adjust the default text segment based on the user's determined language proficiency. In some examples, as described with respect to FIG. 2 , a text segment for output may be adjusted based on a user context associated with the user. In another example, as described with respect to FIG. 3 , a text segment for output may be adjusted by word replacement or sentence restructuring to reduce the complexity of the text segment. For example, the adjustment may include how rare individual words are included in text segments, the types of verbs used (eg compound verbs, verb tenses), the linguistic structure of the text segment (eg number of dependent clauses). , the number of divisions between related words, the degree to which phrases are superimposed, and the like. Also, in another example, the adjustment may be based on the linguistic measures and a reference measure for linguistic features (eg, average distinction between subject and verb, distinction between adjective and noun, etc.). In the example above, the reference measures may represent averages, or may include ranges or examples for different complexity levels.
In some implementations, determining a text segment for output may include selecting text segments having scores that best match reference scores that describe the language proficiency level of the user. In another implementation, individual words or phrases may be scored for complexity, and the most complex words may be replaced, deleted, or restructured so that the overall complexity fits a level appropriate to the user.
Process 500 may include generating 530 audio data comprising a synthesized utterance of a text segment.
Process 500 may include providing 540 audio data for a client device.
6 is a block diagram of computing devices 600 , 650 that may be used to implement the systems and methods described herein, either as a client or as a server or a plurality of servers. Computing device 600 is intended to represent various types of digital computers such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. Computing device 650 is intended to represent various types of mobile devices, such as personal digital assistants, cellular telephones, smart phones, and other similar computing devices. Additionally, computing device 600 or 650 may include USB flash drives. USB flash drives can store operating systems and other applications. USB flash drives may include input/output components such as a USB connector that may be inserted into a USB port of a wireless transmitter or other computing device. The components shown herein, their connections and relationships, and their functionality, are meant to be illustrative only and are not meant to limit the implementation of the inventions described and/or claimed herein.
The computing device 600 includes a high-speed interface 608 and a low-speed bus 614 and a storage device coupled to the processor 602 , memory 604 , storage device 606 , memory 604 , and high-speed expansion port 610 . and a low-speed interface 612 coupled to 606 . Each of components 602 , 604 , 606 , 608 , 610 and 612 may be interconnected using various buses and mounted on a common motherboard or in other suitable manner. Processor 602 processes instructions for execution within computing device 600 , including instructions stored in memory 604 or storage device 606 , such as display 616 coupled to high-speed interface 608 . Graphical information about the GUI can be displayed on an external input/output device. In other implementations, multiple processors and/or multiple buses may be used, as appropriate, with multiple memories and multiple types of memory. Additionally, multiple computing devices 600 may be coupled with each device providing portions of the required operation (eg, a server bank, blade server group, or multi-processor system).
Memory 604 stores information within computing device 600 . In one implementation, memory 604 is a volatile memory unit or units. In another implementation, memory 604 is a non-volatile memory unit or units. Also, memory 604 may be another form of computer-readable medium, such as a magnetic or optical disk.
The storage device 606 may provide large storage for the computing device 600 . In one implementation, the storage device 606 is a device including a floppy disk device, a hard disk device, an optical disk device or tape device, a flash memory or other similar solid state memory device, or a device of a storage area network or other configuration. It may be or include a computer-readable medium, such as an array. The computer program product may be tangibly embodied in an information carrier. A computer program product may also include instructions that, when executed, perform one or more methods as described above. The information carrier is a computer or machine readable medium, such as memory 604 , storage device 606 , or memory on processor 602 .
High speed controller 608 manages bandwidth intensive operations for computing device 600 , while low speed controller 612 manages low bandwidth intensive operations. The assignment of these functions is exemplary only. In one implementation, high-speed controller 608 includes memory 704, display 716 (eg, via a graphics processor or accelerator), and a high-speed expansion port (not shown) that can accommodate various expansion cards (not shown). 610) is connected. In an implementation, the low-speed controller 612 is coupled to the storage device 606 and the low-speed expansion port 614 . A low-speed expansion port, which may include a variety of communication ports (eg, USB, Bluetooth, Ethernet, wireless Ethernet), may include one or more input/output devices such as a keyboard, pointing device, microphone/speaker pair, scanner, or network, for example. It may be coupled to a networking device such as a switch or router via an adapter. Computing device 600 may be implemented in a number of different forms as shown in the figures. For example, it may be implemented as a standard server 620 or many in a group of such servers. It may also be implemented as part of a rack server system 624 . It may also be implemented in a personal computer, such as a laptop computer 622 . Alternatively, components from computing device 600 may be combined with other components within a mobile device (not shown), such as device 650 . Each of the devices may include one or more of the computing devices 600 , 650 , and the overall system may be comprised of multiple computing devices 600 , 650 communicating with each other.
Computing device 600 may be implemented in a number of different forms as shown in the figures. For example, it may be implemented as a standard server 620 or many in a group of such servers. It may also be implemented as part of a rack server system 624 . It may also be implemented in a personal computer, such as a laptop computer 622 . Alternatively, components from computing device 600 may be combined with other components within a mobile device (not shown), such as device 650 . Each of the devices may include one or more of the computing devices 600 , 650 , and the overall system may be comprised of multiple computing devices 600 , 650 communicating with each other.
Computing device 650 includes processor 652 , memory 664 , input/output devices such as display 654 , communication interface 666 , and transceiver 668 , among other components. Device 650 may also be provided with a storage device, such as a micro drive or other device, to provide additional storage. Each of components 650 , 652 , 664 , 654 , 666 and 668 are interconnected using various buses, and several components may be mounted on a common motherboard or in other suitable manners.
Processor 652 may execute instructions in computing device 650 including instructions stored in memory 664 . The processor may be implemented as a chipset of chips comprising separate and multiple analog and digital processors. Additionally, a processor may be implemented using any number of architectures. For example, the processor 610 may be a Complex Instruction Set Computers (CISC) processor, a Reduced Instruction Set Computer (RISC) processor, or a Minimal Instruction Set Computer (MISC) processor. The processor may provide coordination of other components of the device 650 , such as user interfaces, applications executed by the device 650 and wireless communication by the device 650 , for example.
The processor 652 may communicate with a user via a control interface 658 and a display interface 656 coupled to the display 656 . Display 654 may include, for example, a TFT LCD (Thin Film Transistor Liquid Crystal Display) or OLED (Organic Light Emitting Diode) display or other suitable display technology. Display interface 786 may include suitable circuitry for driving display 654 to provide graphics and other information to a user. Control interface 658 may receive commands from a user and translate them for submission to processor 652 . Additionally, an external interface 662 may be provided for communication with the processor 652 to enable short-range communication of the device 650 with other devices. External interface 662 may be provided, for example, for wired communication in some implementations or for wireless communication in other implementations, and multiple interfaces may also be used.
Memory 664 stores information within computing device 650 . Memory 664 may be implemented in one or more of a computer-readable medium or media, a volatile memory unit or units, and a non-volatile memory unit or units. Expansion memory 774 may also be provided and connected to device 650 via expansion interface 772 , which may include, for example, a Single In Line Memory Module (SIMM) card interface. The expansion memory 674 may provide additional storage space for the device 650 , or may store applications or other information about the device 650 . In particular, the extended memory 674 may include instructions to perform or supplement the processes described above, and may also include security information. Thus, for example, expansion memory 674 may be provided as a secure module for device 650 , and may be programmed with instructions to allow secure use of device 650 . Additionally, secure applications may be provided with additional information via SIMM cards, such as placing identification information on the SIMM card in an unhackable manner.
The memory may include, for example, flash memory and/or NVRAM memory, as described below. In one implementation, the computer program product is tangibly embodied in an information carrier. The computer program product also includes instructions that, when executed, perform one or more methods as described above. The information carrier is a computer or machine readable medium, such as memory 664 , expansion memory 674 , or memory on processor 652 , which may be received via transceiver 668 or external interface 662 , for example.
Device 650 may communicate wirelessly via communication interface 666 , which may include digital signal processing circuitry as desired. Communication interface 666 may be provided for communication under various modes or protocols, such as GSM voice calls, SMS, EMS or MMS messaging, CDMA, TDMA, PDC, WCDMA, CDMA2000 or GPRS, among others. Such communication may occur, for example, via radio frequency transceiver 668 . Additionally, short-range communications may occur, such as using Bluetooth, Wi-Fi, or other transceivers (not shown). Additionally, a Global Positioning System (GPS) receiver module 670 may provide the device 650 with additional navigation and location related wireless data that may be suitably used by applications running on the device 780 .
Device 650 may also communicate audibly using audio codec 660 that may receive information spoken from a user and convert it into usable digital information. Audio codec 660 may likewise generate an audible sound for a user, such as through a speaker in a handset of device 650 , for example. Such sound may include sound from voice phone calls, recorded sound (eg, voice messages, music files, etc.), and also sound generated by applications running on device 650 . may include.
Computing device 650 may be implemented in a number of different forms as shown in the figures. For example, it may be implemented as a cellular phone 480 . It may also be implemented as part of a smartphone 682 , a personal digital assistant (PDA), or other similar mobile device.
Various implementations of the systems and methods described herein may be implemented in digital electronic circuitry, integrated circuits, specially designed application specific integrated circuits (ASICs), computer hardware, firmware, software, and/or combinations thereof. . These various implementations may include implementation in one or more computer programs executable and/or interpretable on a programmable system comprising at least one programmable processor, which may be dedicated or general purpose, and may be a storage system , may be coupled to receive data and instructions from, and transmit data and instructions to, at least one input device and at least one output device.
These computer programs (also known as programs, software, software applications or code) contain machine instructions for a programmable processor and may be implemented in high-level procedural language and/or object-oriented programming language and/or assembly/machine language. . As used herein, the terms "machine readable medium", "computer readable medium" include a machine readable medium that receives machine instructions as a machine readable signal, including programmable machine instructions and/or data. refers to any computer program product, apparatus and/or device used to provide a processor, eg, magnetic disk, optical disk, memory, programmable logic device (PLD). The term "machine-readable signal" refers to any signal used to provide machine instructions and/or data to a programmable processor.
In order to provide interaction with a user, the systems and techniques described herein provide a user with a display device such as, for example, a cathode ray tube (CRT) or liquid crystal display (LCD) monitor, and a display device to display information to the user. may be implemented in a computer with a keyboard and pointing device, such as a mouse or trackball, capable of providing input to the computer. Other types of devices may also be used to provide interaction with the user. For example, the feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and input from the user may be received in any form, including acoustic, voice, or tactile input. can
The systems and techniques described herein are compatible with implementations of the systems and techniques described herein, for example, by a user computer or user having a graphical user interface or a backend component such as a data server, a middleware component such as an application server. It may be implemented in a computing system comprising a front-end component, such as a web browser, with which it can interact, or any combination of one or more of the above-mentioned back-end, middleware or front-end components. The components of the system may be interconnected by any form or medium of digital data communication, for example, a communication network. Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
A computing system may include users and servers. A user and a server are usually remote from each other, and they usually interact through a communication network. The relationship between the user and the server arises by virtue of computer programs running on the respective computers and having a user-server relationship with each other.
A number of embodiments have been described. Nevertheless, it will be understood that various modifications may be made without departing from the spirit and scope of the present invention. Additionally, the logic flows depicted in the figures do not necessarily require a particular illustrated order, or time-series order, to achieve desired results. Additionally, other steps may be provided, steps may be omitted from the described flow, and other components may be added to or removed from the described system. Accordingly, other embodiments are also within the scope of the following claims.
7 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| JP2010145873A | Cites | Japan | Search report |
| US2015332665A1 | Cites | United States of America | Search report |
| JP2810750B2 | Cites | Japan | Search report |
36 members in 6 offices
Priority claims3
| Document | Office | Kind | Date |
|---|---|---|---|
| 15009432 | United States of America | – | |
| 201615009432 | United States of America | A | |
| 2016069182 | United States of America | W |
Members36
| Document | Office | Kind | |
|---|---|---|---|
| US2017221471A1 | United States of America | A1 | |
| US2017221472A1 | United States of America | A1 | |
| WO2017131924A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US9799324B2 | United States of America | B2 | |
| US2017316774A1 | United States of America | A1 | |
| US9886942B2 | United States of America | B2 | |
| KR20180098654AThis record | Republic of Korea | A | |
| EP3378059A1 | European Patent Office (EPO) | A1 | |
| CN108604446A | China | A | |
| US10109270B2 | United States of America | B2 | |
| US2019019501A1 | United States of America | A1 | |
| JP2019511034A | Japan | A | |
| US10453441B2 | United States of America | B2 | |
| US2020013387A1 | United States of America | A1 | |
| KR20200009133A | Republic of Korea | A | |
| KR20200009134A | Republic of Korea | A | |
| JP6727315B2 | Japan | B2 | |
| JP2020126262A | Japan | A | |
| US10923100B2 | United States of America | B2 | |
| KR102219274B1 | Republic of Korea | B1 | |
| KR20210021407A | Republic of Korea | A | |
| US2021142779A1 | United States of America | A1 | |
| JP6903787B2 | Japan | B2 | |
| JP2021144759A | Japan | A | |
| EP3378059B1 | European Patent Office (EPO) | B1 | |
| EP4002353A1 | European Patent Office (EPO) | A1 | |
| JP7202418B2 | Japan | B2 | |
| CN108604446B | China | B | |
| US11670281B2 | United States of America | B2 | |
| CN116504221A | China | A | |
| US2023267911A1 | United States of America | A1 | |
| EP4002353B1 | European Patent Office (EPO) | B1 | |
| EP4478349A2 | European Patent Office (EPO) | A2 | |
| US12198671B2 | United States of America | B2 | |
| EP4478349A3 | European Patent Office (EPO) | A3 | |
| US2025131909A1 | United States of America | A1 |
7 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Full renewal or maintenance fee paidU11 | U11 | |
| Written decision to grantGRNT | GRNT | |
| Divisional application of patentA107 | A107 | |
| Decision to grant or registration of patent rightE701 | E701 | |
| Notification of reason for final refusalE90F | E90F | |
| Notification of reason for refusalE902 | E902 | |
| Request for examinationA201 | A201 |
Numbers
- Publication
- 10-2018-0098654
- Application
- 1020187021923
Titles4
- Korean
- 적응적 텍스트-투-스피치 출력
- English
- Adaptive text-to-speech output
- Unlabeled
- 적응적 텍스트-투-스피치 출력
- Unlabeled
- Adaptive text-to-speech output
Classification
- CPC, 7
- G10L13/043
- G10L13/08
- G06F40/253
- G10L13/00
- G06F17/274
- G06F40/289
- G06F17/2775
- IPC, 3
- G10L13 04
- G06F17 27
- G10L13 08