EP2804113A2

Hybrid, offline/online speech translation system

Abstract

A hybrid speech translation system whereby a wireless-enabled client computing device can, in an offline mode, translate input speech utterances from one language to another locally, and also, in an online mode when there is wireless network connectivity, have a remote computer perform the translation and transmit it back to the client computing device via the wireless network for audible outputting by client computing device. The user of the client computing device can transition between modes or the transition can be automatic based on user preferences or settings. The back-end speech translation server system can adapt the various recognition and translation models used by the client computing device in the offline mode based on analysis of user data over time, to thereby configure the client computing device with scaled-down, yet more efficient and faster, models than the back-end speech translation server system, while still be adapted for the user's domain.

EP2804113A2, drawing sheet 1
Sheet 1 of 9

Term

Projected expiry 7 May 2034.

  1. Priority
  2. Filed
  3. Published
  4. Today
  5. Projected expiry

15 claims: 11 independent, 4 dependent

  1. 1
    A speech translation system comprising:a back-end speech translation server system;and a client computing device that is configured for communicating with the back-end speech translation server system via a wireless network, wherein the client computing device comprises: a microphone;a processor connected to the microphone;a memory connected to the processor that stores instructions to be executed by the processor;and a speaker connected to the processor, wherein: the client computing device is for outputting via the speaker a translation of input word phrases from a first language to a second language;and the memory stores instructions such that: in a first operating mode, when the processor executes the instructions, the processor translates the input word phrases to the second language for output to a user;and in a second operating mode: the client computing device transmits to the back-end speech translation server system, via the wireless network, data regarding the input word phrases in the first language received by the client computing device;the back-end speech translation server system determines the translation to the second language of the input word phrases in the first language based on the data received via the wireless network from the client computing device;and the back-end speech translation system transmits data regarding the translation to the second language of the input word phrases in the first language to the client computing device via the wireless network such that the client computing device outputs the translation to the second language of the input word phrases in the first language;wherein the client computing device has a user interface that permits a user to switch between the first operating mode and the second operating mode and/or wherein the client computing device automatically selects whether to use the first operating mode or the second operating mode based on a connection status for the wireless network or on a user preference setting of the user for the client computing device.
  2. 3
    The speech translation system of claims 1 or 2, wherein the client computing device outputs the translations audibly via the speaker.
  3. 4
    The speech translation system of any of claims 1 to 3, wherein:the client computing device stores in memory a local acoustic model, a local language model, a local translation model and a local speech synthesis model for, in the first operating mode, recognizing the speech utterances in the first language and translating the recognized speech utterances to the second language for output via the speaker of the client computing device;the back-end speech translation server system comprises a back-end acoustic model, a back-end language model, a back-end translation model and a back-end speech synthesis model for, in the second operating mode, determining the translation to the second language of the speech utterances in the first language based on the data received via the wireless network from the client computing device;the local acoustic model is different from the back-end acoustic model;the local language model is different from the back-end language model;the local translation model is different from the back-end translation model;and the local speech synthesis model is different from the back-end speech synthesis model.
  4. 5
    The speech translation system of any of claims 1 to 4, wherein the back-end speech translation server system is programmed to:monitor over time speech utterances received by the client computing device for translation from the first language to the second language;and update at least one of the local acoustic model, the local language model, the local translation model and the local speech synthesis model of the client computing device based on the monitoring over time of speech utterances received by the client computing device for translation from the first language to the second language, wherein updates to the at least one of the local acoustic model, the local language model, the local translation model and the local speech synthesis model of the client computing device are transmitted from the back- end speech translation server system to the client computing device via the wireless network.
  5. 6
    The speech translation system of any of claims 1 to 5, wherein the local acoustic model, the local language model, the local translation model and the local speech synthesis model of the client computing device are updated based on analysis of translation queries by the user.
  6. 7
    The speech translation system of any of claims 1 to 6, wherein:the client computing device comprises a GPS system for determining a location of the client computing device;and the back-end speech translation server system is programmed to update at least one of the local acoustic model, the local language model, the local translation model and the local speech synthesis model of the client computing device based on the location of the client computing device, wherein updates to the at least one of local acoustic model, the local language model, the local translation model and the local speech synthesis model of the client computing device are transmitted from the back end speech translation server system to the client computing device via the wireless network.
  7. 8
    The speech translation system of claims 1 to 7, wherein:the back-end speech translation server system is one of a plurality of a back-end speech translation server systems, and the client computing device is configured for communicating with the each of the plurality of back-end speech translation server systems via a wireless network;and in the second operating mode: each of the plurality of back end speech translation server systems is for determining a translation to the second language of the speech utterances in the first language based on the data received via the wireless network from the client computing device;and one of the plurality of back-end speech translation server systems selects one of the translations from the plurality of back-end speech translation server systems for transmitting to the client computing device or merges two or more of the translations from the plurality of back-end speech translation server systems to generate a merged translation for transmitting to the client computing device.
  8. 9
    A speech translation method comprising:in a first operating mode: receiving by a client computing device a first input word phrase in a first language;translating by the client computing device the first input word phrase to a second language;and outputting by the client computing device the first input word phrase in the second language;transitioning by the client computing device from the first operating mode to the second operating mode;in the second operating mode: receiving by a client computing device a second input word phrase in a first language;transmitting, by the client computing device, via a wireless network, data regarding the second input word phrase to a back-end speech translation server system;receiving, by the client computing device, from the back-end speech translation server system via the wireless network, data regarding a translation by the back-end speech translation server system of the second input word phrase from the first language to the second language;and outputting by the client computing device the second input word phrase in the second language.
  9. 11
    The speech translation method of any of claims 1 to 10, further comprising downloading by the client computing device application software for a language translation pair that comprises the first and second languages, in particular wherein downloading the application software for the language translation pair comprises downloading the application software for the language translation pair when suitable connectivity between the client computing device and the back-end speech translation server system is available via the wireless network.
  10. 13
    The speech translation method of any of claims 9 to 12, wherein:the client computing device comprises a graphical user interface have a first language display section and a second language display section that are simultaneously displayed;and each of the first and second language display sections comprise a user-accessible listing of a plurality of languages, and the method further comprises the step of receiving by the client computing device via the graphical user interface a selection of the first language from the listing in the first language display section and the second language in the second language display section, such that the client computing device is thereby configured to translate the input speech utterances from the first language to the second languages.
  11. 15
    The speech translation method of any of claims 9 to 14, wherein transitioning by the client computing device from the first operating mode to the second operating mode is responsive to an input via a user interface of the client computing device to transition from the first mode to the second mode.