Hybrid, offline/online speech translation system.
Abstract
A hybrid speech translation system whereby a wireless-enabled client computing device can, in an offline mode, translate input speech utterances from one language to another locally, and also, in an online mode when there is wireless network connectivity, have a remote computer perform the translation and transmit it back to the client computing device via the wireless network for audible outputting by client computing de vice. The user of the client computing device can transition between modes or the transition can be automatic based on user preferences or settings. The back-end speech translation server system can adapt the various recognition and translation models used by the client computing device in the offline mode based on analysis of user data over time, to thereby configure the client computing device with scaled-down, yet more efficient and faster, models than the back-end speech translation server system, while still be adapted for the user's domain.

Term
7.6 yearsleft in the term
Expires 1 May 2034.
- Priority
- Filed
- Granted
- Today
- Expires
15 claims: 3 independent, 12 dependent
- 1CLAIMS REIVINDICACIONES IMPI MAID INSTITUTO MEXICANO DE LA PROHEDAI > INDUSTRIAL INSTITUTO MEXICANO DE LA PROHEDAI> INDUSTRIAL 1. A speech translation system characterized by comprising:1. Un sistema de traducción de habla caracterizado porque comprende: a translation server system;and a client device that is configured to communicate with the translation server system, wherein the client device comprises: un sistema de servidor de traducción;y un dispositivo de cliente que se configura para comunicarse con el sistema de servidor de traducción, en donde el dispositivo de cliente comprende: a microphone;un micrófono;a processor connected to the microphone;un procesador conectado al micrófono;a memory connected to the processor;and a speaker connected to the processor, where: una memoria conectada al procesador;y un altavoz conectado al procesador, en donde: el dispositivo de cliente es para producir mediante el altavoz una traducción de frases verbales de entrada de un primer idioma a un segundo idioma;y la memoria provoca que el procesador pueda: the client device is to produce by the loudspeaker a translation of input verb phrases from a first language to a second language;and memory causes the processor to: determinar el segundo idioma para las frases de entrada recibidas en el dispositivo de cliente desde un usuario del dispositivo de cliente;determining the second language for input phrases received at the client device from a user of the client device;receive from the user a translation setting mode for the client device for the translation of input verb phrases into the second specified language, the translation setting mode includes a privacy preference for use of the translation server only if a network of wireless security is available;recibir del usuario un modo de configuración de traducción para el dispositivo de cliente para la traducción de las frases verbales de entrada al segundo idioma determinado, el modo de configuración de traducción comprende una preferencia de privacidad de uso del servidor de traducción sólo si una red de seguridad inalámbrica está disponible;In response to determining that a wireless network is not available, automatically select to carry out the translation on the client device, by means of the processor that translates the verbal phrases of input words from the first language to the second language and produce at user in the second language a local translation of the input verb phrases;and in response to determining that a wireless security network is available, automatically select to carry out the translation on the translation server: en repuesta a determinar que una red de se inalámbrica no está disponible, seleccionar automáticamente el llevar a cabo la traducción en el dispositivo de cliente, por medio del procesador que traduce las frases verbales de palabras de entrada del primer idioma al segundo idioma y producir al usuario en el segundo idioma una traducción local de las frases verbales de entrada;y en respuesta a determinar que una red de seguridad inalámbrica está disponible, seleccionar automáticamente llevar a cabo la traducción en el servidor de traducción: el dispositivo de cliente transmite al servidor de traducción información asociada con las frases de palabras de entrada en el primer idioma recibido por el dispositivo de cliente;the client device transmits to the translation server information associated with the input word phrases in the first language received by the client device;el servidor de traducción determina la traducción al segundo idioma de las frases de palabras de entrada en el primer idioma basándose en los datos recibidos mediante la red inalámbrica desde el dispositivo de cliente;y el sistema de traducción transmite datos en relación con la traducción al segundo idioma de las frases de palabras de entrada en el primer idioma al the translation server determines the second language translation of the input word phrases in the first language based on data received via the wireless network from the client device;and the translation system transmits data regarding the second language translation of the input word phrases in the first language to the IMPI MAID INSTITUTO MEXICANO DE LA PROPIEDAD MEXICAN INSTITUTE OF PROPERTY .. · Χ · J | · X -I _ I _1 client device so that the client device produces the translation to according to 3 or Ί di or ma effe the input word phrases in the first language;.. ·Χ· J |· X -I _ I _1 dispositivo de cliente de modo que el dispositivo cliente produce la traducción al según 3 o Ί di o ma efe las frases de palabras de entrada en el primer idioma;el servidor de traducción monitorea a través del tiempo expresiones orales recibidas por el dispositivo de cliente para traducción del primer idioma al segundo idioma;the translation server monitors over time oral expressions received by the client device for translation from the first language to the second language;el servidor de traducción determina, basándose en las expresiones orales monitoreadas, vocabulario empleado por el usuario;y el servidor de traducción actualiza, basándose en el vocabulario determinado, al menos uno de un modelo acústico local, un modelo de idioma local, un modelo de traducción local y un modelo de síntesis de habla local del dispositivo de cliente, en donde el actualizar al menos uno del modelo acústico local, el modelo de idioma local, el modelo de traducción local y el modelo de síntesis de habla local del dispositivo de cliente se transmiten desde el servidor de traducción al dispositivo de cliente mediante la red inalámbrica. the translation server determines, based on the monitored oral expressions, vocabulary used by the user;and the translation server updates, based on the determined vocabulary, at least one of a local acoustic model, a local language model, a local translation model and a local speech synthesis model of the client device, wherein the update at least one of the local acoustic model, the local language model, the local translation model and the local speech synthesis model of the client device are transmitted from the translation server to the client device via the wireless network.
- 7A speech translation method characterized by comprising:7. Un método de traducción de habla caracterizado porque comprende: receiving at a client device, from a user of the client device, an input verb phrase in a first language;recibir en un dispositivo de cliente, de un usuario del dispositivo de cliente, una frase verbal de entrada en un primer idioma;determinar un segundo idioma para traducción de la frase determine a second language for translation of the sentence IMPIég ^ IMPIég^ INSTITUTO MEXICANO de La propiedad verbal de entrada;wowmu» *» recibir del usuario un mode-d-e-’CürTfTgiiracion de traducción para el dispositivo de cliente para la traducción de las frases verbales de entrada al segundo idioma determinado, el modo de configuración de traducción comprende una preferencia de privacidad de uso del servidor de traducción sólo si una red de seguridad inalámbrica está disponible;INSTITUTO MEXICANO de The verbal property of entry;wowmu »*» receive from the user a translation-mode-'CürTfTrotate for the client device for the translation of the input verb phrases to the second specified language, the translation configuration mode includes a privacy preference for server use translation only if a wireless security network is available;In response to determining that a wireless security network is not available, automatically select, carry out the translation on the client device by: en repuesta a determinar que una red de seguridad inalámbrica no está disponible, seleccionar automáticamente, llevar a cabo la traducción en el dispositivo de cliente por: traducir por el dispositivo de cliente la frase verbal de entrada del primer idioma al segundo idioma;y producir en el segundo idioma una traducción local de la frase verbal de entrada;y en respuesta a determinar que una red de seguridad inalámbrica está disponible, seleccionar automáticamente el llevar a cabo la traducción en el servidor de traducción por: translating by the client device the input verb phrase from the first language to the second language;and produce a local translation of the input verb phrase in the second language;and in response to determining that a wireless security network is available, automatically select to carry out the translation on the translation server by: transmitir del dispositivo de cliente al servidor de traducción información asociada con la frase verbal de entrada;transmitting information associated with the input verb phrase from the client device to the translation server;receive, on the client device, data associated with a translation of the verb phrase of recibir, en el dispositivo de cliente, datos asociados con una traducción de la frase verbal de entrada del primer idioma al segundo idioma por el servidor de traducción;y producir en el segundo idioma el servidor de traducción de la frase verbal de entrada;input from the first language to the second language by the translation server;and producing the translation server of the input verb phrase in the second language;monitor, over time, oral expressions received by the client device for translation from the first language to the second language;monitorear, a través, del tiempo expresiones orales recibidas por el dispositivo de cliente para traducción del primer idioma al segundo idioma;determinar, basándose en las expresiones orales monitoreadas, vocabulario empleado por el usuario;y actualizar, basándose en el vocabulario determinado, al menos uno de un modelo acústico local, un modelo de idioma local, un modelo de traducción local y un modelo de síntesis de habla local del dispositivo de cliente, en donde el actualizar al menos uno del modelo acústico local, el modelo de idioma local, el modelo de traducción local y el modelo de síntesis de habla local del dispositivo de cliente se transmiten desde el servidor de traducción al dispositivo de cliente mediante la red inalámbrica. determine, based on the monitored oral expressions, vocabulary used by the user;and updating, based on the determined vocabulary, at least one of a local acoustic model, a local language model, a local translation model, and a local speech synthesis model of the client device, wherein updating at least one of the local acoustic model, local language model, the local translation model and the local speech synthesis model of the client device are transmitted from the translation server to the client device via the wireless network.
- 9The method in accordance with the medical report, further characterized in that it comprises:...... 9. El método de conformidad con la reTvmdica'C^ caracterizado además porque comprende: ...... determinar por el dispositivo de cliente una ubicación del dispositivo de cliente;y descargar por el dispositivo de cliente la aplicación para la combinación de idiomas de traducción basándose en la ubicación determinada del dispositivo de cliente y cuando haya conectividad adecuada disponible entre el dispositivo de cliente y el servidor de traducción mediante la red inalámbrica. determining by the client device a location of the client device;and downloading by the client device the application for the translation language combination based on the determined location of the client device and when adequate connectivity is available between the client device and the translation server via the wireless network.
Independent claims3
240 paragraphs in 25 sections, as filed
(54) Title: SPEECH TRANSLATION SYSTEM OFFLINE / ONLINE, HYBRID.
(54) Title: HYBRID, OFFLINE / ONLINE SPEECH TRANSLATION SYSTEM.
(57) Summary
A hybrid speech translation system whereby a wirelessly enabled client computing device can, in an offline mode, translate input spoken expressions from one language to another locally, and also, in an online mode, where it exists a wireless network connectivity , has a remote computer that performs the translation and transmits it back to the client computing device over the wireless network for audible output by the client computing device. The user of the client computing device may transition between modes or the transition may be automatic based on user preferences or settings. The output terminal speech translation server system can adapt the various recognition and translation models used by the client computing device in the offline mode based on an analysis of user data over time, thereby Therefore, configure the client computing device with scaled-down models, even more efficient and faster, than the output terminal speech translation server system, while still adapting to the user's domain.
(57) Abstract
A hybrid speech translation system whereby a wireless-enabled Client computing device can, in an offline mode, transiate input speech utterances from one language to another locally, and also, in an Online mode when there is wireless network connectivity, have a remote Computer perform the translation and transmit it back to the Client computing device via the wireless network for audible outputting by Client computing de vice. The user of the Client computing device can transition between modes or the transition can be automatic based on user preferences or settings. The back-end speech translation server system can adapt the various recognition and translation models used by the Client computing device in the offline mode based on analysis of user data over time, to thereby configure the Client computing device with scaleddown, yet more efficient and faster , models than the back-end speech translation server system, while still be adapted for the user's domain.
IMPI ¿<· <·> '** ·
PATENT TITLE No. 348169
Owner (s): FACEBOOK, INC.
Address: 1601 Willow Road, Menlo Park, California, 94025, USA
Denomination: HYBRID OFFLINE / ONLINE SPEAKING TRANSLATION SYSTEM.
Classification: CIP: G06F17 / 28; G10L13 / 00; G10L15 / 30
CPC: G06F17 / 289; Ó10L1W; Ó10L13 / 00; G10L15 / 30 ', I
Inventor (s): NAOMI AOKI WAIBEL; ALEXANDER WA1BEL; CHRISTIAN FUEGEN; KAY
ROTTMANN
REQUEST
Number:
n International:
• J · l V - <
MX / a / 201M) 15799
Country:
US US
Validity: V & full years
Date of VtMlóilmientth Tde May <4. 2034 lili fu
H ^ ro: 31 / 822,629 W "l 'i.
'L>'
F chadeExpIMfQlón: 31ttenwo ^^ / f Γ * '7 <*
The patent of references, grants with foundation in losrtelos 1 ?, 2 * fracc * V, 6 ° flaca # W, / 59 of the Law of the ^ ropewrindustrial.
_. ......— *. . t. ..TO . . ... .
Pursuant to Article 23 of the Industrial Property Law, the Pfe ^ nte de veirtfe aJÍSs ^ extendable, counted from the date of submission of the Application and will be subject to the payment of the tariff p ^ w «to keep the right ^>
Whoever subscribes this title does so based on the provisions of the article 6 "shares III and 7 * bis 2" of the Industrial Property Law (Official Gazette of the Federation (DOF) 06/27/1991, amended on 02 / 08/1994, | 5 / W19e «26 ^ 2/1997. T7 / Q ^ 1999, 26/01/2004, 16/06/2005, 25/01/2006, 06/05 / 2009,06 / 01 / 2010, 188) 80118. I6 / 06 / W10, 27 / 0WQ12 and & 9/84/20 Articles 1<sup>or</sup>, 3 * frpptl6'n / jlltiso a), 4th and 12th fractions I and III of the Regulations of the Mexican Institute of the PnjyHÍdpd, Iri ^ sMil (DOF TtflfSTWS refoijn «E> et 01 / (S * / 200Z'15 / 07/2004, 07/28/2004 and 09/07/2007); articles 1, 3, 4, 5 fraction V subsection a), 16 ftpcáppfls by 18 and «| Mel f stgMp¡ £ tampp ^ de Industrial Property (DOF
12/27/1999, amended 10/10/2002, 07/29/2004, 044) 8 / ^ 904 y ^ 09β6 (&), Λ *> 8th Agreement that delegates powers to the Directors
Deputy Generals, Coordinator, DivisioMes Directors, Titiftrap. <K> 8¾ Official Offices, Divisional Deputy Directors, Departmental Coordinators, and other subordinates of the Mexican Institute of Civil Law. (DOF 12/15/1999, amended on 02/04/2000, 07/29/2004, 08/04/2004 and 09/13/2007).
This document is signed with an advanced electronic signature (FIEL), based on articles 7 BIS 2 of the Industrial Property Law; 3 of its Regulations, and 1 section III, 2 section V, 26 BIS and 26 TER of the Agreement establishing the guidelines for the use of the Payment and Electronic Services Portal (PASE) of the Mexican Institute of Industrial Property, in the procedures indicated
THE DIVISIONAL DIRECTOR OF PATENTS
NAHANNY CANAL REYES
<img file="MX348169B_D0001.tif" />
Original string:
NAHANNY MARISOL CANAL REYES | 00001000000403252793 | Tax Administration Service | 1695 || MX / 2017/43602 | MX / a / 2015/015799 | PCT patent title | 1488 | IAR | Page (s) | n51 HqORVLuMKnQOYI EKlsqWgV0
Digital stamp:
LRsWj4Mq9eq9ojVqgytZxu62nTolYcpNHqh / 6DObFu3rPY6d7l6KVS / OvQWXR2V5TI + UzoOObuG2MkFHsnTZOvAc + r
V1mCDXI7OMjPIGmk7hmxOGjurHFGHiZ6DRPFiqM1doanEBRLmvSuDesgm5OWfk2ul8TcK5WUUBqDMeTaliz0SizEt2 6K6F2o7uXsesoSmTQIgU9VcwoDPdZimBoSI / yDndEBJBWIZT + N4ehoOnHgQbuVcdxJcsx9O9Sa1jgwCFsl7gKCm791 UbEkBfT / SqwgWIZN / ke5gwtZ2 / sWzTAIVBT1qQEb + == YMRw3TLL9UjOTPU3pGeSw1EKZrwzGlw
Arenal No 550, Piso 1, Pueblo Santa María Tepepan, Xochimilco, 16020, Mexico City (55) 53340700 ww gob mx / impi
IHIIIIlillll
MX / 2017/43602
-MAID
INSTITUTO MEXICANO DI LA PROPERTY INDUSTRIAL
<img file="MX348169B_D0002.tif" />
OFFLINE / ONLINE SPEECH TRANSLATION SYSTEM.
HYBRID
Background of the invention
Speech-to-speech translation systems (STS) are usually delivered in one of two different ways: online via the Internet or integrated offline into a user device (e.g. smartphone or smartphone). other suitable computing device). The online version has the advantage that it can take advantage of significant processing sources on a large server (the cloud), and it provides a data feed to the service provider that makes enhancements and customization possible. However, online processing requires continuous network connectivity, which cannot be guaranteed 15 in all locations or is not desired in some cases due to roaming costs or privacy / security concerns. As an alternative implementation, speech-to-speech translators, such as Jibbigo's speech translation applications, can be supplied as software that runs integrated locally on the same smartphone, and no network connectivity is needed after the initial download of the translation app. Such built-in offline speech translation capability is the preferred implementation for many, if not most, of practical situations where language support is needed, as networks may not be available, they are
<img file="MX348169B_D0003.tif" />
intermittent or very expensive. Many travelers experience such intermittent or absent connectivity, for example, during airline flights, remote geographic locations, buildings, or simply because data roaming was disabled to avoid associated roaming charges while traveling in a foreign country.
The way in which speech translation software or devices are delivered also has implications as to whether the software can / should operate in a domain-dependent or independent manner and whether it can adapt the context of the user. STS systems will usually work quite well for one domain and not so well for another domain (domain dependency) if they have been optimized and tightly tuned for a specific domain of use, or if they attempt domain independence by working more or less equally well for all domain. Either solution limits performance for all specific situations.
A user commonly runs an online client program on his computing device. This device typically digitizes and possibly encodes speech, then transmits samples or coefficients over a communication line to a server. The server then performs high computing power speech recognition and / or translation and sends the result back to the user via a communication line, and the result is displayed on the user device. Different inline designs have been proposed that move different parts of a chain of
MAID
INSTITUTO MEXICANO D £ LA PROPERTY INDUSTRIAL processing outside the server and do more or less computation work on the device. In speech recognition, translation, and translation systems, the user device can be as simple as a microphone, or an analog-to-digital converter, or provide more complex functions such as noise suppression, encoding as coefficients, one or more phases. speech recognition or one or more stages of language processing. In contrast, an offline design runs the entire application on the same device as an integrated application. All computing is done locally on the device and no transmission is required between a client and a server during use.
Typically an online design has the advantage that it needs only a very simple client and thus an application can run on a very simple computing device or mobile phone, while all the high computing and processing capabilities are done on a computing server. larger. For speech and machine translation this may mean that more advanced but computationally intensive algorithms can be used, and up-to-date background information can be used. It also has the advantage that the developer or operator of the Service can maintain / improve the service or capacity on the server, without requiring the user to download or update new versions of the system.
The downside to an online design is the fact that it is critically dependent on network connectivity. As a user moves or travels to remote locations, however, connectivity can be intermittent and / or very expensive (roaming), and therefore, in many ways, unavailable. For speech and speech translation systems this requirement is often unacceptable. Unlike text or email transmissions, voice cannot allow for a time lag or connectivity just as it cannot allow a corresponding interruption of speech transmission without losing information or real-time performance. An online design can, therefore, ensure continuous real-time transmission and thus continuous connectivity during use.
Brief description of the invention
In a general aspect, the present invention describes a hybrid speech translation system whereby a wirelessly enabled client computing device (e.g., a smartphone or tablet computer) can translate input word phrases (e.g. example, input spoken expressions or input text) from one language to another locally, for example in an "offline" mode, and also, in an "online" mode where there is wireless network connectivity, have a remote computer, for example an output terminal speech translation server system, perform the translation and transmit it back to the client computing device over the wireless network for production by the customer computing device (e.g. audible through a loudspeaker and / or through a
<img file="MX348169B_D0004.tif" />
MAID
INSTITUTE MBUCANO
OF THE INDUSTRIAL ntONEOAD field of text display). In various modes, the client computing device user may transition between modes or the transition may be automatic and transparent to the user based on user preferences or settings. Furthermore, the output terminal speech translation server system can adapt the various speech translation models used by the client computing device in the offline mode based on analysis of user data over time, to configure accordingly. this mode the client computing device with scaled-down models, even more efficient and faster, than the output terminal speech translation server system, while still adapting to the user's domain.
The embodiments according to the invention are particularly described in the appended claims which focus on a language translation system and method, in which any feature mentioned in a category of the claims, for example the method, can also be used in another category of claims, for example a system, too. The dependencies or references in the appended claims and the modalities listed below were chosen for purely formal reasons. However, any subject in question that results from a reference mentioned corresponding to any previous claim or previous modality (in particular multiple units) can also be claimed, so that any combination of claims and their characteristics is described and can be claimed independently of the dependencies chosen in the appended claims. TamoieN btí - describes any combination of characteristics of the modalities listed below, without taking into account the above references mentioned herein.
In one embodiment according to the invention, a speech translation system comprises:
- an output terminal speech translation server system; and
- a client computing device that is configured to communicate with the output terminal speech translation server system via a wireless network, wherein the client computing device comprises:
a microphone;
a processor connected to the microphone;
a memory connected to the processor that stores instructions to be executed by the processor; and a speaker connected to the processor, where:
the client computing device is to produce by the loudspeaker a translation of input word phrases from a first language to a second language; and
- memory stores instructions so that:
In a first mode of operation, when the processor executes the instructions, the processor translates the input word phrases into the second
<img file="MX348169B_D0005.tif" />
language to produce a user; and in a second mode of operation:
- the client computing device transmits to the outgoing terminal speech translation server system, via the wireless network, data regarding the input word phrases in the first language received by the client computing device;
- the output terminal speech translation server system determines the second language translation of the input word phrases in the first language based on the data received via the wireless network from the client computing device; and
- the output terminal speech translation system transmits data regarding the second language translation of the input word phrases in the first language to the client computing device via the wireless network so that the client computing device produces the second language translation of input word phrases in the first language.
The client computing device may have a user interface that allows a user to switch between the first mode of operation.
<img file="MX348169B_D0006.tif" />
operation and the second mode of operation.
The client computing device can automatically select whether to use the first mode of operation or the second mode of operation based on a connection status for the wireless network.
Alternatively, the client computing device may automatically select whether to use the first mode of operation or the second mode of operation based on a user preference setting of the user for the client computing device.
In a further embodiment according to the invention, the input word phrases are entered into the client computing device through one of:
- input oral expressions captured by the microphone of the client computing device; or text input using a text input field in a user interface of the client computing device.
The client computing device can output the translations audibly through the loudspeaker.
In the speech translation system of the present invention, the client computing device can store in memory a local acoustic model, a local language model, a local translation model and a local speech synthesis model for, in the first mode of operation, recognize oral expressions in the first language and translate the recognized oral expressions into
<img file="MX348169B_D0007.tif" />
second language to be produced by the speaker of the client computing device.
The output terminal speech translation server system may comprise an output terminal acoustic model, an output terminal language model, an output terminal translation model, and an output terminal speech synthesis model. , in the second mode of operation, determines the second language translation of the oral expressions in the first language based on the data received via the wireless network from the client computing device.
Preferably, the local acoustic model may be different from the output terminal acoustic model;
the local language model may be different from the output terminal language model;
the local translation model may be different from the output terminal translation model; and the local speech synthesis model may be different from the output terminal speech synthesis model.
In addition, the output terminal speech translation server system can be programmed to: monitor oral expressions received by the client computing device over time for translation from first language to second language and update at least one of the local acoustic model, local language model, local translation model and local speech synthesis model of the client computing device based on monitoring at <sup>10</sup> ungodly
INSTITUTO MEXICANO VÍZíZiZ
OF THE INDUSTRIAL PROPERTY through the time of oral expressions received-onr the client computing device for translation from the first language to the second language, where the updates in at least one of the local acoustic model, the local language model, the local translation model and the local speech synthesis model of the client computing device are transmitted from the output terminal translation server system to the client computing device via the wireless network.
The local acoustic model, the local language model, the local translation model, and the local speech synthesis model of the client computing device can be updated based on the analysis of translation queries by the user.
The client computing device may comprise a GPS system for determining a location of the client computing device.
The output terminal speech translation server system may be further programmed to update at least one of the local acoustic model, the local language model, the local translation model, and the local speech synthesis model of the client computing device based on at the location of the client computing device, where updates to at least one of the local acoustic model, the local language model, the local translation model and the local speaking synthesis model of the client computing device are transmitted from the output terminal translation server system to the device
<img file="MX348169B_D0008.tif" />
<img file="MX348169B_D0009.tif" />
client computing over the wireless network.
Furthermore, the output terminal speech translation server system may be one of a plurality of output terminal speech translation server systems, and the client computing device may be configured to communicate with each of the plurality of outbound terminal speech translation server systems using a wireless network.
In the second mode of operation, each of the plurality of output terminal speech translation server systems may be to determine a second language translation of oral expressions in the first language based on data received via the wireless network. from the client computing device; and one of a plurality of output terminal speech translation server systems may select one of the translations of the plurality of output terminal speech translation server systems to transmit to the client computing device.
Alternatively, one of the plurality of output terminal speech translation server systems merges two or more of the translations of the plurality of output terminal speech translation server systems to generate a merged translation for transmission to the computing device. customer.
According to another aspect of the present invention, a speech translation method is described, the speech translation method comprises:
IMPI Mexican institute of industrial PROPERTY
<img file="MX348169B_D0010.tif" />
in a first mode of operation:
- receiving by a client computing device a first sentence of words in a first language;
- translating by the client computing device the first sentence of words into a second language; and
- produce by the client computing device the first sentence of words in a second language;
- making a transition by the client computing device from the first mode of operation to the second mode of operation;
- in the second mode of operation:
- receiving by a client computing device a second sentence of words in a first language;
- transmitting, by the client computing device, via a wireless network, data related to the second sentence of input words to an output terminal speech translation server system;
receiving, by the client computing device, from the output terminal speech translation server system, via the wireless network, data related to a translation by the output terminal speech translation system of the second sentence of input words from the first language to the second language; and
- produce by the client computing device the second sentence of input words in the second language. In a further embodiment of the invention, the device<sup>13</sup>
INSTHUTO MEXICANO DE LA PROPIEDAD INDUSTBIA The client computer stores in memory a local acoustic model, a local language model, a local translation model and a local speech synthesis model, in order, in the first mode of operation, to recognize oral expressions input in the first language, and translating the recognized input spoken expressions into the second language to be produced by the loudspeaker and the preferred output terminal speech translation server system comprises an output terminal acoustic model, an output terminal language model, a output terminal translation model and an output terminal speech synthesis model for, in the second mode of operation, determining the second language translation of the input spoken expressions in the first language based on the data received over the wireless network from the client computing device.
The method can also comprise the stages of:
- monitoring, by the output terminal speech translation server system, over time oral expressions received by the client computing device for translation from the first language to the second language; and
- updating by the output terminal speech translation server system at least one of the local acoustic model, the local language model, the local translation model and the local speech synthesis model of the client computing device based on the monitoring over time of oral expressions received by the client's computing device for translation of the
MAID
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<img file="MX348169B_D0011.tif" />
first language to the second language, where updates in at least one of the local acoustic model, the local language model, the local translation model, and the local speech synthesis model of the client computing device are transmitted from the client computing system. Outbound terminal translation server to client computing device via wireless network.
The speech translation method may comprise a step for downloading, by the client computing device, application software for a translation language combination comprising the first and second languages.
The step of downloading the application software for the translation language combination may comprise downloading the application software for the translation language combination when adequate connectivity is available between the client computing device and the speech translation server system of output terminal via the wireless network.
In a further embodiment of the invention, the speech translation method may comprise:
- determining by the client computing device a location of the client computing device; and
- download by the client computing device the application software for the translation language combination based on the determined location of the client computing device and where adequate con ctivity is available between the device
<img file="MX348169B_D0012.tif" />
<img file="MX348169B_D0013.tif" />
output terminal via the wireless network.
Preferably, the client computing device may comprise a graphical user interface having a first language display section and a second language display section that are displayed simultaneously; and
- each of the first and second language display sections may comprise a user-accessible list of a plurality of languages.
The method may further comprise the step of receiving by the client computing device, via the graphical user interface, a selection of the first language from the list in the first language display section and the second language in the second display section of language, so that the client computing device is thus configured to translate input spoken expressions from the first language to the second language.
The languages available in the first mode of operation can be visually designed differently in the first and second language display sections of the graphical user interface from languages that are not available in the first mode of operation.
The stage of transition by the client computing device from the first mode of operation to the second mode of operation responds to an input through a user interface d I
IMPI 'Mexican STnvro'
Dt INDUSTRIAL PROPERTY
<img file="MX348169B_D0014.tif" />
client computing device for transition from first mode to second mode.
These and other benefits of the present invention will be apparent from the following description.
Brief description of the drawings
Various embodiments of the present invention are described herein by way of example in conjunction with the following figures, wherein
<td>Figures 1 and 8</td><td>are block diagrams of the hybrid speech translation system in accordance with various embodiments of the present invention;</td>
<td>Figures 2A-2B and 4A-4B</td><td>provide sample graphical user interface screenshots through which a user of the client computing device can select a desired translation language combination, and transition from offline mode to online mode and vice versa, from agreement with several</td>
Figure 3 embodiments of the present invention; is a block diagram of the client computing device of
<img file="MX348169B_D0015.tif" />
in accordance with various embodiments of the present invention;
<td>The</td><td>figure</td><td> 5</td>
<td>The</td><td>figure</td><td> 6</td>
<td>The</td><td>figure</td><td> 7</td>
is a flowchart diagramming a client computing device process to automatically transition between offline and online modes in accordance with various embodiments of the present invention;
is a flowchart diagramming a process for updating various models and the offline speech translation table of the client computing device in accordance with various embodiments of the present invention; y is a flow chart diagramming the speech translation process in offline and online modes according to various embodiments of the present invention;
Detailed description of the invention
The present invention generally describes a speech translation system where a wirelessly enabled Mexican iNSTrrvro De la prohedao INDUSTRIAL client computing device (for example, a smartphone or tablet computer) has online capabilities (for example, speech translation processing by a remote computer system) and offline (for example, speech translation processing built into the client computing device). FIG. 1 is a block diagram of an exemplary embodiment of speech translation system 10 in accordance with various embodiments of the present invention. As illustrated in Figure 1, the system 10 comprises a wirelessly enabled client computing device 12, a wireless network 14, a data communication network 15 (eg, the Internet), a translation server system it speaks of an exit terminal 16 and an application storage server system 18 ("application store"). Client computing device 12 is used by a user to translate oral expressions received by client computing device 12 from a first language to a second (or even another) language. Client computing device 12 can be any suitable computing device, such as a desktop or laptop computer, but more preferably a mobile portable computing device, such as a smartphone or tablet computer. Further details relating to an example client computing device 12 are described below in conjunction with FIG. 3.
The client computing device 12 is also preferably capable of wireless data communication (i.e., the
INSTITUTO MEXICANO DE LA PROPERTY INDUSTRIAL client computing device 12 is “wirelessly enabled”) via wireless network 14. Wireless network 14 can be any suitable wireless network, such as a wireless LAN (WLAN) that uses the IEEE802 WLAN standards. 1 1, such as a WiFi network. Wireless network 14 may also comprise a mobile telecommunications network, such as a 3G or 4G LTE mobile phone communication network, although other suitable wireless networks may also be used. Wireless network 14 preferably provides connection to the Internet 15, such as through an access point or base station. The output terminal speech translation server system 16 and the application store 18 are connected to the Internet 15 and therefore in communication with the client computing device 12 via the wireless network 14.
As described herein, client computing device 12 is provided with software (including models) that enables client computing device 12 to perform offline speech translation, or perform speech translation online, with the server system. output terminal speech translation 16 providing the computationally intensive speech recognition and / or translation processing steps. The output terminal speech translation server system 16 may thus comprise one or more networked computer servers that perform speech translation based on data received from the client computing device 12 via the wireless network 14. The system translation server
<img file="MX348169B_D0016.tif" />
<sup>20</sup> IMPI 'NSTnyTO MEX1CAN <DE ΙΑ output terminal speech 16 industrial PROPERTY can therefore comprise, for example: an automatic speech recognition module 20 (ASR) to recognize speech in the first language in the oral expression data entered; a machine translation module 22 (MT) that converts / translates what is
<img file="MX348169B_D0017.tif" />
recognized in the first language to the second selected language; and a speech synthesis module 24 that synthesizes the translation in the second language to indicate an audible production in the second language. The ASR module 20 can employ, for example, (i) a language model that contains a large list of words and their probability of happening in a given sequence, and (i) an acoustic model that contains a statistical representation of the different sounds that make up each word in the language model. The TM module can use, for example, appropriate translation tables (or models) and language models. The speech synthesis module 24 may employ appropriate speech synthesis models. Similarly, the speech translation software for the client computing device 12 may comprise an ASR module (with language and acoustic models), an MT module (with translation tables / models and language models), and a speech synthesis module (with speech synthesis models). Further details for the ASR, MT, and modules (or motors), for online and offline modes, can be found in US Patents 8,090,570 and 8,204,739, which are incorporated herein by reference in their entirety.
<img file="MX348169B_D0018.tif" />
The user of the client computing device 12 may purchase speech translation software (or application or "app") through the application store 18. In various embodiments, an online version of the translation application, where the output terminal speech translation server system 16 performs most of the speech translation processing assuming a connection to the client computing device 12, which is available for free download via the app store 18. The online translation application provides a user interface to client computing device 12, the ability to collect input word phrases for translations, such as oral expressions (captured by a microphone on client computing device 12) or text ( via a text field provided by the user interface), and producing the translation (via loudspeakers of client computing device 12 and / or verbatim via the user interface). In such an embodiment, the client computing device 12 may transmit to the output terminal speech translation server system 16, via the wireless network 14, data related to the input phrase to be translated, in the first language, recorded by a microphone from the client computing device 12 or entered via the text input field, as the data includes, for example, coded samples, or feature vectors after previously processing the input speech. Based on the received input data, the terminal translation server system 16
MEXICAN INSTITUTE
OF INDUSTRIAL PROPERTY
<img file="MX348169B_D0019.tif" />
Input translates the spoken-words into the second selected language, and transmits data representative of the translation back to the client computing device 12 via the wireless network 14 for processing, synthesis and audible production through the speakers of the client computing device 12.
The speech translation application can also be operated in a connectionless mode, where the client computing device 12 performs speech translation locally, without connection to the output terminal speech translation server system 16.
In various embodiments, the user of the client computing device 12, while having connectivity to the wireless network 14, downloads the application software offline for a language combination of choice (e.g., English-Spanish, etc.), from so that the offline system can run with network connectivity outages. Figures 2A-B illustrate sample user interfaces for display on client computing device 12 using the application that allows the user to select the desired language combination. The illustrated sample user interface also shows that the user can select online or offline mode through a user control. For example, in figure 2A the user changed user control 40 to online, as indicated by the cloud and / or the description "Online translators": in figure 2B the user changed control 40 to offline, as indicated by the diagonal line through the cloud and the description “Offline Translators”. In the
IMPI Instituto Mixicano ...... _ _. . . .
In the examples of Figures 2A-B, the user can scroll the idioirr ^ s in the first and second columns 42 and 44 from the top towards'aTáj'O (Hiuy ”similar to a scroll wheel) until the user gets a combination of desired languages, indicated by the languages in the first and second columns in the highlighted selection area 46. In the example in Figure 2A, the selected language combination is English (International version) and Spanish (Mexico version). In the example in Figure 2B, the selected language combination is English (International version) and Spanish (Spain version).
In online mode, the user can access any combination of languages offered. This can be indicated to the user by displaying the icons (eg, nationality flags) for the available languages in the two columns 42 and 44 in color, as shown in Figure 2A. The user can then scroll the two columns from top to bottom so that the desired language combination is displayed in the selection area 46. When wireless network connectivity is not available (such as due to being disabled by the user through user control 40 or disabled automatically, as described in the following), only the language combination previously installed on the device client computing 12 is available in various forms. Similarly, for the languages available for offline mode, as shown in Figure 2B, they can be indicated to the user by displaying the icons in
<img file="MX348169B_D0020.tif" />
color (eg flags) for firstal'dtius languages in lab! UUT columns 42 and 44, while all languages not installed are grayed out.
Figure 7 is a flow chart illustrating the hybrid online / offline process according to various embodiments. Customer computing device 12 (eg, customer's microphone) captures an input speech in the first language in step 70. If, in step 72, the online mode is used, in step 74 the client computing device 12 transmits, via the wireless network 14, data (for example, samples or coefficients of input speaking) related to speaking input to the output terminal speech translation server system 16, which in step 76 translates the expression into the second language. In step 77, the output terminal speech translation server system 16 transmits data for translation back to the client computing device 12 via the wireless network 14 so that, in step 79, the client computing device 12 (for example, its speakers) may audibly produce the second language translation of the input expression. If in step 72 the offline mode is used, in step 78 the client computing device 12, which runs the downloaded offline speech translation software stored in memory, translates the expression into the second language, and is produced in stage 79.
Figure 3 is a block diagram of the device
<img file="MX348169B_D0021.tif" />
INSTITUTO MgXICANÍ DE LA PROPISDAD INDUSTRIAL
<img file="MX348169B_D0022.tif" />
computer client 12 according to various mnrialiriadAg Cnma shown in the example of figure 3, the device 12 can comprise several processors 202, 204. A baseband processor 202 can handle communication over a mobile telecommunications network ( eg cellular network) according to any suitable communication technology (eg 3G, 4G, etc.). Baseband processor 202 may comprise dedicated random access memory (RAM) 214. In various embodiments, baseband processor 202 may be in communication with a transceiver 206. Transceiver 206 may subsequently be in communication with one or more power amplifiers 208 and an antenna 210. The signals produced for the telecommunications network Mobile can be processed on a baseband by baseband processor 202 and provided to transceiver 206. Baseband transceiver 206 and / or processor 206 can modulate the signal produced at a carrier frequency. One or more of the amplifiers 208 can amplify the produced signal, which can subsequently be transmitted by the antenna 210. Incoming signals for the mobile telecommunications network can be received by the antenna 210, amplified by one or more amplifiers 208 and provided at the transceiver 206. The transceiver 206 or the baseband processor 202 can demodulate the incoming signal to the baseband.
An application processor 204 can run a system
MEXICAN INSTITUTE
DE LA FROFEDAD OteaSSL · ® INDUSTRIAL operating software as well as software applications (for example, downloaded via the app store 18), including the offline and online speech recognition and / or translation functionalities described herein. The application processor 204 may also run the software for the touch screen interface 232. The application processor 204 may also be in communication with the application RAM 212, and the non-volatile data storage (eg, ROM) 216. RAM 212 may store, for execution by processor 204, among other things, application software downloaded via application store 18 for offline and online speech translation, including necessary automatic speech recognition, automatic translation , and speech synthesis modules for offline processing, and software for communicating with the output terminal speech translation server system 16 for online processing.
The application processor 204 may additionally be in communication with other hardware devices such as a combination WI-FI / BLUETOOTH transceiver 218. The WI-FI / BLUETOOTH transceiver 218 can handle radio frequency (RF) communication with a LAN (for example, in accordance with the WI-FI standard, or any suitable standard) or direct RF communication between the device 200 and another device. wireless (for example, according to the BLUETOOTH standard or any suitable standard). In various embodiments, device 200
<img file="MX348169B_D0023.tif" />
MEXICAN INSTITUTE
OF INDUSTRIAL PROPERTY
<img file="MX348169B_D0024.tif" />
It may also comprise a global positioning system 222 (GPS) that is in communication with a satellite GPS system via a GPS 223 antenna to provide the application processor 204 with information describing the geographic location of the device. 12. Touch screen 232 can provide output to user device 12 visually and receive input from the user. The input can be in the form of signals representing screen touches by the user. An audio codec module 224 can provide hardware and / or software to decode and reproduce audio signals. In some embodiments, codec 224 may also comprise an analog-to-digital converter. Audio output signals can be provided to device speaker 16 and / or a connection (not shown) that can receive a set of headphones and / or speakers to reproduce the audio output signal. The audio input signals can be provided by the device microphone (18). The device may also comprise a digital camera 240.
Multiple sensors can be included in certain embodiments. A magnetic sensor 226 can pick up magnetic fields near the device. For example, magnetic sensor 226 can be used by various applications and / or system functionality to implement a compass. An accelerometer 228 and a gyroscope 230 can provide data describing the movement of the device. For example, data from accelerometer 228 and gyroscope 230
<img file="MX348169B_D0025.tif" />
they can be used to orient the display of the touch screen 232 (eg, portrait versus landscape). Device 200 can be powered by a battery 234, which can, in turn, be managed by a power management integrated circuit 236 (PMIC). An I / O transceiver 238 can handle wired communication between the device and other devices, for example according to the Universal Serial Bus (USB) or any other suitable standard. A connector 239 can facilitate hardwired connections. In some embodiments, connections via connector 239 and I / O transceiver 238 can provide power to charge battery 234.
As described above, in various modes the user can switch between online and offline modes, such as by activating user control 40 as shown in the examples of Figures 2A and 2B. Preference online processing provides more extensive vocabularies in language models than online processing, but online processing can improve user privacy and security as data so that input expressions are not transmitted over the wireless network 14 and the internet. The translation application software may also allow the client computing device 12 to automatically switch between online and offline modes in accordance with various modes. For example, the user can provide the settings for the application so that if the wireless network 14 is available (for example,
<img file="MX348169B_D0026.tif" />
network with adequate rate of index / data connection), online mode of operation is used; otherwise the offline mode of operation. Accordingly, for such a mode, as shown in the example flow diagram of FIG. 5, if the client computing device 14 is in a wireless communication mode (e.g., Wi-Fi or cell phone network, such as 3G or 4G) (step 50), the processor of the client computing device 12, which runs the application software stored in memory, can check the data connection / index speed for the WiFi network (step 52), and if it is above a data connection / index speed threshold, the online mode is used (step 54); otherwise the offline mode is used (step 56). In this way, the user automates the ability for continuous translation, and the use of offline or online modes is transparent to the user. The client computing device 12 can visually show which of the modes is to be used in a given period (such as with the cloud and non-cloud icons described above).
In other embodiments, the processor of the client computing device 12, which runs the application software stored in memory, can automatically switch between online and offline modes of operation based on other factors, such as: cost (for example, if roaming charges apply, or if there is no other network connectivity, the offline mode of operation is used, otherwise the online mode is used); quality (by
<img file="MX348169B_D0027.tif" />
example, better translation, acoustic or language models, for example, the specific use of offline speaker, or general domain independent online models); location (eg, based on GPS coordinates as determined by GPS system 222); privacy (for example, use online mode only if a secure wireless network is available); and / or time (for example, specified modes during specified hours of the day). In various embodiments, the user of the client computing device 14 can configure the application through its settings to establish the applicable criteria to automatically transition between online and offline modes of operation. For example, the user could select, according to various modes, to: always use the offline mode (in which case the mode the line is never used); prefer the faster service (in such case online mode is used only if the connection speed of the wireless network exceeds a threshold); the most accurate translations (in which case online mode is used whenever available); limit costs (in such case, for example, offline mode is used when roaming charges may apply). Such user preferences can be influenced by considerations of privacy (data transmission), quality (size and performance for speech translation models), or cost (data roaming).
Language pairs is another aspect of the hybrid offline / online translation system made available at the
<img file="MX348169B_D0028.tif" />
client computing device 12 for offline mode. Due to memory size limitations for client computing device 12, in many cases it is not practical to download all available language combinations on client computing device 12. As such, the user of client computing device 12 preferably downloads, to the client computing device 12, only selected language combinations required by the user. For example, in various modes, the user can choose or purchase offline language combinations available through the app store 18. In various modalities, the user could purchase a package comprising of several language combinations (such as languages for a geographical area, such as Europe, Southeast Asia, etc., or versions of the same language, such as the Spanish versions. from Mexico and Spain, Portuguese versions from Portugal and Brazil, etc.) in which case the application software for all language combinations in the package are available for download to the customer computing device 18. For example, Figure 4A shows a sample screen shot where a user can choose to purchase various language combinations for translation; and Figure 4B shows a sample screenshot for a pack of the language pairs for translation (in this example, a world pack), if the user wants to remove a language pair from the client computing device to save memory, the user can
INSTITUTO MEXICANO DE LA PROPIEDAD INDUSTRIAL do it, in various ways, to eliminate this combination and the models that correspond to it, without losing their availability. That is, the user can download the models again at a later date.
In one embodiment, the user is left with the choice to download a language combination, and the user selects the combinations to be installed on the client computing device for offline translation. If the user wants to install a selected language combination but there is no satisfactory network connectivity, the client computing device stores, requests and issues a reminder message to the user to download the combination when network connectivity is soon available. The reminder message asks the user if they want to download the offline version of the selected language pair and proceeds with the download if confirmed by the user.
In another embodiment, the client computing device 12 may itself manage the translation combinations offline for the user. For example, the client computing device 12 can maintain data regarding languages used around the world, and can automatically download an offline language combination that is relevant to the user's location. For example, if the GPS system 22 shows that the user is in Spain, the Spanish version of Spain can be downloaded, te. Also, the language combinations
<img file="MX348169B_D0029.tif" />
INSTITUTO MEXICANO BE LA mOHEOAI INDUSTRIAL
<img file="MX348169B_D0030.tif" />
they could be downloaded automatically based on, for example, calendar data for the user (for example, a trip) or web search data indicating a user's interest in or planning to travel to a particular region of the world.
Access to a user's location (eg based on GPS data) and / or interests (eg based on internet search data and / or speech translation queries) also offers customization of the translation system speech in their language behavior. Certain words, location names, and types of food may be preferred. In particular, names (location names, person names) will likely be more or less relevant and commonly location dependent (e.g. Kawasaki and Yamamoto for Japan, versus Martinez or Gonzales for Spain, etc.). The modeling parameters of recognition and translation models, more importantly their vocabularies and likely translations, can therefore be adjusted based on user location and interests. In an online mode, this can be done dynamically during use, using established fitting algorithms. But in an offline system, not all words could be stored and memory must be preserved to be effective on a mobile device. Therefore, the system can, in various modes, download custom parametric models even for offline / embedded systems from the output terminal speech translation systems 16, when the <sup>34</sup> IMPI ^
INSTITUTO MSXICANOE LA ntOREDAr. inoustuial network connectivity is available, and exchange vocabulary articles, language models, and modified probabilistic acoustic parameters.
The most memory intensive aspects of a speech translation system are typically provided by the translation tables and language models of a machine translation engine, the acoustic and language models of the recognition engine, and the parameters of a recognition engine. speech synthesis. To reduce the size of the templates for the offline translation application located on the client computing device 12, different techniques may be used depending on the type of the model. Models that have probabilities as model parameters such as acoustic models and language models can be reduced by quantifying the value range of the probabilities so that the value range can be outlined from a continuous to discrete space with a fixed number of only value points. Depending on the quantization factor, the storage requirements can be reduced to a byte or just a few bits. Models that store word phrases, such as translation tables and language models, can efficiently use implemented storage techniques such as prefix trees. Additionally, memory mapping techniques can be used that load only small parts of the models dynamically on demand into RAM 212/214, while unnecessary parts remain intact in non-volatile storage 216.
MAID
MEXICAN INSTITUTE OF INDUSTRIAL PROPERTY
<img file="MX348169B_D0031.tif" />
Another more sophisticated procedure for reducing language models and / or translation models for a given size so that they work on an offline client computing device is to use special fit / expand heuristics that remove vocabularies and word N-gram or Expands a base model by adding additional information. Removal can be done in a timely manner so that a user's most likely words and expressions are still rendered despite resource limitations, for example, by limiting the vocabulary to only a specific subset of the user and selecting those parts of a general model that are covered by this vocabulary or by automatically collecting user-specific information from user requests and selecting those parts of the general model that relate closely with user requests. On the other hand, expansion can be done by selecting only user specific information, for example based on speech style, and / or domain specific information, for example for tourist use or humanitarian use and / or situation specific information. , for example, based on GPS location, and / or general information not related to any of the specific information above on the server, by transferring this information (delta) only from the server to the device and by applying this information to a base model stored on the device.
For example, referring to the flow chart of Figure 6, in step 60 the speech translation server system of <sup>00</sup> IMPI ^ Mexican institute
OF PROPERTY ΓίΓΤΐ INDUSTRIAL output terminal 16 can analyze user data to determine, in step 62, if the user offline language model and / or translation table should be updated to, for example, remove used words or expressions seldom, while maintaining commonly used user words and expressions or extracting commonly used translations and expressions on the server and applying them to a base model stored on the client computing device. As mentioned above, the output terminal speech translation server system 16 can analyze user translation requests (for example, expressions to be translated) and / or Internet browsing history, to determine the words and expressions commonly (and not commonly) used. As such, in various embodiments, the offline mode user's translation requests can be saved and stored by the client computing device 12, and updated for the output terminal speech translation server system 16 for a period of time. network connectivity so that they can be analyzed by the output terminal speech translation server system 16. Similarly, the user's Internet browsing history (for example, cookie data) may be updated for the output terminal speech translation server system 16 during a period of network connectivity so that it can be analyzed by the 16 output terminal speech translation server system to determine the words and expressions
<img file="MX348169B_D0032.tif" />
commonly (and not commonly) used by the user ... yes, through analysis of user data), the output terminal speech translation server system 16 determines that the speech patterns and / or the translation tables of the offline processing software of the client computing device should be updated, the updated software (e.g. models) is downloaded to the client computing device (e.g. from the output terminal speech translation server system 16) at step 64. Instead of downloading the full models, it is also possible to download only the information (delta) that is required to update the model on the computing device of client.
Similarly, user-specific information could also be helpful in reducing the size of the acoustic model by, for example, replacing a more general acoustic model with a smaller user-specific one. Depending on the amount of data espe user-specific data, this can be achieved, for example, by using acoustic model fitting techniques such as MLLR or by completely refitting the acoustic model using the new additional data. Thus, for example, referring back to Figure 6, if at step 66 the output terminal speech translation server system 16 determines that the acoustic model in offline mode for the client computing device 12 user data should be modified based on analysis of user data, updated software (e.g. acoustic model) is downloaded to client computing device (e.g. from
<img file="MX348169B_D0033.tif" />
of the output terminal speech translation server system 16) in step 68.
In a speech-to-speech translation system, the most limiting element of speed is the speech recognition algorithms, since they perform searches on many acoustic hypotheses and many time intervals of the speech signal. The speed of algorithmic searches is predominantly affected by the size of the acoustic model set. In order to maintain the speed of the offline system when performing speech-to-speech translation at the client computing device 12, various techniques can be used. For example, in one embodiment, depending on the size of the model, the lookup tables can be used to calculate the Mahalanobis distances between the model and the input speech instead of calculating the distances on demand. Additionally, Gaussian selection techniques can be used in offline mode to reduce the total number of model parameters that need to be evaluated. As soon as user-specific information is available, more efficient and smaller user-specific models can be used instead, as described above in conjunction with Figure 6.
Furthermore, according to various embodiments, during the online mode, the output terminal speech translation system 16 may use and combine multiple speech recognition and translation engines (modules). These output terminal motors can be provided by the same supplier of<sup>39</sup> IMPI n ^ ITUTO MEXICANO CE LA PROPERTY industrial speech translation and run on the same server, for example, or in other modalities could be enhanced by independent speech translation providers in different locations, as illustrated in the example in Figure 8, showing three separate and independent output terminal speech translation systems 16, although fewer or more output terminal speech translation systems 16 could be used in other embodiments. In such an embodiment, each of the output terminal speech translation systems 16 may be requested to perform the translation simultaneously via the Internet 15. In such embodiment, the output terminal speech translation systems 16 are they are in communication (eg via the Internet 15) and one of the output terminal speech translation systems 16 chooses the best of the translations or combines them. To decide between multiple systems / translations and / or how to weigh each system in combination strongly, confidence measures in ASR and confidence measures for MT can be used. Such confidence measures are used to determine the reliability of an ASR or a MT hypothesis. If two or more ASR or MT engines are to be merged in such a mode, combinations of systems can be used, such as the “ROVER” methods of combining ASR outputs (see, for example, J, G. Fiscus, “A post-processing system to yield reduced error word rates: Recognizer output voting error reduction (ROVER),” IEEE Workshop on Automatic Speech Recognition and Understanding, pp. 347-354, 1997), transversally adapt a system against
<img file="MX348169B_D0034.tif" />
IMP ^ 3
INSTITUTO MEXICANO de la mohedal industuiai otro, or TM system combination techniques (see, for example, Rostí et al ,, Combining Outputs from Multiple Machine Translation Systems, '' Proc. Of NAACL HLT, pp. 228-235, 2007 and K, Heafield et al ,, Combining Machine Translation Output with Open Source, Prague Bulletin of Mathematical Linguistics, No. 93, pp. 27-36, 2010). In such an embodiment, the selected and combined hypothesis can be completed at the output terminal to produce the best output for the user. Once this is done in online mode, the system will remember the best choice arrived in this way for entry into the offline system. For offline system learning, combined online systems 16 can retain the multiple ASR engine recognition hypothesis and / or multiple MT engine translation hypotheses in memory and use the combination, or the best of these hypotheses , to adapt or train new offline systems. Such retrofitted or retrofitted systems can be subsequently swapped back into the connectionless systems, when the wireless network is available.
In a general sense, therefore, the present invention is directed to hybrid speech translation systems and methods for offline and online speech translation. According to various embodiments, the system may comprise an output terminal speech translation server system and a client computing device that is configured for communication with the output terminal speech translation server system via a network. wireless. The client computing device may comprise a microphone, a processor connected to the microphone, a memory connected to the processor that stores instructions to be executed by the processor, a speaker connected to the processor. The client computing device is to produce, for example, by means of the loudspeaker or a text display field, a translation of input word phrases for translation (for example, oral expressions or input text) from a first language to a second language. The memory stores instructions so that, in a first mode of operation (offline mode), when the processor executes the instructions, the processor translates the input word phrases into the second language for output (eg, through the speaker). In a second mode of operation (online mode): (I) the client computing device transmits to the output terminal speech translation server system, via the wireless network, data regarding the input word phrases in the first language received by the microphone; (I) the output terminal speech translation server system determines the second language translation of the input word phrases in the first language based on data received via the wireless network from the client computing device; and (ii) the output terminal speech translation system transmits data regarding the second language translation of the input word phrases into the first language to the client computing device via wireless network conducts the output so that the cli computing device translation into second language of oral expressions in first language.
According to various implementations, the client computing device has a user interface that allows a user to switch between the first mode of operation and the second mode of operation. Alternatively, the client computing device automatically selects whether to use the first mode of operation or the second mode of operation based on a user preference connection setting of the user for the client computing device. Furthermore, the client computing device can store in memory a local acoustic model, a local language model, a local translation model, and a local speech synthesis model to, in the first mode of operation, recognize the spoken expressions in the first language and translate the recognized oral expressions into the second language to be produced through the loudspeaker. Also, the output terminal speech translation server system comprises an output terminal acoustic model, an output terminal language model, an output terminal translation model, and an output terminal speech synthesis model. output, in the second mode of operation, which determine the second language translation of the oral expressions in the first language based on the data received via the wireless network from the client computing device. The local models are different from the output terminal models (for
MAID
INSTmnt> M »! CANO DE LA RROHEDAD INDUSTRIAL (example, a subset of another variation).
Additionally, the output terminal speech translation server system can be programmed to: (i) monitor over time oral expressions received by the client computing device for translation from a first language to the second language (ii) update at least one of the local acoustic model, the local language model, the local translation model and the local speech synthesis model of the client computing device based on monitoring over time of oral expressions received by the client computing device for translation from the first language to the second language. The client computing device may also comprise a GPS system for determining a location of the client computing device. In such an embodiment, the output terminal speech translation server system may also be programmed to update at least one of the local acoustic model, the local language model, the local translation model, and the device's local speech synthesis model. client computing device based on the location of the client computing device. Any such update in at least one of the client computing device models can be transmitted from the output terminal speech translation server system to the client computing device via the wireless network.
Additionally, the client computing device can be configured for downloaded application software (including models) for a combination of languages for translation that
<img file="MX348169B_D0035.tif" />
MAID
INSTITUTO MEXICANO M LA INOUSTÍIAL PROPERTY
<img file="MX348169B_D0036.tif" />
It understands the first and second language, particular euendo. Proper connectivity between the client computing device and the output terminal speech translation server system is available via the wireless network. Also, for modalities where the client computing device comprises a GPS system, The client computing device can be configured to download the application software for the language combination for translation at the determined location of the client computing device and when adequate connectivity between the client computing device and the speech translation server system of output terminal is available via wireless network.
In addition, the client computing device may comprise a graphical user interface having a first language display section and a second language display section that are displayed simultaneously. Each of the first and second language display sections may comprise a user-accessible list of a plurality of languages, such that after a user of the client computing device selects the first language from the listing in the first section language display and the second language from the second language display section, the client computing device is thus configured to translate input spoken expressions from the first language to the second language. The languages available in the first mode of operation (without<sup>45</sup> IMPI ^
INSTITUTO MíXICANC ·
OF PROPERTY <V ~ * INDUSTRIAL connection) can be designed differently-axi_la_j3 first and second language display sections from languages not available in the first mode of operation.
In addition, in various embodiments, the output terminal speech translation server system is one of a plurality of output terminal speech translation server systems, and the client computing device is configured to communicate with each of the plurality of output terminal speech translation server systems via a wireless network. In the second (online) mode of operation, each of the plurality of output terminal speech translation server systems determines a second language translation of the input word phrases in the first language based on the received data. over the wireless network from the client computing device. In such circumstances, one of the plurality of output terminal speech translation server systems selects one of the translations from the plurality of output terminal speech translation server systems to transmit to the client computing device, or two or more of the translations from the plurality of output terminal speech translation server systems are merged to generate a merged translation for transmission to the client computing device.
In a general aspect, the speech translation method comprises, a first mode of operation (without connection): (i) receiving by a client computing device a first phrase of
<img file="MX348169B_D0037.tif" />
<img file="MX348169B_D0038.tif" />
words in a first language; (¡I) translate p'oT ~ '6l ~~ tl + s- ^ GS4tiua client computer the first sentence of words into a second language; and (iii) produce by the client computing device the first oral expression in the second language (for example, audibly through a loudspeaker and / or visually through a text display field). The method further comprises making a transition by the client computing device from the first mode of operation to the second operating mode, and then, in the second (online) mode of operation: (iv) receiving by the client computing device a second phrase of words in a first language; (v) transmitting, by the client computing device, via a wireless network, data regarding the second sentence of input words to an output terminal speech translation server system; and (vi) receiving, by the client computing device, from the output terminal speech translation server system via the wireless network, data regarding a translation by the terminal speech translation server system. output of the second sentence of input words from the first language to the second language; and produce the first oral expression in the second language by the client computing device.
It will be apparent to one of ordinary skill in the art that at least some of the embodiments described herein can be implemented in many different modes of software, firmware, and / or hardware. The software and firmware code can
MAID
INSTITUTO MEXICANO DE LA PROPIEDAD industrial be executed by a processor circuit or any other similar computing device. The specialized control hardware or software code that can be used to implement the modalities is not limiting. For example, the modalities described herein can be implemented in computer software using any type of suitable computer software language, using, for example, conventional or object-oriented techniques. Such software may be stored on any type of suitable computer-readable medium or media, such as, for example, magnetic or optical storage medium. The operation and behavior of the modes can be described without specific reference to specific software code or specialized hardware components. The absence of such specific references is feasible, because it is clearly understood that technicians of ordinary experience may be able to design control software and hardware to implement the modalities based on the present disclosure without more than reasonable effort and without undue experimentation.
On the other hand, the process associated with these modalities can be executed by programmable equipment, such as computers or computer systems, mobile devices, smartphones and / or processors, the software that can cause the programmable equipment to execute processes can be stored on any device. storage, such as, for example, a computer system (non-volatile) memory, RAM,
<img file="MX348169B_D0039.tif" />
ROM, Flash Memory, etc. In addition, at least some of the processes can be programmed when the computer system is manufactured or stored on various types of computer-readable media.
A "computer", "computer system", "main computer", "server", or "processor" can be, for example and without limitation, a processor, a microcomputer, a mini-computer, server, central computer, laptop, personal data assistant (PDA), wireless email device, cell phone, smartphone, Tablet, mobile device, pager, processor, fax, scanner, or any other programmable device configured to transmit and / or receive data over a network. The computer systems and computer-based devices described herein may include memory to store certain software modules or engines used to obtain, process, and communicate Information. It can be appreciated that such memory can be internal or external with respect to the operation of the described modes. Memory can also include any medium for storing software, including a hard disk, an optical disk, a floppy disk, a ROM (read-only memory), a RAM (random access memory), a PROM (programmable ROM), a EEPROM (Electrically Erasable PROM) and / or any other computer-readable media. The software modules and engines described herein can be run by the processor (or processors as the case may be) of the computer devices that have access to the memory that stores the modules.
<img file="MX348169B_D0040.tif" />
In various embodiments described in IIf pitíbeiiler, —a single component can be replaced by multiple components and multiple components can be replaced by a single component to perform a given function or functions. Except where such substitution is not operative, such substitution is within the intended scope of the modalities. Any servers described herein, for example, can be replaced by a "server tower" or other cluster of network servers (such as blade servers) that are located and configured for cooperative functions. It can be appreciated that a tower of servers can serve to distribute the workload among individual components of the tower and can accelerate computational processes by employing the collective and cooperative power of multiple servers. Such server towers may employ load balancing software that accomplishes tasks such as, for example, tracking the demand for processing power of different machines, prioritizing and scheduling tasks based on network demand, and / or providing backup contingency in case of component failure or reduced operability.
Although various embodiments have been described herein, it should be apparent that various modifications, alterations, and adaptations to those embodiments may occur to those skilled in the art with the achievement of at least some of the advantages. The modalities described, therefore, are intended to
MEXICAN INSTITUTE
OF THE PROPERTY include all the modifications, alterations and atfá'pwc depart from the scope of the modalities c herein.
Contents25
56 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41 Sheet 42 Sheet 43 Sheet 44 Sheet 45 Sheet 46 Sheet 47 Sheet 48 Sheet 49 Sheet 50 Sheet 51 Sheet 52 Sheet 53 Sheet 54 Sheet 55 Sheet 56
23 members in 11 offices
Priority claims11
| Document | Office | Kind | Date |
|---|---|---|---|
| 201361822629 | United States of America | P | |
| 61822629 | United States of America | – | |
| 13915820 | United States of America | – | |
| 201313915820 | United States of America | A | |
| 2014036454 | United States of America | W | |
| 13915820 | – | – | – |
| 61822629 | – | – | – |
| PCTUS2014036454 | – | – | – |
| US201313915820 | – | – | – |
| US201361822629P | – | – | – |
| WO2014US36454 | – | – | – |
Members23
| Document | Office | Kind | |
|---|---|---|---|
| US2014337007A1 | United States of America | A1 | |
| EP2804113A2 | European Patent Office (EPO) | A2 | |
| CA2907775A1 | Canada | A1 | |
| WO2014186143A1 | World Intellectual Property Organization (WIPO) | A1 | |
| EP2804113A3 | European Patent Office (EPO) | A3 | |
| AU2014265782A1 | Australia | A1 | |
| CN105210056A | China | A | |
| KR20160006682A | Republic of Korea | A | |
| MX2015015799A | Mexico | A | |
| US9430465B2 | United States of America | B2 | |
| JP2016527587A | Japan | A | |
| US2016364385A1 | United States of America | A1 | |
| KR101729154B1 | Republic of Korea | B1 | |
| IL242230A | Israel | A | |
| AU2014265782B2 | Australia | B2 | |
| MX348169BThis record | Mexico | B | |
| JP6157725B2 | Japan | B2 | |
| BR112015028622A2 | Brazil | A2 | |
| AU2017210631A1 | Australia | A1 | |
| CN105210056B | China | B | |
| CA2907775C | Canada | C | |
| AU2017210631B2 | Australia | B2 | |
| US10331794B2 | United States of America | B2 |
1 legal event, as the office reported them to INPADOC
Events
| Event | Code | |
|---|---|---|
| Grant or registrationFG | FG |
Numbers
- Publication
- 348169
- Publication, DOCDB
- 348169
- Publication, EPODOC
- MX348169
- Application
- 2015015799
- Application, DOCDB
- 2015015799
- Application, EPODOC
- MX202015015799
Titles2
- English
- HYBRID OFFLINE / ONLINE SPEECH TRANSLATION SYSTEM.
- Spanish
- SISTEMA DE TRADUCCION DE HABLA SIN CONEXION/EN LINEA, HIBRIDO.
Classification
- CPC, 4
- G06F40/58
- G10L13/00
- G10L15/30
- G10L13/02