System and method for language translation in a hybrid peer-to-peer environment.
Abstract
An improved system and method are disclosed for peer-to-peer communications. In one example, the method enables an endpoint to send and/or receive audio speech translations to facilitate communications between users who speak different languages.
Term
5 yearsleft in the term
Expires 16 September 2031.
- Priority
- Filed
- Today
- Expires
17 claims: 3 independent, 14 dependent
- 1NOVEDAD DE LA INVENCIÓN REIVINDICACIONES 5 1. Un método para la comunicación, por un primer punto final, en una sesión en curso de comunicación de pares a pares entre el primer punto final y un segundo punto final en una red híbrida de pares a pares que comprende:establecer, por el primer punto final, una ruta de comunicaciones directamente entre el primer punto final y el segundo punto 10 final, en donde las comunicaciones de señalización se envían directamente desde el primer punto final hacia el segundo punto final a través de una ruta de señalización proporcionada por la ruta de comunicaciones;recibir, por el primer punto final, una entrada de audio de voz en un primer idioma hablado desde un usuario del primer punto final;determinar, por el primer punto final, 15 si la entrada de audio de voz ha de traducirse desde el primer idioma hablado hacia un segundo idioma hablado;enviar, por el primer punto final, la entrada de audio de voz hacia un componente de traducción de idiomas accesible al primer punto final a través de la red híbrida de pares a pares si la entrada de audio de voz ha de traducirse desde el primer idioma hablado hacia el 20 segundo idioma hablado, en donde el primer punto final no envía la entrada de audio de voz directamente hacia el segundo punto final si la entrada de audio de voz ha de traducirse desde el primer idioma hablado hacia el segundo idioma hablado;y enviar, por el primer punto final, la entrada de audio de voz directamente hacia el segundo punto final a través de la ruta de comunicaciones si la entrada de audio de voz no ha de traducirse desde el primer idioma hablado hacia el segundo idioma hablado.
- 2El método de conformidad con la reivindicación 1,
- 35 caracterizado además comprende adicionalmente enviar, por el primer punto final, los datos que no son de audio directamente hacia el segundo punto final a través de la ruta de comunicaciones independientemente de si la entrada de audio de voz ha de traducirse desde el primer idioma hablado hacia el segundo idioma hablado. 10 3. El método de conformidad con la reivindicación 1, caracterizado además porque comprende adicionalmente:recibir, por el primer punto final, una traducción de la entrada de audio de voz desde el componente de traducción de idiomas;y proporcionar la traducción de la entrada de audio de voz al usuario del primer punto final. 15 4. El método de conformidad con la reivindicación 1, caracterizado además porque enviar la entrada de audio de voz hacia el componente de traducción de idiomas incluye: establecer, por el primer punto final, una ruta de señalización con un módulo de voz a texto del componente de traducción de idiomas;y enviar la entrada de audio de voz hacia el módulo 20 de voz a texto. 5. El método de conformidad con la reivindicación 1, caracterizado además porque comprende adicionalmente: realizar, por el primer punto final, un proceso de autenticación con un servidor de acceso en la red híbrida de pares a pares;y recibir, por el primer punto final, un perfil desde el servidor de acceso después del proceso de autenticación, en donde el perfil identifica el segundo punto final como un punto final con el cual el primer punto final tiene permiso para comunicarse e identifica que el segundo 5 punto final se asocia con el segundo idioma hablado.
- 46. El método de conformidad con la reivindicación 5, caracterizado además porque comprende adicionalmente recibir, por el primer punto final, una lista de idiomas desde el servidor de acceso, en donde la lista de idiomas incluye el primer y segundo idiomas hablados e identifica los 10 idiomas hablados que pueden traducirse por el componente de traducción de idiomas.
- 57. El método de conformidad con la reivindicación 1, caracterizado además porque comprende adicionalmente recibir, por el primer punto final, una notificación directamente desde el segundo punto final de que 15 el segundo punto final se asocia con el segundo idioma hablado.
- 68. El método de conformidad con la reivindicación 1, caracterizado además porque comprende adicionalmente:recibir, por el primer punto final desde el componente de traducción de idiomas, el audio de voz que se origina desde el segundo punto final, en donde el audio de voz que 20 se origina desde el segundo punto final se traduce desde el segundo idioma hablado hacia el primer idioma hablado por el componente de traducción de idiomas antes de recibirse por el primer punto final;y producir, por el primer punto final, el audio de voz recibido desde el componente de traducción de idiomas como un sonido audible.
- 79. El método de conformidad con la reivindicación 1, caracterizado además porque comprende adicionalmente:establecer, por el primer punto final, una segunda ruta de comunicaciones directamente entre el 5 primer punto final y un tercer punto final, en donde las comunicaciones de señalización se envían directamente desde el primer punto final hacia el tercer punto final a través de una ruta de señalización proporcionada por la segunda ruta de comunicaciones;identificar, por el primer punto final, que el tercer punto final se asocia con el primer idioma hablado;y enviar, por el primer
- 810 punto final, la entrada de audio de voz directamente hacia el tercer punto final a través de la segunda ruta de comunicaciones. 10. Un método para la comunicación, por un primer punto final en una red híbrida de pares a pares, en una sesión en curso de comunicación con un segundo y tercer puntos finales a través de un puente 15 que comprende:identificar, por el primer punto final, que el primer punto final se asocia con un primer idioma hablado, el segundo punto final se asocia con un segundo idioma hablado, y el tercer punto final se asocia con un tercer idioma hablado;enviar, por el primer punto final, una solicitud hacia el puente para que un puerto de entrada y un puerto de salida se proporcionen en el 20 puente para el primer punto final;notificar, por el primer punto final, a un componente de traducción de idiomas en la red híbrida de pares a pares acerca del puerto de entrada, en donde el componente de traducción de idiomas es accesible al primer punto final a través de la red híbrida de pares a pares, y en donde la notificación instruye al componente de traducción de idiomas para que envíe el audio recibido desde el primer punto final hacia el puerto de entrada;enviar hacia el componente de traducción de idiomas, por el primer punto final, la entrada de audio de voz recibida por el primer punto 5 final desde un usuario del primer punto final, en donde la entrada de audio de voz enviada por el primer punto final está en el primer idioma hablado;y recibir, por el primer punto final, el audio de voz desde el segundo y tercer puntos finales directamente desde el puerto de salida en el puente, en donde el audio de voz recibido por el primer punto final a través del puerto de salida 10 se envió por el segundo y tercer puntos finales en el segundo y tercer idiomas hablados, respectivamente, y en donde el audio de voz recibido por el primer punto final directamente desde el puerto de salida se recibe en el primer idioma hablado.
- 911. El método de conformidad con la reivindicación 10, 15 caracterizado además porque comprende adicionalmente:identificar, por el primer punto final, que un cuarto punto final se asocia con el primer idioma hablado;y enviar, por el primer punto final, el audio de voz recibido desde el usuario del primer punto final directamente hacia el cuarto punto final sin usar el puente. 20
- 1012. El método de conformidad con la reivindicación 10, caracterizado además porque comprende adicionalmente:identificar, por el primer punto final, que un cuarto punto final se asocia con el primer idioma hablado;y enviar, por el primer punto final, el audio de voz recibido desde el usuario del primer punto final hacia el puerto de entrada en el puente.
- 1113. El método de conformidad con la reivindicación 10, caracterizado además porque comprende adicionalmente enviar, por el primer punto final, los datos que no son de audio directamente hacia al menos uno 5 del segundo y tercer puntos finales, en donde los datos que no son de audio no pasan a través del puente.
- 1214. El método de conformidad con la reivindicación 10, caracterizado además porque comprende adicionalmente recibir, por el primer punto final, los datos que no son de audio directamente desde al menos uno 10 del segundo y tercer puntos finales, en donde los datos que no son de audio no pasan a través del puente.
- 1315. El método de conformidad con la reivindicación 10, caracterizado además porque comprende adicionalmente notificar directamente al segundo y tercer puntos finales, por el primer punto final, 15 acerca de los puertos de entrada y de salida, en donde la notificación no usa el puente.
- 1416. El método de conformidad con la reivindicación 10, caracterizado además porque comprende adicionalmente recibir, por el primer punto final, notificaciones directamente desde cada uno del segundo y tercer 20 puntos finales acerca de los puertos de entrada y de salida correspondientes al segundo y tercer puntos finales.
- 1517. El método de conformidad con la reivindicación 10, caracterizado además porque comprende adicionalmente enviar una indicación, por el primer punto final, hacia el componente de traducción de idiomas de que la entrada de audio de voz ha de traducirse hacia el segundo y tercer idiomas.
- 1618. Un dispositivo de punto final que comprende:una interfaz 5 de red;un procesador acoplado a la interfaz de red;y una memoria acoplada al procesador y que contiene una pluralidad de instrucciones para su ejecución por el procesador, las instrucciones que incluyen instrucciones para: realizar un proceso de autenticación con un servidor de acceso en una red híbrida de pares a pares, en donde el proceso de autenticación autoriza ai 10 primer punto final para acceder a la red híbrida de pares a pares;recibir un perfil desde el servidor de acceso que identifica un segundo punto final como un punto final dentro de la red híbrida de pares a pares con el cual el primer punto final tiene permiso para comunicarse;determinar que un usuario del primer punto final ha designado un primer idioma humano para usarse por el 15 primer punto final;establecer una ruta de comunicaciones directamente entre el primer punto final y el segundo punto final, en donde las comunicaciones de señalización se envían directamente desde el primer punto final hacia el segundo punto final a través de la ruta de comunicaciones;determinar que el segundo punto final ha designado un segundo idioma humano para usarse por 20 el segundo punto final;recibir una entrada de audio de voz en el primer idioma humano desde el usuario del primer punto final;y enviar la entrada de audio de voz hacia un módulo de traducción de idiomas para su traducción hacia el segundo idioma humano.
- 1719. El dispositivo de punto final de conformidad con la reivindicación 18, caracterizado además porque comprende adicionalmente instrucciones para enviar los datos que no son de audio directamente hacia el segundo punto final a través de la ruta de comunicaciones. 5 20. El dispositivo de punto final de conformidad con la reivindicación 18, caracterizado además porque las instrucciones para determinar que el segundo punto final ha designado que el segundo idioma humano se use por el segundo punto final puede incluir instrucciones para obtener del perfil un identificador que representa el segundo idioma humano.
Independent claims17
187 paragraphs in 6 sections, as filed
(54) Title: SYSTEM AND METHOD FOR THE TRANSLATION OF A LANGUAGE IN A HYBRID ENVIRONMENT FROM COUPLE TO COUPLE.
(54) Title: SYSTEM AND METHOD FOR LANGUAGE TRANSLATION IN A HYBRID PEER-TO-PEER ENVIRONMENT.
(57) Summary
An improved system and method for peer-to-peer communications is disclosed; in one example, the method allows an endpoint to send and / or receive audio language translations to facilitate communications between users who speak different languages.
(57) Abstract
An improved system and method are disclosed for peer-to-peer Communications. In one example, the method enables an endpoint to send and / or receive audio speech translations to facilitate Communications between users who speak different languages.
SYSTEM AND METHOD FOR THE TRANSLATION OF A LANGUAGE IN
HYBRID COUPLE-TO-COUPLE ENVIRONMENT
CROSS REFERENCE TO RELATED APPLICATION
This application claims the benefit of United States patent application number 12 / 890,333, filed on September 24, 2010, and titled LANGUAGE TRANSLATION SYSTEM AND METHOD IN A HYBRID ENVIRONMENT OF COUPLES TO COUPLES, the description of which is incorporated in the present.
BACKGROUND OF THE INVENTION
Current packet-based communication networks 15 can generally be divided into peer-to-peer networks and client / server networks. Traditional peer-to-peer networks support direct communication between multiple endpoints without the use of an intermediary device (for example, a host or server). Each endpoint can initiate requests directly to other endpoints and respond to requests from other endpoints using a credential and address information stored on each endpoint. However, because traditional peer-to-peer networks include the distribution and storage of endpoint information (for example, addresses and credentials) throughout the network at various non-secure endpoints, such networks inherently have an increased security risk. Although a client / server model addresses the security problem inherent in the peer-to-peer model by locating the storage of credentials and address information on a server, a disadvantage of client / server networks is that the server may be unable to adequately support the number of clients trying to communicate with it. Since all communications (even between two clients) must pass through the server, the server can quickly become a bottleneck in the system.
Consequently, what is needed is a system and method that addresses these problems.
BRIEF DESCRIPTION OF THE DRAWINGS
For a more complete understanding, reference is now made to the description that follows taken together with the accompanying drawings in which:
Fig. 1 is a simplified network diagram of an embodiment 20 of a hybrid pair-to-pair system.
Fig. 2A illustrates one embodiment of an access server architecture that can be used within the system of Fig. 1.
Fig. 2B illustrates one embodiment of an endpoint architecture that can be used within the system of Fig. 1.
Fig. 2C illustrates an embodiment of components within the endpoint architecture of Fig. 2B that can be used for connectivity of a cellular network.
FIG. 3 is a sequence diagram illustrating an exemplary process whereby an end point in FIG. 1 can authenticate and communicate with another end point.
Fig. 4 is a simplified diagram of one embodiment of a peer-to-peer environment in which voice audio translations can be performed.
Fig. 5 is a sequence diagram illustrating an embodiment of a process that can be executed within the peer-to-peer environment of Fig. 4.
Fig. 6 is a simplified diagram of another embodiment of a peer-to-peer environment in which voice audio translations can be performed.
Fig. 7 is a sequence diagram illustrating an embodiment of a process that can be executed within the peer-to-peer environment of Fig. 6.
Fig. 8 is a simplified diagram of a more specific embodiment of the peer-to-peer environment of Fig. 4.
Fig. 9 is a simplified diagram of a more specific embodiment of a portion of the peer-to-peer environment of Fig. 8.
Fig. 10 is a sequence diagram illustrating an embodiment of a process that can be executed within the peer-to-peer environment of Fig. 8.
Fig. 11 is a flowchart illustrating a modality of a method that can be executed by an endpoint within the system of the
Fig. 8.
Fig. 12 is a flowchart illustrating an embodiment of a method that can be executed by a Language translation component within the system of Fig. 8.
Fig. 13A is a simplified diagram of one embodiment of a peer-to-peer environment in which voice audio translations can be performed when the endpoints are coupled across a bridge.
Fig. 13B is a simplified diagram of an embodiment of the peer-to-peer environment of Fig. 13A in which signaling is performed directly between the end points.
Fig. 14A is a sequence diagram illustrating one embodiment of a process that can be executed within the peer-to-peer environment of Fig. 13A.
Fig. 14B is a sequence diagram illustrating another embodiment of a process that can be executed within the peer-to-peer environment of Fig. 13A.
Fig. 14C is a sequence diagram illustrating another embodiment of a process that can be executed within the peer-to-peer environment of Fig. 13A.
Fig. 15 is a simplified diagram of another embodiment of a peer-to-peer environment in which voice audio translations can be performed when the endpoints are coupled across a bridge.
Fig. 16 is a sequence diagram illustrating an embodiment of a process that can be executed within the peer-to-peer environment of Fig. 15.
Fig. 17 is a simplified diagram of another embodiment of a peer-to-peer environment in which voice audio translations can be performed when the endpoints are coupled across a bridge.
Fig. 18 is a sequence diagram illustrating an embodiment of a process that can be executed within the peer-to-peer environment of Fig. 17.
Fig. 19 is a simplified diagram of one embodiment of a computer system that can be used in embodiments of the present disclosure.
DETAILED DESCRIPTION OF THE INVENTION
The present disclosure is directed to a system and method for hybrid peer-to-peer communications. The following description is understood to provide many different embodiments or examples. Specific examples of components and arrangements are described below to simplify the present description. These are, of course, mere examples and are not intended to be limiting. Furthermore, in the present description, the letters and / or reference numbers may be repeated in the various examples. This repetition is done for the sake of simplicity and clarity and not to dictate in itself a relationship between the various modalities and / or configurations discussed.
Referring to FIG. 1, one embodiment of a hybrid peer-to-peer system 100 is illustrated. System 100 includes an access server 102 that is coupled to endpoints 104 and 106 via a packet network 108 . Communication between access server 102, endpoint 104, and endpoint 106 is accomplished using predefined and publicly available (i.e. non-proprietary) communication standards or protocols (for example, those defined by the Engineering Task Force). Internet (IETF) or the Telecommunication Standardization Sector of the International Telecommunication Union (ITU-T)). For example, signaling communications (eg, establishment, administration, and session closure) can use a protocol such as Session Initiation Protocol (SIP), while actual data traffic can communicate using a protocol such as real-time transport protocol (RTP). As will be seen in the examples that follow, the use of standard protocols for communication allows endpoints 104 and 106 to communicate with any device that uses the same standards. Communications may include, but are not limited to, voice calls, instant messages, audio and video, emails, and any other type of resource transfer, where a resource represents any digital data. In the description that follows, media traffic is generally based on the User Datagram Protocol (UDP), while authentication is based on the Transmission Protocol / Internet Protocol (TCP / IP). However, it is understood that these are used for example purposes and that other protocols may be used in addition or instead of UDP and TCP / IP.
Connections between Access Server 102, the endpoint
104, and endpoint 106 can Include wired and / or wireless communication channels. In the description that follows, the direct term is understood to mean that there is no endpoint or access server on the communication channel (s) between endpoints 104 and 106, or between any endpoint and the access server. Consequently, access server 102, endpoint 104, and endpoint 106 connect directly even if other devices (eg, routers, firewalls, and other network elements) postclone to each other. Additionally, connections to endpoints, locations, or services can be subscription-based, with an endpoint only accessible if the endpoint has a current subscription.
Furthermore, the description that follows may use the terms user and endpoint interchangeably, although it is understood that a user may be using any one of a plurality of endpoints. Accordingly, if an endpoint Logs into the network, it is understood that the user is logging in through the endpoint and that the endpoint represents the user in the network using the user's identity.
Access server 102 stores profile information for a user, a session table to keep track of users who are currently online, and a routing table that matches the address of an endpoint with each user online. Profile information includes a friend list for each user that identifies other users (friends) who have previously agreed to communicate with the user. Online users in the friends list will be displayed when a user logs in, and friends who log in later will directly notify the user that they are online (as described with respect to Fig. 3). Access server 102 provides the relevant profile information and routing table to each of the endpoints 104 and 106 so that the endpoints can communicate directly with each other. Accordingly, in the current mode, a function of the access server 102 is to serve as a storage location for the information that an endpoint needs in order to communicate with other endpoints and as a temporary storage location for requests, emails voice, etc., as will be described in more detail later.
With additional reference to Fig. 2A, one embodiment of an architecture 200 is illustrated for access server 102 of Fig. 1. Architecture 200 includes functionality that can be provided by hardware and / or software, and which can be combined on a single hardware platform or spread across multiple hardware platforms. For purposes of illustration, the access server in the examples that follow is described as a single device, but the term is understood to apply equally to any type of environment (including a distributed environment) in which at least one is present. portion of the functionality attributed to the access server.
In the current example, the architecture includes web services 202 (for example, based on functionality provided by XML, SOAP, .NET, MONO), web server 204 (using, for example, Apache or IIS), and database data 206 (using, for example, mySQL or SQLServer) to store and retrieve routing tables 208, profiles 210, and one or more session tables 212. The functionality for a STUN server (Simple UDP Tour over NAT (Network Address Translation)) 214 is also present in architecture 200. As is known, STUN is a protocol to assist devices behind a firewall or NAT router with its packet routing. Architecture 200 may further include a redirect server 216 to handle requests originating outside of system 100. One or both of the STUN server 214 and the redirect server 216 can be incorporated into the access server 102 or can be independent devices. In the current mode, both server 204 and redirect server 216 are attached to database 206.
Referring to FIG. 2B, one embodiment of an architecture 250 is illustrated for endpoint 104 (which may be similar or identical to endpoint 106) of FIG. 1. It is understood that the term endpoint can refer to many different devices that have some or all of the described functionality, including a computer, a VolP phone, a personal digital assistant, a cell phone, or any other device that has an IP battery. on which the necessary protocols can be executed. Such devices generally include a network interface, a controller coupled to the network interface, a memory coupled to the controller, and instructions executable by the controller and stored in memory to perform the functions described in the current application. The data that an endpoint needs can also be stored in memory. Architecture 250 includes an endpoint engine 252 positioned between a graphical user interface (GUI) 254 and an operating system 256. GUI 254 provides user access to the endpoint engine 252, while operating system 256 provides the functionality underlying, as known to those of skill in the art.
Endpoint engine 252 can include multiple components and layers that support the functionality required to perform endpoint 104 operations. For example, endpoint engine 252 includes a software switch 258, a management layer 260, a module encryption / decryption 262, one function layer 264, one protocol layer
266, a speech to text engine 268, a text to speech engine 270, a language conversion engine 272, an out-of-network connectivity module 274, a connection module from other networks 276, a negotiation engine-p (eg, inter-point trading) 278 which includes a trading agent p and a trading broker-p, and a cellular network interface module 280.
Each of these components / layers can be further divided into multiple modules. For example, the software switch 258 includes a call control module, an instant messaging (IM) control module, a resource control module, a CALEA agent (communications assistance for the law enforcement act) , a media control module, a peer control module, a signaling agent, a fax control module, and a routing module.
Management layer 260 includes modules for presence (i.e. network presence), peer management (detecting peers and notifying peers that it is online), firewall management (navigation and management), media management, management resource management, profile management, authentication, roaming, fax management, and media playback / recording management.
The encryption / decryption module 262 provides encryption for outgoing packets and decryption for incoming packets. In the current example, the encryption / decryption module 262 provides application-level encryption at the source, rather than the network. However, it is understood that the encryption / decryption module 262 can provide encryption on the network in some ways.
Feature layer 264 provides support for various features such as voice, video, IM, data, voicemail, file transfer, file sharing, class 5 features, short message service (SMS), interactive voice response (IVR) ), faxes, and other resources. Protocol layer 266 includes the protocols supported by the endpoint, which include SIP, HTTP, HTTPS, STUN, RTP, SRTP, and ICMP. It is understood that these are only examples, and that fewer or more protocols may be supported.
The speech-to-text engine 268 converts the voice received by the endpoint (for example, via a microphone or the network) into text, the text-to-speech engine 270 converts the text received by the endpoint into speech ( for example, for output through a loudspeaker), and the language conversion engine 272 can be configured to convert incoming or outgoing information (text or voice) from one language to another language. Out-of-network connectivity module 274 can be used to manipulate connections between the endpoint and external devices (as described with respect to Fig. 12), and the connection module from other networks 276 manipulates incoming connection attempts from external devices. The cellular network interface module 280 can be used to interact with a wireless network.
With additional reference to Fig. 2C, cellular network interface module 280 is illustrated in more detail. Although not shown in Fig. 2B, software switch 258 of endpoint architecture 250 includes a network interface for communication with the 280 cellular network interface module. Additionally, the cellular network interface module 280 includes various components such as a call control module, a signaling agent, a media manager, a protocol stack, and a device interface. It should be noted that these components may correspond to layers within the endpoint architecture 250 and may be incorporated directly into the endpoint architecture in some embodiments.
Referring again to Fig. 2B, in operation, the software switch 258 uses the functionality provided by the underlying layers to manipulate the connections to other endpoints and access server 102, and to manipulate the services that the point needs. end 104. For example, incoming and outgoing calls can use multiple components within the endpoint 250 architecture.
The figures that follow are sequence diagrams illustrating various exemplary functions and operations by which Access Server 102 and endpoints 104 and 106 can communicate. It is understood that these diagrams are not exhaustive and that various stages may be excluded from the diagrams to clarify the aspect described.
Referring to Fig. 3 (and using endpoint 104 as an example), a sequence diagram 300 illustrates an exemplary process whereby endpoint 104 can authenticate with access server 102 and then communicate with the endpoint 106. As will be described, after authentication, all communication (both signaling and media traffic) between endpoints 104 and 106 occurs directly without any intervention by access server 102. In the current example, neither endpoint is understood to be online at the beginning of the sequence, nor are endpoints 104 and 106 friendly. As described above, friends are endpoints that both of you have previously agreed to communicate with each other.
In step 302, endpoint 104 sends a registration and / or authentication request message to access server 102. If endpoint 104 is not registered with access server 102, the access server will receive the request for log in (for example, user ID, password, and email address) and it will create a profile for the endpoint (not shown). The user ID and password will then be used to authenticate endpoint 104 during subsequent logins. It is understood that the user ID and password can allow the user to authenticate from any endpoint, rather than just endpoint 104.
Upon authentication, access server 102 updates a session table residing on the server to indicate that the user ID currently associated with endpoint 104 is online. Access server 102 further retrieves a friend list associated with the user ID currently used by endpoint 104 and identifies which of the friends (if any) are online using the session table. Since endpoint 106 is currently offline, the friends list will reflect this status. Access server 102 then sends the profile information (for example, the friend list) and a routing table to endpoint 104 at step 304. The routing table contains address information for the online members of the friends list. Steps 302 and 304 are understood to represent a start and break connection that breaks after endpoint 104 receives the profile information and routing table.
In steps 306 and 308, endpoint 106 and access server 102 repeat steps 302 and 304 as described for endpoint 104. However, because endpoint 104 is online when endpoint 106 If authenticated, the profile information sent to endpoint 106 will reflect the online status of endpoint 104, and the routing table will identify how to contact you directly. Accordingly, in step 310, endpoint 106 sends a message directly to endpoint 104 to notify endpoint 104 that endpoint 106 is now online. This further provides the endpoint 104 with the address information necessary to communicate directly with the endpoint 106. In step 312, one or more communication sessions can be established directly between the endpoints 104 and 106.
NAT routing and routing can be performed as described, for example, in US Pat.
7,570,636, filed on August 30, 2005, and titled SISTEMA Y
METHOD FOR TRAVELING A NAT DEVICE FOR HYBRID COUPLE-TO-COUPLE COMMUNICATIONS, and the United States patent application with serial number 12 / 705,925, filed on February 2010, entitled SYSTEM AND METHOD FOR STRATEGIC ROUTING IN A COUPLE TO COUPLE.
Referring to FIG. 4, in another embodiment, an environment 400 is illustrated in which endpoint 104 (eg, endpoint 104 in FIG. 1) is associated with a user who speaks a different language than an endpoint 106 user. This speech difference can make audible communications between endpoint 104 and 106 end users difficult or impossible. For example, if the endpoint user 104 only speaks English and the endpoint user 106 only speaks Spanish, then the two users will not be able to communicate by speaking easily and may not be able to communicate at all by speaking. Accordingly, one or both of the endpoints 104 and 106 may perform one-way or two-way translations or may have such translations performed by system components of a hybrid peer-to-peer network (not shown) as described above. .
In the current embodiment, a language translation component 402 is used to translate the voice before sending the translated voice to the other endpoint. However, other communications, such as signaling, are made directly between endpoint 104 and endpoint.
106 as indicated by arrow 404. Other media, such as data, video, and / or audio that do not need to be translated (eg, music or other unspoken audio), also communicate directly between endpoint 104 and endpoint 106. The original (ie untranslated) voice audio is sent to language translation component 402 as indicated by arrow 406, which translates the voice and sends the translated voice to endpoint 106 as indicated by arrow 408.
It is understood that the functionality provided by the language translation component 402 may be provided by a system component of a hybrid peer-to-peer network (not shown) as illustrated in FIG. 4 or some or all of the functionality provided by language translation component 402 can be provided by endpoint 104 itself. For example, endpoint 104 may contain the speech to text engine 268, the text to speech engine 270, and / or the language conversion engine 272 of Fig. 2B. Furthermore, the system component illustrated in Fig. 4 may be comprised of various physical or logical components that operate with each other to provide the functionality described herein, rather than a single component as shown.
Although not shown, it is understood that endpoint 104 can authenticate with the peer-to-peer hybrid network through access server 102 of FIG. 1 or with a similar authorization system. During the authentication process, endpoint 104 may obtain language information, such as one or more languages used by endpoint 106, the languages available through language translation component 402, and / or similar information that may assist the endpoint user 104 to communicate with the endpoint user 106. For example, endpoint 104 may receive language information about endpoint 106 in a profile received from access server 102.
Referring to FIG. 5, a sequence diagram illustrates an embodiment of a message sequence 500 that can occur in the environment of FIG. 4 in which endpoint 104 uses the translation functionality provided by the translation component. of languages 402. In the current example, endpoints 104 and 106 are friends and are able to communicate freely as described above.
In step 502, endpoints 104 and 106 exchange signaling information directly with each other for the purpose of establishing a communication session as described in previous embodiments. This information may include information that identifies the languages that are available and / or preferred, a unique identifier such as a caller ID that is used to identify the communication session being established, encryption keys, and / or similar information. Media such as data, video, and / or audio that do not need to be translated (eg, music or other unspoken audio) can also communicate directly between endpoint 104 and endpoint 106.
In step 504, voice input is received by endpoint 104 from a user. For example, the user can speak into a microphone and endpoint 104 can detect audio received through the microphone as voice. In other embodiments, endpoint 104 can be configured to recognize speech regardless of origin. In still other embodiments, endpoint 104 may allow a user to designate a file or media stream as voice. For example, a file containing voice may be labeled as voice by the user and treated by endpoint 104 as needing translation.
At step 506, endpoint 104 sends the voice to language translation component 402 for translation. Endpoint 104 may inform language translation component 402 of information such as the original language of the voice input, the language into which the translation is to be made, the caller ID, or other information that the translation component needs. of 402 languages to translate and process the received voice. At step 508, the language translation component 402 translates the voice from one language into another language as requested. At step 510, the language translation component 402 sends the translated voice to endpoint 106.
It is understood that, in the modalities where the translation functionality is internal to endpoint 104, steps 506, 508, and 510 may involve sending the voice to an internal module of endpoint 104 instead of to an external component as shown in Fig. 5. The internal module can then perform the translation and return the translated voice to another module at end point 104 to send it to end point 106 or it can send the translated voice directly to end point 106.
Referring to FIG. 6, in another embodiment, an environment 600 is illustrated in which endpoint 104 (eg, endpoint 104 in FIG. 1) is associated with a user who speaks a different language than a Endpoint 106 user. This difference in speech can make audible communications difficult or impossible between endpoint 104 and 106 user. For example, if the endpoint user 104 only speaks English and the endpoint user 106 only speaks Spanish, then the two users will not be able to communicate by speaking easily and may not be able to communicate at all by speaking. Accordingly, one or both of endpoints 104 and 106 may perform one-way or two-way translations or may have such translations performed by system components of a hybrid peer-to-peer network (not shown) as described above. .
In the current embodiment, the language translation component 402 of Fig. 4 is used to translate the voice before sending the translated voice to the other end point. Other communications, such as signaling, are performed directly between endpoint 104 and endpoint 106 as indicated by arrow 602. Other media, such as data, video, and / or audio that do not need to be translated (for example, music or other unspoken audio), also communicate directly between endpoint 104 and endpoint 106. The original voice audio ( that is, untranslated) is sent to language translation component 402 as indicated by arrow 604, which translates the voice and sends the translated voice to endpoint 104 as indicated by arrow 606. Endpoint 104 then sends the translated voice directly to endpoint 106 as indicated by arrow 608.
Referring to FIG. 7, a sequence diagram illustrates an embodiment of a message sequence 700 that can occur in the environment of FIG. 6 in which endpoint 104 uses the translation functionality provided by the translation component. of languages 402. In the current example, endpoints 104 and 106 are friends and are able to communicate freely as described above.
In step 702, endpoints 104 and 106 exchange signaling information directly with each other for the purpose of establishing a communication session as described in previous embodiments. This information may include information that identifies the languages that are available and / or preferred, a unique identifier such as a caller ID that is used to identify the communication session being established, encryption keys, and / or similar information. Media such as data, video, and / or audio that do not need to be translated (eg, music or other unspoken audio) can also communicate directly between endpoint 104 and endpoint 106.
At step 704, the voice input is received by the endpoint
104 from a user. For example, the user can speak into a microphone and endpoint 104 can detect audio received through the microphone as voice. In other embodiments, endpoint 104 can be configured to recognize speech regardless of origin. In still other embodiments, endpoint 104 may allow a user to designate a file or media stream as voice. For example, a file containing voice may be labeled as voice by the user and treated by endpoint 104 as needing translation.
At step 706, endpoint 104 sends the voice to language translation component 402 for translation. Endpoint 104 may inform language translation component 402 of information such as the original language of the voice input, the language into which the translation is to be made, the caller ID, and other information that the voice component needs. 402 language translation to translate and process the received voice. At step 708, the language translation component 402 translates the voice from one language into another language as requested. At the stage
710, the language translation component 402 sends the translated voice to endpoint 104. In step 712, endpoint 104 sends the translated voice directly to endpoint 106.
It is understood that, in modalities where the translation functionality is internal to endpoint 104, steps 706, 708, and 710 may involve sending the voice to an internal module of endpoint 104 rather than to an external component as shown in Fig. 7. The internal module can then perform the translation and return the translated voice to another module at end point 104 to send it to end point 106 or it can send the translated voice directly to end point 106.
Referring to FIG. 8, in another embodiment, an environment 800 is illustrated in which endpoint 104 (eg, endpoint 104 in FIG. 1) is associated with a user who speaks a different language than a Endpoint user 106. In environment 800, the language translation component 402 represents multiple system components of the hybrid peer-to-peer network. Although shown as part of the Languages 402 translation component, it is understood that the various components / functions illustrated in Fig. 8 may be distributed and need not be part of a single component and / or may be part of endpoint 104 and / or or from endpoint 106.
The language translation component 402 includes an 802 speech to text module (STT), an 804 text to speech module (TTS), and an 806 language translation module. The STT 802 module is configured to receive voice audio and convert voice audio to text. The TTS module is configured to receive text and convert text to voice audio. The 806 language translation module is configured to receive text from the STT 802 module and translate the text from one language to another language (for example, from English to Spanish).
Although the term module is used for description purposes, it is understood that some or all of the STT 802 module, the TTS 804 module, and the language translation module 806 may represent servers, server arrangements, distributed systems, or may be configured as any other way necessary to provide the described functionality. For example, the STT 802 module may be an isolated server or it may be a server array that performs load balancing as known in the art. Consequently, the different STT 802 modules can be involved in manipulating the voice received from endpoint 104 in a single communication session. An identifier such as a caller ID can be used to distinguish the communication session that occurs between the endpoint 104-106 from other communication sessions. Similarly, the TTS 804 module can be an isolated server or it can be a server arrangement that performs load balancing, and the different TTS 804 modules can be involved in manipulating the voice sent to the endpoint.
106 in a single communication session.
Each of the 802 speech-to-text module (STT), 804 text-to-speech module (TTS), and language translation module 806 may be comprised of existing and / or custom hardware and / or software components. It is understood that even with existing components, some adaptation may take place with the aim of adapting the components to the hybrid peer-to-peer network in which the endpoints 104 and 106 operate. For example, in the current mode, an RTP 808 layer is positioned between the STT 802 module and endpoint 104, and an RTP 810 layer is positioned between TTS 804 module and endpoint 106. RTP layers 808 and 810 can , among other functions, support encryption functionality between the STT 802 module and endpoint 104 and between TTS 804 module and endpoint 106.
In operation, the STT 802 module, the TTS 804 module, and the 806 language translation module are used to translate voice from endpoint 104 before voice is sent to endpoint 106. Although not shown, it is understood that the process can be reversed for voice flowing from endpoint 106 to endpoint 104. Other communications, such as signaling, are performed directly between endpoint 104 and endpoint 106 as indicated by arrow 812. During initial signaling, endpoint 104 may start a media connection to endpoint 106 for the communication session as indicated by arrow 813, but may never finish establishing the media connection due to the need for translate the voice. For example, the end point
104 You can start establishing the media path with endpoint 106, but you can hold audio packets until signaling ends. Signaling may indicate that the translation is requested using a translation indicator or other identifier and thus endpoint 104 may not finish establishing the media connection as it would for a normal (eg untranslated) call.
Other media, such as data, video, and / or audio that does not need to be translated (for example, music or other unspoken audio), also communicate directly between endpoint 104 and endpoint 106 as indicated by arrow 826 ( which may be the same as arrow 813 in some modalities). Signaling and other means may use a route as described above (eg, a private, public, or relay route) and arrows 812 and 826 may represent a single different route or routes.
In the current example, the original (ie, untranslated) voice audio is sent to the STT 802 module through RTP layer 808 as indicated by arrow 814. For example purposes, arrow 814 represents a route RTP. Consequently, endpoint 104 receives the voice audio as input, packages the voice audio using RTP, and sends the RTP packets to the STT 802 module. The RTP 808 layer is then used to convert the RTP voice audio to a form that the STT 802 module is capable of using. The STT 802 module converts the voice audio to text and sends the text to the language translation module. 806 as indicated by arrow 816. In some embodiments, after the text is translated, the language translation module 806 may send the text to the STT 802 module as indicated by 818, which forwards the text to endpoint 104 as indicated by arrow 820. Endpoint 104 can then display the translated text to the endpoint 104 user. For example, text can be translated from a source language (for example, English) to a target language (for example, Spanish), and then from the target language back to the source language, and the text of the target language. Translated source can be sent to endpoint 104 so that the user can determine if the translation in the source language is correct. The language translation module 806 sends the translated text to the TTS module 804 as indicated by arrow 822. The TTS module 840 converts the translated text to speech audio and forwards the translated speech audio to endpoint 106. The Endpoint 106 can then play back the voice audio in the translated language.
With further reference to Fig. 9, an embodiment of a portion of environment 800 of Fig. 8 is illustrated with each endpoint 104 and 106 that is associated with an STT module and a TTS module. More specifically, endpoint 104 is associated with STT 802a module and TTS 804a module. Endpoint 106 is associated with STT 802b module and TTS 804b module. Each endpoint 104 and 106 sends its outgoing voice to its respective STT 802a and 802b. Similarly, each endpoint 104 and 106 receives its incoming voice from its respective TTS module 804a and 804b.
In some embodiments, the functionality of the STT 802 module, TTS 804 module, and / or language translation module 806 may be included in endpoint 104. For example, endpoint 104 may translate audio from speech to text using an internal STT module 802, send the text to an external language translation module 806, receive the translated text, convert the translated text to voice audio, and send the translated voice audio to endpoint 106. In other embodiments, endpoint 104 can translate speech-to-text audio using an internal 802 STT module, send the text to an external language translation module 806, receive the translated text, and send the translated text to the endpoint. 106. Endpoint 106 would then display the translated text or convert the translated text to voice audio for playback. In still other embodiments, endpoint 104 may have an internal language translation module 806.
Referring to FIG. 10, a sequence diagram illustrates an embodiment of a message sequence 1000 that can occur in the environment of FIG. 8 in which endpoint 104 uses the translation functionality provided by the language translation 806. In the current example, endpoints 104 and 106 are friends and are able to communicate freely as described above.
In step 1002, endpoints 104 and 106 exchange signaling information directly with each other for the purpose of establishing a communication session as described in previous embodiments. This information may include information that identifies the languages that are available and / or preferred, a unique identifier such as a caller ID, which is used to identify the communication session being established, one or more encryption keys (for example , an encryption key for each endpoint 104 and 106), and / or similar information. Data such as files (eg, documents, spreadsheets, and images), video, and / or audio that does not need to be translated (eg, music or other unspoken audio) can also communicate directly between endpoint 104 and the end point 106. It is understood that step 1002 may continue during the following steps, with signaling and other means being passed directly between end points 104 and 106 during message sequence 1000.
In step 1004, voice input is received by endpoint 104 from a user. For example, the user can speak into a microphone and endpoint 104 can detect audio received through the microphone as voice. In other embodiments, endpoint 104 can be configured to recognize speech regardless of origin. In still other embodiments, endpoint 104 may allow a user to designate a file or media stream as voice. For example, a file containing voice may be labeled as voice by the user and treated by endpoint 104 as needing translation.
In step 1006, endpoint 104 sends the voice to the STT 802 module. Endpoint 104 can obtain the network address (eg, IP address) of the STT 802 module from the profile received during authentication, or it can obtain the network address in other ways, such as by querying access server 102 (Fig. 1). Endpoint 104 may send information to the STT 802 module such as the source language, the target language into which the source language is to be translated, the call ID of the communication session, and / or other information to be used in the translation process. It is understood that this information may be passed on to the 402 language translation component as needed. For example, the Caller ID can be used at each stage to distinguish the communication session for which translation is performed from other communication sessions undergoing translation.
In step 1008, the STT module 802 converts the received speech to text and, in step 1010, sends the text to the language translation module 806 along with any necessary information, such as the source and destination languages. In step 1012, the language translation module translates the text from the source language into the target language and, in step 1014, returns the translated text to the STT 802 module.
In the current example, the STT 802 module sends the translated to endpoint 104 at step 1016. This text can then be displayed to the endpoint 104 user to view the translation and, in some modes, to approve or reject the translation and / or request a new translation. At step 1018, the STT 802 module sends the translated text to the TTS module
804. The TTS module 804 converts the translated text to speech audio at step 1020 and, at step 1022, sends the speech audio to endpoint 106.
Referring to Fig. 11, a flowchart illustrates one embodiment of a method 1100 that can represent a process whereby an endpoint such as endpoint 104 in Fig. 8 obtains a translation for voice audio to be sent to another endpoint such as endpoint 106 during a communication session. It is understood that, during the execution of method 1100, endpoint 104 may be receiving the translated voice audio corresponding to the communication session with endpoint 106. In the current example, the communication session has been established before the step 1102 of method 1100. The communication session establishment can identify that the endpoints 104 and 106 will be using different languages and the endpoint 104 can retrieve or obtain the Information necessary to translate the voice audio based on the information exchanged during the establishment of the communication session. call.
It is further understood that signaling and non-audio / translated media may be sent by endpoint 104 directly to endpoint 106 and received by endpoint 104 directly from endpoint 106 during execution of method 1100.
At step 1102, endpoint 104 receives an audio input representing the voice. As previously described, voice audio can be labeled as such by a user, can be identified based on its source (eg, a microphone), can be identified by software and / or hardware configured to identify the voice, and / or can identify yourself in any other way.
In step 1104, it is determined whether translation is required. For example, if the endpoint user 104 only speaks English and the endpoint user 106 only speaks Spanish, then a translation is required. The need for a translation can be set by the user (eg by selecting an option during call set-up or later), it can be automatic (eg based on profiles of endpoints 104 and 106), and / or it can identify yourself in any other way. If no translation is required, the voice audio is sent directly to the endpoint at step 1106. If translation is required, the method
1100 it moves to step 1108. In step 1108, the voice audio is sent to the STT 802 module as described with respect to Fig. 8.
In step 1110, it is determined whether the translated text is expected from the STT 802 module and, if expected, whether the translated text has been received. For example, endpoint 104 may be configured to receive a text translation of the voice audio sent to the STT 802 module in step 1108. If so, step 1110 can be used to determine if the translated text has been received from the STT 1102 module. If not received, method 1100 can wait until the translated text is received, until a timeout is exceeded, or until another defined event occurs. If translated text is not expected, method 1100 ends.
In step 1112, if in step 1110 it was determined that the translated text is expected and the translated text has been received, it is determined whether endpoint approval 104 is required. For example, the STT 802 module may require the approval from endpoint 104 before sending the translated text to the TTS 804 module. Accordingly, in some embodiments, endpoint 104 displays the translated text to the user of endpoint 104 and may then wait for user input. In other modes, approval can be granted automatically by endpoint 104. For example, the user may approve some translations until the translation quality is satisfied, and can then select an option to automatically approve future translations.
If approval is not required as determined in step
1112, method 1100 ends. If approval is required, method 1100 proceeds to step 1114, where it is determined whether approval is granted (for example, whether the user has provided an entry indicating approval of the translated text). If approval is granted, method 1100 moves to step 1116, where approval is sent to the module
STT 802. If the approval is rejected, method 1100 moves to step 1118, where it is determined whether to request a new translation.
If a new translation is requested, method 1100 returns to step 1108. It is understood that endpoint 104 may request a new translation, and the STT 802 module may then use a different language reference, may select an alternative translation, or may perform other actions to achieve the requested result. It is further understood that new translations may not be available in some modalities. If a new translation is not requested, method 1100 can continue to step 1120, where a reject is sent to the module
STT 802.
Referring to FIG. 12, a flowchart illustrates an embodiment of a method 1200 that may represent a process whereby the language translation component 402 of FIG. 4 receives, translates, and sends voice audio. It is understood that the language translation component 402 is the one illustrated in Fig. 8 for example purposes, but can be configured differently. The endpoint user 104 communicates in English in the current example and the endpoint user 106 communicates in Spanish.
In step 1202, the STT 802 module receives English voice audio from the endpoint 104. The voice audio in the current example is received through the RTP 808 interface, but can be received in any format supported by the endpoint 104 and / or the STT 802 module. The STT 802 module may further receive information regarding the received text, such as the original language (for example, English), the target language (for example, Spanish), the call ID associated with the communication session, and / or other information.
In step 1204, the STT 802 module converts English speech audio to English text and sends the text to the 806 language translation module. In step 1206, the 806 language translation module converts the English text. into Spanish and sends the translated text to the STT 802 module. At step 1208, it can be determined whether endpoint 104 needs to approve the translation as described above with respect to Fig. 11. If approval is required, the STT 802 module sends the Spanish text back to endpoint 104 at step 1210. At step 1212, it can be determined whether approval has been received. If approval has not been received, method 1200 ends. If approval has been received as determined by step 1212 or if approval is not needed as determined in step 1208, method 1200 continues to step 1214.
In step 1214, the TTS 804 module receives the Spanish text from the STT 802 module and converts the text into Spanish voice audio. In step 1216, the TTS module 804 sends the Spanish voice audio to the endpoint 106 through the RTP layer 810. Although RTP is used for example purposes, it is understood that the Spanish voice audio can be sent in any format compatible with endpoint 106 and / or module
TTS 804.
Referring to FIG. 13A, in another embodiment, an environment 1300 is illustrated in which end point 104 (eg, end point 104 in FIG. 1), end point 106, an end point 1302, and an endpoint 1303 is on a conference call using a conference bridge 1304. Users of endpoints 104 and 1303 speak the same Language, while users of endpoints 104/1303, 106, and 1302 each speak a different language. For example, endpoint users 104/1303 can speak only English, endpoint user 106 can speak only German, and endpoint user 1302 can speak only Spanish. Although users of endpoints 104 and 1303 can easily communicate due to their Common Language, speech differences between users of endpoints 104/1303, 106, and 1302 can make audible communications difficult or impossible. For purposes of illustration, Fig. 13A shows the outgoing voice from the perspective of endpoint 104 and the incoming voice from the perspective of endpoints 106, 1302, and 1303. It is understood that each endpoint 104, 106, 1302, and 1303 can send and receive voice in a similar or identical way.
To aid in communications between endpoint users 104, 106, 1302, and 1303, each endpoint can access an STT module and a TTS module as described above with respect to Fig. 8. For example, the endpoint 104 is associated with a STT 802a module and a TTS 804a module, endpoint 106 is associated with a STT 802b module and a TTS 804b module, endpoint 1302 is associated with a module
STT 802c and a TTS 804c module, and endpoint 1303 is associated with a STT 802d module and a TTS 804d module. It is understood that STT 802a-802d modules and / or TTS 804a-804d modules can be the same module or can be part of one or more servers, server arrays, distributed systems, endpoints, or can be configured in any other way required to provide the functionality described as described above. In the current example, each pair of STT / TTS modules is illustrated as providing outgoing translation functionality only for their respective endpoints 104, 106, 1302, and 1303. Accordingly, each pair of STT / TTS modules is illustrated with the STT module positioned closer to the respective end point 104, 106, 1302, and 1303. However, it is understood that the same STT / TTS modules or other STT / TTS modules can provide inbound translation functionality for an endpoint, in which case a TTS module can communicate directly with the endpoint as illustrated by the TTS module. 804 of Fig. 8.
Environment 1300 includes the 806 language translation module that couples to the STT 802a module and the TTS 804a module. STT 802b-802d modules and TTS 804b-804d modules can be attached to the 806 language translation module or can be attached to another language translation module (not shown).
Conference bridge 1304 provides bridging capabilities for a multi-party conference call that includes endpoints 104, 106, 1302, and 1303. In the current example, the conference bridge
1304 provides a separate pair of ports for each endpoint 104, 106,
1302, and 1303. Accordingly, endpoint 104 communicates with conference bridge 1304 through an input port (from the perspective of conference bridge 1304) 1306 and an output port 1308, endpoint 106 communicates with bridge conference 1304 through an input port 1310 and an output port 1312, endpoint 1302 communicates with conference bridge 1304 through an input port 1314 and an output port 1316, and endpoint 1303 communicates with conference bridge 1304 through an input port 1318 and an output port 1320.
It is understood that conference bridge 1304 can be configured differently, with more or fewer input ports and / or output ports. The input and / or output ports can also be configured differently. For example, instead of a single output port for each endpoint 104, 106, 1302, and 1303, conference bridge 1304 may have one output port for each language (for example, English, German, and Spanish), and an endpoint can connect to the port associated with a desired language. In another example, all endpoints can send voice audio to a shared port or ports on the conference bridge. In some embodiments, conference bridge 1304 can also bridge the conference call with devices that are not endpoints, such as telephones and computers that do not have the endpoint functionality described herein.
With additional reference to Fig. 13B, endpoints 104,
106, 1302, and 1303 can communicate through conference bridge 1304 for voice, but can communicate directly for signaling and / or non-audio / untranslated means as described above. For illustration purposes, Fig. 13B shows direct communication routes 1350, 1352, and 1354 from the perspective of endpoint 104, but it is understood that such routes can exist directly between each pair of endpoints made up of endpoints 104, 106, 1302, and 1303. Accordingly, endpoint 104 can send the voice to endpoints 106, 1302, and 1303 through conference bridge 1304, but can send the signaling information and / or a file or data stream (for example, a word processing document, spreadsheet, image, or video stream) directly to endpoints 106,
1302, and 1303 through routes 1350, 1352, and 1354, respectively.
Alternatively or additionally, conference bridge 1304 may provide forwarding or other distribution capabilities for those communications, in which case routes 1350, 1352, and / or 1354 may not exist or may carry less traffic (for example, only signaling).
Referring again to Fig. 13A, in operation, endpoint 104 receives voice input, such as helium in English. Endpoint 104 sends the voice input to the STT 802a module as illustrated by arrow 1318. The STT 802a module converts speech to text and sends the text to language translation module 806 as illustrated by arrow 1320 . The language translation module 1304 converts the text from English to German (eg Guten Tag) and Spanish (eg Hello) before sending the translated text back to the STT 802a module as illustrated by arrow 1322 The STT 802a module sends the translated text to the TTS 804a module as illustrated by arrow 1324, although the language translation module 806 can send the translated text directly to the TTS 804a module in some ways. The module
TTS 804a converts the translated German and Spanish text to German and Spanish voice audio and outputs the audio to input port 1306 on conference bridge 1304 as illustrated by arrow 1326.
Conference bridge 1304 identifies endpoint 106 as corresponding to German and endpoint 1302 as corresponding to Spanish. For example, conference bridge 1304 may have a list of all endpoints that are connected to a call, and the list may include the language associated with each endpoint. When conference bridge 1304 receives packets from endpoint 104 that identifies their language as German or Spanish, conference bridge 1304 can query endpoints that have German or Spanish listed as their language and send packets to appropriate output ports. Consequently, conference bridge 1304 sends German voice audio to output port 1312 associated with endpoint 106 as illustrated by arrow 1328 and sends Spanish voice audio to output port 1316 associated with the end point 1302 as illustrated by the arrow
1330. It is understood that voice audio can be buffered or in any other way temporarily bypassed by the conference bridge
1304 and that voice audio may not move directly from input port 1306 to output ports 1312 and 1316. German voice audio is then sent to endpoint 106 through output port 1312 as illustrated by arrow 1332 and Spanish voice audio is sent to endpoint 1302 through output port 1316 as illustrated by arrow 1334.
In the current example, the translated voice audio is sent directly to each endpoint 106 and 1302 instead of to the STT or TTS associated with each endpoint. Since the English voice audio has been converted to text, translated, and converted back to voice audio by the STT 802a module and the TTS 804a module associated with the endpoint 104, the translated voice audio can be sent from the bridge Conference 1304 directly to endpoints 106 and 1302. In other modes, the voice can be sent as text from the STT 802a module or the 806 language translation module to the conference bridge 1304, from the conference bridge 1304 to the TTS module of each endpoint 106 and 1302 and become voice audio, and go to the associated endpoint 106 and 1302. If the text is sent untranslated from the STT 802a module to conference bridge 1304, the text can be sent to language translation module 806 from conference bridge 1304 or modules
SST / TTS of endpoints 106 and 1302.
Since the endpoint user 1303 uses the same language as the endpoint user 104, there is no need to translate the voice before sending it to the endpoint 1303. Consequently, the original voice audio is sent from the TTS module 804a to conference bridge 1304 as illustrated by arrow 1330. For example, the STT 802a can receive the original voice audio and, in addition to converting it to text for translation, forward the original voice audio to the TTS 804a module. . In other embodiments, endpoint 104 can send the original voice audio directly to TTS module 804a for forwarding. The TTS 804a module can convert the translated voice text into German and Spanish as described above, and can also receive the original voice audio as forwarded by the STT 802a. The TTS 804a module can then forward the original voice audio to input port 1306 on conference bridge 1304.
Conference bridge 1304 identifies endpoint 1303 as corresponding to English, and when conference bridge 1304 receives packets from endpoint 104 that identify their language as English, conference bridge 1304 can query endpoints that have Listed English as your language and send packets to the appropriate outgoing ports. Accordingly, conference bridge 1304 sends the English voice audio to the output port 1320 associated with the end point 1303 as illustrated by arrow 1340. The English voice audio is then sent to the end point 1303 to through exit port 1320 as illustrated by arrow 1342.
With reference to Figs. 14A-14C, the sequence diagrams illustrate an embodiment of message sequences 1400, 1450, and 1480 that can occur in the environment of Fig. 13A in which endpoint 104 uses the translation functionality provided by the translation module of language 806. Some communications between endpoints 104 and 1303 are detailed with respect to Fig. 13B, rather than Fig. 14A.
Referring specifically to Fig. 14C, the end points
104, 106, 1302, and 1303 are all coupled to conference bridge 1304 and can send signaling and / or non-audio / non-translated media directly from one to the other, although these can be sent through conference bridge 1304 in other ways. . In the current mode, signaling occurs directly (eg, does not pass through conference bridge 1304) between endpoint 104 and endpoints 106, 1302, and 1303 as illustrated in steps 1482, 1484, and 1486. , respectively. Although Fig. 13C is from the perspective of endpoint 104, signaling may further occur between the other endpoints in a similar or identical manner, with each endpoint sending signals directly to the other endpoints. It is understood that this signaling may continue while a conference call is in session, and may include setup and maintenance signaling for the conference call. For example, endpoints 104, 106, 1302, and 1303 may send signals directly to establish the conference call, and the signaling may contain the parameters necessary to establish the conference call, such as source_language for an endpoint, target_language_2, target_language_2 , target_language_x and identifiers of which endpoint is associated with each language, such as endpoint 104: English, endpoint 106: German, endpoint 1302: Spanish, and end point 1303: English. Signaling between endpoints can be used to establish a communication route, such as one or more of the private, public, and / or relay routes described above.
Non-audio / untranslated media can also be transferred directly between endpoint 104 and endpoints 106, 1302, and 1303 as illustrated in steps 1488, 1490, and 1492, respectively. Although Fig. 13C is from the perspective of the end point
104, such transfers may further occur between the other endpoints in a similar or identical manner, with each endpoint directly transferring the non-audio / untranslated media to the other endpoints. It is understood that non-audio / untranslated media transfers can continue while a conference call is in session, and can begin before the conference call is established and continue after the conference call ends. Non-audio / untranslated media may be transferred through one or more of the private, public, and / or relay routes as described for signaling.
Referring specifically to Fig. 14A (which is directed to endpoints 104, 106, and 1302), at steps 1402, 1404, and 1406, endpoints 104, 106, and 1302, respectively, contact the conference bridge 1304 and get one or more ports. In the current example, conference bridge 1304 assigns each endpoint 104, 106, and 1302 an input port (from the perspective of conference bridge 1304) and an output port as illustrated in Fig. 13A.
At step 1408, the voice input is received by the endpoint
104 from a user. For example, the user can speak into a microphone and endpoint 104 can detect audio received through the microphone as voice. In other embodiments, endpoint 104 can be configured to recognize speech regardless of origin. In still other embodiments, endpoint 104 may allow a user to designate a file or media stream as voice. For example, a file containing voice may be labeled as voice by the user and treated by endpoint 104 as needing translation.
In step 1410, endpoint 104 sends the voice to the STT 802a module. Endpoint 104 can obtain the network address (for example, the IP address) of the STT 802a module from the profile received during authentication, or it can obtain the network address in other ways, such as by querying access server 102 (Fig. . one). Endpoint 104 may include information such as the source language (for example, English), the target language (s) (for example, German and Spanish) towards which language (s) of source to be translated, the call ID of the communication session, and / or other information to be used in the translation process. It is understood that this information may be passed on to the 402 language translation component as needed. For example, the Caller ID can be used at each stage to distinguish the communication session for which translation is performed from other communication sessions.
In step 1412, the STT module 802a converts the received speech to text and, in step 1414, sends the text to the language translation module 806 along with any necessary information, such as the source and destination languages. In step 1416, the language translation module 806 translates the text from the source language into the target language and, in step 1418, returns the translated text to the STT 802a module. In some embodiments, the 806 language translation module can send the translated text directly to the TTS 804a module.
In step 1420, the STT 802a module sends the translated text to the TTS 804a module. The TTS 804a module converts the translated text to voice audio in step 1422. More specifically, the TTS 804a module converts the German translated text to German speech audio and converts the Spanish translated text to Spanish voice audio. In step 1424, the TTS module 804a sends the translated voice audio to the conference bridge 1304. In the current example, the TTS 804a module sends both languages simultaneously to input port 1306 (Fig. 13A). For example, the TTS 804a module can send languages through RTP and each language sequence can be identified by an SSRC or another identifier. In other modes, the TTS 804a module can send the audio for each language to a different port. For example, the conference bridge 1304 can set a separate input port for each language, and the TTS 804a module can send German voice audio to a German input port and Spanish voice audio to an input port. in Spanish. At step 1426, conference bridge 1304 sends the German voice audio to endpoint 106 and, at step 1428, sends Spanish voice audio to endpoint 1302. As described with respect to FIG. . 13A, conference bridge 1304 sends the voice audio directly to endpoints 106 and 1302 in the current example, bypassing the SST / TTS module pairs associated with each endpoint.
With reference to Fig. 14B (which is directed to the end points
104 and 1303), a sequence diagram illustrates an embodiment of a message sequence 1450 that can occur in the environment of Fig. 13A in which endpoint 104 communicates the original voice audio to endpoint 1303. The actual translation The original voice audio from English to German and Spanish for endpoints 106 and 1302 is not detailed in the current example, but may occur in a manner that is similar or identical to that described with respect to Fig. 13A. Accordingly, since Fig. 14B and Fig. 14A have many identical or similar stages, only stages specifically directed to endpoint 1303 are described in detail in the current example.
In the current mode, after the ports in steps 1452 and 1454 are configured and endpoint 104 receives a voice input in step 1456, endpoint 104 sends the original English voice audio to the STT 802a module in step 1458. The STT 802a module converts the original speech audio to text in step 1460, translated it in steps 1462,
1464, and 1466, and sends the translated text and the original speech audio to the TTS 804a module in step 1468. The TTS 804a module converts the speech to text in step 1470 and sends the audio in English, German, and Spanish to conference bridge 1304 in step 1472. Conference bridge 1304 then sends the English voice audio to endpoint 1303 in step 1474. Accordingly, English voice audio originating from endpoint 104 can be sent without translation using the same mechanism (eg, SST 802a module and TTS 804a module).
Referring to Fig. 15, in another embodiment, an environment 1500 is illustrated in which endpoint 104, endpoint 106, endpoint 1302, and endpoint 1303 are on a conference call using the bridge. lecture 1304 of Fig. 13A. Environment 1500 may be similar or identical to environment 1300 of Fig. 13A except for an additional communication path between endpoint 104 and the conference bridge
1304 for the original voice audio.
In the current example, endpoint 104 sends the original voice audio to the SST 802A module as described with respect to Fig. 13A. The SST 802a module can then manipulate speech to text conversion, translation, and forwarding to the TTS 804a module. The TTS module
804á can then manipulate the text-to-speech conversion and forward the audio in German and Spanish to conference bridge 1304. Endpoint 104 also sends the original speech audio directly to conference bridge 1304 as illustrated by the arrow. 1502. Consequently, instead of sending the original voice audio only to the SST 802a module for both translation and forwarding, endpoint 104 sends the original voice audio to the SST 802a module for translation and to the conference bridge. 1304 for forwarding.
With reference to Fig. 16 (which is directed to the end points
104 and 1303), a sequence diagram illustrates an embodiment of a message sequence 1600 that can occur in the environment of Fig. 15 in which endpoint 104 communicates the original voice audio to endpoint 1303. The actual translation The original voice audio from English to German and Spanish for endpoints 106 and 1302 is not detailed in the current example, but may occur in a manner that is similar or identical to that described with respect to Fig. 13A.
In the current mode, after the ports are configured in steps 1602 and 1604 and endpoint 104 receives a voice input in step 1606, endpoint 104 sends the original English voice audio to the STT 802a module for its translation at step 1608. Additionally, endpoint 104 sends the original voice audio directly to conference bridge 1304 at step 1610. Conference bridge 1304 then sends the original voice audio to endpoint 1303 at step 1612.
Referring to Fig. 17, in another embodiment, an environment 1700 is illustrated in which endpoint 104, endpoint 106, endpoint 1302, and endpoint 1303 are on a conference call using the bridge. lecture 1304 of Fig. 13A. Environment 1700 may be similar or identical to environment 1300 of Fig. 13A except that translation occurs after the original voice audio passes through conference bridge 1304. Consequently, conference bridge 1304 may not be configured to identify the languages associated with specific ports or endpoints in some modes. In such modalities, conference bridge 1304 can send the voice audio to a particular network address regardless of whether the address is for an endpoint, an STT module, or another network component. In other embodiments, conference bridge 1304 can be configured to recognize languages for the purpose of determining whether to send voice audio to a particular address, such as an endpoint or an STT module associated with an endpoint.
In operation, endpoint 104 receives voice input, such as helium in English. Endpoint 104 sends the voice input as the original English voice audio to input port 1306 on conference bridge 1304 as illustrated by arrow 1702. Conference Bridge 1304 sends the original English voice audio to output port 1312 associated with endpoint 106 as illustrated by arrow 1704, to output port 1316 associated with endpoint 1302 as illustrated by arrow 1706, and toward exit port 1320 associated with endpoint 1303 as illustrated by arrow 1708.
The original English voice audio is sent to the STT module
802b associated with endpoint 106 through output port 1312 as illustrated by arrow 1710. STT 802b module converts speech to text and sends text to language translation module 806 as illustrated by arrow 1712 The 806 language translation module converts the text from English to German (eg Guten Tag) before sending the translated text back to the STT 802b module as illustrated by the arrow.
1714. The STT 802b module sends the translated text to the TTS 804b module as illustrated by arrow 1716, although the language translation module 806 can send the translated text directly to the TTS 804b module in some ways. The TTS 804b module converts the translated German into German voice audio and outputs the audio to endpoint 106 as illustrated by arrow 1718.
The original English voice audio is also sent to the SST 802c module associated with endpoint 1302 through output port 1316 as illustrated by arrow 1720. The STT 802c module converts speech to text and sends text to the 806 language translation module as illustrated by arrow 1722. The 806 language translation module converts the text from English to Spanish (eg Hello ”) before sending the translated text back to the STT 802c module as illustrated by arrow 1724. The STT 802c module sends the text translated into the TTS module
804c as illustrated by arrow 1726, although the 806 language translation module can send the translated text directly to the TTS module
804c in some modalities. The TTS 804c module converts the translated Spanish into Spanish voice audio and sends the audio to the endpoint
1302 as illustrated by arrow 1728.
The original English voice audio is also sent to endpoint 1303 through output port 1320 as illustrated by arrow 1730. As endpoint 104 and endpoint 1303, both are associated with users who speak the same language. , the original voice audio is not passed to the STT 802d module in the current mode. In other embodiments, the original voice audio may be passed to STT 802d and forwarded to endpoint 1303, rather than sent directly to endpoint 1303.
With reference to Fig. 18 (which is directed to the end points
104, 106, and 1303), a sequence diagram illustrates an embodiment of a 1800 message sequence that can occur in the environment of Fig. 17 in which endpoint 104 uses the translation functionality provided by the language translation 806. End point 1302 of Fig. 17 is not illustrated in Fig. 18, but the message sequence that would occur with respect to endpoint 1302 may be similar or identical to the message sequence for endpoint 106 except that the language translation would be from English to Spanish rather than English to German.
Although not shown, it is understood that the messaging of Fig. 14C can be applied to endpoints 104, 106, and 1303 of the current example. In steps 1802, 1804, and 1806, endpoints 104, 106, and 1303, respectively, contact conference bridge 1304 and obtain one or more ports. In the current example, conference bridge 1304 assigns each endpoint 104, 106, and 1303 an input port (from the perspective of conference bridge 1304) and an output port as illustrated in
Fig. 13A.
In step 1808, voice input is received by endpoint 104 from a user, and in step 1810, endpoint 104 sends the original English voice audio to the input port on the conference bridge
1304. In step 1812, conference bridge 1304 sends the original English voice audio to the STT 802b associated with the endpoint 106 and, in step 1814, sends the original English voice audio to the endpoint 1303. As Endpoint 1303 does not need to translate the original voice audio into English, Endpoint 1303 does not need to use the associated 802d STT.
In step 1816, the STT 802b module converts the received speech to text and, in step 1818, sends the text to the language translation module 806 along with any necessary information, such as the source and destination languages. In step 1820, the Languages translation module 806 translates the text from the source English language into the target German language and, in step 1822, returns the translated text to the STT 802b module. In some embodiments, the 806 language translation module can send the translated text directly to the TTS 804b module. At step 1824, the STT 802b module sends the translated text to the TTS module
804b. The TTS 804b module converts the translated text to German speech audio in step 1826. In step 1828, the TTS 804b module sends the translated speech audio to the endpoint 106.
Referring to Fig. 19, one embodiment of a 1900 computing system is illustrated. The 1900 computing system is a possible example of a system component or device such as an endpoint, an access server, or a shadow server. The computer system 1900 may include a central processing unit (CPU) 1902, a memory unit 1904, an input / output (l / O) device 1906, and a network interface 1908. Components 1902, 1904, 1906, and 1908 are interconnected by a transportation system (eg, a bus) 1910. A 1912 power supply (PS) can provide power to components of the 1900 computer system, such as the CPU 1902 and memory unit 1904. It is understood that computer system 1900 may be configured differently and that each of the listed components may actually represent several different components. For example, CPU 1902 may actually represent a multiprocessor or a distributed processing system; memory unit 1904 can include different levels of cache memory, main memory, hard drives, and remote storage locations; device 1/01906 may include monitors, keyboards, and the like; and network interface 1908 may include one or more network cards that provide one or more wired and / or wireless connections to packet network 108 (FIG. 1). Therefore, a wide range of flexibility is anticipated in the configuration of the 1900 computer system.
The 1900 computer system can use any operating system (or multiple operating systems), including various versions of the operating systems provided by Microsoft (such as WINDOWS), Apple (such as Mac OS X), UNIX, and LINUX, and may include Operating systems developed specifically for handheld devices, personal computers, and servers depending on the use of the 1900 computer system. The operating system, as well as other instructions (eg, for the endpoint engine 252 in Fig. 2B if it is an endpoint), can be stored in memory unit 1904 and executed by processor 1902. For example, if the 1900 computer system is an endpoint (for example, one of the endpoints 104, 106, 1302, and 1303), the SST 802 module, the TTS 804 module, the 806 language translation module, or With conference bridge 1304, memory unit 1904 may include 10 instructions for performing the functions as described herein with respect to the various modalities illustrated in the sequence diagrams and flow diagrams.
In another embodiment, a method for communicating, by a first endpoint, in an ongoing peer-to-peer communication session between the first endpoint and a second endpoint in a peer-to-peer hybrid network comprises establishing, by first end point, a communication path directly between the first end point and the second end point, wherein signaling communications are sent directly from the first end point to the second end point through a signaling route provided by the communication route; receiving, by the first endpoint, a voice audio input in a first language spoken from a user of the first endpoint; determining, by the first endpoint, whether the voice audio input is to be translated from the first spoken language to a second spoken language; send, by the first endpoint, the voice audio input to a Language translation component accessible to the first endpoint through a hybrid peer-to-peer network if the voice audio input is to be translated from the first Language spoken towards the second spoken language, wherein the first endpoint does not send the voice audio input directly to the second endpoint if the voice audio input is to be translated from the first spoken Language to the second spoken Language; and sending, by the first endpoint, the voice audio input directly to the second endpoint through the communications path if the voice audio input is not to be translated from the first spoken language to the second spoken language. The method may further comprise sending, via the first endpoint, the non-audio data directly to the second endpoint via the communications path regardless of whether the voice audio input is to be translated from the first spoken Language. towards the second spoken language. The method may further comprise receiving, by the first endpoint, a translation of the voice audio input from the Language translation component; and providing the translation of the voice audio input to the user of the first endpoint. Sending the voice audio input to the language translation component may include establishing, by the first endpoint, a signaling path with a speech to text module of the language translation component; and send the voice audio input to the voice to text module. The method may further comprise performing, by the first endpoint, an authentication process with an access server in the hybrid peer-to-peer network; and receive, by the first endpoint, a profile from the access server after the authentication process, where the profile identifies the second endpoint as an endpoint with which the first endpoint has permission to communicate and identifies that the Second end point is associated with the second spoken language. The method may further comprise receiving, by the first endpoint, a language list from the access server, where the language list includes the first and second spoken languages and identifies the spoken languages that can be translated by the translation component of Languages. The method may further comprise receiving, by the first endpoint, a notification directly from the second endpoint that the second endpoint is associated with the second spoken language. The method may further comprise receiving, by the first endpoint from the language translation component, voice audio originating from the second endpoint, wherein voice audio originating from the second endpoint is translated from the second language spoken to the first language spoken by the language translation component before being received by the first endpoint; and produce, by the first endpoint, the voice audio received from the language translation component as an audible sound. The method may further comprise establishing, by the first endpoint, a second communication path directly between the first endpoint and a third endpoint, where signaling communications are sent directly from the first endpoint to the third endpoint to through a signaling route provided by the second communication route; identify, by the first endpoint, that the third endpoint is associated with the first spoken language; and send, by the first endpoint, the voice audio input directly to the third endpoint through the second communication path.
In another embodiment, a method for communicating, by a first endpoint in a hybrid peer-to-peer network, in an ongoing communication session with the second and third endpoints through a bridge comprises identifying, by the first endpoint , that the first end point is associated with a first spoken language, the second end point is associated with a second spoken language, and the third end point is associated with a third spoken language; send, on the first endpoint, a request to the bridge for an input port and an output port to be provided on the bridge for the first endpoint; notify, by the first endpoint, a language translation component on the hybrid peer-to-peer network about the port of entry, where the language translation component is accessible to the first endpoint through the hybrid network peer-to-peer, and where the notification instructs the language translation component to send the received audio from the first endpoint to the input port; send to the language translation component, by the first end point, the voice audio input received by the first end point from a user of the first end point, where the voice audio input sent by the first end point is in the first spoken language; and receive, by the first endpoint, voice audio from the second and third endpoints directly from the output port on the bridge, where the voice audio received by the first endpoint through the output port was sent by the second and third endpoints in the second and third spoken languages, respectively, and where the voice audio received by the first endpoint directly from the output port is received in the first spoken language. The method may further comprise identifying, by the first endpoint, that a fourth endpoint is associated with the first spoken language; and sending, by the first endpoint, the voice audio received from the user of the first endpoint directly to the fourth endpoint without using the bridge. The method may further comprise identifying, by the first endpoint, that a fourth endpoint is associated with the first spoken language; and sending, by the first endpoint, the voice audio received from the user of the first endpoint to the input port on the bridge. The method may further comprise sending, through the first endpoint, the non-audio data directly to at least one of the second and third endpoints, where the non-audio data does not pass through the bridge. The method may further comprise receiving, by the first endpoint, the non-audio data directly from at least one of the second and third endpoints, where the non-audio data does not pass through the bridge. The method may further comprise directly notifying the second and third endpoints, by the first endpoint, about the input and output ports, where the notification does not use the bridge. The method may further comprise receiving, by the first end point, notifications directly from each of the second and third end points about the input and output ports corresponding to the second and third end points. The method may further comprise sending an indication, by the first endpoint, to the language translation component that the voice audio input is to be translated into the second and third languages.
In another embodiment, a method of translating voice audio into a hybrid peer-to-peer network comprises receiving, by a voice-to-text module, a first voice audio medium from a first endpoint through a hybrid peer-to-peer network. in pairs, where the first voice audio medium is in a first human language; convert, by the speech to text module, the first voice audio medium into original text; send the original text via the speech to text module to a language translation module; translate, by the language translation module, the original text into translated text in a second human language; send, via the language translation module, the translated text to a text-to-speech module; converting, by the text-to-speech module, the translated text into a second voice audio medium, wherein the second voice audio medium is in the second human language;
and sending the second voice audio medium to a second endpoint in the peer-to-peer hybrid network. Sending of the second voice audio medium can be done by the text-to-speech module. The method may further comprise sending, via the speech to text module, the translated text to the first endpoint. The method may further comprise waiting, by the speech-to-text module, for an approval of the translated text from the first endpoint before sending the original text to the language translation module. Sending the second voice audio medium to the second end point may include sending the second voice audio medium to a port on a bridge identified by the first end point. The method may further comprise translating, through the language translation module, the original text into text translated into a third human language; send, via the language translation module, the translated text to the text-to-speech module;
converting, by the text-to-speech module, the translated text into a third voice audio medium, wherein the third voice audio medium is in a third human language; and sending the third voice audio medium to a third end point by sending the third voice audio medium to a port on the bridge identified by the first end point. The ports for the second and third media can be the same port. The method may further comprise receiving, by the speech-to-text module, instructions from the first endpoint identifying the first and second human languages.
In another embodiment, an endpoint device comprises one)
network interface; a processor coupled to the network interface; and a memory coupled to the processor and containing a plurality of instructions for execution by the processor, the instructions including instructions for;
performing an authentication process with an access server on a hybrid peer-to-peer network, where the authentication process authorizes the first endpoint to access the hybrid peer-to-peer network; receiving a profile from the access server that identifies a second endpoint as an endpoint within the hybrid peer-to-peer network with which the first endpoint has permission to communicate; determining that a user of the first endpoint has designated a first human language to be used by the first endpoint; establishing a communication route directly between the first end point and the second end point, where signaling communications are sent directly from the first end point to the second end point through the communication route; determine that the second endpoint has designated a second human language to be used by the second endpoint; receive a voice audio input in the first human language from the user of the first endpoint; and send the voice audio input to a language translation module for translation into the second human language. The endpoint device may further comprise instructions for sending non-audio data directly to the second endpoint via the communications path. Instructions to determine that the second endpoint has designated the second human language to be used by the second endpoint may include instructions to obtain from the profile an identifier representing the second human language. Instructions for determining that the second endpoint has designated the second human language to be used by the second endpoint may include instructions for obtaining an identifier representing the second human language directly from the second endpoint.
Although the foregoing description shows and describes one or more embodiments, it will be understood by those skilled in the art that various changes in form and detail may be made therein without departing from the spirit and scope of the present disclosure. For example, the various stages illustrated within a particular sequence diagram or flow chart can be combined or further divided. Additionally, the steps described in a sequence diagram or flow chart can be incorporated into another sequence diagram or flow chart. Furthermore, the described functionality can be provided by hardware and / or software, and can be distributed or combined on a single platform. Furthermore, the functionality described in a particular example can be achieved in a different way than that illustrated, but is still covered within the present description. Therefore, the claims should be interpreted broadly, consistent with the present description.
Contents6
2 priority claims, no other members on record
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 89033310 | United States of America | A | |
| 2011051877 | United States of America | W |
1 legal event, as the office reported them to INPADOC
Events
| Event | Code | |
|---|---|---|
| Grant or registrationFG | FG |
Numbers
- Application
- 2013003277
Titles2
- English
- SYSTEM AND METHOD FOR LANGUAGE TRANSLATION IN A HYBRID PEER-TO-PEER ENVIRONMENT.
- Spanish
- SISTEMA Y METODO PARA LA TRADUCCION DE UN IDIOMA EN UN ENTORNO HIBRIDO DE PARES A PARES.
Classification
- CPC, 3
- G06F40/58
- G06F40/40
- G10L15/005
- IPC, 3
- H04L12 18
- G06F15 16
- G06F17 28