A mass-scale, user-independent, device-independent, voice message to text conversion system
Abstract
A large-scale vocal messaging system, independent of the user and independent of the device, which allows to convert an unstructured vocal message into text for a visual presentation on a screen; characterized in that the system comprises (i) subsystems implemented by computer, as well as (ii) a network connection to provide transcription and quality control to human operators; the system being adapted to optimize the effectiveness of human operators, including: a grid subsystem implemented by computer to generate a grid of possible sequences of phrases or words and to allow a human operator to guide a conversion subsystem by presenting one or more converted candidate words or phrases from the grid and to allow the operator select the candidate word or phrase or by entering one or more characters for a different converted word, Operationally start the conversion subsystem to propose an alternative word or phrase.
Term
0.4 yearsto projected expiry
Projected expiry 12 February 2027, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
17 claims: 3 independent, 14 dependent
- 1REIVINDICACIONES 1. Un sistema de mensajería vocal a gran escala, independiente del usuario e independiente del dispositivo, que permite convertir un mensaje vocal no estructurado en texto para una presentación visual en una pantalla; caracterizado porque el sistema comprende (i) subsistemas puestos en práctica por ordenador, así como (ii) una conexión de red para proporcionar una transcripción y un control de calidad a operadores humanos; estando el sistema adaptado para optimizar la eficacia de los operadores humanos comprendiendo, además:un subsistema de retícula puesta en práctica por ordenador para generar una retícula de posibles secuencias de frases o palabras y para permitir a un operador humano guiar un subsistema de conversión presentando una o más palabras o frases convertidas candidatas a partir de la retícula y para permitir al operador seleccionar la palabra o la frase candidata o bien, introduciendo uno o varios caracteres para una palabra convertida diferente, iniciar operativamente el subsistema de conversión para proponer una palabra o frase alternativa.
- 2El sistema según la reivindicación 1, en donde el subsistema de retícula está configurado para recibir entradas procedentes de un subsistema que gestiona la información de registro histórico de llamada del par.
- 3El sistema según la reivindicación 1, en donde el subsistema de retícula está configurado para recibir entradas procedentes de recursos de conversión.
- 4El sistema según la reivindicación 1, en donde el subsistema de retícula está configurado para recibir entradas procedentes de un subsistema de contexto que tiene conocimiento del contexto de un mensaje.
- 5El sistema según la reivindicación 1, en donde el subsistema de retícula está configurado para aprender, a partir de las entradas del operador humano, palabras o frases probables que corresponden a un modelo sonoro.
- 6El sistema según la reivindicación 1, en donde el operador humano debe seleccionar solamente una sola tecla para aceptar una palabra o una frase.
- 7El sistema según la reivindicación 1, en donde el subsistema de retícula está configurado para proporcionar automáticamente mayúsculas iniciales y signos de puntuación.
- 8El sistema según la reivindicación 1, en donde el subsistema de retícula está configurado para proponer números, nombres reales, direcciones web, direcciones de correo electrónico, direcciones físicas, información de localización u otras coordenadas candidatas.
- 9El sistema según la reivindicación 1, en donde el subsistema de retícula está configurado para realizar automáticamente la distinción entre las partes del mensaje que son susceptibles de ser importantes y las que son susceptibles de no tener importancia.
- 10El sistema según la reivindicación 1, en donde las partes sin importancia del mensaje son confirmadas por el operador, como perteneciente a una clase propuesta por el subsistema de retícula y a continuación, se convierten únicamente por un motor de reconocimiento vocal ASR del aparato (1).
- 11El sistema según la reivindicación 1, en donde el operador humano puede pronunciar la palabra correcta en el destino del sistema de conversión, que está configurado para su transcripción automática más adelante.
- 12El sistema según la reivindicación 3, en donde los recursos de conversión analizan una palabra o una frase convertida con respecto a un cuerpo de conocimiento en línea.
- 13El sistema según la reivindicación 12, en donde el cuerpo de conocimiento en línea es Internet, accesible por un motor de búsqueda.
- 14El sistema según la reivindicación 13, en donde el cuerpo de conocimiento en línea es una base de datos de motor de búsqueda.
- 15El sistema según una cualquiera de las reivindicaciones precedentes, en donde el mensaje es uno de los mensajes siguientes:(a) un correo de voz destinado a un teléfono móvil y el sistema está configurado para convertir el mensaje vocal en texto y enviar el mensaje vocal a ese teléfono móvil o (b) un mensaje vocal destinado a un servicio de mensajería instantánea y el sistema está configurado para convertir el mensaje vocal en texto y enviar el mensaje vocal a un servicio de mensajería instantánea para su presentación visual en una pantalla o (c) un mensaje vocal destinado a un servicio web y el sistema está configurado para convertir el mensaje vocal en texto y enviar el mensaje vocal a un servidor para una presentación visual como parte del servicio web.
- 16El sistema según cualquier reivindicación precedente, en donde el mensaje es uno de los mensajes siguientes:5 (a) un mensaje vocal destinado a convertirse al formato de texto y enviarse bajo la forma de mensaje de texto o (b) un mensaje vocal destinado a convertirse al formato de texto y enviarse en tanto como mensaje de correo electrónico o 10 (c) un mensaje vocal destinado a convertirse al formato de texto y enviarse bajo la forma de una nota o de un memorándum, por correo electrónico o texto, a un expedidor del mensaje.
- 17Un método que permite proporcionar un sistema de mensajería vocal a gran escala, independiente del usuario e 15 independiente del dispositivo, que convierte un mensaje vocal no estructurado en texto para una presentación visual en una pantalla, caracterizado por cuanto que el sistema comprende:(i) subsistemas puestos en práctica por ordenador así como (ii) una conexión de red para proporcionar una transcripción y un control de calidad a operadores humanos;optimizando el método la eficacia de los operadores humanos comprendiendo las etapas de: 20 un subsistema de retícula puesto en práctica por ordenador, que genera una retícula de posibles secuencias de palabras o frases y que permite a un operador humano guiar un subsistema de conversión presentándole una o más palabras o frases convertidas candidatas a partir de la retícula y que permite al operador seleccionar la palabra o la frase candidata o, introduciendo uno o varios caracteres para una palabra o una frase convertida diferente, para iniciar operativamente el subsistema de conversión para proponer una palabra o frase alternativa. Calidad de la voz y confianza en el Reconocedor Motor de ASR Modelos de Reconocedor de voz Control de calidad Procesamiento lenguaje postconversión Voz Correo de voz Clasificación de voz y unidad decisión sobre estrategia de conversión Mejora de la voz Figura 1 Entrada mensaje de voz Aplicación de Control de Calidad QC Gestor de cola de espera (TAT) Preprocesamiento Adaptación de canal, lenguaje, ruido ASR Multi-motor, independiente del usuario que habla Figura 2 Postprocesamiento Frases de lenguaje natural Salida de texto convertido Corregir: hi jonathan i will be in the stag and hounds at seven forty see you soon andy Salida en pantalla | hi john it's tam i will be into stagecoach after four to meet you soon amy Entrada: accept_word Salida: hi | john it's tam i will be into stagecoach after four to meet you soon amy Entrada: 3 * accept_char Salida: hi jo | hn it's tam i will be into stagecoach after four to meet you soon amy Entrada: n Salida: hi jon | athan i will be into stagecoach after four to meet you soon amy Entrada: 4 * accept_word 3 * accept_char Salida: hi jonathan i will be in | to stagecoach after four to meet you soon amy Entrada: space Salida: hi jonathan i will be in | the stadium from after four to meet you soon amy Entrada: accept_word 4 * accept_char Salida: hi jonathan i will be in the sta | dium from after four to meet you soon amy Entrada: g Salida: hi jonathan i will be in the stag | ecoach after four to meet you soon amy Entrada: " Salida: hi jonathan i will be in the stag | and hounds after four to meet you soon amy Entrada: 3 * accept_word 2 * accept_char Salida: hi jonathan i will be in the stag and hounds a | fter four to meet you soon amy Figura 3 Entrada: t Salida: hi jonathan i will be in the stag and hounds at | seven forty see you soon amy Entrada: 6 * accept_word 2 * accept_cbar Salida: hi jonathan i will be in the stag and hounds at seven forty see you soon a | my Entrada: n Salida: hi jonathan i will be in the stag and hounds at seven fourty see you soon a | ndy Entrada: accept_utterance Salida en pantalla: | HEY john it's tam i will be into stagecoach after four to SOON amy Figura 4
Independent claims17
548 paragraphs in 4 sections, as filed
A large-scale system, independent of the user and independent of the device for converting the vocal message to text.
BACKGROUND OF THE INVENTION
1. Field of the Invention
The invention relates to a large-scale voice messaging system, independent of the user and independent of the device, which converts unstructured voice messages into text for visual display on a screen. It should be noted initially the operational challenges faced by a large-scale, voice-independent voice messaging system that can convert unstructured voice messages into text. First, 'large-scale' means that the system must be capable of scaling to very large numbers, for example, 500,000 or more subscribers (usually these are subscribers to a mobile operator) and yet, allow effective and fast processing times, being a message, in general, only useful if it is received within 2 to 5 minutes from when it is left. This is a much stricter requirement than most of the implementations of the ASR recognition system. Second, the expression 'independent of usu user ' means that there is absolutely no need for a user to enable the system to recognize their expression or voice models (as opposed to conventional voice dictation systems). Third, the expression 'device independent' means that the system is not required to receive inputs from a particular input device; Some prior art systems require entry from, by way of example, a touch tone telephone. Fourth, the term 'unstructured' means that the messages have no predefined structure, unlike the response to vocal requests. Fifth, the expression 'vocal messages' is
It refers to a very specific and quite narrow field of applications that poses different challenges to those who have to face numerous conventional automated voice recognition (ASR) systems. As an example, voicemail messages, for a mobile phone, usually include hesitations, 'ers' and 'ums'. A conventional ASR method would have to faithfully convert all oral expressions, even meaningless sounds. The set of verbal or neat transcription characterizes the method of the majority of participants in the field of automatic ASR voice recognition. However, in reality, it is not suitable, at all, for the domain of voice messaging. In the domain of vocal messaging, the creative challenge is not an exact or neat transcription at all, but instead captures the meaning in the most useful way for the intended recipient.
Only through a satisfactory approach to these five requirements is it possible to have a correct implementation.
two. Description of the prior art
The conversion from voice to text (STT) uses automatic voice recognition (ASR) and has been applied, so far, mainly to dictation and command tasks. The use of ASR technology to convert voicemail to text is a new application with several features that are task specific. Reference can be made to WO 2004/095821, which discloses a voicemail system, through Spinvox Limited, which allows voicemail, for a mobile phone, to be converted to SMS text and sent to mobile phone. The management of voicemail in the form of text is an attractive option. It is usually faster to read than to listen to messages and, once in the form of text, voicemail messages can be memorized and searched as easily as an email or SMS short message text. In one embodiment, subscribers to the SpinVox service divert their voicemail to a dedicated SpinVox phone number. Caller subscribers leave voicemail messages as usual for the subscriber. Next, SpinVox converts voice messages to text, with the aim of capturing the full meaning as well as the stylistic and idiomatic elements of the message, but without necessarily converting it word by word. The conversion is done with an important level of input by human operators. The text is then sent to the subscriber as a text of short SMS or email messages. Consequently, subscribers can manage voicemail as easily and quickly as text and email messages and can use client applications to integrate their voicemail - now in the form of archivable and searchable text - with their Other messages
The problem with transcription systems that are significantly based on human operators, however, is that they can be expensive and difficult to establish on a larger scale - eg, to a user base of 500,000 or more subscribers. Consequently, it is not feasible for the main mobile or cellular telephone operators to offer them to their subscriber base because, for the required rapid response times, it is simply too expensive to have human operators listening and transcribing the integrity of each message; in this way, the cost per message transcribed would be prohibitively high. Therefore, the fundamental technical problem is to design a system based on IT information technology that allows the human transcription operator to act with maximum efficiency.
In WO 2004/095821, some degree of ASR front end processing combined with human operators was considered: it was essentially a hybrid system; The present invention develops this inventive idea and defines specific tasks that the IT information technology system can do, which greatly increases the efficiency of the entire system.
Hybrid systems are known in other contexts, but the conventional method for voice conversion is to completely eliminate the human element; This is the operational challenge for experts in ASR techniques, particularly STT techniques. Therefore, we will now consider some of the technical background for STT.
The basic technology of voice to text conversion (STT) is classification. The classification aims to determine to which 'class' some given data belong. The maximum probability estimate (MLE), like numerous statistical tools, uses an underlying model of the data generation process - either the act of throwing a coin or the human voice generation system. The parameters of the underlying model are estimated in order to maximize the probability that the model will generate the data. Classification decisions are then made by comparing the characteristics obtained from the test data with the model parameters obtained from the training data for each class. The test data is then classified as belonging to the class with the best match. The probability function describes how the probability of observing the data varies with the parameters of the model. The maximum probability can be found from the points of investment in the probability function if the function and its derivatives are available or can be estimated. Methods for estimating maximum probability include simple gradient descent as well as faster Gauss-Newton methods. However, if the probability function and its derivatives are not available, algorithms based on the Expectation - Maximization (MS) principles can be used that, starting from an initial estimate, converge to a local maximum of the probability function of the data observed.
In the case of STT, a supervised classification is used in which classes are defined by training data more commonly as three-phase units, which means a particular phoneme spoken in the context of the preceding and following phoneme. (The unsupervised classification, where classes are deducted by the classifier, can be considered as a grouping of data). The classification in STT is required not only to determine to which three-phase class each sound belongs in the voice signal but, it is very important to know which sequence of triphos is the most likely. The latter is usually achieved by creating voice models with a hidden Markov (HMM) model that represents the way in which voice characteristics vary over time. The parameters of the HMM model can be found using the Baum-Welch algorithm, which is a form of MS.
The sorting task, led by the SpinVox system, can be declared in a simplified form such as: “Of all the possible text strings, which could be used to represent the message, which string is the most likely given the voice mail signal from Registered voice and language properties used in voicemail? ”. It is immediately evident that this is a problem of classification of large dimensions and complexity.
Automatic voice recognition (ASR) engines have been under development for more than twenty years in research laboratories worldwide. In recent years, applications to request continuous voice, ASR of wide vocabulary have included systems of dictation and automation of call centers of which they are important
examples "Naturally Speaking" (Nuance) and "How May I Help You" (AT&T). It became clear that satisfactory development
Voice-based systems depend, to a large extent, on the design of the system as is the case with ASR performance and possibly due to this factor, ASR-based systems have not yet been used by most telecommunications users and IT information technology.
ASR engines have three main elements. 1. Feature extraction is performed on the input voice signal approximately every 20 ms to extract a representation of the voice that is compact and as free as possible from interference including phase distortion and variations of the telephone set. Cepstral coefficients of the Mel-frequency algorithm are usually chosen and it is known that linear transformations can be performed on the coefficients before recognition in order to improve their ability to discriminate between the various sounds of the voice. 2. ASR engines employ a set of models, which are usually based on three-phase units, which represent all the various voice sounds and their preceding and following transitions. The parameters of these models are learned by the system prior to their development using examples of adequate voice training voice. The training procedure estimates the probability of occurrence of each sound, the probability of all possible transitions and a set of grammar rules that restrict the sequence of the word and the sentence structure of the ASR output. 3. ASR engines use a model classifier to determine the most likely text given the voice signal at the input. Classifiers of the Hidden Markov model are often preferred since they can classify a sequence of sounds regardless of the speed of speaking and have a very suitable structure for creating voice models.
An ASR engine provides, at the output, the most likely text in the sense that the match between the characteristics of the input voice and the corresponding models is optimized. In addition, however, ASR must also take into account the probability of occurrence of the recognizer's output text in the target language. As a simple example, the English text “see you at the cinema at eight” is a much more likely text than “see you at the cinema add eight”, although the analysis of the voice waveform would most likely detect "Add" that "at" in the use of
ordinary english The study of the occurrence statistics of language elements is referred to as a modeling of the language. It is common in ASR to use acoustic modeling, which refers to the analysis of the voice waveform, as well as language modeling to significantly improve the recognition performance.
The simplest language model is a unigram model that contains the frequency of occurrence of each word in the vocabulary. This model would be constructed by analyzing large texts to estimate the probability of occurrence of each word. A more sophisticated modeling uses n-gram models that contain the frequency of occurrences of strings of n elements in length. It is common to use n = 2 (bigrama) on = 3 (trigram). These language models are of much greater cost in the sense of computer calculus but are capable of calculating the language use much more specifically than the unigram models. As an example, bigramas word models are capable of indicating a high probability that 'degrees' will be followed by 'centigrade' or 'Fahrenheit' and a low probability that they will be followed by 'centipede' or 'foreign'. Language modeling research is ongoing worldwide. The issues include the improvement of the implicit quality of the models, the introduction of syntactic structural limitations in the models and the development of computer-efficient ways to adapt the language models to different languages and accents.
Continuous voice ASR systems independent of the broadest vocabulary of the speaker claim recognition rates greater than 95%, which means less than one word error in twenty. However, this error rate is too high to gain the confidence of the user necessary for a large-scale use of technology. In addition, the performance of the ASR greatly decreases when the voice contains noise or if the characteristics of the voice do not adapt adequately with the characteristics of the data used to train the recognizer models. A specialized or colloquial vocabulary is also not well recognized without additional training.
To build and develop satisfactory ASR-based voice systems, a specific optimization of the technology for the application and the added reliability and robustness obtained at the system level are clearly needed.
To date, no one has thoroughly investigated the practical design requirements for a large-scale, user-independent, hybrid voice messaging system that can convert unstructured voice messages into text. The key applications are the conversion of voicemail sent to a mobile phone into text and email; Other applications, where a user wishes to communicate a message orally instead of entering it through a keyboard (of any format) are also possible, such as instant messaging, where a user orally communicates a response that is captured as part of a message. called IM thread; orally communicate a text, where a user communicates a message that is intended to be sent as a text message, such as a source communication or a response to a voice message or text or some other communication; or, the system called 'speak-a-blog', where a user communicates the words he wishes to appear on a blog and those words are then converted to text and added to the blog. Actually, where there is a requirement, or potential advantage to be obtained by allowing a user to communicate a message orally instead of having to directly enter that message as text and have that message converted to text and appear on the screen, in that case , the user-independent, large-scale hybrid voice messaging systems of the class described in this technical specification can be used.
SUMMARY OF THE INVENTION
In a first aspect of the inventive idea, a voice messaging system, independent of the user and independent of the device, is disclosed on a large scale in accordance with the provisions of claim 1. In a second aspect of the inventive idea, a method for providing voice messaging using a voice messaging system, independent of the user and independent of the device, on a large scale, in accordance with the provisions of claim 17 is disclosed.
In one embodiment, a large-scale vocal messaging system, independent of the user and independent of the device, is converted that converts unstructured vocal messages into text for visual display on a screen; the system comprising (i) subsystems implemented by computer and also (ii) a network connection to human operators that provide transcription and quality control; the system being adapted to optimize the efficiency of human operators, also comprising: 3 basic subsystems, namely (i) a pre-processing front end that determines an appropriate conversion strategy; (ii) one or more conversion resources and (iii) a quality control subsystem.
Other embodiments are provided in Annex III. The invention is a contribution to the field of designing a voice messaging system, independent of the user and independent of the device, on a large scale, which converts unstructured vocal messages into text for visual display on a screen. As explained above, this field presents numerous different operational challenges for the system designer compared to other areas in which the ASR system has been previously developed.
BRIEF DESCRIPTION OF THE DRAWINGS
The invention will be described with reference to the accompanying drawings, in which Figures 1 and 2 are schematic views of a voice messaging system, independent of the user and independent of the device, on a large scale, which converts unstructured vocal messages into text for visual presentation on a screen, as defined by this invention. Figures 3 and 4 are embodiments, by way of example, of how the system presents the possible choices of words and phrases to a human operator for acceptance or rejection.
DETAILED DESCRIPTION OF THE INVENTION
The designers of the SpinVox system had to face numerous operational challenges:
Automatic voice and language recognition models
First and foremost, it was clear to designers that established ASR technology, by itself, was not enough to provide a reliable STT for voicemail (and other user-independent voice messaging applications, to large scale). ASR relies on the assumptions established from theoretical models of voice and language including, by way of example, language models that contain probabilities of previous words and grammar rules. Numerous, if not all, these assumptions and rules are not generally valid for voicemail voice. Factors found in the voicemail STT application, which are beyond the capabilities of standard ASR technology include:
<dl><dt>-</dt><dd> voice quality is subject to environmental noise, variation of coding-decoding and telephone apparatus, network interference, including noise and vocal fading; </dd></dl>
<dl><dt>-</dt><dd>users do not know that they are speaking an ASR system and feel comfortable leaving a message that uses natural and sometimes poorly structured language; </dd></dl>
<dl><dt>-</dt><dd> the language itself and accents used in voicemail are not restricted or predictable; </dd></dl>
<dl><dt>-</dt><dd> Vocabulary variations occur quickly even within the same language so that, by way of example, language statistics may vary due to important facts about current affairs. </dd></dl>
IT Information Infrastructure
The design of the IT infrastructure to maintain the availability and quality of the SpinVox service establishes demands for accuracy in computing power, network and storage bandwidth and server availability. The load on the SpinVox system is subject to unpredictable highs as well as more predictable cyclic variations.
Non-convertible messages
It is expected that a fraction of the messages will be non-convertible. These could be empty messages, such as 'slam-downs', messages in an unsupported language or calls unintentionally dialed.
Quality assessment
The quality assessment at each stage of the SpinVox system is, in itself, an operational challenge. Signal processing provides numerous analysis techniques that can be applied to the vocal signal, ranging from direct SNR measurement to more sophisticated techniques that include explicit detection of common interference. However, direct measurements, such as these, are not significant in themselves, but need to be evaluated in terms of their impact on the subsequent conversion process. Similarly, the confidence of ASR can be measured in terms of the probability of exit of alternative recognition hypotheses but, as before, it is important to measure the quality in terms of impact on the conversion of global text and the complexity of the quality control necessary to achieve it. .
User experience and human factors
The value of the system for customers is influenced, to a large extent, by the level of success with which human factors are supported by design. Users will quickly lose their confidence in the system if they receive distorted messages or find that the system is not transparent or easy to use.
The previous difficulties have been overcome in the design of the SpinVox system as follows:
System design. A simplified block diagram of the SpinVox system design showing the main functional units is shown in Figure 1. The basic core consists of the ASR 1 engine. The SpinVox system makes a clear distinction between ASR and a complete STT conversion. ASR 1 is a subsystem that generates 'gross' text, given the voice signal at the input. This is a key element used for STT conversion, but it is only one of several important subsystems that are needed to achieve a reliable STT conversion. The front-end pre-processing subsystem 2 can perform a broad classification of the voice signal that can be used to determine the conversion strategy based on the choice of a combination of ASR engine, model set and processing improvement the voice. The quality control subsystem 3 measures the quality of the input voice and the confidence in the ASR, from which information a quality control strategy can be determined. The quality control subsystem 4 operates at the output of ASR. Its purpose is to generate semantically correct, meaningful and idiomatic text, to represent the message within the restrictions of the text format. Knowledge of the context, including the caller ID, the specific language models of the recipient and the party making the call, based on time, can be used to improve the quality of the conversion, to a large extent, compared to gross ASR output. The converted text is provided, finally, from the post-conversion language processing subsystem 5 to an SMS text and email gateway.
Main features. The main features of the method adopted by SpinVox are:
<dl><dt>-</dt><dd>Significant Message Conversion </dd></dl>
Text conversion captures the message, its meaning, style and languages, but it is not necessarily a word-by-word conversion of voicemail.
- Round trip time The round trip time for the conversion of a message is guaranteed.
<dl><dt>-</dt><dd>Reliability </dd></dl>
The system can never send a distorted text message. Subscribers are notified regarding non-convertible messages, which can be heard in the conventional manner.
<dl><dt>-</dt><dd>Standard language </dd></dl>
The messages are sent in standard language and are not 'textured'.
<dl><dt>-</dt><dd>Wide availability </dd></dl>
The system works entirely in the infrastructure and does not make any operational requirements on the telephone or network other than call forwarding.
<dl><dt>-</dt><dd> Adaptive operation </dd></dl>
The system can optimize its performance using built-in quality control strategies, driven by the knowledge gained over time regarding the modeling of voicemail languages, in general, as well as the modeling of specific languages of the calling user . In addition, the system can choose from several possible voice to text conversion strategies, based on the characteristics of the voicemail message. Voicemail message data and corresponding text conversions are continuously analyzed in order to update and adapt the SpinVox STT system.
<dl><dt>-</dt><dd>Quality Supervision </dd></dl>
The quality of the voice to text conversion can be monitored at each stage and thus, quality control can be carried out efficiently, by human agents or by automatic agents.
<dl><dt>-</dt><dd>Language processing </dd></dl>
Post-conversion language processing 5 can be performed to improve the quality of the text of the converted message, eliminate obvious redundancies and validate the message elements such as the greeting structures that are usually used.
<dl><dt>-</dt><dd> ASR of modern technology </dd></dl>
Commercial ASR engines can be used in order to gain a competitive advantage of the most modern ASR technology. Different ASR engines can be requested to manage different messages or even different parts of the same message (with decision unit 2 decided as to which engine to use). Human operators, by themselves, could also be considered as an operational instance of an ASR engine, which is suitable for some tasks, but not for others.
<dl><dt>-</dt><dd>Stable and safe </dd></dl>
The service is performed on secure and very stable Unix servers and can be adapted to the demand for several languages since different maximum experience in time zones are dispensed through each 24-hour period.
Quality control The SpinVox system has developed a detailed knowledge of the expectations and wishes of the users regarding telephony-based messaging systems. They have identified a zero tolerance of users for the conversion of voice to text without meaning; in particular when there is obvious evidence that errors have been introduced by a machine other than by human error. The quality control of the converted text is therefore of key importance. Three alternative quality strategies can be used; decision unit 2 selects the optimum. (i) messages for which the confidence in the conversion of ASR is sufficiently high can be checked automatically by the quality assessment subsystem 3 for its conformity with the quality standards.
(ii) Messages for which the confidence in the conversion of ASR is not high enough, can be routed to a human agent 4 for verification and, if necessary, its correction. (iii) The messages for which the confidence in the conversion of ASR is very low are marked as non-convertible and the user is informed of the receipt of a non-convertible message. Non-convertible messages can be heard by the user, if desired, using a single key press. The result of these strategies is that the SpinVox system is designed so that a conversion failure is favored by generating a conversion that contains errors. User confidence in the system is therefore protected. SpinVox statistics indicate that a significant percentage of voicemails are converted successfully.
One of the important tools used by SpinVox to improve the quality of the converted messages is the knowledge of the language (common phrases, common greetings and transmission closures, etc.) used in voicemail messages. From the data accumulated over time, specific statistical language models can be developed for the voice of the mail and then used to guide the STT conversion process. This greatly improves the accuracy of the conversion for non-standard language constructions.
The most obvious feature of SpinVox is that it provides a service that many users did not realize they needed, but soon find that they cannot perform any management without such a service. This is the first real-time system that provides a voice-to-text conversation of voicemails. Its impact for network operators is an increase in network traffic, improved call continuity, both for voice and data. The objective success of the SpinVox system has been achieved by adopting a design method that is oriented first by quality of service and secondly by technology. The system design is based on a detailed knowledge of the expectations of the users of the service and, more importantly from an engineering perspective, the strengths and weaknesses of ASR technology. Exploiting the strengths of ASR and remedying the weaknesses through rigorous quality control, SpinVox is an effective development that meets the practical design requirements for an unstructured, user-independent, and large-scale hybrid voice messaging system.
SpinVox has demonstrated its success in providing a service based on voice processing by concentrating its conversion technology on a very specific objective application, which is the conversion of voicemail. The indication is that the design of the system that is objective for a very well defined application is a more productive method than the search, seemingly unlimited, for improvements, increasingly minor, in gross performance measures of, by way of example , the ASR engines. This method opens up the possibility of new areas of application in which technology components and knowledge of SpinVox's privately owned systems design could be developed.
SpinVox has developed an important technological knowledge of its own as a systems architect for voice-based applications with its own experience that covers the fields of voice recognition, telecommunications applications, cellular networks and human factors. The opportunities for growth development, in advanced messaging technologies, are likely to be aimed at allowing the integration of voice and text messaging, thus facilitating the operations of search, management and archiving of voice mails with all the Same advantages previously enjoyed by email and SMS text messaging, including operational simplicity and automatic documentation. These developments are in parallel with the convergence of voice and data in other telecommunications systems.
The SpinVox system is moving from what, from the outside, is a problem independent of the person speaking with respect to a problem dependent on the speaker, which is a great discernment in the performance of voice work in telephony. The reason is that the fact that calls, messaging and other communication are driven by community use is being used; As an example, 80% of voice mails from only 7 to 8 people. SMS messages only from 5 to 6 people. IM instant messaging only 2 to 3 people. SpinVox uses the historical record of 'peer call' to perform several operations:
<dl><dt>1. </dt><dd>Develop a profile of what a given caller says each time he calls - a speaker model depending on the user who speaks - how the caller speaks (intonation, etc.); </dd></dl>
<dl><dt>2. </dt><dd>Establish a language model of what the caller says to someone - a language model dependent on the user who speaks - of what the caller says (words, grammar, phrases, etc.); </dd></dl>
<dl><dt>3. </dt><dd>In sections 1 and 2, we are really building a language model of how A speaks to B. This is a more refined model than the one based on how A speaks in general. It is atypical for the type of messaging (that is, how it speaks in voicemail) and it is also atypical of how it speaks to B (eg, the way a person</dd></dl>
communicating a message to his mother is very different in intonation / grammar / phrases / accent / etc, than when communicating to his wife).
<dl><dt>4. </dt><dd>SpinVox is building models of speaker-receiver pairs that range from the independence of the language / user that speaks, in general, to which it depends without any user input or training; </dd></dl>
<dl><dt>5. </dt><dd>SpinVox has the ability to use the language of both parties with each other (eg, how I can make a redial and leave a message) to further refinement of the relevant words (eg, dictionary), grammar / phrase, etc. </dd></dl>
Further details on these aspects of the SpinVox voice message conversion system are provided in annex I below.
Annex I SpinVox - Voice Message Conversion System
The SpinVox Voice Message Conversion System (VMCS) focuses on a single objective: the conversion of spoken messages into their meaningful text equivalent. This objective is unique as are the advanced methods and technologies used here.
Concept
A new method of converting vocal messages into text using multi-stage automatic recognition techniques and quality control processes and quality assurance techniques with human assistance. Automated and human elements interact directly with each other to generate real-time / direct feedback that is essential for the system to always be able to learn from direct data to remain in operational tuning and to provide stable quality. It is also designed to take advantage of the inherent limits of AI (ASR) and greatly improve accuracy through the use of contextual links, human guidance and language data directly from the Internet.
Problematic issue
Traditional methods for voice conversion have been designed, to a large extent, at the recognizer level and with the creation of high-quality automatic voice recognition in laboratory conditions, where the inputs are highly controlled and guarantee a high level of accuracy
The problem is that, in the real world, voice recognition has numerous other elements to overcome with:
<dl><dt>-</dt><dd> Random calling users - anyone can use it; </dd></dl>
<dl><dt>-</dt><dd> Noisy input - background noise and poor speaker quality; </dd></dl>
<dl><dt>-</dt><dd>Poor and variable transmission quality with loss of understanding and defective mobile phone connections; </dd></dl>
<dl><dt>-</dt><dd> Grammatically incorrect voice, idioms or very localized expressions; </dd></dl>
<dl><dt>-</dt><dd> Contextual sensitive grammar or implicit meaning from a united context between the creator of the message and its recipient; </dd></dl>
<dl><dt>-</dt><dd> Context changes within a message - context limits - that invalidate the use of normal grammar rules </dd></dl>
to cite only a few and all of them constantly varying in time, so the real origin input is not a problem defined in time, but a problem in constant evolution.
Solution
The key is to correctly define the problem: conversion of vocal messages into their equivalent in meaningful text.
This does not mean a perfect neat transcription, rather than the most important elements of the message presented in an easily understandable way. Accuracy measures are quantitative and qualitative since the ultimate rating is a quality precision of the user's voice where the VMCS of SpinVox achieves a constant rating of 97%.
There are two main parts:
<dl><dt>- </dt><dd>Use of a constant direct feedback mechanism for the learning system, oriented to the human operator. </dd></dl>
<dl><dt>- </dt><dd>Use of contextual information to better define each conversion problem. </dd></dl>
The use of contextual information helps the system to better estimate the probability that something included in a message given the characteristics of:
<dl><dt>-</dt><dd> Type </dd></dl>
<dl><dt>-</dt><dd> Length </dd></dl>
<dl><dt>-</dt><dd> Time of the day </dd></dl>
<dl><dt>-</dt><dd> Geography </dd></dl>
<dl><dt>-</dt><dd> Context of the calling user - caller and recipient (historical record of par-call) </dd></dl>
<dl><dt>-</dt><dd> Recent operational events </dd></dl>
<dl><dt>-</dt><dd> Etc. </dd></dl>
and the structure of known language that is most likely to occur in some types of messages - natural language, as described below.
Natural language
When vocal messages and spoken text messages are analyzed, regular models occur in what the user says, how he says it and in what order - natural language. This clearly varies by context or type of message, so ordering a pizza would be different.
As an example, in voicemail, how users greet each other can be well defined with 35 or a similar number of the most common expressions - "Hello, it's me", "Hello, here Daniel", "What's up, what? What are you doing? ”,“ How are you partner? ”,“ Correct? ”, Etc. and similarly, the farewell expression can be well defined by common expressions - "Agree, goodbye", "Greetings", "Greetings colleague", "Thanks now, goodbye, greetings", etc ...
Obviously, different parts of a spoken message have implicit meaning and therefore, with the use of this key we can improve the accuracy of recognition using this context to select the most probable classification of what was actually said.
Its construction in related, statistically broad models, is what is defined as our Natural Language Model, one for each language, including dialects within any language.
Context vectors
Natural language is usually regulated by its context, so when the body of a message is converted, the context of what has been said can be used to better estimate what was actually said - eg, a call to a reservation line it will contain, more likely, expressions related to the request and some specific company, product and price names in front of a call to a home phone number, where friendly expressions regarding greetings, 'how are you', 'call me later', etc., are much more likely.
Within vocal messages, we can use context vectors to better estimate the likely content and established natural language that applies:
<dl><dt>-</dt><dd> CLI (or any part identifier) is a very powerful context vector </dd></dl>
<dl><dt>- </dt><dd>It can probably refer to the geographical area of the number and the language / dialect and regional real names of the most likely use. </dd></dl>
<dl><dt>- </dt><dd>It can be indicated if the number is a known commercial number and therefore a better predicted type of message - eg, calls from 0870 numbers are commercial and therefore, with a high possibility that it is a business message while that the 07 range is a personal mobile phone, so the time of day will motivate messages of a more probable type among business, personal, social or other content. </dd></dl>
<dl><dt>-</dt><dd> It allows you to get your correct number better if you communicate within the message. </dd></dl>
<dl><dt>-</dt><dd> It is a key from which you can establish a historical record and known dictionary / grammar - eg, it always says in English 'dat's wicked man' in an accent 'on the street'. </dd></dl>
<dl><dt>-</dt><dd> You can build a user-dependent recognition system that speaks - that is, we can adjust the ASR as a particular calling user and obtain a much higher recognition accuracy, its own vocabulary, grammar, phraseology, vocabulary and general natural language. </dd></dl>
<dl><dt>-</dt><dd> Historical peer call log - deeper use of CLI (or any part identifier) </dd></dl>
<dl><dt>- </dt><dd>You can train the system much more accurately for a historical record of messages of the peer call type. </dd></dl>
<dl><dt>- </dt><dd>You can train for the voice of part A (calling user) ignoring part B (recipient). </dd></dl>
<dl><dt>- </dt><dd>You can train for the content and language area of part A that you use with part B. </dd></dl>
<dl><dt>- </dt><dd>It can train multiple relationships of parts A and B and develop the system for higher accuracy and speed. </dd></dl>
<dl><dt>-</dt><dd> Time of day, day of the week </dd></dl>
<dl><dt>-</dt><dd> Voicemail traffic rates, average message length and type of content vary with the time of day in each language market, from very business messages during peak demand hours (8 am to 6 pm ) to more personal (7 to 10 in the afternoon) to very social (11 at night to 1 in the morning) to very functional (2 to 6 in the morning). This situation also varies by the day of the week, with Wednesday being the busiest day and containing the highest levels of business messages, but on Saturday and Sunday they have a very different profile of message type (largely personal conversation messages ) that need to be treated differently. </dd></dl>
<dl><dt>-</dt><dd> International numbers </dd></dl>
<dl><dt>-</dt><dd> By analyzing the country code (eg, 44, 33, 39, 52, 01) we can better determine the language and dialect. -Customer data available -Customer name, address and possibly your workplace.</dd></dl>
Implicit context between parts A and B Likewise, there are many other very important clues that can help us better estimate the likely content of a message, particularly those that relate to who the two parties are, what is the likely purpose of the message and where or from where it is being called.
In the conversion of voicemails to spoken text and text, we know that the fact of having the number of the user who calls-allows us to better estimate any number left inside the message.
<dl><dt>-</dt><dd> Establish the historical record of words, expressions, phrases, etc., known between the two parties. </dd></dl>
<dl><dt>-</dt><dd> Probable language (eg, a call from +33 to +33 will probably be in French, but calls from +33 to +44 may have a 50% chance of being in French). </dd></dl>
<dl><dt>-</dt><dd> Names and their correct spelling. </dd></dl>
If you know the historical record of the calls / messages of part A and its historical record of messages for part B, you can establish a profile dependent on the user who speaks and obtain significant improvements from your recognizer and grammar.
Conversion quality
To solve this problem, the definition of the actual required result is essential since it makes a big difference in its method of solving how to convert voice messages (voice mails, verbal SMS messages, instant messaging, etc.) to text and how to apply, optimally, the conversion resource available to you.
When someone leaves us a voice message, the purpose is a message, and not a formal written communication element, so a less accurate conversion will be tolerated as long as the message's meaning is correctly transmitted.
In addition, there is an asymmetry, so that the depositor of the message is not comparing with what was said with the converted text. With the context of the caller, the recipient is reading the converted output with the aim of finding out what the message is about, so the requirement is an excellent extraction of the message for conversion and not a neat conversion (word for word , expression by expression). In reality, on the contrary, a neat conversion, unless it is well dictated, is usually perceived as a low quality message since it contains abundant unwanted and non-elegant elements of the language of vocal messages (eg, uhmms, ahhs, repetitions , word spelling, etc.).
Therefore, quality in this context refers to the extraction of the important elements of a message.
<dl><dt>-</dt><dd> Smart conversion </dd></dl>
In its simplest form, there are three key elements that provide maximum meaning and therefore are essential to achieve a good message quality:
<dl><dt>1. </dt><dd>Whose is it - with great courage to understand the meaning from this context. </dd></dl>
<dl><dt>2. </dt><dd>What is the purpose of the message - eg, call me urgently, late completion, change plans / times, call me at this number, just say hello, etc. </dd></dl>
<dl><dt>3. </dt><dd>Any specific facts, the most common being: </dd></dl>
<dl><dt>to. </dt><dd> Name (s </dd></dl>
<dl><dt>b. </dt><dd> Numbers, telephone numbers </dd></dl>
<dl><dt>c. </dt><dd> Hour </dd></dl>
<dl><dt>d. </dt><dd> Direction </dd></dl>
Other information in the message is largely in support of the transmission of these key elements and usually helps to provide a better context for these key elements.
Variation of the sensitivity of the quality within the message
What is also very important to understand is that we need to recognize that each main part of any message has a different function in the delivery of the message and therefore, we can attribute another quality dimension to each one in order to achieve it during the conversion.
Messages can be broken down into:
<dl><dt>- </dt><dd> Greetings (top) </dd></dl>
<dl><dt>- </dt><dd> Message (body) </dd></dl>
<dl><dt>- </dt><dd> Farewell (final) </dd></dl>
The percentage of messages that contain any body is obviously a function of the length of the deposited message, so we know that short messages (eg, less than 7 seconds) usually contain only a greeting and a farewell. First of all, the probability of a significant message body increases exponentially. This fact also helps us to better estimate the likely conversion strategy that we should use. Greetings and farewells
The way someone greets you can be classified into about 50 frequently recognized greetings (eg, "hello everyone", "Hello, it's me", "Hello, this is a call to X from Y", "Hello, I just call you for…. ”, etc.). Similarly, the 'farewell' element of a message can be classified in an order of similar magnitude of farewells of common recognition (eg, “Thank you very much,” “See you,” “Ta,” “Goodbye,” “I greet you,” "See you later", etc.).
Two issues dictate our demand for conversion quality:
<dl><dt>1. </dt><dd>Greetings and farewells are for message protocol and usually contain little value for the main message, so our tolerance for low accuracy is high, provided they make sense. </dd></dl>
<dl><dt>2. </dt><dd>We can classify the vast majority of greetings and farewells into some 50 categories of frequent recognition </dd></dl>
each. Therefore, the quality requirement during a greeting, a reception or farewell is much less than with respect to the content in the body of the message, usually the main issue of a message or a key fact - eg, call me at 020 7965 2000.
Message body Of course, the message body has a higher quality requirement, but it is likely that it can often found that it contains regular natural language models that relate to the context and
that, therefore, we can also apply a rating to help us better get the answer correct. A suitable example is: "Hi Dan, I'm John" - Top of the message (or greeting) “You can call me back at 0207965200 when you can” - Message body "Thank you very much partner, greetings, goodbye" - End of the message (or farewell). In this case, the body of the message is a well structured element of the Voicemail Language that you have learned.
SpinVox conversion system. You can then break down the message and make a correct assignment to the result of the message body. The elements that apply in this case are: -Parts A and B known
- The phone number is John CLI or seen before in his calls to others
- Message length - less than 10 seconds, making common expression more likely
<dl><dt>-</dt><dd> Time of day - work schedule - John usually does not leave detailed messages in the work schedule but simply simple and short messages. </dd></dl>
SpinVox vocal message conversion system Having correctly indicated our problem and identified some very important characteristics of the voice and how it relates to the text equivalent, the SpinVox system (see Figure 2) was designed to obtain the maximum competitive advantage of these characteristics:
SpinVox voice message conversion system
This diagram illustrates the three key stages that allow us to optimize our ability to correctly convert vocal messages (voicemail, spoken SMS message, instant messages, voice clips, etc.) to text. A key concept is that the system uses the term Agent for any conversion resource, be it human,
either computer / machine based. Preprocessing It can be broken down into two characteristics:
<dl><dt>1. </dt><dd>Optimize the quality of the audio signal for our conversion system by eliminating noise, debugging known defects, normalizing signal / volume energy, eliminating silent / empty sections, etc. </dd></dl>
<dl><dt>2. </dt><dd>Classify the type of message for optimal routing of the message for conversion or not. </dd></dl>
Message type classification is done using a range of 'Detectors':
<dl><dt>-</dt><dd> Language </dd></dl>
<dl><dt>-</dt><dd>Pe, English from the United Kingdom / United States / Australia / New Zealand / South Africa / Canada and then, the types of dialects within those languages (eg, within the United Kingdom - S. East, Cockney, Birmingham, Glasgow, Northern Ireland , etc.). </dd></dl>
<dl><dt>-</dt><dd> It allows us to determine if we support the language. </dd></dl>
<dl><dt>-</dt><dd> It allows you to select which conversion path to use: QC / QA profile, TAT rules (SLA), which ASR stages strategy (engines) to load and what post-processing strategy to apply. </dd></dl>
Methods: -Identification of statistical language
- Previous technique: -various known automatic language identification methods
<dl><dt>-</dt><dd> SpinVox solution: </dd></dl>
<dl><dt>-</dt><dd> basic decision on context: knowledge about registration, location and historical record of calls of the calling user and the receiver </dd></dl>
<dl><dt>-</dt><dd> Identification of the language based on the signal -Problem with the prior art: -High precision methods require voice recognition with great vocabulary or at least phonetic recognition, so it is expensive to obtain and use </dd></dl>
<dl><dt>-</dt><dd> requirement for a reliable and fast method based exclusively on records (labeled with the language but nothing else) </dd></dl>
<dl><dt>-</dt><dd> SpinVox solution: </dd></dl>
<dl><dt>1. </dt><dd>Automatic grouping of voice data for each language (vector quantization) </dd></dl>
<dl><dt>2. </dt><dd>combination of grouping centers </dd></dl>
<dl><dt>3. </dt><dd>use of the statistical sequence grouping model for each language to find a better match </dd></dl>
<dl><dt>4. </dt><dd>establish a relationship model between qualification differences between models and the expected accuracy </dd></dl>
<dl><dt>5. </dt><dd>combination of several versions of 1-4 (based on variable training data, feature extraction methods, etc.) until the desired accuracy is achieved. </dd></dl>
<dl><dt>-</dt><dd> Noise - SNR detector </dd></dl>
<dl><dt>-</dt><dd> If the amount of noise in a message exceeds a certain threshold, then it becomes increasingly difficult to correctly detect the signal of the message and proceed to its conversion. More significant is if the ratio of the signal to noise becomes lower than a certain level, in which case, it will have a high degree of confidence and not being able to convert the message.</dd></dl>
<dl><dt>-</dt><dd> SpinVox users value the fact that when they receive a notification that the message was inconvertible, the source audio is so poor that more than 87% of the time it is called or the text, the person returns directly and continues the 'conversation '. </dd></dl>
<dl><dt>-</dt><dd> Voice Quality Estimator </dd></dl>
<dl><dt>-</dt><dd> If someone's voice quality is probably too low for the use of the conversion system or agent. Or the content that a user would hear for himself - eg, someone calls them with the sound of a happy birthday.</dd></dl>
<dl><dt>-</dt><dd> SpinVox's solution includes: </dd></dl>
<dl><dt>1. </dt><dd>find deletions (voice packets lost during transmission) based primarily </dd></dl>
at zero cross counts
<dl><dt>2. </dt><dd>also estimate noise levels </dd></dl>
<dl><dt>3. </dt><dd>calculate the overall measure of voice quality and use an adaptive threshold to reject messages of the lowest quality. </dd></dl>
<dl><dt>-</dt><dd> Off-hook operation detector ('slam-down') </dd></dl>
<dl><dt>-</dt><dd> Messages where someone called, but left no significant audio content. Normally, short messages with verbal expressions in the background.</dd></dl>
<dl><dt>-</dt><dd> Inadvertent Call Detector </dd></dl>
<dl><dt>-</dt><dd>Normally, a call from the press of a redial button while in someone's pocket and leaving a long rumble message without any significant audio content. </dd></dl>
<dl><dt>-</dt><dd> Standard messages </dd></dl>
<dl><dt>-</dt><dd>Pre-recorded messages, more common in the United States, from an automatic dialing system or service or call notifications. </dd></dl>
<dl><dt>-</dt><dd> Greetings and farewell </dd></dl>
<dl><dt>-</dt><dd> If the message only contains these parts, in that case, we can use a dedicated ASR element to correctly convert these messages. </dd></dl>
<dl><dt>-</dt><dd> Message length and voice density </dd></dl>
<dl><dt>- </dt><dd>The length allows us to initially estimate the probability of the type of message - eg, the short call usually consists simply of 'hello, I am X, call me later please' in front of a long call that will contain something more complex to convert. </dd></dl>
<dl><dt>- </dt><dd>Voice density will allow you to adjust your estimate of the probability that a message length is a good indicator of the type - low density and short message is likely to be simply 'hello, I'm X, call me later, please', but a Short, high-density message will bias the user in need of a higher level of conversion resources since the complexity of the message will be higher. </dd></dl>
Obviously, pre-processing allows us in some cases (eg, slan-down (off-hook), inadvertent call, foreign language / not supported) to immediately route the message as classified and send the notification text
correct to the recipient (eg, 'this person called, but left no message'), saving any additional use of
valuable resources of the conversion system.
Automatic voice recognition (ASR)
It is a dynamic process. The optimal use of conversion resources is determined at a message level.
We assume the input from the pre-processing stage in the message classification and from any context vectors and use it to choose the optimal conversion strategy. This means that this stage is using the best ASR technology for the particular task. The reason is that different types of ASR are very suitable for specific tasks (eg, one is excellent for greetings, another for telephone numbers, another for addresses in French).
This stage is designed to use a range of conversion agents, whether ASR or human and only discerns between them depending on how the conversion logic is configured at that time. As you learn the system, this situation is adapted and different strategies, conversion resources and sequences can be used.
This strategy not only applies to the entire message level, but can also be applied within a message.
Top 'n' final message
One strategy is to subdivide the sections of greetings (top), body and farewell (end) of the message, send them to different ASR elements that are optimal for that element of a message. Once this operation has been completed, they are assembled as a single message.
Number routing
Another strategy is to subdivide any simple elements where telephone numbers, coins or the obvious use of numbers communicated in the message. These sections are sent to a specific ASR element
or the agent for optimal conversion and then, reassembled with the rest of the converted messages.
Address Routing
Similarly, the subdivision of elements where an address is being communicated can be sent to a specific ASR or agent element and to an address coordinate validator to ensure that all elements of a converted address are real. As an example, if you cannot detect the street name, but you have a clear zip code, you can complete the most likely street name. The accuracy of finding the street name is improved by a new address processing, but with its street name estimated a priori by readjusting its ASR classification variables to a much more limited set and ascertaining whether or not there is a high coincidence
Real Name Routing
Real names are renamed to obtain an unreliable ASR. Again, concentrating on only this part and applying a more specialized, but more expensive, computer resource, you can estimate the real name much better.
Mail processing
The ASR stage contains its own dictionaries and grammar, but this is not enough to correctly convert the most complex forms with which we speak. ASR is very oriented to the conversion at the level of words and very short sequences of words to estimate the probability of sequences of words and basic grammar (grid techniques and n-grams). One problem is that, from the mathematical point of view, when you expand the set of words trying to estimate the possible combinations beyond 3 or 4, the permutations become so great that your ability to capture the correct one decreases faster than any advantage obtained by expanding the number of sequential words, so it is currently an unreliable strategy.
A very suitable method is to look for broader sentences or sentence structures that occur in natural language. By focusing on the problem from the macro level, you can estimate better solutions for errors or words / sections where the confidence in ASR is low.
However, this method also has its weaknesses. As indicated above, the human voice contains a lot of noise, distortions and due to the intimate relationship between parts A and B it is prone to great contextual limitations. Some things make no sense to someone other than the person who has a broad set of context in which seemingly random phrases or words of unreliable sounds make a lot of sense.
As an example, the English phrase "see you by the tube at Piccadilly opposite the Trocadero and mine's a skinny mocha when you get here" would not make sense to someone unless they knew the possible meaning of 'tube', they had been in London and they would know that Piccadilly has a building called the 'Trocadero' very close and would understand that in nearby Starbucks they serve a drink called 'mocha' and is low in calories, that is, 'skinny'.
Real world bodies - context checking
One solution is to look for a very broad body of English words, real names, phrases, idioms, ordinary expressions to verify that your conversion could have contained these word sequences.
The problem is that in the normal voice there are great possible combinations of the above elements and critically, this lacks no real-world context check. How do you know that the combinations of real names Piccadilly, Trocadero, mocha and skinny are valid, allowing only good conversions of your source audio signals? The only absolute thing is a real-world check and unfortunately, by definition, human beings are the only ones, at this time, able to discern whether or not something has a validity in the real world - we do it after all Fixed program computers and databases are reliable in this regard.
With intelligence at the human level, one can check, with the greatest precision, whether these seemingly disconnected elements have any probable real-world context. However, human beings also lack complete knowledge of everything that a significant percentage of Londoners have that they would not find comfortable knowing if this phrase was likely, or not, given their knowledge of Piccadilly.
One solution is to use the largest bodies on the planet of human knowledge. The billions of pages and databases created by human publishers available on the Internet. A simple query as to whether the sentence, or any element of its conversion, is cited on the Internet provides you with a very qualified real-world test of whether it is something that human beings have probably experienced and registered and so Therefore, it could be real. Therefore, in the previous example, we found that Google, Yahoo !, MSN and other important search engines are able to provide many successes of pages with these elements because we have a much improved confidence that our conversion is correct, in reality .
In addition, using the Internet, we can very often find the correct spelling of phonetic approximations of words, real names and place names that ASR tries with new or unknown words if it is crossed. Currently, this operation is carried out through a high-cost and time-consuming manual programming of ASR dictionaries.
The other very valuable advantage of this solution is that the Internet is a 'live' system that reflects, with great accuracy, the
Current language, which is a dynamic and evolving subject and may vary with a unique news heading, so you are not relying on constantly updating your ASR dictionaries with a limited subset of natural language, but you have access to probably the source more current and of greater magnitude of natural language on the planet.
Example SpinVox converts the following audio signals: Message from an English person Audio: In English the phrase “The cat sat on Sky when Ronaldo scored against Cacá” Converted text: The cat sat on sky when Rownowdo / Ron Al Doh / Ronaldow / Ronahldo score against Caka / Caca / Caker. Problems:
- 'sat on sky' is grammatically incorrect, 'you can't sit on' sky 'in the context of the dictionary.
- Rownowdo / Ron Al Doh / Ronaldow / Ronahldo are possible solutions for an unusual real name.
-Caka / Caca / Caker are guesses of a very unusual real name. Google search for difficult items of this class indicate: -The sky is a brand name - number 1 for 'in the sky'. Therefore, it is a real name for an object,
so 'The cat sat on sky' is possible grammatically and validly.
<dl><dt>-</dt><dd> The first name is more likely Ronaldo simply from the unique spelling checks of </dd></dl>
all versions (in Google "Did you mean: Ronaldo?").
<dl><dt>-</dt><dd> Ronaldo is in close correlation with the 'Ronaldo scored' as being a very famous soccer player and the search provides a large number of exact matches for this phrase. </dd></dl>
<dl><dt>-</dt><dd> The middle name is more likely Cacá, because Cacá has more hits for 'score against Cacá'. </dd></dl>
<dl><dt>- </dt><dd>In addition, we rely on the search for 'soccer Cacá'-football is a word that is derived from the context of' Ronaldo scored 'and we obtain a large number of search results in very close correlation. Given that 'qualified Ronaldo' has already obtained a large number of satisfactory searches, we are more confident that 'Cacá' is the most likely solution.</dd></dl>
<dl><dt>- </dt><dd>In addition, the real-world nature of Google data initiation means that terms that are currently being used, current terms, obtain higher rankings than less current terms, which is essential for getting voice recognition for Work with the current language and its context. </dd></dl>
Queue Manager The queue manager is responsible for: -Determine what should happen to a vocal message at each stage - conversion strategy. -Management of the decision of each automated stage when it requires human assistance
<dl><dt>-</dt><dd> If at any stage of the automated conversion, the confidence intervals or other measures indicate that any part of a message is not good enough, in that case, the queue manager directs it to the correct human agent for assistance. </dd></dl>
<dl><dt>-</dt><dd> Guarantee our service level agreement with any client, ensuring that we convert any message within an agreed time - Round Trip Time (TAT) </dd></dl>
<dl><dt>-</dt><dd> TAT usually has an average value of 3 minutes, 95% within 10 minutes and 98% within 15 minutes. </dd></dl>
<dl><dt>-</dt><dd> Decision making calculating compromise solutions between conversion time and quality. This is one</dd></dl>
function of what the SLA agreement allows, in particular, to address the maximum demands for the use of abnormal language or
of unforeseen traffic and performance to date.
The above is achieved using large state machines that, for any given queue of language, can decide how to best process messages through the system. It interacts with all parties and is the operating core of the VMCS SpinVox.
Quality Control Application
Annex II contains a more complete description of this grid method as used within the SpinVox quality control application.
As illustrated in the diagram in Figure 2 of the voice message conversion system (VMCS), human agents interact with messages in several stages. They do it using the quality control application.
They also use a variant of this tool to randomly inspect messages to ensure that the system is converting messages correctly as one of the problems with AI is that it is unable to make sure that it is really accurate.
A key inventive stage is the use of human agents to 'guide' the conversion of each message. This lies in the
VMCS SpinVox databases, which contain a large body of possible matches, ASR and a human introduction of some words to create predictive typing solutions. In its extreme case, no human agent is required and the conversion is completely automatic.
Question
ASR is only acceptable in word level matches. To convert a meaningful message, phrases, sentences and grammar for vocal messaging is necessary. ASR will provide a measure of statistical confidence for each match at the word level and at the phrase level where available. It is unable to use the rules of natural language or context of use to perform the process of a meaningful and correct conversion.
Which automated systems are good in their spelling and grammar base - coherence.
Which human agents are good in terms of meaning, context, natural language, spoken grammar, ambiguous input management and make sense of them. Human agents tend to be inconsistent with spelling, grammar and speed.
Business issue
The use of human agents costs money, so something can be done to use them only for matters such as that matter and therefore, economic value is essential.
SpinVox VMCS uses the concept of Agent Conversion Relationship (ACR) - the ratio of time the agent actually uses to process a message with respect to the length of the voice message. Something that reduces the ACR ratio and improves the quality of message conversion is a commercial factor such as a 1% reduction in ACR that results in at least a 1% improvement in gross margin. In fact, the sensitivity is even higher when not only the direct cost of the merchandise sold is reduced, but the management overhead and the operational availability of the service and its scalability benefit all of them from the lower number of human agents required. .
Solution
Grid method: use human agent to guide the system to capture the correct set of words, phrases, sentences, messages from a predetermined list of the most likely options.
SpinVox VMCS databases maintain an abundant historical record of message data such as large statistical models (dictionaries and grammar with context vectors that relate to them) that can be extracted in two main ways:
Reticle Method
i. The VMCS language model uses the context (eg, call pair history, language, time of day, etc. - see section of Context Vectors) to capture the most likely conversion (the proposed conversion) to show to the agent.
ii. When the message is played, the agent selects a letter to choose an alternative (it can be simply
the first letters of the correct word) or enter 'accept' to accept the proposed section of text and
Scroll to the next section.
iii. As the types of agents change, the system uses it as input to capture the most likely new conversion and as feedback (learning) so that, on the next occasion, it is more likely to get the correct match at the first time.
<dl><dt>iv. </dt><dd>What would normally require an agent to write a complete message with the appropriate characters (eg, 250), it would only take a few keystrokes to complete and in real time or faster. </dd></dl>
<dl><dt>v. </dt><dd>The agent's exit is now restricted to the correct spelling, grammar and phraseology or rules about them that control the quality and better meaning of the message. </dd></dl>
The above can be presented to the agent in two main ways:
1. Proposed conversion assisted by ASR
In this case, ASR is used first to better predict what text the proposed conversion should be for the agent. It uses what is really in the audio signal of the vocal message to reduce the set of possible conversion options to a minimum, thereby improving the accuracy and speed of the agent.
<dl><dt>to. </dt><dd>ASR can be used for the initial proposed conversion. </dd></dl>
<dl><dt>b. </dt><dd>ASR can then be used continuously when the agent enters selections to further readjust the remaining sections of the proposed conversion. Prior art: human agents correct transcripts with choice of word alternatives.</dd></dl>
Problem with the prior art: -The corrections are still time consuming.
<dl><dt>-</dt><dd> The ASR engine could have made a better decision (later in an oral expression) if the user's correction had been known during decoding. </dd></dl>
two. Fully predictive text typing Similar to that described in section 1 above, but where no ASR is used to select the proposed conversion shown to an agent. This is different for standard predictive text editors since
a specific historical record is trusted (VMCS language models and use of context vectors - eg, call pair history) and works at the phrase level and above. Prior technique: predict the most frequent word (list of alternatives) given the partial human input. Problem with the prior art: -The most frequent word is not usually what the user wants. -Predictions only for one word. In either case, the SpinVox VMCS language models are simply trained by human agents or
by a combination of ASR and human agents. In the extreme case, the system is fully trained and capable of always capturing the correct proposed conversion for the first time and only requires human assistance for quality assurance for random sampling.
and to check if the VMCS is self-regulating correctly. Annex II - Grid method Various observations and assumptions
<dl><dt>1. </dt><dd>Given the breadth of the vocabulary and the variable audio quality, it seems impossible to achieve a fairly high voice recognition accuracy for a fully automatic conversion for more than a small fraction of oral expressions. Reliably detecting this fraction, that is, deciding that no human verification is needed is a long-term research problem, of great interest, but probably not a realistic option in the short term.</dd></dl>
<dl><dt>2. </dt><dd>Although a good operator has a target ACR of 3-4, the most likely average is 6-8. </dd></dl>
<dl><dt>3. </dt><dd>The correction of an oral expression that is already 90% correct takes approximately 1.2 (source: SpinVox 2005 Operational Research). </dd></dl>
<dl><dt>4. </dt><dd>75% of the correction time is spent finding and selecting errors (Wald et al). </dd></dl>
<dl><dt>5. </dt><dd>Word selection lists (alternatives) reduce listening time (Burke 2006). </dd></dl>
<dl><dt>6. </dt><dd>Errors tend to be grouped (Burke 2006). </dd></dl>
<dl><dt>7. </dt><dd>A double speed reproduction maintains intelligibility and users seem to prefer it after a short training (Arons 97). </dd></dl>
<dl><dt>8. </dt><dd>The elimination of pauses and 50% faster reproduction provides a real time factor of 1/3 (Arons 97). </dd></dl>
<dl><dt>9. </dt><dd>According to Bain et al 2005, normal typing has ACR 6.3 which is equal to ACR to edit the ASR output with 70% accuracy. The transcription 'ghost' is mentioned as 'viable' for a live subtitling.</dd></dl>
Focus
The main objective has to be to reduce the agent conversion ratio (ACR) using voice technology to support the agent. This can be achieved in several ways:
<dl><dt>1. </dt><dd>Allow the agent to make the decisions we cannot provide to make mistakes, that is, the overall meaning of the message or individual phrases. The machine can fill in the details.</dd></dl>
<dl><dt>2. </dt><dd>Offer predictions while the agent introduces / edits the oral expression. This could not only save typing time, but also help avoid spelling mistakes.</dd></dl>
<dl><dt>3. </dt><dd>Provide capitalization (simplified) and punctuation automatically, so that the agent does not need to address these issues. </dd></dl>
Call Management Stages
<dl><dt>1. </dt><dd>The agent listens to the message at high speed (eg, 1/2 real time). </dd></dl>
<dl><dt>2. </dt><dd>The agent presses a button to select the category (eg, "redial", "redial only", ... "general"). </dd></dl>
<dl><dt>3. </dt><dd>In some cases, oral expression is accepted immediately. This will happen if the message follows a simple model defined for the message category, the voice quality was good and there is no important part, but easily confused in the message (eg, times in English).</dd></dl>
<dl><dt>4. </dt><dd>The system proposes a converted string, the agent edits while the system continuously updates (and instantly) the proposed oral expression using predictions based on the voice recognition results. </dd></dl>
<dl><dt>5. </dt><dd>The agent presses a key to accept the oral expression as soon as the displayed oral expression is correct. Figure 3 illustrates, by way of example, the call management stage 4. As an example, the agent would need 35 keystrokes to edit an oral expression with 17 words and 78 </dd></dl>
characters:
<dl><dt>-</dt><dd> 15 * <accept_word> (eg, tab) </dd></dl>
<dl><dt>-</dt><dd> 14 * <accept_char> (eg, right arrow) </dd></dl>
<dl><dt>-</dt><dd> 6 * normal input </dd></dl>
<dl><dt>-</dt><dd> 1 * <accept_utterance> (eg, Enter). </dd></dl>
Most of them must be very fast because the same key has to be pressed a few times. Only 6 of they require selecting a normal key. It should be noted that only 6 of the 17 words (35%) were correct in oral expression (utterance)
originally proposed by the system. Implementation Processing stages
<dl><dt>1. </dt><dd>A speech recognition engine (eg, HTK) converts the oral expression voice file into a reticle (that is, a word hypothesis graphic - a directed graphic, a cyclic graphic that represents a large number of possible word sequences) . </dd></dl>
<dl><dt>2. </dt><dd>The grid is reclassified to take into account the specific information of the telephone number (pair) (eg, names, frequent phrases in previous calls, etc.). </dd></dl>
<dl><dt>3. </dt><dd>The reticle is increased to allow a very fast search during the editing phase (eg, the most likely route at the end of the oral expression is calculated for each node and the arc that starts this route is memorized, sub are added</dd></dl>
character trees to each node representing decision points). "Families", that is, several arches that differ
Only in their start and end times are they combined with some limits.
Four. When the agent selects a specific category (stage 2 in “call management stages”), a grammar model and corresponding language are selected for syntactic analysis and dynamic requalification. When the
category is "general", an unrestricted "grammar" is used.
<dl><dt>5. </dt><dd>The highest grade route is selected through the grid that matches the selected category grammar (if appropriate). </dd></dl>
<dl><dt>6. </dt><dd>The result found in this way will be accepted immediately if: </dd></dl>
<dl><dt>to. </dt><dd>The category is not "general." </dd></dl>
<dl><dt>b. </dt><dd>The difference in qualification is the highest qualification on an unrestricted route that is within a given margin. This margin can be used as a parameter to dynamically control the compromise solution between speed and precision.</dd></dl>
<dl><dt>c. </dt><dd>Depending on the grammar used to find the route, the oral expression does not contain crucial parts that are easily confused (eg, times in English). </dd></dl>
<dl><dt>7. </dt><dd>When the user accepts words or characters, the system moves along the selected path through the grid. </dd></dl>
<dl><dt>8. </dt><dd>When characters and words are accepted or entered, their color or print source changes. </dd></dl>
<dl><dt>9. </dt><dd>When the agent types something, the system selects the highest grade route (again, taking into account the current grammar and possibly another, eg, statistical information) that starts with the characters typed. This new route is displayed at this time.</dd></dl>
<dl><dt>10. </dt><dd>When an agent types a word not found in the grid, it is automatically checked by spelling and a correction is offered if appropriate. </dd></dl>
<dl><dt>11. </dt><dd>After the agent presses <accept_utterance>, the text is processed to add capitalization and punctuation, correct spelling errors, replace number words with digits, etc. This operation uses a solid probabilistic analyzer that uses semi-automatically grammars derived from training data.</dd></dl>
Audio signal playback The nodes in the grid contain timing information and therefore, the system can keep a record of the part of the message the agent is editing. The agent can configure how many seconds the system will play. If the agent so wishes, the system reproduces the oral expression (utterance) from the word before the current node.
Refinement Options
Mark important and unimportant parts
i. Depending on the grammar and relevant category, specific parts of the visualized oral expression text that are considered crucial are highlighted while the particularly unimportant parts (eg, greeting phrases) are shaded.
Use of phrase classes for unimportant parts
Parts of the message are displayed as phrase classes instead of individual words. The agent only needs to confirm the class while the choice of the individual class is left to the ASR engine because an error in this area is not considered important. As an example, the class "HEY" could mean "hi, there, hey, find, hello" and the previous example could be visualized as illustrated in Figure 4. In this version, the <accept_word> key applied to a phrase would accept the entire sentence. The typing of a character changes back to the word mode, eg, the phrase class marker is replaced by individual words.
Limit prediction display
The display of erroneous predictions could really confuse the agent and it might be worthwhile to visualize only those (partial) of which the system has relative certainty. As an alternative, trust
relative in several predictions could be color-coded in some way (see “trusted shading” section), eg, some uncertain ones (usually far from the cursor) are printed in a very light gray color while
that the most reliable appear darker and bold.
Segmentation of oral expression
Longer periods of silence are detected and used to break the message into segments. The user interface reflects the segmentation and an extra key is assigned to "<accept_segment>". This allows the confirmation of longer sentences with a single keystroke and also resynchronization if the agent writes a word that does not extend the current path through the reticle.
Cursor maintenance on the left side of the screen
Have a wide area, in the middle part of the screen, that shows the current phrase in large letters. As editing continues, the text scrolls (keeping the cursor in the same position). Only a few words are shown to the left of the cursor. As words are removed from the middle zone, they move to the top (smaller, gray type of print). More phrases are displayed below, again in small and gray characters.
Show phrase alternatives
Always or after a key press, show alternate phrase realizations as a drop-down menu to the right of the cursor and allow selection with arrow keys. This means that the agent does not need to think about the first characters of the correct word that would help with difficult words.
Move cursor with voice
As the message is played, the spoken word is automatically highlighted and the cursor moves to the beginning of the word.
Play the highlighted area
The agent can select a zone (eg, left and right mouse button) and the system maintains segment playback between the markers until the agent is moved.
Transcription 'ghost' for individual words or phrases
The words are highlighted as they are reproduced to the agent and in addition to their typing, the agent can simply say the word to replace the currently highlighted one (and the rest of the sentence). The system dynamically establishes a grammar from the alternative candidates in the grid (words and phrases) and uses ASR to select the correct solution. This is a technically difficult option because ASR needs to be used from within the QC application and models dependent on the proper way of speaking need to be trained and selected at the time of execution.
Accuracy Considerations
Most likely reserve
The accuracy of the highest rated result, after the voice recognition stage, is expected to be quite low (eg, 25%). Therefore, the result initially displayed only rarely will be correct. IBM reports a word error rate of 28% for voicemail in Padmanabhan et al 2002.
When a category of oral expression can be identified (conjecture: 20% of cases), the possibility of obtaining a correct overall result should be reasonably high (eg, 70%) if the “phrase class” method is used, that is, if the errors in the exact phrases used for the top and bottom are accepted and there is no difficult part in the message or can be verified using other information (phone owners, previous calls). An approximate conjecture would be that an overall percentage of approximately 10% of oral expressions could be managed with simply a keystroke (the one required for category selection).
Error correction
It has been observed that errors in speech recognition tend to occur in clusters, eg, the average number of subsequent words that contain an error is approximately 2 (ALL: find reference). This is usually due to:
<dl><dt>-</dt><dd>Segmentation errors - the first incorrect word is shorter or longer than the correct one and therefore the </dd></dl>
<dl><dt>Next word must also be wrong. </dt><dd /></dl>
<dl><dt>-</dt><dd>The influence of the language model. </dd></dl>
<dl><dt>-</dt><dd>Possibly coarticulation modeling. </dd></dl>
This observation motivates the expectation that a correction of a word, during the editing process, will normally correct more than one error in the assumption of oral expression.
In very broad terms, a character's keyboard limits the number of contestants for the next word by a factor of 1/26. Two characters limit it to 1/676 and almost certainly, you must exclude all the wrong words of highest qualification. This motivates another prediction: an ASR error should, on average, require no more than one keystroke for correction.
Best route through the grid
A very important factor for the operational success of the system is the percentage of grids that contain the correct path even if it has a comparatively low rating. If the correct route is not in the grid, the system will not be able, at some point, to follow a route through the grid and therefore it will be difficult to obtain new predictions. The system may need to wait for the agent to type two or three words before finding the appropriate points in the remaining grid again to create more predictions. The size of the grid and therefore, the possibility of obtaining the correct oral expression, can be controlled by parameters (number of so-called tokens and pruning) and in theory, the total search space could be included. This would generate large reticles, however, that could not be transmitted to the client within an acceptable time frame. In addition, we have to address the occasional occurrence of previously unseen words that, consequently, would not be in the vocabulary. After a few months of operation (and therefore, data collection) a rate of approximately 95% seems achievable.
If the “oral expression segmentation” version described above is used, the segments would provide easy points to restart the prediction.
Linguistic Post-Processing
It might be worth defining a syntax of "SpinVox Messages" for each language. Message services
Short SMSs are not usually intended to contain adequate complete sentences and the attempt to add a lot of punctuation (and often do it wrongly) could, however, be worthwhile for use on rare occasions, but in a consistent manner.
Initial capitalization This operation is comparatively easy in English but more difficult in other languages (eg, German).
Expected benefits
<dl><dt>1. </dt><dd>When performing the conversion or editing, the system keeps a record of where, in the oral expression, the agent is currently and therefore, the audio reproduction can be better controlled. </dd></dl>
<dl><dt>2. </dt><dd>For a given proportion of oral expressions, where the agent only needs to determine the category, the ACR could be less than one (theoretically 1/3 with rapid reproduction and elimination of silence). </dd></dl>
<dl><dt>3. </dt><dd>A significant number of messages, for which the ASR performance is high, will require only a quick check and very few keystrokes to make corrections, providing an ACR of approximately 2. </dd></dl>
<dl><dt>4. </dt><dd>Most messages will still need an important issue. To what extent these cases will be advantageous from the predictions have yet to be determined.</dd></dl>
<dl><dt>5. </dt><dd>The management of initial capitalization and punctuation should automatically reduce the ACR by a small percentage and also improve consistency. </dd></dl>
Issues / Editions
1. When prediction with editions controlled by ASR become more time consuming than simple typing? To take advantage of the predictions, the agent needs to read them. If the following words are predicted correctly, their simple acceptance should be faster than typing them but if the next word is wrong, the additional time required for verification is simply wasted. On the contrary, the agent needs to listen in some way and could also use the time to check the predictions.
Combination of prediction methods
It seems promising to use statistical predictions as an operational backup for ASR-based prediction. Since the statistical prediction model is static (not dependent on calls) and since it does not need to be transmitted to the QC application with each message, it can be quite global. The reticles have to be transmitted for each message and therefore do not have to be kept within certain size limits and some of the necessary assumptions are likely to be missing.
Statistical models and prediction models based on ASR would be represented as graphs and the task of combining the predictions would imply going through both graphs separately and then choosing the most reliable prediction or combining them based on some interpolation formula.
This method could be extended to more graphs of prediction models, by way of example, based on call pairs.
Statistical predictions
These predictions are based on n-gram language models. These models memorize the conditional probabilities of word sequences. As an example, a 4-gram model could memorize the probability of the word "to" after the context of three words "I am going", English. These models can be very large and efficient ways to memorize it are also required to allow a rapid generation of predictions.
Put into practice
N-gram models are usually memorized in a graphic structure where each node represents a context (words already transcribed) and each outgoing link is annotated with a word and the corresponding conditional probability.
Since there will always be words that were never found (or very rarely) after a given context, but are still required at runtime, the model needs a way to treat previously unseen words in a given context. This is achieved with a "reserve copy" for the corresponding shorter context. In our example, if the word "to" had not been observed after "I am going" in English, the model would search for "to" after "am going." If it was never observed, it would look for the context node for "going" and finally, in the empty context node where all the words in the vocabulary are represented. This “backup” is implemented by adding a special link to each context node that points to the node with the corresponding shorter context and is annotated with “back-off penalty” that can be interpreted as the (logarithm of) the mass of probability not distributed through all other links that leave the node.
The overall log probability of "to" after "I am going" could, by way of example, be calculated as back_off ("I am going") + back_off ("am going") + link_prob ("to" @ context node "Going").
Word Graphics Expansion
It would be high cost from the point of view of the calculation, to search for the word (link) most likely to start with a given sequence of characters each time the user presses a key. One way to accelerate this operation is based on expanding the word chart in a character chart where outgoing links, in each node, are classified by decreasing probability. It should be noted that the maximum number of outbound links, on each node, is the number of characters in the language plus two for the backup reservation link and the final word link. Since searching through this list would require a maximum of 100 character comparisons for the English language with the expected cost of less than approximately 50 comparisons taking into account that the most likely words would be tried first.
The above ignores the cost of searching for words not found in the current context node. When this circumstance is required, it might be better to accept that predictions cannot be generated very quickly and use the normal security link to search in the security reserve context nodes. The alternative of memorizing backup links in each character node would require too much memory space.
Expansion from the word chart to the character chart can be implemented as follows:
1. For each context node (word level):
<dl><dt>-</dt><dd> Sort all outbound links by their probability (decreasing). </dd></dl>
<dl><dt>-</dt><dd> For each link (in order): </dd></dl>
<dl><dt>-</dt><dd> set the pointer for the current node (word) </dd></dl>
<dl><dt>-</dt><dd> for each character: </dd></dl>
<dl><dt>-</dt><dd> if it is already an annotated link with this character, set the pointer to the node to which the link points </dd></dl>
<dl><dt>-</dt><dd> In any other case: add a new link to the node to which the pointer points and create a new node as the destination for the link. Put the pointer to this new node.</dd></dl>
<dl><dt>-</dt><dd> Add a new link to the target pointer, pointing to the destination of the current word link. </dd></dl>
After the expansion, all word links (including their probabilities) can be deleted, except for backup reservation links. It should be noted that this will always allow finding the most likely phrase prediction, but not the list of the least likely. If the latter is required (later), the sequence in which the character expansion was performed would have to be memorized in some way.
Prediction
Taking a word node identifier, a character node identifier, the current word substring and a
character as input, the "prediction" method would be:
<dl><dt>1. </dt><dd>Go to the character node [character_node_id] </dd></dl>
<dl><dt>2. </dt><dd>Find the link annotated with the input character (use the linear search in links classified by probability). </dd></dl>
<dl><dt>3. </dt><dd>If it is found, follow the link and starting from its target node, follow the first link leaving each node until some stop condition is reached. At each transition, add the character found in the link to the result string. Resend the node identifier of the initial target node and the result string.</dd></dl>
<dl><dt>4. </dt><dd>In other cases: use the security reserve link from the word nodes [word_node_id] and the current word chain to find predictions in the security reserve nodes. This operation is not intended to find real-time predictions if the user types quickly.</dd></dl>
Annex III
Basics The following concepts are covered. Each basic concept A - I can be combined with any other basic concept in an implementation.
The following text also describes several subsystems that, inter alia, implement characteristics of the basic concepts. These subsystems do not need to be separated from each other. As an example, a subsystem can be part of another subsystem. Nor do the subsystems have to be discrete in some other way, the code implementation functions of one subsystem can be part of the same computer program as the code implementation functions of another subsystem.
Basic concept A
A large-scale, voice-independent voice messaging system that converts unstructured voice messages into text for visual display on a screen; the system comprises (i) subsystems implemented by computer and also (ii) a network connection for human operators that provide transcription and quality control; the system being adapted to optimize the effectiveness of human operators including, in addition:
3 basic subsystems, namely (i) a preprocessing front end that determines an appropriate conversion strategy; (ii) one or more conversion resources and (iii) a quality control subsystem.
Other features:
<dl><dt>- </dt><dd>Conversion resources include one or more of the following: one or more ASR engines; signal processing resources; Human operators</dd></dl>
- The signal processing resources optimize the quality of the audio signal for conversion by performing one or more of the following functions: elimination of noise, debugging of known defects, normalization of signal / volume energy, elimination of empty sections / in silence.
<dl><dt>- </dt><dd>Human operators perform random quality assurance tests on converted messages and provide informative feedback to the front end of pre-processing and / or conversion resources. </dd></dl>
Basic concept B
Context vectors
A voice messaging system, independent of the user and independent of the device, on a large scale, which converts unstructured vocal messages into text for visual display on a screen; the system comprising (i) subsystems implemented by computer and also (ii) a network connection for human operators that provide transcription and quality control; the system being adapted to optimize the effectiveness of human operators including, in addition:
a context subsystem implemented by the computer adapted to use information about the context of a message or a part of a message to improve the accuracy of the conversion.
Other features:
<dl><dt>-</dt><dd> Context information is used to limit the vocabulary used in any ASR engine or to refine the search or match processes used by the ASR engine. </dd></dl>
<dl><dt>- </dt><dd>Context information is used to select a particular conversion resource a combination of conversion resources, such as a particular ASR engine. </dd></dl>
<dl><dt>- </dt><dd>The context information includes one or more of the caller IDs, caller ID, if the calling user or the recipient is a company or other classifiable entity, or not, specific language of the calling user, historical record of call pairs; call time; day of the call; georeference or other location data of the user that called or of the called user; PIM data (personal information management data, including address book, diary) of the calling user or of the called user; the type of messages, including if the message is a voicemail, spoken text, an instant message, a blog entry, an email, a memo or a memo; message length; discoverable information using an online body of knowledge; presence data; voice density of the message; Message voice quality. </dd></dl>
<dl><dt>- </dt><dd>The context subsystem includes a recognizer trust subsystem that automatically determines the level of trust associated with a conversion of a specific message or part of a message, using the context information. </dd></dl>
<dl><dt>- </dt><dd>The context subsystem includes or is connected to a recognizer trust subsystem that determines caller the level of trust associated with a conversion of a specific message, or part of a message, using the output of one or more ASR engines. </dd></dl>
<dl><dt>-</dt><dd> The recognizer's trust subsystem can dynamically weigh how it uses the output of different ASR motors depending on their accuracy or probable effectiveness. </dd></dl>
<dl><dt>-</dt><dd>The context knowledge of a message is extracted by a subsystem and sent, directly, to a downstream subsystem that uses that context information to improve the conversion performance. </dd></dl>
<dl><dt>-</dt><dd> A downstream subsystem is a control and guarantee and / or quality supervision subsystem. </dd></dl>
Basic concept C
Historical record of call pairs
A voice messaging system, independent of the user and independent of the device, on a large scale, which converts unstructured vocal messages into text for visual display on a screen; the system comprising (i) subsystems implemented by computer and also (ii) a network connection for human operators that provide transcription and quality control; the system being adapted to optimize the effectiveness of human operators including, in addition:
a call pair subsystem implemented by computer, adapted to use historical record of call pairs to improve conversion accuracy.
Other features:
<dl><dt>-</dt><dd> The historical record of call pairs allows the system to be independent of the user but to acquire over the course of time, without explicit user training, user-dependent data that improve the conversion performance. </dd></dl>
<dl><dt>-</dt><dd> The historical record of call pairs is associated with a pair of numbers, including the associated numbers </dd></dl>
with mobile phones, landlines, IP addresses, email addresses or unique addresses
provided by a network.
<dl><dt>-</dt><dd> The historical record of call pairs includes information related to one or more of: language or dialect probably used; country called from or called to; temporary zones; call time; day of the call; specific phrases used; specific language of the calling user; intonation; PIM data (personal information management data, including address book, diary).</dd></dl>
<dl><dt>-</dt><dd> A dynamic language model subsystem implemented by computer, adapted to build a model </dd></dl>
dynamic language using one or more of: dependency of the calling user; call dependency
pair; called user dependency.
<dl><dt>-</dt><dd> The caller is anyone who deposits a vocal message, ignoring the attempt to make a vocal call and the called user is anyone who reads the converted message, ignoring whether they planned to receive a vocal call. </dd></dl>
<dl><dt>-</dt><dd>A personal profile subsystem implemented by computer adapted to build a personal profile of a calling user to improve the accuracy of the conversion. </dd></dl>
<dl><dt>-</dt><dd> The personal profile includes words, phrases, grammar or intonation of the caller. </dd></dl>
Basic concept D
3-part message taxonomy
A voice messaging system, independent of the user and independent of the device, on a large scale, which converts unstructured vocal messages into text for visual display on a screen; the system comprising (i) subsystems implemented by computer and also (ii) a network connection for human operators that provide transcription and quality control; the system being adapted to optimize the effectiveness of human operators including, in addition:
a subsystem of selection of limits implemented by computer, adapted to process a message looking for the boundaries between sections of the message that transmit different types of content or transmit different types of message.
Other features:
<dl><dt>-</dt><dd> The limits selection subsystem implemented by computer analyzes one or more of the following component parts: a greeting part; a body part; A parting farewell.</dd></dl>
<dl><dt>- </dt><dd>Different conversion strategies are applied to each part, being the optimal applied strategy for the conversion of that part. </dd></dl>
<dl><dt>- </dt><dd>Different parts of the message have different quality requirements and a quality assessment subsystem applies different standard levels to different parts. </dd></dl>
<dl><dt>-</dt><dd> A voice quality estimator detects boundaries between sections of the message that transmit different types of content or transmit different types of messages. </dd></dl>
<dl><dt>-</dt><dd> Limits are detected or inferred in zones in the message where the voice density changes. </dd></dl>
<dl><dt>- </dt><dd>Limits are detected or inferred in a message pause. </dd></dl>
<dl><dt>- </dt><dd>Limits are inferred as arising in a predefined proportion of the message. </dd></dl>
<dl><dt>-</dt><dd> A limit of greetings is inferred by approximately 15% of the length of the entire message. </dd></dl>
Basic concept E
Front end of preprocessing
A voice messaging system, independent of the user and independent of the device, on a large scale, which converts unstructured vocal messages into text for visual display on a screen; the system comprising (i) subsystems implemented by computer and also (ii) a network connection for human operators that provide transcription and quality control; the system being adapted to optimize the effectiveness of human operators including, in addition:
a preprocessing front end subsystem implemented by a computer that determines an appropriate conversion strategy used to convert voice messages.
Other features:
<dl><dt>- </dt><dd>The preprocessing front end optimizes the quality of the audio signal for conversion by performing one or more of the following functions: noise elimination, debugging of known defects, signal / volume energy normalization, removal of empty sections / in Silence and classify the type of message for optimal routing of the message for conversion, or not. </dd></dl>
<dl><dt>- </dt><dd>The front end of preprocessing determines the language that is being used by the calling user, based on one or more of the following: knowledge about registration, location and historical call log of the calling user and / or the receiver. </dd></dl>
<dl><dt>- </dt><dd>The preprocessing front end selects a particular ASR engine to convert a message or part of a message. </dd></dl>
<dl><dt>-</dt><dd> Different conversion resources, such as ASR engines, are used for different messages. </dd></dl>
<dl><dt>-</dt><dd> Human operators are treated as ASR engines. </dd></dl>
<dl><dt>-</dt><dd> The preprocessing front end uses or is connected to a recognizer trust subsystem to automatically determine the level of trust associated with a conversion of a specific message, or part of a message, and a particular conversion resource, such as a ASR engine, then develops dependent on that level of trust. </dd></dl>
<dl><dt>-</dt><dd> The conversion strategy involves the selection of a conversion strategy from a set of conversion strategies that include the following: (i) messages for which an ASR conversion confidence that is sufficiently high is checked by a subsystem of quality assessment for compliance with quality standards; (ii) messages for which the confidence in the conversion of ASR is not high enough is routed to a human operator for verification and, if necessary, its correction; (iii) messages for which the confidence in the conversion of ASR is very low are indicated as non-convertible and the user is informed of the receipt of a non-convertible message.</dd></dl>
Basic concept F
Queue Manager
A voice messaging system, independent of the user and independent of the device, on a large scale, which converts unstructured vocal messages into text for visual display on a screen; the system comprising (i) subsystems implemented by computer and also (ii) a network connection for human operators that provide transcription and quality control; the system being adapted to optimize the effectiveness of human operators including, in addition:
a queue manager subsystem implemented by a computer that intelligently manages the load and calls on resources that are required to ensure that converted message delivery times meet a predefined standard.
Other features:
<dl><dt>-</dt><dd> The queue manager subsystem determines what should happen to a voice message at each stage of processing through the system. </dd></dl>
<dl><dt>-</dt><dd> If at any stage of the automated conversion, the confidence intervals or other messages indicate that any part of a message is not good enough, in that case, the queue manager directs it to the correct human operator for assistance. </dd></dl>
<dl><dt>-</dt><dd>The queue manager subsystem makes decisions calculating compromise solutions between conversion time and quality. </dd></dl>
<dl><dt>-</dt><dd> The queue manager subsystem uses state machines that, for any given language queue, can decide how to best process messages through the system. </dd></dl>
Basic concept G
Reticle
A voice messaging system, independent of the user and independent of the device, on a large scale, which converts unstructured vocal messages into text for visual display on a screen; the system comprising (i) subsystems implemented by computer and also (ii) a network connection for human operators that provide transcription and quality control; the system being adapted to optimize the effectiveness of human operators including, in addition:
a grid subsystem implemented by computer that generates a grid of possible word sequences
or phrases and allows a human operator to guide a conversion subsystem by displaying one or more candidate converted words or phrases from the grid and allowing the operator to select that candidate word or phrase or, enter one or more characters for a different converted word or phrase, to operatively start the conversion subsystem to propose an alternative word or phrase.
Other features:
<dl><dt>- </dt><dd>The conversion subsystem receives inputs from a subsystem that manages the information in the historical record of call pairs. </dd></dl>
- The conversion subsystem receives inputs from conversion resources.
<dl><dt>- </dt><dd>The conversion subsystem receives inputs from a context subsystem that has knowledge of the context of a message. </dd></dl>
<dl><dt>-</dt><dd>The conversion subsystem learns, from the inputs of human operators, the words that probably correspond to a sound model. </dd></dl>
<dl><dt>-</dt><dd> The human operator is obliged to manage only a single key to accept a word or phrase. </dd></dl>
<dl><dt>- </dt><dd>The conversion subsystem automatically provides initial capitalization and punctuation. </dd></dl>
<dl><dt>- </dt><dd>The conversion subsystem can propose candidate numbers, real names, web addresses, email addresses, physical addresses, location information or other coordinates. </dd></dl>
<dl><dt>- </dt><dd>The conversion subsystem automatically differentiates between parts of the message that are probably important and those that are probably not important. </dd></dl>
<dl><dt>-</dt><dd> The unimportant parts of the message are confirmed by the operator as belonging to a class proposed by the conversion subsystem and then converted exclusively by a machine ASR engine. </dd></dl>
<dl><dt>-</dt><dd> The human operator can speak the correct word to the conversion system, which then transcribes it automatically. </dd></dl>
Basic concept H
Inline body
A voice messaging system, independent of the user and independent of the device, on a large scale, which converts unstructured vocal messages into text for visual display on a screen; the system comprising (i) subsystems implemented by computer and also (ii) a network connection for human operators that provide transcription and quality control; the system being adapted to optimize the effectiveness of human operators including, in addition:
a search subsystem implemented by computer that analyzes a converted message regarding an online body of knowledge.
Other features:
<dl><dt>- </dt><dd>The body of knowledge online is the Internet, as accessed by a search engine. </dd></dl>
<dl><dt>- </dt><dd>The online body of knowledge is a search engine database, such as Google. </dd></dl>
- The analysis of the converted message allows the accuracy of the conversion to be evaluated by a human operator and / or a trust subsystem of the recognizer.
<dl><dt>- </dt><dd>The analysis of the converted message allows the resolution of ambiguities in the message by a human operator and / or an ASR engine. </dd></dl>
Basic concept I
Detectors
A voicemail system, independent of the user and independent of the device, on a large scale, which converts voice messages unstructured in text for visual display on a screen; the system that comprises (i) subsystems implemented by computer and also (ii) a network connection for human operators that provide transcription and quality control; the system being adapted to optimize the effectiveness of the human operators including, in addition: a detector subsystem implemented by computer that is adapted to detect off-hook operations.
Other features:
<dl><dt>-</dt><dd> The pick-up operation detector is implemented as part of a preprocessing front end. </dd></dl>
Other detectors that can also be used:
<dl><dt>-</dt><dd> A computerized detector subsystem that adjusts to detect different spoken languages such as English, Spanish, French, etc. </dd></dl>
-The language detector can detect changes in the language part through a message. -The language detector can use inputs from a subsystem that has historical record information
of call pairs that record how language changes occurred in previous messages. -A detector subsystem implemented by computer that is adapted to estimate voice quality.
<dl><dt>-</dt><dd> The voice quality estimator finds voice fades, estimates noise levels and calculates a global measure of voice quality and uses an adaptive threshold to reject lower quality messages. </dd></dl>
<dl><dt>-</dt><dd>A detector subsystem implemented by a computer that is adapted to detect off-hook operations. </dd></dl>
<dl><dt>-</dt><dd> The pick-up operation detector is implemented as part of a preprocessing front end. </dd></dl>
<dl><dt>-</dt><dd> A detector subsystem implemented by a computer that is adapted to detect inadvertent calls. </dd></dl>
<dl><dt>-</dt><dd> The inadvertent call detector is implemented as part of a preprocessing front end. </dd></dl>
<dl><dt>- </dt><dd>A detector subsystem implemented by a computer that is adapted to detect and convert preregistered messages. </dd></dl>
<dl><dt>- </dt><dd>A detector subsystem implemented by a computer that is adapted to detect and convert spoken numbers. </dd></dl>
- A detector subsystem implemented by a computer that is adapted to detect and convert spoken addresses.
<dl><dt>- </dt><dd>A detector subsystem implemented by a computer that is adapted to detect and convert real names, numbers, web addresses, email addresses, physical addresses, location information and other coordinates. </dd></dl>
Message Types
<dl><dt>- </dt><dd>The message is a voicemail intended for a mobile phone and the voice message is converted to text and sent to that mobile phone. </dd></dl>
<dl><dt>- </dt><dd>The message is a voice message intended for an instant messaging service and the voice message is converted to text and sent to an instant messaging service for a visual presentation on a screen. </dd></dl>
<dl><dt>- </dt><dd>The message is a vocal message intended for a web blog and the voice message is converted to text and sent to a server for visual presentation as part of the web blog. </dd></dl>
<dl><dt>- </dt><dd>The message is a vocal message intended to be converted to text format and sent as a text message. </dd></dl>
<dl><dt>- </dt><dd>The message is a vocal message intended to be converted to text format and sent as an email message. </dd></dl>
<dl><dt>- </dt><dd>The message is a vocal message intended to be converted to text format and sent as a note or memo, by email or text, to a message sender. </dd></dl>
Other elements of the value chain
<dl><dt>-</dt><dd> A mobile telephone network that is connected to the system of any preceding claim. </dd></dl>
<dl><dt>-</dt><dd> A mobile phone when displaying a message converted by the system of any preceding claim. </dd></dl>
<dl><dt>-</dt><dd> A computer visual display screen when displaying a message converted by the system of any preceding claim. </dd></dl>
<dl><dt>-</dt><dd> A method of providing voice messaging, comprising the step of a user sending a voice message to a messaging system as set forth in any preceding claim.</dd></dl>
Contents4
94 members in 10 offices
Priority claims4
| Document | Office | Kind | Date |
|---|---|---|---|
| 0602682 | United Kingdom | A | |
| 0700376 | United Kingdom | A | |
| 0700377 | United Kingdom | A | |
| 2007000483 | United Kingdom | W |
Members94
| Document | Office | Kind | |
|---|---|---|---|
| GB0602682D0 | United Kingdom | D0 | |
| GB0700376D0 | United Kingdom | D0 | |
| GB0700377D0 | United Kingdom | D0 | |
| GB0700379D0 | United Kingdom | D0 | |
| GB0702706D0 | United Kingdom | D0 | |
| GB0702707D0 | United Kingdom | D0 | |
| US2007127688A1 | United States of America | A1 | |
| GB0708658D0 | United Kingdom | D0 | |
| GB2435147A | United Kingdom | A | |
| AU2007213532A1 | Australia | A1 | |
| CA2641853A1 | Canada | A1 | |
| WO2007091096A1 | World Intellectual Property Organization (WIPO) | A1 | |
| GB0717246D0 | United Kingdom | D0 | |
| GB0717247D0 | United Kingdom | D0 | |
| GB0717249D0 | United Kingdom | D0 | |
| GB0717250D0 | United Kingdom | D0 | |
| GB0800315D0 | United Kingdom | D0 | |
| GB0800318D0 | United Kingdom | D0 | |
| GB0800319D0 | United Kingdom | D0 | |
| GB0800320D0 | United Kingdom | D0 | |
| GB0800321D0 | United Kingdom | D0 | |
| US2008049906A1 | United States of America | A1 | |
| US2008049907A1 | United States of America | A1 | |
| US2008049908A1 | United States of America | A1 | |
| US2008052070A1 | United States of America | A1 | |
| US2008052071A1 | United States of America | A1 | |
| US2008063155A1 | United States of America | A1 | |
| US2008109221A1 | United States of America | A1 | |
| US2008133219A1 | United States of America | A1 | |
| US2008133231A1 | United States of America | A1 | |
| US2008133232A1 | United States of America | A1 | |
| US2008162132A1 | United States of America | A1 | |
| GB2445666A | United Kingdom | A | |
| GB2445667A | United Kingdom | A | |
| GB2445668A | United Kingdom | A | |
| GB2445669A | United Kingdom | A | |
| GB2445670A | United Kingdom | A | |
| AU2008204402A1 | Australia | A1 | |
| AU2008204404A1 | Australia | A1 | |
| CA2674971A1 | Canada | A1 | |
| CA2674987A1 | Canada | A1 | |
| WO2008084207A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2008084207A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2008084209A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2008084209A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2008084211A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2008084211A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2008084213A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2008084213A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2008084215A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2008084215A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2008084209A3 | World Intellectual Property Organization (WIPO) | A3 | |
| WO2008084209A3 | World Intellectual Property Organization (WIPO) | A3 | |
| WO2008084211A3 | World Intellectual Property Organization (WIPO) | A3 | |
| WO2008084211A3 | World Intellectual Property Organization (WIPO) | A3 | |
| EP1992154A1 | European Patent Office (EPO) | A1 | |
| WO2008084215A3 | World Intellectual Property Organization (WIPO) | A3 | |
| WO2008084215A3 | World Intellectual Property Organization (WIPO) | A3 | |
| EP2116035A1 | European Patent Office (EPO) | A1 | |
| EP2118830A2 | European Patent Office (EPO) | A2 | |
| EP2119205A2 | European Patent Office (EPO) | A2 | |
| EP2119208A1 | European Patent Office (EPO) | A1 | |
| EP2119209A2 | European Patent Office (EPO) | A2 | |
| CN201355842Y | China | Y | |
| MX2009007396A | Mexico | A | |
| MX2009007397A | Mexico | A | |
| BRMU8702846U2 | Brazil | U2 | |
| AU2007213532B2 | Australia | B2 | |
| BRPI0806206A2 | Brazil | A2 | |
| BRPI0806207A2 | Brazil | A2 | |
| EP2523441A1 | European Patent Office (EPO) | A1 | |
| EP2523442A1 | European Patent Office (EPO) | A1 | |
| EP2523443A1 | European Patent Office (EPO) | A1 | |
| AU2008204402B2 | Australia | B2 | |
| AU2012258326A1 | Australia | A1 | |
| US8374863B2 | United States of America | B2 | |
| EP1992154B1 | European Patent Office (EPO) | B1 | |
| AU2008204404B2 | Australia | B2 | |
| US2013165086A1 | United States of America | A1 | |
| ES2420559T3This record | Spain | T3 | |
| EP2523441B1 | European Patent Office (EPO) | B1 | |
| EP2523443B1 | European Patent Office (EPO) | B1 | |
| US8654933B2 | United States of America | B2 | |
| US8750463B2 | United States of America | B2 | |
| US8903053B2 | United States of America | B2 | |
| US8934611B2 | United States of America | B2 | |
| US8953753B2 | United States of America | B2 | |
| US8976944B2 | United States of America | B2 | |
| US8989713B2 | United States of America | B2 | |
| AU2012258326B2 | Australia | B2 | |
| US2015180809A1 | United States of America | A1 | |
| US9191515B2 | United States of America | B2 | |
| CA2641853C | Canada | C | |
| CA2674971C | Canada | C |
Numbers
- Publication
- 2420559
- Application
- 7712713
Titles2
- Spanish
- Un sistema a gran escala, independiente del usuario e independiente del dispositivo de conversión del mensaje vocal a texto
- English
- A large-scale system, independent of the user and independent of the device for converting the vocal message to text
Classification
- CPC, 6
- H04M3/53333
- H04M3/533
- H04M3/4936
- H04M3/5183
- H04M2201/60
- G10L15/26
- IPC, 1
- H04M3 533