Voice recognition rejection scheme
Abstract
A method of capturing a speech unit in a voice recognition system (10), comprising the steps of: comparing (18) the speech unit with a first stored word to generate a first score; compare (18) the speech unit with a second stored word to generate a second score; and determine (18) a difference between the first score and the second score; process (20) the speech unit based on the first score and the difference determined by: comparing the first score with a first slope threshold value and reject the speech unit if the first score is greater than the first slope threshold value; otherwise, compare the first score with a second slope threshold value and apply an N-best algorithm to verify the speech unit if the first score is greater than the second slope threshold value; otherwise, accept the speech unit; in which the first and second slope values vary with the difference determined.

Term
Term ended
Projected expiry passed 4 February 2020, 6.6 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
21 claims: 2 independent, 19 dependent
- 1ES 2 286 014 T3 REIVINDICACIONES 1. Un procedimiento de captura de una unidad de habla en un sistema (10) de reconocimiento de voz, que comprende las etapas de:comparar (18) la unidad de habla con una primera palabra almacenada para generar una primera puntuación;comparar (18) la unidad de habla con una segunda palabra almacenada para generar una segunda puntuación;y determinar (18) una diferencia entre la primera puntuación y la segunda puntuación;procesar (20) la unidad de habla basándose en la primera puntuación y la diferencia determinada por: comparar la primera puntuación con un primer valor umbral de pendiente y rechazar la unidad de habla si la primera puntuación es mayor que el primer valor umbral de pendiente;en caso contrario, comparar la primera puntuación con un segundo valor umbral de pendiente y aplicar un algoritmo N-best para verificar la unidad de habla si la primera puntuación es mayor que el segundo valor umbral de pendiente;en caso contrario, aceptar la unidad de habla;en el que el primer y segundo valor umbral de pendiente varían con la diferencia determinada.
- 2El procedimiento según la reivindicación 1, en el que la etapa de comparación de la primera puntuación con un segundo valor umbral de pendiente compara además la diferencia determinada con un umbral de diferencia y se aplica el algoritmo N-best para verificar la unidad de habla si la primera puntuación es mayor que el segundo valor umbral de pendiente y la diferencia es menor que el umbral de diferencia.
- 3El procedimiento según la reivindicación 1 ó 2, en el que la pendiente de dichos primeros valores umbral de pendiente y la pendiente de dichos segundos valores umbral de pendiente son la misma.
- 4El procedimiento según la reivindicación 1, en el que la diferencia corresponde a un cambio de puntuación entre la primera puntuación y la segunda puntuación.
- 5El procedimiento según la reivindicación 1, en el que la primera palabra almacenada comprende un candidato mejor en un vocabulario de un sistema (10) de reconocimiento de voz, y la segunda palabra almacenada comprende un candidato mejor siguiente en un vocabulario de un sistema (10) de reconocimiento de voz.
- 6El procedimiento según la reivindicación 1, en el que la primera puntuación comprende un resultado de comparación más ajustado, y la segunda puntuación comprende un resultado de comparación más ajustado siguiente.
- 7El procedimiento según la reivindicación 1, en el que la primera puntuación y la segunda puntuación comprenden coeficientes de codificación predictiva lineal.
- 8El procedimiento según la reivindicación 1, en el que la primera puntuación y la segunda puntuación comprenden coeficientes cepstral.
- 9El procedimiento según la reivindicación 1, en el que la primera puntuación y la segunda puntuación comprenden salidas de filtro paso banda.
- 10El procedimiento según la reivindicación 1, en el que la diferencia comprende una diferencia entre un resultado de comparación más ajustado y un resultado de comparación más ajustado siguiente.
- 11Un sistema (10) de reconocimiento de voz, que comprende:medios para comparar (18) la unidad de habla con una primera palabra almacenada para generar una primera puntuación;medios para comparar (18) la unidad de habla con una segunda palabra almacenada para generar una segunda puntuación;y medios para determinar (18) una diferencia entre la primera puntuación y la segunda puntuación;medios para procesar (20) la unidad de habla basándose en la primera puntuación y la diferencia determinada, que pueden operar para: ES 2 286 014 T3 comparar la primera puntuación con un primer valor umbral de pendiente y rechazar la unidad de habla si la primera puntuación es mayor que el primer valor umbral de pendiente;en caso contrario, comparar la primera puntuación con un segundo valor umbral de pendiente y aplicar un algoritmo N-best para verificar la unidad de habla si la primera puntuación es mayor que el segundo valor umbral de pendiente;en caso contrario, aceptar la unidad de habla;en el que el primer y segundo valor umbral de pendiente varían con la diferencia determinada.
- 12El sistema de reconocimiento de voz según la reivindicación 11, en el que los medios para procesar (20) pueden operar para comparar la primera puntuación con un segundo valor umbral de pendiente y para comparar la diferencia determinada con un umbral de diferencia, y en el que el algoritmo N-best se aplica para verificar la unidad de habla si la primera puntuación es mayor que el segundo valor umbral de pendiente y la diferencia es menor que el umbral de diferencia.
- 13El sistema de reconocimiento de voz según la reivindicación 11 ó 12, en el que la pendiente de dichos primeros valores umbral de pendiente y la pendiente de dichos segundos valores umbral de pendiente son la misma.
- 14El sistema (10) de reconocimiento de voz según la reivindicación 11, que comprende:medios (14) para extraer parámetros del habla a partir de muestras del habla digitalizadas de la unidad de habla, en el que los medios (18) para comparar la unidad de habla con una primera palabra almacenada, los medios para comparar (18) la unidad de habla con una segunda palabra almacenada, los medios (18) para determinar una diferencia, los medios (20) para determinar una relación y los medios (20) para procesar son todos partes de un medio único.
- 15El sistema (10) de reconocimiento de voz según la reivindicación 14, en el que:los medios para extraer (14) comprenden un procesador (14) acústico;y el medio único comprende un procesador acoplado al procesador (14) acústico.
- 16El sistema (10) de reconocimiento de voz según la reivindicación 11, en el que la primera palabra almacenada comprende un candidato mejor en un vocabulario del sistema (10) de reconocimiento de voz, y la segunda palabra almacenada comprende un candidato mejor siguiente en un vocabulario del sistema (10) de reconocimiento de voz.
- 17El sistema (10) de reconocimiento de voz según la reivindicación 11, en el que la primera puntuación comprende un resultado de comparación más ajustado, y la segunda puntuación comprende un resultado de comparación más ajustado siguiente.
- 18El sistema (10) de reconocimiento de voz según la reivindicación 11, segunda puntuación comprenden coeficientes de codificación predictiva lineal.
- 19El sistema (10) de reconocimiento de voz según la reivindicación 11, segunda puntuación comprenden coeficientes cepstral.
- 20El sistema (10) de reconocimiento de voz según la reivindicación 11, segunda puntuación comprenden salidas de filtro paso banda. en el que la primera puntuación y la en el que la primera puntuación y la en el que la primera puntuación y la
- 21El sistema (10) de reconocimiento de voz según la reivindicación 11, en el que la diferencia comprende una diferencia entre un resultado de comparación más ajustado y un resultado de comparación más ajustado siguiente.
Independent claims21
36 paragraphs in 2 sections, as filed
ES 2 286 014 T3
DESCRIPTION
Speech recognition rejection scheme.
Background of the invention
I. Field of the invention
The present invention belongs generally to the field of communications, and more specifically to speech recognition systems.
II. Background
Voice recognition (VR) represents one of the most important techniques for equipping a machine with simulated intelligence to recognize user or user voice commands and to facilitate human interface with the machine. VR also represents a key technique for understanding human speech. Systems that employ techniques to retrieve a linguistic message from an acoustic speech signal are called speech recognizers. A speech recognizer typically comprises an acoustic processor, which extracts a sequence of information-bearing characteristics information, or vectors, necessary to achieve the VR of the incoming raw speech, and a word decoder, which decodes the sequence of characteristics, or vectors, to produce a coherent and desired output format such as a sequence of linguistic words corresponding to the input speech unit. To increase the performance of a given system, training is required to equip the system with valid parameters. In other words, the system needs to learn before it can function optimally.
The acoustic processor represents an input speech analysis subsystem in a speech recognizer. In response to an input speech signal, the acoustic processor provides an appropriate representation to characterize the time-varying speech signal. The acoustic processor should discard irrelevant information such as background noise, channel distortion, speaker characteristics, and manner of speaking. Efficient acoustic processing provides speech recognizers with improved acoustic discrimination power. For this purpose, a useful characteristic to analyze is the short-time spectral envelope. Two commonly used spectral analysis techniques to characterize the short-time spectral envelope are linear predictive coding (LPC) and filterbank-based spectral modeling. Exemplary LPC techniques are described in US Patent No. 5,414,796, which is assigned to the assignee of the present invention, and in LB Rabiner & RW Schafer, Digital Processing of Speech Signals 396-453 (1978).
The use of VR (also commonly referred to as speech recognition) is becoming increasingly important for safety reasons. For example, VR can be used to replace the manual task of pressing buttons on a cordless telephone keypad. This is especially important when a user is initiating a phone call while driving a car. When using a phone without VR, the driver must take one hand off the wheel and look at the phone keypad while pressing the buttons to dial the call. These acts increase the possibility of a car accident. A speech-enabled phone (that is, a phone designed for speech recognition) would allow the driver to make phone calls while continually looking at the road. A hands-free car kit system would additionally allow the driver to keep both hands on the wheel during call initiation.
Speech recognition devices are classified as speaker-dependent or speaker-independent devices. Speaker independent devices can accept voice commands from any user. The more common speaker-dependent systems are trained to recognize commands from particular users. A speaker-dependent VR typically operates in two phases, a training phase and a recognition phase. In the training phase, the VR system prompts the user to say each of the words in the system's vocabulary once or twice so that the system can learn the characteristics of the user's speech for these particular words or phrases. Alternatively, for a phonetic VR device, training is done by reading one or more short articles specifically written to cover all phonemes in the language. An exemplary vocabulary for a hands-free car kit might include the digits above the keypad; the keywords "call", "send", "dial", "cancel", "delete", "add", "delete", "history", "schedule", "yes", and "no"; and the names of a predefined number of commonly called coworkers, friends, or family members. Once the training is finished, the user can initiate calls in the recognition phase by saying the trained keywords. For example, if the name “John” were one of the trained names, the user could initiate a call to John by saying the phrase “call John”. The VR system would recognize the words "call" and "John", and would dial the number that the user had previously entered as John's phone number.
The overall performance of a VR system can be defined as the percentage of cases in which a user goes through a recognition task successfully. A reconnaissance task typically comprises multiple stages. For example, in voice dialing with a cordless phone, overall performance refers to the average percentage of times a user successfully completes a phone call with the VR system. The number of steps required to achieve a successful VR phone call can vary from call to call. In general, the overall performance of a VR system depends mainly on two factors, the accuracy of
ES 2 286 014 T3 recognition of the VR system, and the human-machine interface. A subjective human user perception of VR system performance is based on overall performance. Therefore, there is a need for a VR system with high recognition precision and a human-machine interface to increase overall performance.
US Patent No. 4,827,520 and EP 0 867 861 A describe procedures that compare a speech unit with stored templates. Scores are determined as a result of the comparisons and, in particular circumstances, differences between the scores can be assessed.
Summary of the invention
The present invention is directed to a VR system with high recognition precision and an intelligent human-machine interface to increase the overall performance. Accordingly, in one aspect of the invention, a method for capturing a speech unit in a voice recognition system advantageously includes the steps of claim 1.
In another aspect of the invention, a speech recognition system advantageously includes the features of claim 11.
Brief description of the drawings
Figure 1 is a block diagram of a speech recognition device.
Figure 2 is a graph of score versus score change for a VR system rejection scheme, illustrating rejection, N-best, and acceptance regions.
Detailed description of the preferred embodiments
According to one embodiment, as illustrated in Figure 1, a speech recognition system 10 includes an analog-digital (A / D) converter 12, an acoustic processor 14, a VR template database 16, logic 18 of pattern comparison, and decision logic. The VR system 10 may reside in, for example, a cordless phone or a hands-free car kit.
When the VR system 10 is in the speech recognition phase, a person (not shown) says a word or phrase, generating a speech signal. The speech signal is converted to an electrical speech signal s (t) with a conventional transducer (not shown either). The speech signal s (t) is provided to the A / D 12, which converts the speech signal s (t) to digitized speech samples s (n) according to a known sampling procedure such as, for example, pulse modulation. encoded (PCM).
Speech samples s (n) are provided to acoustic processor 14 for parameter determination. The acoustic processor 14 produces a set of extracted parameters that models the characteristics of the signal s (t) of the input speech. The parameters can be determined according to any of a number of known speech parameter determination techniques including, for example, speech coder encoding and the use of cepstrum coefficients based on fast Fourier transform (FFT), as described. in the aforementioned US Patent No. 5,414,796. Acoustic processor 14 can be implemented as a digital signal processor (DSP) . The DSP may include a speech coder. Alternatively, acoustic processor 14 can be implemented as a speech coder.
Parameter determination is also performed during system 10 VR training, in which a set of templates for all vocabulary words in the system 10 VR are routed to the VR template database 16 for permanent storage therein. . The VR template database 16 is advantageously implemented as any conventional form of non-volatile storage medium, such as, for example, flash memory. This allows the templates to remain in the VR template database 16 when power is removed to the VR system 10.
The set of parameters is provided to the pattern comparison logic 18. Pattern comparison logic 18 advantageously detects the start and end points of a speech unit, calculates dynamic acoustic characteristics (such as, for example, time derivatives, second time derivatives, etc.), compresses acoustic characteristics by selecting important frames, and quantifies static and dynamic characteristics. Various known methods of end point detection, derivation of dynamic acoustic characteristics, pattern compression, and pattern quantification are described, for example, in Lawrence Rabiner & Biing-Hwang Juang, Foundations of Speech Recognition (1993).
The pattern comparison logic 18 compares the parameter set with all the templates stored in the VR template database 16. The results of the comparison, or distances, between the parameter set and all the templates stored in the VR template database 16 are provided to the decision logic 20. The decision logic 20 may (1) select from the VR template database 16 the templates that most closely correspond to the set of parameters, or it may (2) apply an "N-best" selection algorithm. ", Which chooses the closest matching N within a predefined matching threshold; or
ES 2 286 014 T3 can (3) reject the set of parameters. If an N-best algorithm is used, then the person is asked what choice was planned. The output of decision logic 20 is the decision of which vocabulary word is said. For example, in an N-best situation, the person might say, "John Anders," and the VR 10 system might respond, "Did you say John Andrews?" The person would then respond, "John Anders." The 10 VR system might then respond, "Did you say John Anders?" The person would then respond, "Yes," at which point the VR 10 would initiate dialing a phone call.
Pattern comparison logic 18 and decision logic 20 can be advantageously implemented as a microprocessor. Alternatively, the pattern comparison logic 18 and decision logic 20 can be implemented as any conventional form of processor, controller, or state machine. The VR system 10 can be, for example, an application-specific integrated circuit (ASIC). The recognition accuracy of the VR system 10 is a measure of how well the VR system 10 correctly recognizes spoken words or phrases in the vocabulary. For example, a recognition accuracy of 95% indicates that the VR 10 system correctly recognizes vocabulary words ninety-five times out of 100.
In one embodiment, a scoring plot is segmented against a change in score into regions of acceptance, N-best, and rejection, as illustrated in Figure 2. The regions are separated by lines according to known techniques of linear discriminative analysis. , which are described in Richard O. Duda & Peter E. Hart, Pattern Classification and Scene Analysis (1973). Each speech unit input to the VR system 10 is assigned a comparison result for, or distance from, each template stored in the VR template database 16 by pattern comparison logic 18, as described above. . These distances, or "scores", can advantageously be Euclidean distances between vectors in an N-dimensional vector space, summed over multiple frames. In one embodiment the vector space is a twenty-four dimensional vector space, the score is accumulated over twenty frames, and the score is an integer distance. Those skilled in the art would understand that the score could equally be expressed as a fraction or other value. Those skilled in the art would also understand that other metrics can be substituted for Euclidean distances, such that the scores could be, for example, measures of probability, measures of possibility, etc.
For a given speech unit and a given VR template from the VR template database 16, the lower the score (that is, the smaller the distance between the speech unit and the VR template), the tighter the correspondence between speech unit and VR template. For each speech unit, the decision logic 20 analyzes the score associated with the best match in the VR template database 16 in relation to the difference between that score and the score associated with the second best match in the base 16 of VR template data (that is, the second lowest score). As depicted in the graph of Figure 2, "score" is plotted versus "score change" and three regions are defined. The rejection region represents an area where a score is relatively high and the difference between that score and the next lowest score is relatively small. If a speech unit falls within the rejection region, the decision logic 20 rejects the speech unit. The acceptance region represents an area in which a score is relatively low and the difference between that score and the next lowest score is relatively large. If a speech unit falls within the acceptance region, the decision logic 20 accepts the speech unit. The N-best region lies between the reject region and the accept region. The N-best region represents a zone in which either a score is less than a score in the reject region or the difference between that score and the next lowest score is greater than the difference for a score in the region of rejection. The N-best region also represents a zone in which either a score is greater than a score in the acceptance region or the difference between that score and the next lowest score is less than the difference for a score in the region acceptance, since the difference for the score in the Nbest region is greater than a predefined threshold score change value. If a speech unit falls within the N-best region, the decision logic 20 applies an N-best algorithm to the speech unit, as described above.
In the embodiment described with reference to Figure 2, a first line segment separates the rejection region from the N-best region. The first line segment intersects the "score" axis at a predefined threshold score value. The slope of the first line segment is also predefined. A second line segment separates the N-best region from the acceptance region. The slope of the second line segment is predefined since it is the same as the slope of the first line segment, so the first and second line segments are parallel. A third line segment extends vertically from a predefined threshold change value on the "score change" axis to meet an end point of the second line segment. Those skilled in the art would appreciate that the first and second line segments need not be parallel, and could have any arbitrarily assigned slopes. Also, the third line segment does not need to be used.
In one embodiment the threshold score value is 375, the threshold change value is 28, and if the end point of the second line segment were to extend, the second line segment would intersect the "score" axis at the value 250, from So the slopes of the first and second line segments are each 1. If the score value is greater than the change in score value plus 375, the speech unit is rejected. Otherwise, if either the score value is greater than the score change value plus 250 or the score change value is less than 28, an N-best algorithm is applied to the speech unit. Otherwise, the speech unit is accepted.
In the embodiment described with reference to Figure 2, two dimensions are used for linear discriminative analysis. The dimension "score" represents the distance between a given speech unit and a given VR template,
ES 2 286 014 T3 as derived from the outputs of multiple band pass filters (not shown). The dimension "change in score" represents the difference between the lowest score, that is, the most closely matched score, and the next lowest score, that is, the score for the next best matched speech unit. In another embodiment the dimension "score" represents the distance between a given speech unit and a given VR template, as derived from the cepstral coefficients of the speech unit. In another embodiment the dimension "score" represents the distance between a given speech unit and a given VR template, as derived from the linear predictive coding (LPC) coefficients of the speech unit. Techniques for deriving LPC coefficients and cepstral coefficients from a speech unit are described in the aforementioned US Patent No. 5,414,796.
In alternative embodiments the linear discriminative analysis is not restricted to two dimensions. Accordingly, a first score based on bandpass filter outputs, a second score based on cepstral coefficients, and a change in score are analyzed in relation to each other. Alternatively, a first score based on bandpass filter outputs, a second score based on cepstral coefficients, a third score based on LPC coefficients, and a change in score are analyzed in relation to each other. As would be readily appreciated by those skilled in the art, the number of dimensions for "scoring" need not be restricted to any particular number. Experts would appreciate that the number of scoring dimensions is limited only by the number of words in the vocabulary of the VR system. Those of skill would also appreciate that the types of scores used need not be limited to any particular type of scoring, but may include any scoring procedure known in the art. Furthermore, and also readily appreciated by those skilled in the art, the number of dimensions for "score change" need not be restricted to one, or any particular number. For example, in one embodiment, a score is analyzed in relation to a change in score between the best match and the next closest match, and the score is also analyzed in relation to a change in score between the best match and the next best match. third tightest correspondence. Those skilled in the art would appreciate that the number of dimensions of change in punctuation is limited only by the number of words in the vocabulary of the VR system.
Therefore, a novel and improved speech recognition rejection scheme based on linear discriminative analysis has been described. Those skilled in the art would understand that the various illustrative logic blocks and algorithm steps described with respect to the embodiments described herein can be implemented or performed with a digital signal processor (DSP), an application-specific integrated circuit (ASIC). , discrete gates and transistor logic, discrete hardware components such as registers and FIFOs, a processor executing a firmware instruction set, or any conventional programmable software module and processor. The processor can advantageously be a microprocessor, but alternatively, the processor can be any conventional processor, controller, microcontroller, or state machine. The software module could reside in RAM memory, flash memory, registers, or any other form of writable storage medium known in the art. Those skilled in the art would further appreciate that the data, instructions, commands, information, signals, bits, symbols, and code elements that may be referred to throughout the foregoing description are advantageously represented by voltages, currents, electromagnetic waves, magnetic particles or fields, optical particles or fields, or any combination thereof.
Therefore, preferred embodiments of the present invention have been shown and described. However, it should be apparent to one skilled in the art that numerous alterations can be made to the embodiments described herein without departing from the scope of the invention. Therefore, the present invention is not limited except according to the following claims.
Contents2
2 sheets
Sheet 1 Sheet 2
18 members in 11 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 19990248513 | United States of America | – | |
| 24851399 | United States of America | A | |
| 24851399 | United States of America | A | |
| 00914513248513 | – | – | – |
| US19990248513 | – | – | – |
Members18
| Document | Office | Kind | |
|---|---|---|---|
| WO0046791A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU3589300A | Australia | A | |
| KR20010093327A | Republic of Korea | A | |
| EP1159735A1 | European Patent Office (EPO) | A1 | |
| CN1347547A | China | A | |
| US2002055841A1 | United States of America | A1 | |
| JP2002536691A | Japan | A | |
| US6574596B2 | United States of America | B2 | |
| CN1178203C | China | C | |
| HK1043423B | Hong Kong, China | B | |
| KR100698811B1 | Republic of Korea | B1 | |
| EP1159735B1 | European Patent Office (EPO) | B1 | |
| AT362166T | Austria | T | |
| ATE362166T1 | Austria | T1 | |
| DE60034772D1 | Germany | D1 | |
| ES2286014T3This record | Spain | T3 | |
| DE60034772T2 | Germany | T2 | |
| JP4643011B2 | Japan | B2 |
Numbers
- Publication
- 2286014
- Publication, DOCDB
- 2286014
- Publication, EPODOC
- ES2286014T
- Application
- 914513
- Application, DOCDB
- 00914513
- Application, EPODOC
- ES20000914513T
Titles2
- Spanish
- ESQUEMA DE RECHAZO DE RECONOCIMIENTO DE VOZ.
- English
- VOICE RECOGNITION REJECTION SCHEME.
Classification
- CPC, 3
- G10L15/10
- G10L15/00
- G10L15/22
- IPC, 6
- G10L15 00
- G10L15 10
- G06F3 16
- G10L15 08
- G10L15 22
- G10L15 28