Speech recognition system and method for recognising given speech patterns, specially for voice control
6 claims: 2 independent, 4 dependent
- 1Verfahren zur Spracherkennung vorgegebener Sprachmuster zur Sprachsteuerung von Kraftfahrzeugsystemen, bei dem in einem Spracherkennungssystem eine mit einer Aufnahmevorrichtung erfaßte Sprachäußerung mit den vorgegebenen Sprachmustern verglichen wird, für die Übereinstimmung eine Wahuscheinlichkeit ermittelt wird und bei dem das vorgegebene Sprachmuster, das mit der erfaßten Sprachäußerung mit der größten Wahrscheinlichkeit übereinstimmt, als erkanntes Sprachmuster ausgewählt wird, wobei bei Auftreten einer definierten Unsicherheitssituation im Hinblick auf die Spracherkennung vom Spracherkennungssystem eine automatische Nachtrainingsroutine ermöglicht wird und wobei im Falle einer Durchführung der Nachtrainingsroutine im ersten Schritt ein nachzutrainierendes vorgegebenes Sprachmuster ausgewählt wird, im zweiten Schritt die daraufhin folgende Sprachäußerung erfaßt wird und im dritten Schritt diese erfaßte Sprachäußerung dem ausgewählten vorgegebenen Sprachmuster zugeordnet wird und als neues Sprachmuster im Spracherkennungssystem hinterlegt wird, dadurch gekennzeichnet, dass im Falle einer Durchführung der Nachtrainingsroutine im ersten Schritt das nachzutrainierende vorgegebene Sprachmuster automatisch auflistend, beginnend mit dem Sprachmuster mit der größten Wahrscheinlichkeit, vom Spracherkennungssystem ausgewählt wird.
- 2Verfahren nach Patentanspruch 1, dadurch gekennzeichnet, daß bei erneutem Auftreten einer definierten Unsicherheitssituation für eine aktuell erfaßte Sprachäußerung bei einem Vergleich zunächst nur mit den vorgegebenen Sprachmustern anschließend die aktuell erfaßte Sprachäußerung mit dem in der Nachtrainingsroutine neu hinterlegten Sprachmuster verglichen und bei ausreichender Ähnlichkeit ausgewählt wird.
- 3Verfahren nach Patentanspruch 1, dadurch gekennzeichnet, daß die in den Nachtrainingsroutinen neu hinterlegten Sprachmuster zusammen mit den vorgegebenen Sprachmustern bei einem Vergleich mit einer aktuell erfaßten Sprachäußerung sofort herangezogen werden, ohne eine definierte Unsicherheitssituation bei einem ersten Vergleich nur mit den vorgegebenen Sprachmustern abzuwarten.
- 4Spracherkennungssystem zur Spracherkennung vorgegebener Sprachmuster zur Sprachsteuerung von Kraftfahrzeugsystemen, bei dem Mittel vorgesehen sind, die eine mit einer Aufnahmevorrichtung erfaßte Sprachäußerung mit den vorgegebenen Sprachmustern vergleichen, die für die Übereinstimmung eine Wahrscheinlichkeit ermitteln und die das vorgegebene Sprachmuster das der erfaßten Sprachäußerung mit der größten Wahrscheinlichkeit übereinstimmt, als erkanntes Sprachmuster auswählen, wobei Mittel vorgesehen sind, die bei Auftreten einer definierten Unsicherheitssituation im Hinblick auf die Spracherkennung eine automatische Nachtrainingsroutine ermöglichen und wobei Mittel vorgesehen sind, die im Falle einer Durchführung der Nachtrainingsroutine im ersten Schritt ein nachzutrainierendes vorgegebenes Sprachmuster auswählen, im zweiten Schritt die daraufhin folgende Sprachäußerung erfassen und im dritten Schritt diese erfaßte Sprachäußerung dem ausgewählten vorgegebenen Sprachmuster zuordnen und als neues Sprachmuster im Spracherkennungssystem hinterlegen, dadurch gekennzeichnet, dass die vorgesehenen Mittel auch derart ausgestaltet sind, dass im Falle einer Durchführung der Nachtrainingsroutine im ersten Schritt das nachzutrainierende vorgegebene Sprachmuster automatisch auflistend, beginnend mit dem Sprachmuster mit der größten Wahrscheinlichkeit, vom Spracherkennungssystem ausgewählt wird.
- 5Spracherkennungssystem nach Patentanspruch 4, dadurch gekennzeichnet, daß Mittel vorgesehen sind, die bei erneutem Auftreten einer definierten Unsicherheitssituation für eine aktuell erfaßte Sprachäußerung bei einem Vergleich zunächst nur mit den vorgegebenen Sprachmustern anschließend die aktuell erfaßte Sprachäußerung mit dem in der Nachtrainingsroutine neu hinterlegten Sprachmuster vergleichen und bei ausreichender Ähnlichkeit auswählen.
- 6Spracherkennungssystem nach Patentanspruch 4, dadurch gekennzeichnet, daß Mittel vorgesehen sind, die die in den Nachtrainingsroutinen neu hinterlegten Sprachmuster zusammen mit den vorgegebenen Sprachmustern bei einem Vergleich mit einer aktuell erfaßten Sprachäußerung sofort heranziehen, ohne eine definierte Unsicherheitssituation bei einem ersten Vergleich nur mit den vorgegebenen Sprachmustern abzuwarten.
Independent claims6
17 paragraphs, as filed
p0001The invention relates to a speech recognition system and method for speech recognition of predetermined speech patterns, in particular for voice control of motor vehicle systems, according to the preamble of claim 1 and 4 respectively.
p0002Such a speech recognition system is known for example from US 5,774,841. This uncertainty situations are treated poorly in speech recognition.
p0003For voice control of vehicle functions, such as navigation systems and car phones, voice recognition systems are already in use. These speech recognition systems is a with a recording device, eg. As a hands-free microphone, detected utterance with predetermined speech patterns compared. Such speech patterns are explained using the example of a car telephone, commands like "Dialing", "store name", "correction", "Delete phone book" and individual numbers. Here, the predetermined speech pattern that is the detected voice utterance is most similar, as a detected valid. To determine the similarity are speech, z. B. HMM recognizer, known which detect a speech utterance and determine probability of matching a detected voice utterance with the predetermined speech patterns. The predetermined voice pattern with the highest probability is selected. For a more detailed explanation of the operation of HMM recognizers for example made to EP 0559349 A1.
p0004Can not be assigned because, for example, two speech patterns the same probability was determined a speech utterance, the utterance or the command is rejected or the voice recognition system asks in dialogue with the user according to ( "what?"). Since this speech strongly depend on the pronunciation of the user (dialect sociolect, accent, etc.) can lead to frequent rejections of recognition in terms of the training population in intensively exposed speech. This leads to the reduction of customer acceptance of such language controls that are to be expanded, especially for motor vehicle functions in the future (eg. As heating / cooling control, radio / CD control, turn signal control, lighting control, etc.).
p0005It is an object of the invention to improve a speech recognition system and a speech recognition method of the initially mentioned type such that a greater independence from the speech is obtained.
p0006This object is achieved by the features of claim 1 or 4 procedurally or device standpoint. Advantageous developments of the invention are the subjects of the dependent claims.
p0007According to the invention, in a method for speech recognition of predetermined speech patterns, in particular for voice control of vehicle systems, wherein in a voice recognition system a detected with a receiving device voice utterance with the predetermined speech patterns are compared, and wherein the predetermined speech pattern that is the detected voice utterance is most similar, as detected applies speech pattern and is selected cherheitssituation in view of the speech recognition at a defined uncer occurrence from the voice recognition system enables an automatic Nachtrainingsroutine.
p0008A defined uncertainty situation can be a defined uncertainty value. For example, there is an uncertainty value, when a maximum probability of a given speech pattern is detected, is below a predetermined threshold. Also there may be a situation of uncertainty, when at least two predetermined speech patterns with identical or very similar probabilities (neighbors) are detected and thus a selection is not possible. Likewise, there is an uncertainty situation when the detected voice statement is repeatedly rejected or if a detected utterance repeatedly requests are issued by the speech recognition system.
p0009The operator or the user of the speech recognition system can Nachtrainingsroutine preferably accept or reject. If it is approved, a nachzutrainierendes predetermined speech pattern is selected manually by the operator or automatically by the speech recognition system for performing Nachtrainingsroutine in the first step. In the second step then the following utterance is detected. In the third step this detected voice utterance is associated with the selected given voice pattern and filed as a new voice pattern in the voice recognition system.
p0010In a first alternative, the currently detected voice statement is to re-occur a defined uncertainty situation for a currently detected voice utterance when comparing initially only with the standard voice patterns then compared to the in Nachtrainingsroutine newly recorded voice pattern and with sufficient similarity selected.
p0011Thus, the speech recognition system is in principle not changed. But for example, the primary user can utterances that are often dismissed retrain.
p0012In a second alternative in the Nachtrainingsroutinen newly recorded voice patterns are immediately taken with the predetermined speech patterns when compared with a currently detected voice statement, without waiting for a defined uncertainty situation at a first comparison with the predetermined speech patterns. Here, although the voice recognition system will be changed, but not the users in this case can retrain.
p0013The Nachtrainingsroutinen can be carried out as often. Either the detected in the Nachtrainingsroutinen utterances all be stored as a voice pattern for the comparison or only the last a certain predetermined speech patterns associated speech utterance is stored as a newly detected speech patterns for comparison.
p0014The drawing shows an embodiment of the invention is shown. It shows a rough functional sequence within a speech recognition system according to a preferred alternative method.
p0015In a not shown microprocessor controlled speech recognition system with microphone, a speech utterance is first detected and routed to an HMM speech recognizer. 1 In HMM speech 1 the predetermined voice pattern, eg. As also the possible commands a voice operated car phone stored. The predetermined speech patterns are compared with the detected utterance. In probability calculation block 2, the probabilities P are determined, with the matches the detected voice utterance with each predetermined speech patterns. The voice pattern with the highest probability P<sub>Max</sub> for a match is selected and forwarded to the decision block. 3 In decision block 3 is queried whether the highest probability P<sub>Max</sub> is smaller than a predetermined threshold. In this case, an absolute threshold or a relative threshold are given which defines the probability distance to the next probable speech pattern. If the maximum probability P<sub>Max</sub> greater than the predetermined threshold, the associated probability of being selected given voice pattern in the selection block 4 and -if it a command performed actual. If the maximum probability is less than the predetermined threshold, whereby an uncertainty situation is, going on to Nachtrainingsroutineblock. 5 Lies in Nachtrainingsroutineblock 5 no nachtrainiertes newly deposited speech patterns before, the user is asked whether he wants to carry out the Nachtrainingsroutine. Affirmed this, the implementation of Nachtrainingsroutine begins. Here the first step is a nachzutrainierendes predetermined speech pattern, for example, manually by the operator, eg., Via the non-voice-operated controls, or automatically auflistend, beginning with the voice pattern with the highest probability is selected from the speech recognition system. In the second step then the following utterance is detected, the third step this detected voice statement is associated with the selected predetermined speech patterns and stored as a new voice pattern in Nachtrainingsroutineblock 5 the speech recognition system.
p0016Lies in Nachtrainingsroutineblock 5 ever nachtrainiertes newly deposited speech patterns before, the detected voice statement is compared with this speech patterns and likely (z. B. also P> S) or in excessive similarity selected. Otherwise restarts a Nachtrainingsroutine.
p0017In this embodiment, therefore, a comparison with a newly stored voice pattern is performed only when the probability of an originally given voice pattern is too low or if two probabilities have too small a distance.
1 sheet
Sheet 1
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| WO9940570A | Cites | World Intellectual Property Organization (WIPO) | – |
| US4618984A | Cites | United States of America | – |
| US5774841A | Cites | United States of America | – |
| KOSAKA T.; SAGAYAMA S.: "Tree-structured speaker clustering for fast speaker adaptation", PROCEEDINGS OF THE CONFERENCE ON ACOUSTICS, SPEECH, AND SIGNAL PROCESSING, 1994, ADELAIDE, SA, AUSTRALIA, 19 April 1994 (1994-04-19), NEW YORK, NY, USA, IEEE, pages I-245 - I-248 | Non-patent | – | Examiner |
| KOSAKA T.; SAGAYAMA S.: 'Tree-structured speaker clustering for fast speaker adaptation' PROCEEDINGS OF THE CONFERENCE ON ACOUSTICS, SPEECH, AND SIGNAL PROCESSING, 1994, ADELAIDE, SA, AUSTRALIA 19 April 1994, NEW YORK, NY, USA, IEEE, Seiten I-245 - I-248 | Non-patent | – | – |
6 members in 2 offices; this record represents the family
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 19933323 | Germany | – | |
| 19933323 | Germany | A |
Members6
| Document | Office | Kind | |
|---|---|---|---|
| EP1069551A2 | European Patent Office (EPO) | A2 | |
| EP1069551A3 | European Patent Office (EPO) | A3 | |
| DE19933323A1 | Germany | A1 | |
| DE19933323C2 | Germany | C2 | |
| EP1069551B1This record | European Patent Office (EPO) | B1 | |
| DE50012632D1 | Germany | D1 |
32 legal events, as 5 offices reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | Office | |
|---|---|---|---|
| Expiry of rightR071 | R071 | DE | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Notification of lapseLapsedST | ST | FR | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Gb: european patent ceased through non-payment of renewal feeCeasedGBPC | GBPC | EP | |
| Ep patent has lapsedLapsedEUG | EUG | SE | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| No opposition filedOpposition26N | 26N | EP | |
| No opposition filed within time limitOppositionORIGINAL CODE: 0009261PLBE | PLBE | EP | |
| Information on the status of an ep patent application or granted ep patentGrantedSTATUS: NO OPPOSITION FILED WITHIN TIME LIMITSTAA | STAA | EP | |
| Fr: translation filedET | ET | EP | |
| Translation of granted ep patentGrantedTRGR | TRGR | SE | |
| Gb: translation of ep patent filed (gb section 77(6)(a)/1977)GBT | GBT | EP | |
| Corresponds to:REF | REF | EP | |
| Designated contracting statesAK | AK | EP | |
| European patent grantedGrantedNOT ENGLISHFG4D | FG4D | GB | |
| (expected) grantORIGINAL CODE: 0009210GRAA | GRAA | EP | |
| Grant fee paidORIGINAL CODE: EPIDOSNIGR3GRAS | GRAS | EP | |
| Despatch of communication of intention to grant a patentORIGINAL CODE: EPIDOSNIGR1GRAP | GRAP | EP | |
| First examination report despatched17Q | 17Q | EP | |
| Designation fees paidDE FR GB SEAKX | AKX | EP | |
| Request for examination filed17P | 17P | EP | |
| Designated contracting statesAK | AK | EP | |
| Request for extension of the european patentAL;LT;LV;MK;RO;SIAX | AX | EP | |
| Search report despatchedORIGINAL CODE: 0009013PUAL | PUAL | EP | |
| Designated contracting statesAK | AK | EP | |
| Request for extension of the european patentAL;LT;LV;MK;RO;SIAX | AX | EP | |
| Public reference made under article 153(3) epc to a published international application that has entered the european phaseORIGINAL CODE: 0009012PUAI | PUAI | EP |
Numbers
- Publication
- 1069551
- Application
- 1121219
Titles3
- German
- Spracherkennungssystem und Verfahren zur Spracherkennung vorgegebener Sprachmuster insbesondere zur Sprachsteuerung
- English
- Speech recognition system and method for recognising given speech patterns, specially for voice control
- French
- Système de reconnaissance de la parole et procédé de reconnaissance des formes déterminées de parole, en particulier pour commande vocale
Classification
- CPC, 3
- G10L15/063
- G10L2015/0631
- G10L2015/0635
- IPC, 2
- G10L15 06
- G10L15 22
Designated states4
- Contracting states, 4
- Germany
- France
- United Kingdom
- Sweden
