Method and device to enter by speech an address of destination in a navigation system in real time
Abstract
The invention relates to a method for voice input of a destination address in a route guidance system in real-time operation, in which input speech utterances of a user are recognized by means of a speech recognition device and classified according to their recognition probability, and that speech utterance with the greatest recognition probability is identified as the speech utterance entered, at least one speech utterance permissible voice command is which activates the operating functions of the route guidance system assigned to this voice command, all permissible utterances being stored on at least one database. According to the invention, at least one operating function of the route guidance system comprises at least one input dialog, wherein after the activation of the at least one operating function of the route guidance system, depending on the at least one input dialog, at least one lexicon is generated in real time from the permitted utterances uttered on at least one database and then the at least one lexicon is used as Vocabulary is loaded into the speech recognition device.

Term
Term ended
Projected expiry passed 3 March 2018, 8.6 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
13 claims: 13 independent, 0 dependent
- 1Method for voice input of a destination address in a route guidance system in real-time operation, in which input speech utterances of a user are recognized by a speech recognition device and classified according to their recognition probability and that speech utterance with the greatest recognition probability is identified as the speech utterance entered, at least one speech utterance being a permissible voice command, which activates the operating functions of the route guidance system assigned to this voice command, all permissible utterances being stored on at least one database,characterized,that at least one operating function of the route guidance system comprises at least one input dialog, wherein after the activation of the at least one operating function of the route guidance system, depending on the at least one input dialog, at least one lexicon is generated in real time from the permitted utterances expressed in at least one database, and then the at least one lexicon is used as a vocabulary is loaded into the speech recognition device. Verfahren zur Spracheingabe einer Zieladresse in ein Zielführungssystem im Echtzeitbetrieb, bei welchem eingegebene Sprachäußerungen eines Benutzers mittels einer Spracherkennungseinrichtung erkannt und gemäß ihrer Erkennungswahrscheinlichkeit klassifiziert werden und diejenige Sprachäußerung mit der größten Erkennungswahrscheinlichkeit als die eingegebene Sprachäußerung identifiziert wird, wobei mindestens eine Sprachäußerung ein zulässiges Sprachkommando ist, welches die diesem Sprachkommando zugeordneten Bedienfunktionen des Zielführungssystems aktiviert, wobei alle zulässigen Sprachäußerungen auf mindestens einer Datenbasis gespeichert sind, dadurch gekennzeichnet, daß wenigstens eine Bedienfunktion des Zielführungssystems mindestens einen Eingabedialog umfaßt, wobei nach der Aktivierung der wenigstens einen Bedienfunktion des Zielführungssystems in Abhängigkeit des mindestens einen Eingäbedialogs aus den auf mindestens einer Datenbasis gespeicherten zulässigen Sprachäußerungen in Echtzeit mindestens ein Lexikon generiert und anschließend das mindestens eine Lexikon als Vokabular in die Spracherkennungseinrichtung geladen wird.
- 2Method according to claim 1,characterized,that at least one lexicon is generated from the permissible utterances uttered on at least one database in an 'off-line editing mode', and after the activation of the at least one operating function of the navigation system, depending on the at least one input dialog, the at least one in the 'off-line Editing mode 'generated lexicon is loaded in real time as vocabulary into the speech recognition device. Verfahren nach Anspruch 1, dadurch gekennzeichnet, daß mindestens ein Lexikon aus den auf mindestens einer Datenbasis gespeicherten zulässigen Sprachäußerungen in einem 'Off-Line-Editiermodus' generiert wird, wobei nach der Aktivierung der wenigstens einen Bedienfunktion des Zielführungssystems in Abhängigkeit des mindestens einen Eingabedialogs das mindestens eine im 'Off-Line-Editiermodus' generierte Lexikon in Echtzeit als Vokabular in die Spracherkennungseinrichtung geladen wird.
- 3Method according to claim 1 or 2,characterized,that at least one utterance is a place name and / or a street name, with all permitted place names being stored in a destination file, with all permitted street names being stored in a street list for at least one permitted place name. Verfahren nach Anspruch 1 oder 2, dadurch gekennzeichnet, daß mindestens eine Sprachäußerung ein Ortsnamen und/oder ein Straßennamen ist, wobei alle zulässigen Ortsnamen in einer Zieldatei abgelegt sind, wobei für wenigstens einen zulässigen Ortsnamen alle zulässigen Straßennamen in einer Straßenliste abgelegt sind.
- 4Method according to claim 1,characterized,that the speech recognition device comprises at least one speaker-independent speech recognizer and at least one speaker-dependent additional speech recognizer, where, depending on the at least one input dialog, the speaker-independent speech recognizer is used for recognizing individually spoken place names and / or street names and / or letters spoken individually and / or in groups and / or parts of words and the speaker-dependent additional speech recognizer for recognizing at least one spoken keyword is used. Verfahren nach Anspruch 1, dadurch gekennzeichnet, daß die Spracherkennungseinrichtung mindestens einen sprecherunabhängigen Spracherkenner und mindestens einen sprecherabhängigen Zusatz-Spracherkenner umfaßt, wobei abhängig von dem mindestens einen Eingabedialog der sprecherunabhängige Spracherkenner zur Erkennung von einzeln gesprochenen Ortsnamen und/oder Straßennamen und/oder von einzeln und/oder in Gruppen gesprochenen Buchstaben und/oder von Wortteilen verwendet wird und der sprecherabhängige Zusatz-Spracherkenner zur Erkennung von mindestens einem gesprochenen Schlüsselwort verwendet wird.
- 5Method according to claim 4,characterized,that the at least one keyword is assigned a specific destination address, the at least one spoken keyword being stored in a personal address register, and wherein a name lexicon is generated from the personal address register and loaded into the speech recognition device. Verfahren nach Anspruch 4, dadurch gekennzeichnet, daß dem mindestens einen Schlüsselwort eine bestimmte Zieladresse zugeordnet wird, wobei das mindestens eine gesprochene Schlüsselwort in einem persönlichen Adreßregister gespeichert wird, und wobei aus dem persönlichen Adreßregister ein Namen-Lexikon generiert und in die Spracherkennungseinrichtung geladen wird.
- 6Method according to claim 2,characterized,that a basic lexicon that was generated in the 'off-line editing mode' contains the 'p' largest places of a state. Verfahren nach Anspruch 2, dadurch gekennzeichnet, daß ein Grund-Lexikon, welches im 'Off-Line-Editiermodus' generiert wurde, die 'p' größten Orte eines Staates enthält.
- 7Method according to claim 6,characterized,that the basic lexicon is stored in an internal non-volatile memory of the navigation system. Verfahren nach Anspruch 6, dadurch gekennzeichnet, daß das Grund-Lexikon in einem internen nicht flüchtigen Speicher des Zielführungssystems gespeichert ist.
- 8Method according to claim 1,characterized,that an environment lexicon which is generated in real time contains 'a' locations in the vicinity of the current vehicle location, the environment lexicon being updated at regular intervals. Verfahren nach Anspruch 1, dadurch gekennzeichnet, daß ein Umgebungs-Lexikon, welches in Echtzeit generiert wird, 'a' Orte im Umkreis des aktuellen Fahrzeugstandortes enthält, wobei das Umgebungs-Lexikon in regelmäßigen Abständen aktualisiert wird.
- 9A method according to claim 8,characterized,that the environment lexicon is stored in an internal non-volatile memory of the route guidance system. Verfahren nach Anspruch 8, dadurch gekennzeichnet, daß das Umgebungs-Lexikon in einem internen nicht flüchtigen Speicher des Zielführungssystems gespeichert ist.
- 10Method according to claim 3,characterized,that after activating an input dialog 'spell destination' a partial word lexicon for spelling recognition is loaded into the speech recognition device, that the user then enters individual letters and / or letter groups as utterances, which are then compared in the speech recognition device with the partial word lexicon, whereby a hypothesis list with word hypotheses is formed from the recognized letters and / or groups of letters, that the first 'n' word hypotheses are then compared with the target file and a whole-word lexicon is generated from the result of the comparison and loaded into the speech recognition device for full-word recognition, and that subsequently a stored acoustic value in the speech recognition device for full-word recognition with the full-word lexicon is compared this acoustic value was generated from a speech uttered as a whole word before loading the partial word lexicon. Verfahren nach Anspruch 3, dadurch gekennzeichnet, daß nach der Aktivierung eines Eingabedialogs 'Zielort buchstabieren' ein Teilwort-Lexikon zur Buchstabiererkennung in die Spracherkennungseinrichtung geladen wird, daß anschließend der Benutzer einzelne Buchstaben und/oder Buchstabengruppen als Sprachäußerungen eingibt, welche anschließend in der Spracherkennungseinrichtung mit dem Teilwort-Lexikon verglichen werden, wobei aus den erkannten Buchstaben und/oder Buchstabengruppen eine Hypothesenliste mit Worthypothesen gebildet wird, daß danach die ersten 'n' Worthypothesen mit der Zieldatei abgeglichen werden und aus dem Ergebnis des Abgleichs ein Ganzwort-Lexikon generiert und zur Ganzworterkennung in die Spracherkennungseinrichtung geladen wird, und daß anschließend ein gespeicherter akustischer Wert in der Spracherkennungseinrichtung zur Ganzworterkennung mit dem Ganzwort-Lexikon verglichen wird, wobei dieser akustische Wert aus einer als Ganzwort gesprochenen Sprachäußerung vor dem Laden des Teilwort-Lexikons erzeugt wurde.
- 11Method according to claim 1,characterized,that after recognizing a 'rough destination' entered through an input dialog 'enter rough destination', 'm' locations in the vicinity of 'rough destination' are calculated in real time by the route guidance system and a fine destination lexicon is generated from these 'm' locations and in the speech recognition device is loaded. Verfahren nach Anspruch 1, dadurch gekennzeichnet, daß nach dem Erkennen eines durch einen Eingabedialog 'Grobziel eingeben' eingegebenen 'Grobziels', durch das Zielführungssystem in Echtzeit 'm' Orte im Umkreis um den Ort 'Grobziel' berechnet werden und aus diesen 'm' Orten ein Feinziel-Lexikon generiert und in die Spracherkennungseinrichtung geladen wird.
- 12Apparatus for carrying out the method according to one of the preceding claims, in which a route guidance system (2) is connected to a speech dialog system (1) via corresponding connections (12), wherein depending on at least one speech uttered by a user in a speech input device (5), if the spoken utterance is recognized as a permissible voice command by a speech recognition device (7), An operating function of the route guidance system (2) assigned to the voice command can be activated by means of a dialog and sequence control (8), wherein all permissible utterances can be stored on at least one database (9;4;3),characterized,that by means of the dialog and sequence control (8), depending on at least one input dialog, which is part of at least one operating function, at least one lexicon can be generated in real time from the permissible utterances expressed on the at least one database (9;4;3) and as a vocabulary can be loaded into the speech recognition device (7). Vorrichtung zum Ausführen des Verfahrens nach einem der vorhergehenden Ansprüche, bei welcher ein Zielführungssystem (2) über entsprechende Verbindungen (12) mit einem Sprachdialogsystem (1) verbunden ist, wobei abhängig von mindestens einer von einem Benutzer in eine Spracheingabeeinrichtung (5) gesprochene Sprachäußerung, wenn die gesprochene Sprachäußerung als ein zulässiges Sprachkommando von einer Spracherkennungseinrichtung (7) erkannt wird, mittels einer Dialog- und Ablaufsteuerung (8) eine dem Sprachkommando zugeordnete Bedienfunktion des Zielführungssystems (2) aktivierbar ist, wobei alle zulässigen Sprachäußerungen auf mindestens einer Datenbasis (9;4;3) speicherbar sind, dadurch gekennzeichnet, daß mittels der Dialog- und Ablaufsteuerung (8) in Abhängigkeit von mindestens einem Eingabedialog, welcher Teil mindestens einer Bedienfunktion ist, in Echtzeit wenigstens ein Lexikon aus den auf der mindestens einen Datenbasis (9;4;3) gespeicherten zulässigen Sprachäußerungen generierbar und als Vokabular in die Spracherkennungseinrichtung (7) ladbar ist.
- 13Device according to claim 12,characterized,that by means of the dialog and sequence control (8), depending on at least one input dialog, at least one lexicon generated in the 'off-line editing mode' and stored on at least one database (9;4;3) in real time as a vocabulary in the Speech recognition device (7) can be loaded. Vorrichtung nach Anspruch 12, dadurch gekennzeichnet, daß mittels der Dialog- und Ablaufsteuerung (8) in Abhängigkeit von mindestens einem Eingabedialog mindestens ein im 'Off-Line-Editiermodus' generiertes Lexikon, welches auf mindestens einer Datenbasis (9;4;3) gespeichert ist, in Echtzeit als Vokabular in die Spracherkennungseinrichtung (7) ladbar ist.
Independent claims13
29 paragraphs, as filed
The invention relates to a method for voice input of a destination address into a route guidance system in real time operation according to the preamble of claim 1 and a device for executing the method according to the preamble of claim 12.
DE 196 00 700 describes a route guidance system for a motor vehicle, in which a fixed switch, a touch panel switch or a speech recognition device can be used as the input device. However, the document does not address the voice input of a destination address in a route guidance system.
EP 0 736 853 A1 also describes a route guidance system for a motor vehicle. The voice input of a destination address in a route guidance system is not the subject of this document.
DE 36 08 497 A1 describes a method for the voice-controlled operation of a telecommunications terminal, in particular a car telephone. A disadvantage of the method is that it does not deal with the special problems with the voice input of a destination address into a route guidance system.
The applicant's older, not previously published patent application P 195 33 541.4-52 discloses a generic method for the automatic control of one or more devices by voice commands or by voice dialog in real-time operation and a device for carrying out the method. According to this method, the entered voice commands are recognized by means of a speech recognition device, which comprises a speaker-independent speech recognizer and a speaker-dependent additional speech recognizer, and identified according to their likelihood of recognition as the entered voice command, and the functions of the device or devices assigned to this voice command are initiated. The voice commands or the speech dialogues are formed on the basis of at least one synatax structure, at least one basic command vocabulary and, if necessary, at least one speaker-specific additional command vocabulary. The syntax structures and the basic command vocabularies are specified in speaker-independent form and are fixed during real-time operation. The speaker-specific additional vocabulary is entered and / or changed by the respective speaker in that during training phases within and outside of real-time operation, an additional speech recognizer working according to a speaker-dependent recognition method is trained by the respective speaker by entering the additional commands at least once for the speaker-specific characteristics of the respective speaker. In real-time operation, the voice dialog is handled and / or the devices are controlled as follows:<ul id="ul0001" list-style="dash" compact="compact"><li>Voice commands entered by the user are forwarded to a speaker-independent speech recognizer working on the basis of phonemes and the speaker-dependent additional speech recognizer, where they are subjected to a feature extraction and examined and classified in the speaker-dependent additional speech recognizer based on the features extracted there for the presence of additional commands from the additional command vocabulary.</li><li>the commands and syntax structures of the two speech recognizers, which are classified as having a certain probability, are then combined to form hypothetical speech commands, and these are examined and classified in terms of their admissibility and recognition probability in accordance with the predetermined syntax structure.</li><li>the permissible hypothetical voice commands are then checked for plausibility according to predetermined criteria, and of the hypothetical voice commands recognized as plausible, the one with the highest probability of recognition is selected and identified as the voice command entered by the user.</li><li>the functions of the devices to be controlled which are assigned to the identified voice command are then initiated and / or responses are generated in accordance with a predetermined voice dialog structure in order to continue the voice dialog. According to this document, the described method can also be used to operate a route guidance system, the destination address being entered by entering letters or groups of letters in a spelling mode and the user making a list for storing destination addresses for the route guidance system under predefined names / abbreviations can be created.</li></ul>
A disadvantage of this described method is that the special properties of a route guidance system are not dealt with and only the voice input of a destination is specified using a spelling mode.
The object of the invention is to develop a generic method in such a way that the special properties of a route guidance system are taken into account and simplified, faster voice input of a destination address into a route guidance system is possible, and the ease of use is thereby improved. A suitable device for carrying out the method is also to be specified.
According to the invention, this object is achieved with the features of claim 1 or 12, the features of the subclaims characterizing advantageous refinements and developments.
The method according to the invention for voice input of destination addresses into a route guidance system uses a voice recognition device known from the prior art for voice recognition, as described for example in the document acknowledged in the introduction to the description, the voice recognition device comprising at least one speaker-independent speech recognizer and at least one speaker-dependent additional speech recognizer. The method according to the invention enables various input dialogues for voice input of destination addresses. In a first input dialog, hereinafter referred to as "destination entry", the speaker-independent speech recognizer is used to recognize isolated spoken destinations and, if the isolated spoken destination was not recognized, to recognize spoken letters and / or groups of letters. In a second input dialog, hereinafter called 'spell destination', the speaker-independent speech recognizer is used to recognize spoken letters and / or groups of letters. In a third input dialog, hereinafter referred to as "rough destination entry", the speaker-independent speech recognizer is used to recognize isolated spoken destinations and, if the isolated spoken destination was not recognized, to recognize spoken letters and / or groups of letters. In a fourth input dialog, hereinafter referred to as “indirect input”, the speaker-independent speech recognizer is used to recognize connected spoken numbers and / or groups of numbers. In a fifth input dialog, hereinafter referred to as 'street input', the speaker-independent speech recognizer is used to recognize isolated spoken street names and, if the isolated spoken street name is not recognized, to recognize spoken letters and / or groups of letters. Through the input dialogs described so far, verified destination addresses, each comprising a destination and a street, are transferred to the route guidance system. In a sixth input dialog, hereinafter called 'retrieve address', in addition to the speaker-independent speech recognizer, the speaker-dependent additional speech recognizer is used to recognize key words spoken in isolation. In a seventh input dialog, hereinafter referred to as 'save address', a keyword entered by the user is assigned a destination address entered by the user, a destination address assigned to the corresponding recognized keyword being transferred to the destination guidance system during the input dialog 'retrieve address'.
The method according to the invention is mainly based on the fact that not the entire permissible vocabulary for a speech recognition device is loaded in the speech recognition device at the time of its activation, but that, depending on the input dialogs required for executing an operating function, at least one necessary lexicon from the entire possible vocabulary during real-time operation generated and loaded into the speech recognition device. There are over 100,000 places in the Federal Republic of Germany that can be used as a vocabulary for the route guidance system. If this vocabulary were loaded into the speech recognition device, the recognition process would be extremely slow and error-prone. A lexicon generated from this vocabulary contains only approx. 1500 words, which makes the recognition process much faster and the recognition rate increases. At least one target file is used as the data basis for the method according to the invention, which contains all possible destination addresses and certain additional information on the possible destination addresses of a route guidance system and is stored on at least one database. From this target file, lexicons are generated which comprise at least parts of the target file, at least one lexicon being generated in real time as a function of at least one activated input dialog. It is particularly advantageous if the destination file contains additional information for each stored destination, for example political affiliation or addition of name, postcode or postcode area, area code, state, number of inhabitants, geographic coding, phonetic description or lexicon affiliation. This additional information can then be used to resolve ambiguities or to speed up the process of finding the desired destination. Instead of the phonetic description, a transcription of the phonetic description in the form of a chain of indices, depending on the realization of the transcription, can be used for the speech recognition device. Furthermore, a so-called automatic phonetic transcription can be provided, which carries out a rule-based conversion of orthographically existing names, including an exception table, into a phonetic description. The entry of the lexicon membership is only possible if the corresponding lexica were generated from the destination file in an 'off-line editing mode', outside the actual operation of the route guidance system, and on the at least one database, for example a CD-ROM or an external one Database in a center, which can be accessed via appropriate communication devices, such as a cellular network, has been stored. The generation of lexicons in the 'off-line editing mode' only makes sense if there is enough storage space on the at least one database, and is particularly useful for lexicons which are required particularly frequently. A CD-Rom or an external database is particularly suitable as the database for the target file, as this enables the target file to be kept up to date at all times. At the moment, not all possible place names of the Federal Republic of Germany have been digitized and stored on a database. Likewise, a corresponding street list is not available for all locations. It is therefore important to be able to update the data basis at any time. An internal non-volatile memory area of the route guidance system can also be used as a database for the at least one lexicon generated in the 'off-line editing mode'.
In order to be able to carry out the voice input of a desired destination address to the route guidance system more quickly, a basic vocabulary is loaded after the initialization phase of the route guidance system or in the case of a sufficiently large non-volatile internal memory area after each change of the database, which contains at least one basic lexicon generated from the destination file includes. This basic lexicon may already have been generated in the 'off-line editing mode'. The basic lexicon can be stored in addition to the target file on the database or can be stored in the non-volatile internal memory area of the route guidance system. As an alternative, the generation of the basic lexicon can only take place after the initialization phase. The dynamic generation of lexicons during the real-time operation of the route guidance system, ie during the runtime, has two major advantages. On the one hand, this gives you the option of compiling any lexicons from the data basis stored on the at least one database and, on the other hand, you save considerable storage space on the at least one database, since not all of the lexicons required for the various input dialogues are saved on the at least one database before the speech recognizer is activated must be filed.
In the embodiment described below, the basic vocabulary comprises two lexica generated in the 'off-line editing mode' and stored on the at least one database and two lexica which are generated after the initialization phase. If the speech recognition device has enough working memory, the basic vocabulary is loaded into the speech recognition device after the initialization phase in addition to the permissible voice commands for the speech dialog system according to P 195 33 541.4-52 . After the initialization phase and pressing the PTT button, the voice dialogue system then allows various information to be entered with the intention of controlling the devices connected to the voice dialogue system, as well as operating the basic functions of a route guidance system and a destination and / or a street as the destination address for the route guidance system to enter. If the speech recognition device has too little working memory, the basic vocabulary is only loaded into the speech recognition device when a corresponding operating function which uses the basic vocabulary has been activated. The basic lexicon stored on the at least one database comprises the 'p' largest cities in the Federal Republic of Germany, the parameter 'p' being set to 1000 in the described embodiment. This will bring about 53 million Germans or 65% of the population directly recorded. The basic lexicon includes all places with more than 15,000 inhabitants. A region lexicon likewise stored on the database comprises 'z' names of regions and areas such as, for example, Lake Constance, Swabian Alb, etc., with the region lexicon comprising approximately 100 names in the embodiment described. The region lexicon is used to record known areas and common region names. These names hide summaries of place names that can be generated and loaded as a new area lexicon after the area or region name has been spoken. An environmental lexicon, which is only generated after initialization, includes 'a' dynamically reloaded place names around the current vehicle location, which means that even smaller places in the immediate vicinity can be addressed directly, with the parameter 'a' in the described embodiment 400 is set. This environment lexicon is also updated at certain intervals during the journey, so that it is always possible to speak to places in the immediate vicinity. The current vehicle location is communicated to the route guidance system by a location method known from the prior art, for example by means of a global positioning system (GPS). The dictionaries described so far are assigned to the speaker-independent language recognizer. A name lexicon which is not generated from the target file and which is assigned to the speaker-dependent speech recognizer comprises approximately 150 user-spoken keywords from the user's personal address register. Each keyword is assigned a specific target address from the target file using the 'Save address' input dialog. This specific destination address is transferred to the route guidance system by voice input of the assigned keyword using the input dialog 'Get address'. This results in a basic vocabulary of approx. 1650 words that are known to the speech recognition device and can be entered as an isolated spoken word (place name, street name, key word).
In addition, it can be provided to transmit addresses from an external data source, for example a personal electronic data system or a portable small computer (laptop), by means of data transmission to the voice dialogue system or to the route guidance system and to integrate them into the basic vocabulary as an address lexicon. Normally, no phonetic descriptions for the address data (name, destination, street) are stored on the external data sources. In order to still be able to transfer this data into the vocabulary for a speech recognition device, an automatic phonetic transcription of this address data, in particular the name, must be carried out. An assignment to the correct destination is then implemented using a table.
For the dialog examples described below, a target file must be stored on the at least one database of the route guidance system, which contains a data record according to Table 1 for each location recorded in the route guidance system. Depending on the storage space and availability, parts of the information given may also be missing. However, this only affects data that are used to resolve ambiguities such as name addition, district affiliation, area codes, etc. If address data from an external data source is used, the address data must be supplemented accordingly. The word subunits are of particular importance for the speech recognition device, which works as a hidden Markov model speech recognizer (HMM recognizer)<tables id="tabl0001" num="0001"><table frame="all"><title>Table 1</title><tgroup cols="2" colsep="1" rowsep="1"><colspec colnum="1" colname="col1" colwidth="78.75mm" /><colspec colnum="2" colname="col2" colwidth="78.75mm" /><thead valign="top"><row rowsep="1"><entry namest="col1" nameend="col1" align="left"><b>Description of the entry</b></entry><entry namest="col2" nameend="col2" align="left"><b>example</b></entry></row></thead><tbody valign="top"><row><entry namest="col1" nameend="col1" align="left">Place names</entry><entry namest="col2" nameend="col2" align="left">Flensburg</entry></row><row><entry namest="col1" nameend="col1" align="left">Political affiliation or additional name</entry><entry namest="col2" nameend="col2" align="left">-</entry></row><row><entry namest="col1" nameend="col1" align="left">Postcode or postcode area</entry><entry namest="col2" nameend="col2" align="left">24900 - 24999</entry></row><row><entry namest="col1" nameend="col1" align="left">Area code</entry><entry namest="col2" nameend="col2" align="left">0461</entry></row><row><entry namest="col1" nameend="col1" align="left">City or county affiliation</entry><entry namest="col2" nameend="col2" align="left">Flensburg, city district</entry></row><row><entry namest="col1" nameend="col1" align="left">state</entry><entry namest="col2" nameend="col2" align="left">Schleswig-Holstein</entry></row><row><entry namest="col1" nameend="col1" align="left">population</entry><entry namest="col2" nameend="col2" align="left">87526</entry></row><row><entry namest="col1" nameend="col1" align="left">Geographic coding</entry><entry namest="col2" nameend="col2" align="left">9.43677, 54.78204</entry></row><row><entry namest="col1" nameend="col1" align="left">Phonetic description</entry><entry namest="col2" nameend="col2" align="left">| fl'Ens | bUrk |</entry></row><row rowsep="1"><entry namest="col1" nameend="col1" align="left">Word subunits for the HMM speech recognizer</entry><entry namest="col2" nameend="col2" align="left">f [LN] le E [LN] n [C] sb [Vb] U [Vb] r k. or 101 79 124 117 12 39 35 82 68</entry></row><row rowsep="1"><entry namest="col1" nameend="col1" align="left">Lexicon affiliation</entry><entry namest="col2" nameend="col2" align="left">3, 4, 78 ...</entry></row></tbody></tgroup></table></tables>
The invention is explained in more detail below on the basis of exemplary embodiments in conjunction with the drawings. Show it:<ul id="ul0002" list-style="none"><li>1: a schematic representation of an overview of the possible input dialogues for voice input of a destination address for a route guidance system;</li><li>2: shows a schematic diagram of a flow chart of a first exemplary embodiment for the input dialog 'destination input';</li><li>3: shows a schematic diagram of a flow chart of a second exemplary embodiment for the input dialog 'destination entry';</li><li>4: shows a schematic diagram of a flow chart for the input dialog 'select from list';</li><li>5: shows a schematic diagram of a flowchart for the input dialog 'resolve ambiguity';</li><li>6: shows a schematic diagram of a flowchart for the input dialog 'spell destination';</li><li>7: shows a schematic diagram of a flow chart for the input dialog 'rough destination input';</li><li>8: shows a schematic diagram of a flow chart for the input dialog 'save address';</li><li>9: shows a schematic diagram of a flowchart for the input dialog 'road input';</li><li>Fig. 10: a schematic representation of a block diagram of an apparatus for performing the method according to the invention.</li></ul>
1 shows an overview of the possible input dialogues for voice input of a destination address for a route guidance system. A speech dialogue between a user and a speech dialogue system according to FIG. 1 begins after the initialization phase with a waiting state 0 in that the speech dialogue system remains until the PTT key (push-to-talk key) is pressed and into which the speech dialogue system after the end of Voice dialog returns. The user activates the voice dialog system by pressing the PTT key in step 100. The voice dialog system responds in step 200 with an acoustic output, for example by a signal tone or by a voice output, which indicates to the user that the voice dialog system is ready to accept a voice command. In step 300, the voice dialog system now waits for a permissible voice command in order to control the various devices connected to the voice dialog system by means of a dialog and sequence control or to start a corresponding input dialog. In detail, however, only the permissible voice commands that relate to the route guidance system are dealt with in detail. The following voice commands relating to the various input dialogs of the route guidance system can be entered here:<ul id="ul0003" list-style="none"><li>- "Destination entry" E1: The input dialog 'Destination entry' is activated with this voice command.</li><li>- "Spell destination" E2: With this voice command, the input dialog 'Spell destination' is activated.</li><li>- "Coarse target entry" E3: The input dialog 'Coarse target entry' is activated with this voice command.</li><li>- "Postal code" E4 or "Area code" E5: These two voice commands activate the input dialog 'Indirect input'.</li><li>- "Street entry" E6: With this voice command, the input dialogue 'Street entry' is activated.</li><li>- "Save address" E7: This voice command activates the input dialog 'Save address'.</li><li>- "Retrieve address" E8: The input dialog 'Retrieve address' is activated with this voice command.</li></ul>
Instead of the above, other terms can of course also be used to activate the various input dialogs. In addition to the above voice commands, general voice commands for controlling the route guidance system, for example "route guidance information", "route guidance start / stop" etc., can also be entered.
After starting an input dialog by speaking the corresponding voice command, the corresponding lexicons are loaded as vocabulary into the speech recognition device. If the destination has been entered successfully as part of the destination entry using one of the input dialogs 'Destination entry' in step 1000, 'Spell destination' in step 2000, 'Rough destination entry' in step 3000 or 'Indirect entry' in step 4000, it is then checked in step 350 whether there is a corresponding street list for the recognized destination or not. If the check produces a negative result, the method branches to step 450. If the check comes to a positive result, a query is made in step 400 as to whether the user wants to enter a street name or not. If the user answers the query 400 with "Yes", the input dialog 'street input' is called. If the user answers the query 400 with "No", the method branches to step 450. The query 400 is therefore only carried out if the street names for the corresponding destination are also recorded in the route guidance system. In step 450, the recognized desired destination is automatically supplemented by the specification center or city center as a street entry, since only a complete destination address can be transferred to the route guidance system, the destination address also a street or a special destination, for example train station, airport, in addition to the destination. City center, etc., includes. In step 500, the destination address is transferred to the route guidance system. The speech dialogue is then ended and the speech dialogue system returns to the waiting state 0. If, at the start of the speech dialog in step 300, the voice command "street entry" E6 was spoken by the user and recognized by the speech recognition device, then the entry dialog 'street entry' is activated in step 5000. After successfully entering the desired destination and the street, the destination address is then transferred to the route guidance system in step 500. If the voice command "Get address" E8 was spoken by the user at the beginning of the voice dialog in step 300 and recognized by the speech recognition device, then the input dialog "Get address" is activated in step 6000. In the input dialog 'Retrieve address' a keyword is spoken by the user and the address assigned to the spoken keyword is transferred to the route guidance system as the destination address in step 500. If the voice command "Save address" E7 was spoken by the user at the beginning of the voice dialog in step 300 and recognized by the speech recognition device, then the input dialog 'save address' is activated in step 7000. Using the input dialog 'Save address', an entered destination address is saved under a keyword spoken by the user in the personal address register. The input dialog 'Retrieve address' is then closed and the system returns to waiting state 0.
2 shows a schematic representation of a first embodiment of the input dialog 'destination input'. After activation of the input dialog "destination entry" in step 1000 by the voice command "destination entry" E1 spoken by the user in step 300 and recognized by the speech recognition device, as can be seen from FIG. 2, the basic vocabulary is loaded into the speech recognition device in step 1010. The loading of the basic vocabulary into the speech recognition device can in principle also take place at another point, for example after the initialization phase or after the PTT key has been pressed. This depends on the speed of the loading process and on the type of speech recognition device used. The user is then asked in step 1020 to enter a destination. In step 1030, the user inputs the desired destination by means of voice input. In step 1040, this voice input is transferred to the speech recognition device as an acoustic value 〈destination_1〉 and compared there with the loaded basic vocabulary, with samples in the time or frequency range or feature vectors being able to be transferred to the speech recognition device as an acoustic value. The type of acoustic value transferred also depends on the type of speech recognizer used. As a result, the speech recognizer supplies a first hypothesis list hypo.1 with place names, which are sorted according to the probability of recognition. If the hypothesis list contains hypo.1 homophonic place names, ie place names that are pronounced identically but written differently, e.g. Ahlen and Aalen, then both place names are given the same recognition probability and both place names are taken into account when the input dialog is continued. Then in step 1050, the place name with the greatest probability of recognition is output as a voice output 〈hypo.1.1〉 to the user with the question whether 〈hypo.1.1〉 corresponds to the desired destination 〈destination_1〉 entered or not. It is irrelevant here whether there are several entries in the first position of the hypothesis list or not, since the place names are spoken identically. If the query 1050 is answered with "yes", then a jump is made to step 1150. If the user answers the question with "No", the acoustic value 〈destination_1〉 of the entered destination is stored in step 1060 for a possible later recognition process with another lexicon. The user is then asked in step 1070 to speak the destination again. In step 1080, the user re-enters the destination by means of voice input. In step 1090, this voice input is transferred to the speech recognition device as an acoustic value 〈destination_2〉 and compared there with the loaded basic vocabulary. As a result, the speech recognition device supplies a second hypothesis list hypo.2 with place names, which are sorted according to the probability of recognition. In step 1100 it is checked whether the ASCII value of the place name or for homophonic place names, the ASCII values of the place names with the greatest probability of detection hypo.1.1 of the hypothesis list hypo.1 with the ASCII value of the place name, or for homophonic place names the ASCII values of the place names with the greatest probability of recognition hypo.2.1 or match. If this is the case, then in step 1110 the place name with the second largest probability of recognition from the second hypothesis list hypo.2 is output as voice output 〈hypo.2.2〉 with the question to the user whether 〈hypo.2.2〉 is the desired destination or not . If the check 1100 comes to a negative result, in step 1120 the place name with the greatest probability of recognition from the second hypothesis list hypo.2 is output to the user as a speech output 〈hypo.2.1〉 with the question whether 〈hypo.2.1〉 is the desired one Destination is or not. If the response from the user shows that the desired destination has still not been recognized, the input dialog 'spell destination' is called in step 1140. If the response of the user reveals that a destination has been recognized, then in step 1150, step 1150 is also reached if the query 1050 was answered with "yes", the ASCII value of the recognized destination, or in the case of homophonic place names, the Compare the ASCII values of the recognized place names (either hypo.1.1, hypo.2.1 or hypo.2.2) with the ASCII values of the place names stored in the data records of the target file. An ambiguity list is then generated from all the place names in which one of the recognized destinations is contained in the orthography. With homophonic place names, the ambiguity list always contains several entries, and the result is therefore not clear.
At this point, however, a non-homophonic place name can lead to an ambiguity list with several entries, the so-called "Neustadt problem", if the orthographic spelling of the entered destination is available several times in the destination file. It is therefore checked in step 1160 whether the destination has been clearly recognized or not. If the destination is not clear, the process branches to step 1170. An eighth input dialog is called there, hereinafter called 'resolve ambiguity'. If the destination is unambiguous, the voice output in step 1180 outputs the destination found with certain additional information, for example postcode, place name and state, with the question to the user whether it is the desired destination or not. If the user answers the question with "No", a branch is made to step 1140, which calls the input dialog 'spell destination'. If the user answers the question with "yes", then the destination is temporarily stored in step 1190 and a jump is made to step 350 (see description of FIG. 1).
Fig. 3 shows a schematic representation of a second embodiment of the input dialog 'destination input'. Process steps 1000 to 1060 have already been dealt with in the description of FIG. 2. In contrast to the first embodiment of the input dialog, after step 1060 the input dialog 'destination entry' is continued with step 1075 and not with step 1070. In step 1075, the place name with the second largest probability of recognition of the first hypothesis list hypo.1 is output as a speech output 〈hypo.1.2〉 with the question to the user whether 〈hypo.1.2〉 corresponds to the desired entered destination 〈Destination_1〉 or not. If the query 1075 is answered with "yes", then a jump is made to step 1150. If the user answers the question with "No", the process branches to step 1140. In step 1140 the input dialog 'Spell destination' is called. The method steps from step 1150 have already been dealt with in the description of FIG. 2 and are therefore no longer described here.
Fig. 4 shows a schematic representation of an embodiment of a flow chart for an eighth input dialog, hereinafter referred to as 'select from list', for selecting an entry from a list. After activating the input dialog 'Select from list' in step 1430 by another input dialog, the user is informed in step 1440 of the number of entries in the list and asked in step 1445 whether the list should be read out or not. If the user answers the question with "yes", a reading flag is set in step 1450 and the program then branches to step 1460. If the user answers the question with "No", the process continues directly with step 1460. In step 1460 the user is shown the list with the request to select an entry from the list, the list being displayed only when the reading flag is not set on a display device of the route guidance system. If the reading flag is set, the list is also read out by voice. The list is divided into pages that contain, for example, up to four entries, with the entries on each page being consecutively numbered starting with one. In step 1460, the user can speak various voice commands to continue the input dialog. With a first voice command E9, for example "continue", the user can turn to the next page of the list in step 1470 and then return to step 1460. With a second voice command E10, for example “back”, the user can scroll back to the previous page of the list in step 1475 and also return to step 1460. With a third voice command E11, for example "number X", the user can select a specific entry from the list, where X stands for the number of the desired entry. After the third voice command E11 has been spoken, a branch is made to step 1480. With a fourth voice command E12, for example "Abort", the user can end the input dialog 'Select from list' if he has not found the desired entry, for example. Therefore, after entering the fourth voice command, the process branches to step 1500. In step 1500, depending on whether the entries in the list consist of street names or place names, the user is informed by voice output that the street name or place name could not be found. The system then returns to waiting state 0. In step 1480, a query to the user again checks whether the “entry_X” is correct or not. In order to make the selection of the desired entry from the list more convenient, it can also be provided that the corresponding list is generated as a list lexicon and loaded into the speech recognition device. As a result, the user can select the entry as described above by speaking the corresponding number as the third voice command E11, or he can read the corresponding entry from the list and enter it as the third voice command E11. For example, the list contains the desired entry: 4. Neustadt an der Weinstrasse, the user can either speak "Number 4" or "Neustadt an der Weinstrasse" as the third voice command E11 and the system recognizes the desired entry in both cases. If the user answers the query 1480 with "Yes", the input dialog 'Select from list' has ended in step 1490 and the selected 'Entry_X' is passed as a result to the calling input dialog. If the user answers the query 1480 with "No", the process branches to step 1500.
5 shows a schematic representation of an embodiment of a flowchart for a ninth input dialog, hereinafter referred to as "resolving ambiguity", for resolving ambiguities, for example for the so-called "Neustadt problem" or for homophonic place names. After activating the input dialog 'resolve ambiguity' in step 1170 by another input dialog, the user is informed in step 1200 how many place names are entered in the ambiguity list. The location with the largest number of inhabitants is then searched for in step 1210 and output in step 1220 as an acoustic value 〈largest_location〉 with the query whether the user desires 〈largest_location〉 as the destination or not. If the user answers the query 1220 with "Yes", the process jumps to step 1410. In step 1410, "largest_location" is passed as the result to the calling input dialog and the input dialog "resolve ambiguity" is ended. If the user answers the inquiry 1220 with "No", then in step 1230 it is checked whether the ambiguity list contains more than k entries or not. If the ambiguity list comprises k or fewer entries, the input dialog 'Select from list' is called in step 1240. The parameter k should not be chosen too large, otherwise the input dialog 'Select from list' takes too long. During testing, k = 5 has proven to be a favorable value. If the check 1230 yields a positive result, an attempt is made in step 1250 with a first query dialog to reduce the number of entries in the ambiguity list. After the first query dialog, it is checked in step 1260 whether the destination is unique or not. If the check 1260 comes to a positive result, then a branch is made to step 1410; if the check 1260 comes to a negative result, a check is carried out analogously to step 1230 in step 1270 whether or not the ambiguity list comprises more than k entries. If the ambiguity list comprises k or fewer entries, the process branches to step 1240; if the ambiguity list includes more than k entries, a second query dialog in step 1280 attempts to reduce the number of entries in the ambiguity list. The processes described are repeated up to an nth query dialog in step 1290. After the nth query dialog, in step 1300 it is checked analogously to step 1260 whether the destination is clear or not. If the destination is unambiguous, the process branches to step 1410; if not, the process branches to step 1240. The input dialog 'Select from list' called in step 1240 returns a clear destination to the input dialog 'Resolve ambiguity'. In step 1410, a clear destination is passed on to the calling input dialog as a result of the input dialog 'resolve ambiguity' and the input dialog 'resolve ambiguity' is ended. Queries for the postcode, the area code, the state or the next largest city can be used as query dialogs. The query dialogs begin with a check as to whether the corresponding query makes sense or not. If, for example, all place names on the ambiguity list are in a federal state, then the query for the federal state is meaningless and the federal state query dialog is ended. Various criteria can be used to evaluate whether a request makes sense or not. For example, an absolute threshold value, for example 8 entries have the query criterion or 80% of the entries have the query criterion. After checking whether the activated request dialog is useful or not, the request is issued to the user, for example "Do you know the state in which the destination is located?" or "Do you know the postcode (or area code, or the next largest city) of the destination?". The input dialog is then continued in accordance with the user's response. If the user does not know the query criterion, the system branches to the next query. If the user knows the query criterion, he will be asked to enter a voice. For the federal state query, a federal state lexicon, if not yet available, can be generated and loaded into the speech recognition device as a vocabulary. In addition to the correct designation of the federal states, common abbreviations, e.g. B. Württemberg instead of Baden Württemberg may be included in the state lexicon. If a query does not result in a reduction of the entries in the ambiguity list, the original ambiguity list is used for the further input dialog 'resolve ambiguity'. If a query results in a reduction of the entries in the ambiguity list, the reduced ambiguity list is used for the further input dialog 'resolve ambiguity'. It is particularly advantageous if the query for the postcode is carried out as the first query dialog, since this criterion leads to a clear result in most applications. This also applies to the query for the area code.
6 shows a schematic representation of an embodiment of the input dialog 'spell destination'. After activating the input dialog 'Spell destination' in step 2000 the user is asked in step 2010 to spell the destination. In step 2020, the voice is input by the user, and the letters of the destination can be spoken in one piece or in groups of letters, separated by short pauses. In addition, certain word endings such as -heim, -berg, -burg, -hausen, -tal etc. or word beginnings such as upper, lower, free, new, bad etc. can be allowed as word input, the word starts and / or the word endings are contained in a partial word lexicon together with the permissible letters, the partial word lexicon being loaded or activated in the speech recognizer at the beginning of the input dialog 'spell destination'. The entered word starts, letters or word endings are fed to the speaker-independent speech recognizer for spelling recognition. In step 2030, a third hypothesis list hypo.3 with words, which were formed from the recognized letters, is returned by the speech recognizer. Then, in step 2040, the hypothesis with the greatest probability of recognition hypo.3.1 from the third hypothesis list hypo.3 is compared with the target file. The result is a new location list, which is also sorted according to the probability of recognition. Step 2050 then checks whether or not an acoustic value for the destination is stored, ie it is checked whether the input dialog 'spell destination' was called or not by another input dialog, for example 'destination entry'. If there is no acoustic value for the destination, then in step 2075 the new location list is adopted as the fourth hypothesis list hypo.4 to continue the input dialog and the system jumps to step 2080. If query 2050 delivers a positive result, then in step 2060 a whole-word lexicon is generated from the new location list and loaded into the speech recognition device for whole-word recognition. There, in step 2070, the stored acoustic value of the desired destination is then compared with the full-word lexicon generated from the location list. The result of the speech recognition device in step 2075 is a fourth hypothesis list hypo.4 sorted according to the recognition probability. This fourth hypothesis list hypo.4 is adopted to continue the input dialog and step 2080 is carried out. In step 2080, the hypothesis with the greatest probability of recognition hypo.4.1 from the fourth hypothesis list hypo.4 is output with the query to the user whether the hypothesis hypo.4.1 corresponds to the desired destination or not. If the user answers the query with "No", the input dialog 'Select from list' is called up in step 2090. The result of the input dialog 'Select from list', which is a unique destination, is then used to carry out step 2100. In step 2100, the system determines an ambiguity list from the target file of all possible locations, in which all locations from the target file, which in the orthography of hypothesis hypo.4.1 from the fourth hypothesis list hypo.4 or the result of the input dialog 'Select from list' correspond, are included. Here it can happen that the entered destination appears several times in the list, for example due to the "Neustadt problem", and is therefore not clear. It is therefore checked in step 2110 whether the destination is unique or not. If the destination occurs only once, the input dialog is continued with step 2130. If the destination is not clear, the input dialog 'resolve ambiguity' is called up according to step 2120. The result of the input dialog 'resolve ambiguity' is then passed to step 2130. In step 2130, the recognized destination is verified with certain additional information, for example postcode and state, in which the user is asked whether the entered destination is correct or not. If the user answers the query 2130 with "No", the user is informed in step 2150 that the destination could not be found and the input dialog is ended. If the user answers the query 2130 with "yes", then the recognized destination is temporarily stored in step 2140 and branched for checking 350 according to FIG. 1.
Fig. 7 shows a schematic representation of an embodiment of the input dialog 'rough destination input'. In this input dialog, the user is asked to speak to a larger city known to him as a rough destination near the actual destination, the rough destination should be contained in the basic lexicon. After the activation of the input dialog 'rough destination' in step 3000, an input dialog 'input rough destination' is called in step 3010. For the input dialog 'Enter rough destination' in step 3010 there is almost the same procedure as for the input dialog 'Destination entry'. In contrast to step 1020 according to FIG. 2 or FIG. 3, the user is not prompted to enter the destination, but rather to enter the rough destination. After the part input dialog 'Enter rough destination' according to step 3010 has been carried out, the result of the input dialog 'Enter rough destination' is passed to step 3300 to continue the input dialog 'Rough destination input'. In step 3300, m place names are calculated in the vicinity of the rough destination which is passed as a result of the input dialog 'Enter rough destination'. The parameter m is dependent on the performance of the speech recognition device used and the transmission capacity between the at least one database and the speech recognition device. In the exemplary embodiment described, the parameter m is set to 1500. From these 1500 place names, which are part of the destination file, a fine destination lexicon is generated in step 3310 and loaded into the speech recognition device as a vocabulary. Then, in step 3320, the input dialog 'destination entry' according to FIG. 2 or 3 is carried out. The difference is that step 1010 is not carried out because the vocabulary required to recognize the destination is already loaded in the speech recognition device. To shorten the input dialog 'rough destination', it is also conceivable that in the input dialog 'enter rough destination' a shortened execution of the input dialog 'destination input' according to FIG. 3 is carried out. In the shortened version of the input dialog 'destination entry' according to FIG. 3, the input dialog 'spell destination' is dispensed with and after query 1075 according to FIG. 3 the input dialog 'spell destination' is not called up but the user is informed by voice output that the rough destination is not could be found and the input dialog 'rough destination entry' is ended. To speed up the input dialog, a rough target lexicon can be generated for the input dialog 'Enter rough target' and then loaded or activated as a vocabulary in the speech recognition device, the rough target lexicon instead of the 1000 largest locations of the basic lexicon only the 400 largest locations of the Federal Republic contains. In step 1150 according to FIG. 3, this results in a much shorter list of ambiguities in most applications. In addition, the input dialog 'resolve ambiguity' can be omitted. Instead of the input dialog 'Resolve ambiguity', the input dialog 'Select from list' is called up to let the user select the desired rough destination or only the two places with the largest population in the ambiguity list are output to the user for the final selection of the rough destination. If the desired rough destination was not found, the input dialog 'rough destination entry' is ended and the system returns to waiting state 0.
8 shows a schematic representation of an embodiment of the input dialog 'save address'. After activating the input dialog 'Save address' in step 7000, it is checked in step 7010 whether a destination address has already been entered. If the check 7010 yields a positive result, the method branches to the query 7020. If the check 7010 yields a negative result, the method branches to step 7030. In step 7020, a query is issued to the user as to whether he wants to save the current destination address or not. If the user answers the query with "yes", the method branches to step 7040. If the user answers the query with "No", the process branches to step 7030. In step 7030 an input dialog 'Enter address' is called. The input dialog 'Enter address' gives the user a query as to which of the four input dialogs 'Destination entry', 'Rough destination entry', 'Spell destination' or 'Indirect entry' the user wants to enter to enter the destination address to be saved. The desired input dialog can be selected in the same way as the input dialog 'Select from list'. It is also conceivable that the input dialog 'Enter address' immediately after activation activates one of the four input dialogues for voice input of a destination ('Destination entry', 'Spell destination', 'Rough destination entry' or 'Indirect entry') without asking the Output users. After the voice input of a destination, analogous to steps 350 and 400 according to FIG. 1 checks whether a street name can be entered and if so, whether a street name should be entered or not. If no street name is entered, step 450 according to FIG. 1 is carried out. The entered destination address is then temporarily saved as a result of the input dialog 'Enter address' to continue the input dialog 'Save address'. In step 7040, the user is asked to speak a keyword which is to be assigned to the entered destination address and under which the destination address is to be stored in the personal address register. The keyword spoken by the user in step 7050 is fed in step 7060 as an acoustic value 〈keyword〉 to the speaker-dependent additional speech recognizer and, if necessary, verified by repeated, in particular twice, speaking. Then in step 7070 the entered destination address is assigned to the keyword and the acoustic value 〈keyword〉 is stored in the personal address register. The user is informed by means of a voice output in step 7080 that the destination address has been stored under the entered 〈keyword〉. Then, in step 7090, the input dialog 'save address' is ended and a branch is made to waiting state 0 according to FIG. 1. Using the input dialog 'Retrieve address', which is not shown in detail, the stored destination address can be called up by voice input of the assigned keyword, which is fed to the speaker-dependent speech recognizer for recognition, and passed on to the route guidance system. If the keywords were included in the list of locations, spelling of the keyword is also possible if the keyword was not recognized as a whole-word entry.
Fig. 9 shows a schematic representation of an embodiment of the input dialog 'street input'. After activation of the input dialog 'street input' in step 5000, it is checked in step 5010 whether a destination has already been entered or not. If the check 5010 yields a positive result, the entered destination is used to continue the input dialog 'street entry' and a branch is made to step 5040. If the check 5010 yields a negative result, a query is output to the user in step 5020 as to whether the street is in the current location or not. The route guidance system determines the current location using a location method known from the prior art, for example using the Global Positioning System (GPS). If the user answers the query with "yes", the current location is used as the destination for the continuation of the input dialog 'street input' and a branch is made to step 5040. If the user answers the query with "No", the process branches to step 5030. In step 5030 an input dialog 'Enter destination' is called. The input dialog 'Enter destination', similar to the input dialog 'Enter address', queries the user which of the four input dialogs 'Destination entry', 'Rough destination entry', 'Spell destination' or 'Indirect entry' of the user to enter the destination address, which one want to save. The desired input dialog can be selected in the same way as the input dialog 'Select from list'. In addition, it is conceivable that the input dialog 'enter destination' immediately after activation, without issuing a query to the user, calls one of the four input dialogues for voice input of a destination. After the input dialog 'enter destination' has been carried out, the entered destination is used as the result for the continuation of the input dialog 'street input' and a branch is made to step 5040. In step 5040, it is checked whether the number of streets of the desired destination is greater than m or not. The parameter m is dependent on the type of speech recognition device. In the described embodiment, m = 1500 was set. If the number of streets of the desired destination is less than m, a street list of the desired destination is passed on to step 5060 in order to continue the input dialog 'street entry'. If the number of streets of the desired destination is greater than m, the process branches to step 5050. In step 5050, an input dialog 'Limit scope' is activated with the aim of reducing the scope of the street list to fewer than m entries of street names. For this purpose, the user can be asked to enter various selection criteria, for example the name of a district, the postal code or the first letter of the desired street name, by means of voice input. Analogous to the input dialog 'resolve ambiguity', the selection criteria can be combined as desired. The input dialog 'Limit perimeter' is ended when the perimeter has been reduced to m or fewer street names. As a result of the input dialog, the reduced street list is transferred to step 5060. In step 5060, a street lexicon is generated from the passed street list and loaded into the speech recognition device. In step 5070, the user is asked to speak the street name. The continuation of the input dialog 'street name' is now analogous to the input dialog 'destination entry' according to FIG. 3. The "street_1" spoken by the user (step 5080) is transferred to the speech recognition device as an acoustic value 〈street_1〉. As a result, the speech recognition device returns a fifth hypothesis list hypo.5 for continuing the input dialogue 'street input' (step 5090). In step 5100, the street name with the greatest probability of recognition hypo.5.1 is output to the user with the query whether 〈hypo.5.1〉 is the desired street name or not. If the user answers the query 5100 with "Yes", the method branches to step 5140. If the user answers the query 5100 with "No", the acoustic value 〈street_1〉 is stored in step 5110. Then in step 5120 the street name with the second largest recognition probability hypo.5.2 is output to the user with the query whether Abfrage hypo.5.2〉 is the desired street name or not. If the query 5120 is answered with "yes", the process branches to step 5140. If the user answers the query with "No", an input dialog 'Spell street' is called up in step 5130. Up to step 2100, the input dialog 'Spell street' is analogous to the input dialog 'Spell destination' and has already been dealt with in the description of FIG. 6. It is only necessary to replace the terms used in the description, destination, new town list, by the terms street name, new street list. Instead of step 2100 according to the input dialog 'spell destination', the input dialog 'spell street' is ended and the result of the input dialog 'spell street' is passed to step 5140 to continue the input dialog 'street entry'. In step 5140, the system determines from the street list, which contains all possible street names of the desired destination, an ambiguity list, in which all street names from the street list, which in the orthography of hypothesis hypo.5.1 or hypothesis hypo.5.2 from the fifth hypothesis list hypo.5 or the result of the input dialog 'Spell street'. In step 5150 it is checked whether the street name entered is unique or not. If the street name entered is unique, the process branches to step 5200. If the query yields a negative result, the method branches to step 5160. In step 5160 it is checked whether the ambiguity list contains more than k entries or not. If the ambiguity list comprises k or fewer entries, the process branches to step 5190. If the ambiguity list contains more than k entries, the method branches to step 5170. In step 5170, it is checked whether the ambiguity can be resolved by entering additional query criteria, for example the postcode or the district. If the check 5170 yields a positive result, an input dialog 'resolve road ambiguity' is called in step 5180. This input dialog runs analogously to the input dialog 'resolve ambiguity' according to FIG. 5. The postal code or the district can be entered as query criteria. The result of the input dialog 'resolve road ambiguity' is then passed on to step 5200 to continue the input dialog 'road input'. If the check 5170 yields a negative result, the method branches to step 5190. In step 5190 the input dialog 'Select from list' is activated and carried out. The result of the input dialog 'Select from list' is passed on to step 5200 in order to continue the input dialog 'Street input'. In step 5200, the input dialogue 'street input' is ended and the result, together with the desired destination, is transferred to step 500 according to FIG. 1 as the destination address.
10 shows a schematic diagram of a block diagram of an apparatus for carrying out the method according to the invention. As can be seen from FIG. 10, the device for carrying out the method according to the invention comprises a voice dialog system 1, a route guidance system 2 and an external database 4, on which, for example, the target file is stored. The speech dialogue system 1 comprises a speech recognition device 7 for recognizing and classifying speech utterances entered by a user via a microphone 5, a speech output device 10 which can deliver speech utterances to a user using a loudspeaker 6, a dialogue and sequence control 8 and an internal database 9, in which, for example, all voice commands are saved. The route guidance system 2 comprises an internal non-volatile memory 3, in which, for example, the basic lexicon is stored, and an optical display device 11. By means of the dialog and sequence control 8, data can be transmitted between the individual persons via corresponding connections 12, which can also be designed as a data bus Components of the device are replaced.
To clarify the input dialogs described, various input dialogs are shown in Tables 2 to 6.<tables id="tabl0002" num="0002"><img file="EP0865014A2_D0001.tif" /></tables><tables id="tabl0003" num="0003"><img file="EP0865014A2_D0002.tif" /></tables><tables id="tabl0004" num="0004"><img file="EP0865014A2_D0003.tif" /></tables><tables id="tabl0005" num="0005"><img file="EP0865014A2_D0004.tif" /></tables><tables id="tabl0006" num="0006"><table frame="all"><title>Table 2a</title><tgroup cols="7" colsep="1" rowsep="1"><colspec colnum="1" colname="col1" colwidth="22.50mm" /><colspec colnum="2" colname="col2" colwidth="22.50mm" /><colspec colnum="3" colname="col3" colwidth="22.50mm" /><colspec colnum="4" colname="col4" colwidth="22.50mm" /><colspec colnum="5" colname="col5" colwidth="22.50mm" /><colspec colnum="6" colname="col6" colwidth="22.50mm" /><colspec colnum="7" colname="col7" colwidth="22.50mm" /><thead valign="top"><row><entry namest="col1" nameend="col7" align="center">Ambiguity list</entry></row><row rowsep="1"><entry namest="col1" nameend="col1" align="center">Current -No.</entry><entry namest="col2" nameend="col2" align="left">POSTCODE.</entry><entry namest="col3" nameend="col3" align="center">Local area code</entry><entry namest="col4" nameend="col4" align="center">state</entry><entry namest="col5" nameend="col5" align="center">name suffix</entry><entry namest="col6" nameend="col6" align="left">circle</entry><entry namest="col7" nameend="col7" align="center">Inhabitant.</entry></row></thead><tbody valign="top"><row><entry namest="col1" nameend="col1" align="left">1:</entry><entry namest="col2" nameend="col2" align="left">66510-66589</entry><entry namest="col3" nameend="col3" align="left">06821</entry><entry namest="col4" nameend="col4" align="left">SL</entry><entry namest="col5" nameend="col5" /><entry namest="col6" nameend="col6" align="left">Neunkirchen</entry><entry namest="col7" nameend="col7" align="left">51863</entry></row><row><entry namest="col1" nameend="col1" align="left">2:</entry><entry namest="col2" nameend="col2" align="left">53819</entry><entry namest="col3" nameend="col3" align="left">02247</entry><entry namest="col4" nameend="col4" align="left">NW</entry><entry namest="col5" nameend="col5" align="left">Seelscheid</entry><entry namest="col6" nameend="col6" align="left">Rhein-SiegKreis</entry><entry namest="col7" nameend="col7" align="left">17360</entry></row><row><entry namest="col1" nameend="col1" align="left">3:</entry><entry namest="col2" nameend="col2" align="left">57290</entry><entry namest="col3" nameend="col3" align="left">02735</entry><entry namest="col4" nameend="col4" align="left">NW</entry><entry namest="col5" nameend="col5" /><entry namest="col6" nameend="col6" align="left">Victories</entry><entry namest="col7" nameend="col7" align="left">14804</entry></row><row><entry namest="col1" nameend="col1" /><entry namest="col2" nameend="col2" /><entry namest="col3" nameend="col3" /><entry namest="col4" nameend="col4" /><entry namest="col5" nameend="col5" /><entry namest="col6" nameend="col6" align="left">Wittgen</entry><entry namest="col7" nameend="col7" /></row><row><entry namest="col1" nameend="col1" /><entry namest="col2" nameend="col2" /><entry namest="col3" nameend="col3" /><entry namest="col4" nameend="col4" /><entry namest="col5" nameend="col5" /><entry namest="col6" nameend="col6" align="left">stone</entry><entry namest="col7" nameend="col7" /></row><row><entry namest="col1" nameend="col1" align="left">...</entry><entry namest="col2" nameend="col2" align="left">...</entry><entry namest="col3" nameend="col3" align="left">...</entry><entry namest="col4" nameend="col4" align="left">...</entry><entry namest="col5" nameend="col5" align="left">...</entry><entry namest="col6" nameend="col6" align="left">...</entry><entry namest="col7" nameend="col7" align="left">...</entry></row><row><entry namest="col1" nameend="col1" align="left">17:</entry><entry namest="col2" nameend="col2" align="left">83317</entry><entry namest="col3" nameend="col3" align="left">08666</entry><entry namest="col4" nameend="col4" align="left">BY</entry><entry namest="col5" nameend="col5" align="left">on the Teisenberg</entry><entry namest="col6" nameend="col6" align="left">Berchtesgadener Land</entry><entry namest="col7" nameend="col7" align="left">0</entry></row><row rowsep="1"><entry namest="col1" nameend="col1" align="left">18:</entry><entry namest="col2" nameend="col2" align="left">95466</entry><entry namest="col3" nameend="col3" align="left">09278</entry><entry namest="col4" nameend="col4" align="left">BY</entry><entry namest="col5" nameend="col5" /><entry namest="col6" nameend="col6" align="left">Bayreuth</entry><entry namest="col7" nameend="col7" align="left">0</entry></row></tbody></tgroup></table></tables><tables id="tabl0007" num="0007"><img file="EP0865014A2_D0005.tif" /></tables><tables id="tabl0008" num="0008"><img file="EP0865014A2_D0006.tif" /></tables><tables id="tabl0009" num="0009"><img file="EP0865014A2_D0007.tif" /></tables><tables id="tabl0010" num="0010"><table frame="all"><title>Table 4</title><tgroup cols="3" colsep="1" rowsep="1"><colspec colnum="1" colname="col1" colwidth="52.50mm" /><colspec colnum="2" colname="col2" colwidth="52.50mm" /><colspec colnum="3" colname="col3" colwidth="52.50mm" /><thead valign="top"><row><entry namest="col1" nameend="col3" align="center">Dialog example 'rough target entry' without ambiguity</entry></row><row rowsep="1"><entry namest="col1" nameend="col1" align="left">system</entry><entry namest="col2" nameend="col2" align="left">user</entry><entry namest="col3" nameend="col3" align="left">annotation</entry></row></thead><tbody valign="top"><row><entry namest="col1" nameend="col1" /><entry namest="col2" nameend="col2" align="left">presses the PTT button</entry><entry namest="col3" nameend="col3" align="left">The user activates the voice dialog system</entry></row><row><entry namest="col1" nameend="col1" align="left">Beep</entry><entry namest="col2" nameend="col2" /><entry namest="col3" nameend="col3" /></row><row><entry namest="col1" nameend="col1" /><entry namest="col2" nameend="col2" align="left">"Rough goal"</entry><entry namest="col3" nameend="col3" /></row><row><entry namest="col1" nameend="col1" align="left">"Please speak rough target"</entry><entry namest="col2" nameend="col2" /><entry namest="col3" nameend="col3" /></row><row><entry namest="col1" nameend="col1" /><entry namest="col2" nameend="col2" align="left">"Stuttgart"</entry><entry namest="col3" nameend="col3" /></row><row><entry namest="col1" nameend="col1" align="left">"Stuttgart is that correct?"</entry><entry namest="col2" nameend="col2" /><entry namest="col3" nameend="col3" align="left">Verification of the recognition result</entry></row><row><entry namest="col1" nameend="col1" /><entry namest="col2" nameend="col2" align="left">"Yes"</entry><entry namest="col3" nameend="col3" /></row><row><entry namest="col1" nameend="col1" align="left">"Loading dictionary for Stuttgart"</entry><entry namest="col2" nameend="col2" /><entry namest="col3" nameend="col3" align="left">The lexicon with 1500 place names in the Stuttgart area is generated and loaded into the speech recognition device. If necessary, the lexicon can also be pre-calculated based on the database. After loading, the desired destination can be entered.</entry></row><row><entry namest="col1" nameend="col1" align="left">"Please speak place name"</entry><entry namest="col2" nameend="col2" /><entry namest="col3" nameend="col3" /></row><row><entry namest="col1" nameend="col1" /><entry namest="col2" nameend="col2" align="left">"Wolfschlugen"</entry><entry namest="col3" nameend="col3" /></row><row><entry namest="col1" nameend="col1" align="left">"Wolfschlugen is that right?"</entry><entry namest="col2" nameend="col2" /><entry namest="col3" nameend="col3" align="left">Verification of the recognition result</entry></row><row><entry namest="col1" nameend="col1" /><entry namest="col2" nameend="col2" align="left">"Yes"</entry><entry namest="col3" nameend="col3" /></row><row><entry namest="col1" nameend="col1" align="left">"Route guidance is programmed to Wolfschlugen in the Esslingen district in Baden-Württemberg"</entry><entry namest="col2" nameend="col2" /><entry namest="col3" nameend="col3" align="left">Since Wolfschlugen is clear, no further queries need to be carried out and the route guidance system can load the possibly existing street names of Wolfschlugen if necessary.</entry></row><row rowsep="1"><entry namest="col1" nameend="col1" align="left">...</entry><entry namest="col2" nameend="col2" align="left">...</entry><entry namest="col3" nameend="col3" align="left">...</entry></row></tbody></tgroup></table></tables><tables id="tabl0011" num="0011"><img file="EP0865014A2_D0008.tif" /></tables><tables id="tabl0012" num="0012"><img file="EP0865014A2_D0009.tif" /></tables><tables id="tabl0013" num="0013"><img file="EP0865014A2_D0010.tif" /></tables><tables id="tabl0014" num="0014"><img file="EP0865014A2_D0011.tif" /></tables><tables id="tabl0015" num="0015"><img file="EP0865014A2_D0012.tif" /></tables><tables id="tabl0016" num="0016"><img file="EP0865014A2_D0013.tif" /></tables>
24 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| EP1071075A2 | Cited by | European Patent Office (EPO) | Search report |
| EP1058236A2 | Cited by | European Patent Office (EPO) | Search report |
| EP1058236A3 | Cited by | European Patent Office (EPO) | Search report |
| WO0205264A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| US6885990B1 | Cited by | United States of America | Applicant |
| EP1071075A3 | Cited by | European Patent Office (EPO) | Search report |
| CN103376115A | Cited by | China | Search report |
| WO02077974A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| EP1273886A1 | Cited by | European Patent Office (EPO) | Search report |
| US7039629B1 | Cited by | United States of America | Applicant |
| EP0346483A1 | Cites | European Patent Office (EPO) | Search report |
| US4866778A | Cites | United States of America | Search report |
| JPH0764480A | Cites | Japan | Examiner |
| JPH08202386A | Cites | Japan | Examiner |
11 members in 6 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 19709518 | Germany | A | |
| 19709518 | Germany | A | |
| 19709518 | Germany | – | |
| 19709518 | – | – | – |
| DE1997109518 | – | – | – |
Members11
| Document | Office | Kind | |
|---|---|---|---|
| DE19709518C1 | Germany | C1 | |
| CA2231851A1 | Canada | A1 | |
| EP0865014A2This record | European Patent Office (EPO) | A2 | |
| JPH1115493A | Japan | A | |
| EP0865014A3 | European Patent Office (EPO) | A3 | |
| US6230132B1 | United States of America | B1 | |
| CA2231851C | Canada | C | |
| EP0865014B1 | European Patent Office (EPO) | B1 | |
| AT303580T | Austria | T | |
| DE59813027D1 | Germany | D1 | |
| DE19709518C5 | Germany | C5 |
47 legal events, as 5 offices reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | Office | |
|---|---|---|---|
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Patent expired after termination of 20 yearsExpiredPE20 | PE20 | GB | |
| Expiry of rightR071 | R071 | DE | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Fee paymentPLFP | PLFP | FR | |
| Fee paymentPLFP | PLFP | FR | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Patent ceasedCeasedPL | PL | CH | |
| No opposition filedOpposition26N | 26N | EP | |
| No opposition filed within time limitOppositionORIGINAL CODE: 0009261PLBE | PLBE | EP | |
| Information on the status of an ep patent application or granted ep patentGrantedSTATUS: NO OPPOSITION FILED WITHIN TIME LIMITSTAA | STAA | EP | |
| Fr: translation filedET | ET | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Corresponds to:REF | REF | EP | |
| Gb: translation of ep patent filed (gb section 77(6)(a)/1977)GBT | GBT | EP | |
| Designated contracting statesAK | AK | EP | |
| European patent takes effect as a national patent in ch/liEP | EP | CH | |
| European patent grantedGrantedNOT ENGLISHFG4D | FG4D | GB | |
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Party data changed (applicant data changed or rights of an application transferred)RAP1 | RAP1 | EP | |
| Information provided on ipc code assigned before grantRIC1 | RIC1 | EP | |
| Information provided on ipc code assigned before grantRIC1 | RIC1 | EP | |
| (expected) grantORIGINAL CODE: 0009210GRAA | GRAA | EP | |
| Grant fee paidORIGINAL CODE: EPIDOSNIGR3GRAS | GRAS | EP | |
| Despatch of communication of intention to grant a patentORIGINAL CODE: EPIDOSNIGR1GRAP | GRAP | EP | |
| First examination report despatched17Q | 17Q | EP | |
| Party data changed (applicant data changed or rights of an application transferred)RAP1 | RAP1 | EP | |
| Designation fees paidAT CH DE ES FI FR GB IT LI SEAKX | AKX | EP | |
| Party data changed (applicant data changed or rights of an application transferred)RAP1 | RAP1 | EP | |
| Request for examination filed17P | 17P | EP | |
| Designated contracting statesAK | AK | EP | |
| Request for extension of the european patentAL;LT;LV;MK;RO;SIAX | AX | EP | |
| Information provided on ipc code assigned before grant7G 01C 21/20 A, 7G 10L 3/00 BRIC1 | RIC1 | EP | |
| Search report despatchedORIGINAL CODE: 0009013PUAL | PUAL | EP | |
| Party data changed (applicant data changed or rights of an application transferred)RAP1 | RAP1 | EP | |
| Party data changed (applicant data changed or rights of an application transferred)RAP1 | RAP1 | EP | |
| Party data changed (applicant data changed or rights of an application transferred)RAP1 | RAP1 | EP | |
| Party data changed (applicant data changed or rights of an application transferred)RAP1 | RAP1 | EP | |
| Designated contracting statesAK | AK | EP | |
| Request for extension of the european patentAL;LT;LV;MK;RO;SIAX | AX | EP | |
| Public reference made under article 153(3) epc to a published international application that has entered the european phaseORIGINAL CODE: 0009012PUAI | PUAI | EP |
Numbers
- Publication
- 0865014
- Publication, DOCDB
- 0865014
- Publication, EPODOC
- EP0865014
- Application
- 98103662
- Application, DOCDB
- 98103662
- Application, EPODOC
- EP19980103662
Titles3
- German
- Verfahren und Vorrichtung zur Spracheingabe einer Zieladresse in ein Zielführungssystem im Echtzeitbetrieb
- English
- Method and device to enter by speech an address of destination in a navigation system in real time
- French
- Procédé et appareil d'indication de destination par commande vocale à un système de navigation en temps réel
Classification
- CPC, 2
- G01C21/3608
- G10L15/22
- IPC, 11
- G01C21 00
- B60R16 02
- G01C21 36
- G06F3 16
- G08G1 0969
- G10L15 00
- G10L15 18
- G10L15 22
- G10L15 24
- G10L15 28
- G10L15 32
Designated states24
- Contracting states, 18
- Austria
- Belgium
- Switzerland
- Germany
- Denmark
- Spain
- Finland
- France
- United Kingdom
- Greece
- Ireland
- Italy
- Liechtenstein
- Luxembourg
- Monaco
- Netherlands (Kingdom of the)
- Portugal
- Sweden
- Extension states, 6
- Albania
- Lithuania
- Latvia
- North Macedonia
- Romania
- Slovenia